Route routine requests without generation
For support queues and agent tools. Select a handler from your action list with Qwen3 Embedding 4B; Qwen3.8 27B handles lower-score requests. Run both on your infrastructure.
“I am having difficulties to verify my identity.”
- Qwen3 Embedding 4B → fixed classifier
- Score 0.599
Score ≥ 0.29 · selected
unable to verify identity
No generation call
Score < 0.29 · alternate
Qwen3.8 27B
Selects one route from your action list
Your app dispatches the chosen handler.
76.8%
of requests skipped the generator in the recorded study
The fixed classifier returns the route above a 0.29 score. Below it, Qwen chooses from the same action list.
Get the same handler with fewer model calls
“reserve a meeting room for 5pm on friday”
Expected: schedule meeting · classifier score 0.484
Examples for “schedule meeting”
- can i book a meeting room from 2:00 pm to 3:00 pm
- can you schedule a meeting with damon for 1
- i want to know if there is meeting room available at 8
- are there meeting rooms available between 7-9
- are there meeting rooms available between 11-12
- i want to find out how do i schedule a meeting
- show me how do i schedule a meeting
- help me schedule a meeting
- schedule a meeting for me
- can you schedule a meeting with matt for 3
- SIE cascade · no generation
- ✓ schedule meeting
- OpenAI GPT-6 Luna · generation
- schedule meeting
- Anthropic Claude Sonnet 5 · generation
- schedule meeting
About $0.19 of compute for 1,000 routes
1 October 2026
CLINC150, BANKING77 and MASSIVE, 1,000 cases each. Correct-route accuracy, with fixed routes and threshold.
SIE: self-host compute at 80% utilization. Rivals: theoretical listed lower bounds, including Batch or ideal cache reads. These are cost scenarios, not a managed bill.
Use one action list for both paths
client.encode(
"Qwen/Qwen3-Embedding-4B",
[{"text": "How can I top up my account with a card?"}],
output_types=["dense"],
... # frozen settings in the full recipe
)Embedding → fixed classifier → Qwen below 0.29. This saved request took the model path.
View complete request
import json
import os
from run import (
ENCODER, GENERATOR, INSTRUCTION, THRESHOLD,
NoRetrySIEClient, routing_body, score_front,
)
from pathlib import Path
from run import load_assets
assets, fronts = load_assets(Path("evidence"))
request = "How can I top up my account with a card?"
labels = assets["banking77"]["labels"]
row = {
"dataset": "banking77", "id": "my-request",
"text": request,
}
with NoRetrySIEClient(
"http://localhost:8080",
api_key=os.environ.get("SIE_API_KEY"),
timeout_s=45, max_connections=1,
) as client:
encoded = client.encode(
ENCODER, [{"text": request}],
output_types=["dense"], output_dtype="float32",
instruction=INSTRUCTION, is_query=True,
wait_for_capacity=False, max_oom_retries=0,
)
front = score_front(
encoded[0]["dense"], fronts["banking77"]
)
if front["conf"] >= THRESHOLD:
selected = front["pred"]
else:
reply = client.chat_completions(
model=GENERATOR,
**routing_body(row, front, assets),
wait_for_capacity=False, max_oom_retries=0,
)
selected = json.loads(
reply["choices"][0]["message"]["content"]
)["route"]
if selected not in labels:
raise ValueError("Unknown route")
print(selected){
"route": "topping up by card"
}- Classifier score 0.221 is below 0.29
- Qwen selects one allowed route; your app dispatches it
Switch to SIE in 5 minutes
Get the frozen route assets and replay the study
# Run inside the pinned request-routing example
python fetch.py --revision "406fe8eb62efd43cf2d2af6fbdaadd4b7dd9b68c"
python score.py --evidence evidence The public recipe supplies the route assets and NumPy classifier. These commands replay saved outputs offline. The routing call uses Qwen3 Embedding 4B and Qwen3.8 27B on your selected deployment.
Give your requests a fast path
Managed Cloud
Full compute toolkit for your agents with zero ops.
- No idle GPUs, pay for what you use
- Fits your stack: SDK, API, CLI, MCP
- Zero lock-in, self-host the same stack
- SOC 2 Type 2, US or EU data residency
Qwen3 Embedding 4B
Encoder component at its published token rate
Qwen3.8 27B
Generator component; each credit figure uses the full balance
no credit card required
Self-host with K8s
Easy & scalable deployment in your own cloud.
- Terraform to your cloud in minutes
- Apache-2.0, same engine as Cloud
- Scales to zero, no bill between jobs
- Per-tenant pools, no noisy neighbors
Deploy SIE to our AWS account with the superlinked/sie/aws Terraform module. Docs: superlinked.com/docs/deployment Serve Qwen/Qwen3-Embedding-4B and Qwen/Qwen3.8-27B-FP8 using the public model configurations.Deploy SIE to our GCP project with the superlinked/sie/google Terraform module. Docs: superlinked.com/docs/deployment Serve Qwen/Qwen3-Embedding-4B and Qwen/Qwen3.8-27B-FP8 using the public model configurations.Deploy SIE to our Azure AKS cluster via helm install. Requirements: superlinked.com/docs/deployment Serve Qwen/Qwen3-Embedding-4B and Qwen/Qwen3.8-27B-FP8 using the public model configurations. Run locally
Serve both routing models on your NVIDIA GPUs.
- NVIDIA GPUs
- One container, no cluster
- Qwen3 Embedding 4B + Qwen3.8 27B, fully offline
- Same SDK and IDs, no code changes
docker run --gpus all -p 8080:8080 \
ghcr.io/superlinked/sie-server:v0.9.0-cuda13-sglang-cu130 \
serve --host 0.0.0.0 --port 8080 \
--models Qwen/Qwen3-Embedding-4B,Qwen/Qwen3.8-27B-FP8