Why did we open-source our inference engine? Read the post

Route routine requests without generation

For support queues and agent tools. Select a handler from your action list with Qwen3 Embedding 4B; Qwen3.8 27B handles lower-score requests. Run both on your infrastructure.

Customer request
“I am having difficulties to verify my identity.”
Selected handler without generation
Qwen3 Embedding 4B → fixed classifier
Score 0.599

Score ≥ 0.29 · selected

unable to verify identity

No generation call

Score < 0.29 · alternate

Qwen3.8 27B

Selects one route from your action list

Your app dispatches the chosen handler.

Recorded request: BANKING77 · PolyAI · CC BY 4.0

Generation is optional

76.8%

of requests skipped the generator in the recorded study

Keep a model for uncertain requests

The fixed classifier returns the route above a 0.29 score. Below it, Qwen chooses from the same action list.

Get the same handler with fewer model calls

Return the route without generation
“reserve a meeting room for 5pm on friday”

Expected: schedule meeting · classifier score 0.484

Examples for “schedule meeting”
  • can i book a meeting room from 2:00 pm to 3:00 pm
  • can you schedule a meeting with damon for 1
  • i want to know if there is meeting room available at 8
  • are there meeting rooms available between 7-9
  • are there meeting rooms available between 11-12
  • i want to find out how do i schedule a meeting
  • show me how do i schedule a meeting
  • help me schedule a meeting
  • schedule a meeting for me
  • can you schedule a meeting with matt for 3
The same route with fewer model calls
SIE cascade · no generation
✓ schedule meeting
OpenAI GPT-6 Luna · generation
schedule meeting
Anthropic Claude Sonnet 5 · generation
schedule meeting

CLINC150 · Larson et al., CC BY 3.0

About $0.19 of compute for 1,000 routes

Modeled cost per 1,000 routes
USD / 1,000 routes ↓
GPT-6 Luna
$0.1307 GPT-6 Luna
SIE cascade
$0.19 SIE cascade
Claude Sonnet 5
Claude Sonnet 5 $3.6972
Requests sent to the right route
Correct routes (%) ↑
GPT-6 Luna
GPT-6 Luna 90.57%
SIE cascade
SIE cascade 89.57%
Claude Sonnet 5
Claude Sonnet 5 89.00%
3,000 paired requests

1 October 2026
CLINC150, BANKING77 and MASSIVE, 1,000 cases each. Correct-route accuracy, with fixed routes and threshold.

SIE: self-host compute at 80% utilization. Rivals: theoretical listed lower bounds, including Batch or ideal cache reads. These are cost scenarios, not a managed bill.

Use one action list for both paths

Workflow input
client.encode(
    "Qwen/Qwen3-Embedding-4B",
    [{"text": "How can I top up my account with a card?"}],
    output_types=["dense"],
    ... # frozen settings in the full recipe
)

Embedding → fixed classifier → Qwen below 0.29. This saved request took the model path.

View complete request
import json
import os
from run import (
    ENCODER, GENERATOR, INSTRUCTION, THRESHOLD,
    NoRetrySIEClient, routing_body, score_front,
)
from pathlib import Path
from run import load_assets

assets, fronts = load_assets(Path("evidence"))
request = "How can I top up my account with a card?"
labels = assets["banking77"]["labels"]

row = {
    "dataset": "banking77", "id": "my-request",
    "text": request,
}
with NoRetrySIEClient(
    "http://localhost:8080",
    api_key=os.environ.get("SIE_API_KEY"),
    timeout_s=45, max_connections=1,
) as client:
    encoded = client.encode(
        ENCODER, [{"text": request}],
        output_types=["dense"], output_dtype="float32",
        instruction=INSTRUCTION, is_query=True,
        wait_for_capacity=False, max_oom_retries=0,
    )
    front = score_front(
        encoded[0]["dense"], fronts["banking77"]
    )
    if front["conf"] >= THRESHOLD:
        selected = front["pred"]
    else:
        reply = client.chat_completions(
            model=GENERATOR,
            **routing_body(row, front, assets),
            wait_for_capacity=False, max_oom_retries=0,
        )
        selected = json.loads(
            reply["choices"][0]["message"]["content"]
        )["route"]
if selected not in labels:
    raise ValueError("Unknown route")
print(selected)
Output
{
  "route": "topping up by card"
}
  • Classifier score 0.221 is below 0.29
  • Qwen selects one allowed route; your app dispatches it

Switch to SIE in 5 minutes

# Existing route list and Route schema stay in your app.
reply = anthropic.messages.parse(
    model="claude-sonnet-5",
    max_tokens=64,
    system=("Route the request. Answer with exactly one "
            "of the routes, copied exactly."),
    messages=[{
        "role": "user",
        "content": "Routes, each with example requests:\n"
        + routes + "\n\nRequest:\n" + request,
    }],
    output_format=Route,
)
selected = reply.parsed_output.route
print(selected)
SIE
import os
from run import NoRetrySIEClient, route

# Frozen route assets loaded by the full recipe.
with NoRetrySIEClient(
    "http://localhost:8080",
    api_key=os.environ.get("SIE_API_KEY"),
    timeout_s=45, max_connections=1,
) as client:
    result = route(client, {
        "dataset": "banking77",
        "id": "my-request", "text": request,
    }, assets, fronts)
if result["error"] is not None:
    raise RuntimeError(result["error"])
selected = result["answer"]
print(selected)
Get the frozen route assets and replay the study
# Run inside the pinned request-routing example
python fetch.py --revision "406fe8eb62efd43cf2d2af6fbdaadd4b7dd9b68c"
python score.py --evidence evidence

The public recipe supplies the route assets and NumPy classifier. These commands replay saved outputs offline. The routing call uses Qwen3 Embedding 4B and Qwen3.8 27B on your selected deployment.

Give your requests a fast path

Self-host with K8s

Easy & scalable deployment in your own cloud.

  • Terraform to your cloud in minutes
  • Apache-2.0, same engine as Cloud
  • Scales to zero, no bill between jobs
  • Per-tenant pools, no noisy neighbors
Agent prompt
Deploy SIE to our AWS account with the superlinked/sie/aws Terraform module. Docs: superlinked.com/docs/deployment Serve Qwen/Qwen3-Embedding-4B and Qwen/Qwen3.8-27B-FP8 using the public model configurations.Deploy SIE to our GCP project with the superlinked/sie/google Terraform module. Docs: superlinked.com/docs/deployment Serve Qwen/Qwen3-Embedding-4B and Qwen/Qwen3.8-27B-FP8 using the public model configurations.Deploy SIE to our Azure AKS cluster via helm install. Requirements: superlinked.com/docs/deployment Serve Qwen/Qwen3-Embedding-4B and Qwen/Qwen3.8-27B-FP8 using the public model configurations.
Deploy guide

Run locally

Serve both routing models on your NVIDIA GPUs.

  • NVIDIA GPUs
  • One container, no cluster
  • Qwen3 Embedding 4B + Qwen3.8 27B, fully offline
  • Same SDK and IDs, no code changes
docker run --gpus all -p 8080:8080 \ ghcr.io/superlinked/sie-server:v0.9.0-cuda13-sglang-cu130 \ serve --host 0.0.0.0 --port 8080 \ --models Qwen/Qwen3-Embedding-4B,Qwen/Qwen3.8-27B-FP8
Quickstart

Contact us

Tell us about your use case and we'll get back to you shortly.

Apply for an inference grant

Free capacity on our hosted cluster for selected projects. Tell us what you run and we reply by email.