# Request routing: give routine requests a path without generation

This page reports one fixed paired recording over **3,000 measured requests**:
1,000 each from CLINC150, BANKING77 and MASSIVE en-US. It compares the recorded
SIE cascade with GPT-6 Luna and Claude Sonnet 5 on the same selected cases.
The separate warm local development check supports fast handler selection on
18 requests that skipped generation. Neither recording establishes general
quality parity, population or rival-service latency, cold or batch embedding
stability, managed availability, observed settlement or self-hosting margin.

## Publication and reproduction

All three evidence folders are published at immutable `superlinked/sie-task-evidence`
revision `406fe8eb62efd43cf2d2af6fbdaadd4b7dd9b68c`:
[primary recording](https://huggingface.co/datasets/superlinked/sie-task-evidence/tree/406fe8eb62efd43cf2d2af6fbdaadd4b7dd9b68c/request-routing/20261001-primary3000-v1),
[serving check](https://huggingface.co/datasets/superlinked/sie-task-evidence/tree/406fe8eb62efd43cf2d2af6fbdaadd4b7dd9b68c/request-routing/20261001-serving-v2),
and [paired local latency](https://huggingface.co/datasets/superlinked/sie-task-evidence/tree/406fe8eb62efd43cf2d2af6fbdaadd4b7dd9b68c/request-routing/20261001-paired-latency-v1).
The [public cascade example](https://github.com/superlinked/sie/tree/99b523b17962167b19e321a9c85c593ff3243574/examples/request-routing)
is merged and pinned at `99b523b17962167b19e321a9c85c593ff3243574`.
The example reproduces the primary recording; its original evidence revision
contains the same primary bytes. The additional latency bundle is a separate
development check, not a rival-service latency claim.
The visible migration uses the example's unchanged `load_assets`,
`score_front` and `routing_body` functions. Both sides read the same BANKING77
action list and the same identity-verification request. Claude receives all
route examples; SIE performs one embedding call and invokes the generator only
below the original 0.29 threshold. The snippet does not refit the classifier
or reproduce a benchmark score. Replay commands remain a separate optional
way to obtain the frozen assets and check the study offline.

The source bundle's `manifest.json` SHA-256 is
`bc9bd3ad71002436506ef13ee2950cd1760187c053af0951e94196e3bd4efe0f`.
The compact website evidence is generated by `import-request-routing.mjs`;
it verifies every manifest file and size before parsing records or results.
The public example's `score.py` independently verified the complete bundle,
recomputed the frozen classifier and paired statistics, and passed all 3,000
cases. The website does not refit a model or implement a second bootstrap.

Before publication, run the importer with `--require-publication-pins`, verify
the public example link without authentication, and verify that a nonexistent
path at the same commit fails. Public example and evidence publication precede
the website merge. Local draft rendering authorizes none of those actions.

## Warm local handler selection

The page leads **52 ms** median time to a validated route for the fixed **18
warm local development pairs that skipped generation**. The corresponding
always-generation workflow took **1,319 ms**. Both paths freshly invoke the
encoder and classifier; always generation uses that front to construct the
same Qwen3.8 27B prompt. This is a complete local routing-workflow comparison,
not a standalone generator or rival-provider timing.

| Path on the fast subset | Median | Paired-bootstrap 95% interval |
| --- | ---: | --- |
| SIE cascade, front only | 52 ms | 49.70 to 59.25 ms |
| Same Qwen3.8 27B workflow, always generation | 1,319 ms | 1299.79 to 1417.09 ms |

Unrounded medians are 51.709479499976396 and 1318.8686520000203 ms. The ratio
of medians is 25.505355396211684 (95% interval 22.29378708557215 to
27.238835745675587); the page makes no unqualified speed-ratio claim.
All 29 pairs completed, including eleven that generated. The all-pair cascade
median is 60 ms (51.82 to 1094.98 ms); always generation is 1,318 ms. Its ratio
interval is wide, 1.2055269282825753 to 25.43064208215262. Fast-subset latency
is not population coverage, and the separate 76.8% weighted generation
avoidance must not be assigned this timing.

Two of eighteen fast-path routes differ from always generation. The development
inputs have no accuracy labels; this latency comparison makes no quality claim.
The fixed quality study and its full uncertainty remain below.

The sanitized [latency method](https://huggingface.co/datasets/superlinked/sie-task-evidence/blob/406fe8eb62efd43cf2d2af6fbdaadd4b7dd9b68c/request-routing/20261001-paired-latency-v1/METHOD.md)
defines the timer, selection and paired bootstrap. Manifest SHA-256:
`54d90fb1901143ac9a3747a971770c51b6aad7c0cfdcdd1ace04f3b7d888d0c5`.
The importer verifies all three files before reading summary statistics and
projects those aggregates, without a second bootstrap or timing calculation.
Intervals use 2,000 paired percentile resamples, NumPy 2.4.6, seed 20261001,
linear percentiles and a fixed fast subset. They describe variation across
this development sample, not independent deployments or production traffic.

The draw used one H100 80GB, eight CPU cores and 64 GiB, with both models
resident on SGLang 0.5.20 and CUDA 13. Requests are sequential, encoder fanout
one, with alternating arm order and no retries. The timer includes SDK and
local loopback, frozen asset loading, front computation, prompt/tokenizer work
and the local before-POST journal flush. It excludes model startup, evidence
Volume commits, cleanup and network travel to the user. The excluded encoder
and generator cold warmups took 9.302 s and 2.520 s respectively.

Normal prefix caching remains enabled. The two arms repeat the same prompt
and share route rosters and training examples; alternating order does not
remove cache effects. Warm vector/front/top10/threshold/body agreement and
escalated-route agreement passed, as did model brackets and owned cleanup.
The warm single-item agreement does not clear the earlier cold/batch stability
limitation. No first-request, remote service, tail latency, sustained
mixed-traffic capacity or maximum-throughput claim is made.

## Separate listed token-fee comparison

| Router | Route accuracy | 95% interval | Modeled USD per 1,000 requests |
| --- | ---: | --- | ---: |
| SIE cascade | 89.57% | 88.43% to 90.67% | $0.1010 upper bound |
| GPT-6 Luna | 90.57% | 89.47% to 91.57% | $0.1307 theoretical lower bound |
| Claude Sonnet 5 | 89.00% | 87.87% to 90.10% | $3.6972 theoretical lower bound |

The exact bounds are $0.10099870980252592, $0.1307053131378542 and
$3.697167107020075 per 1,000 respectively. The cascade's modeled bound is
**97.27% lower than Sonnet's** and **22.73% lower than Luna's**. Display
formatting happens once in the page data module. The cost model includes the
retained unresolved embedding attempt allowance; failure or unknown usage is
not priced as zero.

Paired differences are SIE minus comparator, in percentage points:

| Comparator | Difference | Paired 95% interval |
| --- | ---: | --- |
| Claude Sonnet 5 | +0.57 | -0.33 to +1.43 |
| GPT-6 Luna | -1.00 | -1.80 to -0.20 |

The registered joint quality rule **failed** (`R2_both_pass=false`). Luna's
paired lower bound fails the required one-point margin. Descriptive accuracy
and uncertainty remain published; a Sonnet point estimate above the comparator
does not prove superiority or general noninferiority. Per-dataset estimates
and intervals are shown on the page. All selected cases and strict failures
remain in the aggregates.

Generation avoidance is **76.8%**, using original population weights.
The unweighted selected sample generated 673 of 3,000 requests. These two
figures have different weighting and must not be substituted for each other.
CLINC150 contains 818 in-scope and 182 out-of-scope measured cases. Out-of-scope
recall is 92.31% (87.91% to 96.15%); population-weighted precision is 95.45%
(92.30% to 98.26%), with 176 predicted out-of-scope cases in the sample.

## Fixed method and limitations

Qwen3 Embedding 4B, revision
`5cf2132abc99cad020ac570b19d031efec650f2b`, uses the fixed query instruction
“Given a user request, find the requests that ask for the same action”.
Recorded dense vectors retain the float16 archive roundtrip, then float32
normalization. The front uses the saved logistic-regression coefficients,
C=10, original max_iter=2,000 and frozen route order, without refitting.
Maximum score ≥ 0.29 returns the top route; this score is not a calibrated
probability of correctness.

Below-threshold requests use Qwen3.8 27B FP8, revision
`017b9c7af6b5689d5dd426a76e0bc077eb5ca20a`, with all route names and up to
ten examples for each of the ten shortlisted routes. MASSIVE's “cooking query”
route has four. The schema returns exactly one route; CLINC150 includes “none
of these”. Max output 64, temperature 0, top_p 1, presence_penalty 0 and thinking
disabled remain fixed. The generator context is 8,192 tokens.

The archived rival recordings provide all route names and available training
examples, capped at ten per route. Their prompt shapes differ from the
shortlisted cascade prompt. The same prospectively selected cases are paired
under the recorded source-order extraction rule. No new rival inference was
used to score this selected subset.

Route accuracy gives each dataset equal weight; CLINC150 uses 9/11 in-scope
and 2/11 out-of-scope weights. Paired 95% intervals use 2,000 shared-ID stratified
bootstrap draws, seed 20260930, across four strata, without finite-population
correction. These intervals describe sampled-case uncertainty conditional on
this serving draw, not cross-session or repeated-request stability.

The self-hosted recording used an H100 80GB, SGLang 0.5.20, Transformers 5.12.1,
PyTorch 2.13.0 and CUDA 13.0, collecting models sequentially. Public runtime
source revision: `ee7ca75d733f7a959820ddc43acb35dcf23bed21`.
Embedding collection was interrupted; sixteen successful chunks were retained,
one unknown attempt received a prospectively bounded identical-body replay,
and its original missing model-after observation remains unproven. The
continuation's own model brackets matched. This limitation remains part of the
recording and its conservative attempt cost bound.

The separate development serving checks do not replace this fixed quality
recording. The latency result is qualified warm local workflow time; no rival
speed, population latency, maximum throughput or sustained mixed-traffic
claim is made.

## Self-host compute scenario and serving scope

The page leads with the fixed study's generation avoidance and control over
the action list. The qualified warm local timing remains in its own
methodology disclosure, alongside the full accuracy and cost definitions.
The supporting default-placement scenario is **about $0.19 compute per 1,000
routes**, at concurrency four and **80% assumed utilization**, roughly **95%**
below Sonnet's theoretical listed lower bound. Luna's theoretical listed bound
is lower than this compute estimate, and its measured route accuracy is higher.
The token-fee model above is separate; it establishes neither hosted availability
nor a demonstrated serving margin.

The engineering draw completed **58 development calls**, 29 per model, with
both models resident on one H100 80GB, eight CPU cores and 64 GiB memory.
Each model used one concurrency-one warm-up, twelve measured requests at
concurrency one, four concurrency-four warm-ups and twelve measured requests
at concurrency four. Endpoint groups ran sequentially; no held-out quality
case was used in this engineering draw. Request and group timers exclude
evidence-volume commits. Per-request time includes direct SDK and loopback
transport, queueing, service and usage capture; the group also includes local
executor and validation work. Normal prefix caching remained enabled.

The full sanitized serving bundle's manifest SHA-256 is
`e871db1735596081cc54279b9cc9db35bbae52e012b71f639723a0f17cf028f0`.
The importer verifies METHOD, observations and economics bytes before parsing
them and binds the generation share and comparator price bounds to the fixed
quality recording. Only economics and the displayed completion/hardware facts
are projected. All three evidence folders share the immutable dataset revision above.

The scenario consumes the separately observed embedding and generation group
rates, weighted by the 23.191317452887983% population generation share. It
assumes service demand adds without interference, the development prompt mix
represents routed requests and the chosen utilization can be sustained.
Concurrent mixed traffic was not measured. The exact C4/default/80% estimate is
$0.1856042226499419 per 1,000 routes; display rounds it to about $0.19. The page
also exposes C1/C4, 50%/80%/100% utilization and 1/1.15/1.75 placement sensitivities.
Placement factors apply to the full resource rate; actual placement was not
established and non-preemptible execution is excluded.

The [dated Modal resource price](https://modal.com/pricing) totals $0.00134388
per second for the H100, eight CPU cores and 64 GiB. Reserved CPU and memory
are included. Startup, shutdown, storage, egress, additional gateway hosts,
subscriptions, operations and provider margins are excluded. This is a compute
planning scenario, not an invoice or API offer. Listed token fees have no
demonstrated serving margin in this check. Sampled GPU allocation reached
79.23221226351476% with 16,459 MiB minimum observed free memory; one-second NVML
sampling does not prove sub-second workspace peaks. Owned workers were reaped.


## Price basis, retrieved 1 October 2026

SIE's modeled denomination is 100,000 credits per USD. Embedding is 3/500
credit per input token; generation is 1/40 input and 1/5 output credit per
token. Embeddings round up per individual request, including extra-attempt
allowances. Generation input and output round up together per attempt.
Successful calls use validated tokens; failed generation attempts retain the
full registered prompt/output bounds. Actual embedding collection batched up
to 64 cases, so individual-request ceilings are not observed batch settlement.

Rival bounds allow the cheaper theoretical choice of 50% Batch pricing or ideal
input cache reads with full output pricing. This is a deliberately favorable
rival floor, not a measured real-time invoice or a promise of cache coverage.
The page labels this material comparison setting next to the cost chart.
GPU expense, idle time, ancillary costs and utilization are separate from API
listed-rate bounds; no operating margin follows from the comparison.

- [OpenAI pricing](https://developers.openai.com/api/docs/pricing)
- [Anthropic pricing](https://platform.claude.com/docs/en/about-claude/pricing)
- [Anthropic prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching)

## Sources, licences and display transformations

- [CLINC150](https://github.com/clinc/oos-eval), Larson et al., 2019, CC BY 3.0:
  150 in-scope labels plus “none of these”.
- [BANKING77](https://github.com/PolyAI-LDN/task-specific-datasets),
  Casanueva et al., 2020 / PolyAI, CC BY 4.0: 77 banking intents.
- [MASSIVE](https://huggingface.co/datasets/AmazonScience/massive),
  Amazon Science, CC BY 4.0: 60 assistant intents, en-US partition.

The hero is the first source-ordered short BANKING77 request routed correctly
by the front. The three retained cases are the first short CLINC150 front success,
BANKING77 generation success, and MASSIVE cascade miss answered correctly by
Luna. They illustrate paths and a failure; their selection never changes the
aggregate. The page opens the CLINC150 same-request comparison and shows the
BANKING77 fallback input beside its returned route. The MASSIVE miss remains
in the pinned recording. Text, route names, rival answers and supplied examples are verbatim.
Only numeric display formatting and the card layout change. Required licence
credits remain beside displayed texts. There are no fabricated probabilities,
latency figures, successful calls, route examples or quality passes.
