# Search page sources

The /search page compares SIE against OpenAI text-embedding-3-large, the most
widely deployed embedding API. This file gives the provenance of everything the
page shows and the complete results of the study behind it, every model
included.

## What the page shows

**The hero and the pain cases.** Questions from the held-out look-alike test
below, chosen by one rule from the recorded rankings: SIE Qwen3 Embedding 4B and
Qwen3 Embedding 8B rank the answer first, and OpenAI text-embedding-3-large
ranks the question's named look-alike first. Twelve of the 440 questions meet
it. The hero is `q340` ("Once my boarding pass has been scanned at the gate, can
the airline still bump me off an overbooked flight?"): SIE ranked 14 CFR
250.7(a) first; text-embedding-3-large ranked 14 CFR 250.11(a) first and the
answer second. The pain cases are `q146` (PostgreSQL `multixact_member_buffers`,
default 32, against the offsets cache's 16; PostgreSQL License, source
`config.sgml` at postgres 168fa9dd) and `q353` (45 CFR 164.412(a), a written
law-enforcement statement, against (b), an oral one). Passages are verbatim and
were re-read at their sources on 8 October 2026. Every ranking is in
`lookalike-search/rankings.json` of
[superlinked/sie-task-evidence@b939e410](https://huggingface.co/datasets/superlinked/sie-task-evidence/tree/b939e410/lookalike-search).

**The API pane.** The two encode calls recorded on `https://api.superlinked.com`
on 30 September 2026 for question `q120` and three HHS breach-notification
passages (45 CFR 164.408(b), 164.408(c), 164.406(a)): the question as a query,
then the passages as documents. The ranking is the caller's cosine over the
returned vectors. Requests, gzipped responses and their sha256 are in
`apps/site/tests/fixtures/reference/search/playground/`. That recording sent
`is_query` at the top of `params`, which the Cloud gateway reads; the Copy code
now sends it in `params.options`, where a self-hosted server also reads it.
Both encode the question as a query: the recorded query vector matches the
study's `is_query` vector for `q120` at cosine 0.99986. Run opens the console
on the same question, passages and model through the registered preset
`search-lookalike-q120-qwen3-embedding-4b-2026-09-30`.

**Price.** Real-time list prices per 1M input tokens: SIE Qwen3 Embedding 4B
$0.06 (SIE's vendored rate book); OpenAI text-embedding-3-large $0.13 and
text-embedding-3-small $0.02, read at
[developers.openai.com/api/docs/pricing](https://developers.openai.com/api/docs/pricing)
on 8 October 2026. Qwen3 Embedding 8B has no Cloud rate yet.

**Switching from the OpenAI SDK.** The OpenAI SDK's `embeddings.create` has no
query flag, so the OpenAI pair sends the model's query instruction in front of
the question by hand:
`Instruct: Given a query, retrieve relevant passages that answer the query\nQuery: `.
That gives the same vector as `is_query` on the native API. Without it the
question is encoded as a document; in the study that cost 5.2 points on the
eight MTEB tasks.

## The evaluation

The study ran on 30 September 2026. Its comparison for the page is SIE against
OpenAI text-embedding-3-large; the tables below keep every model we ran.

**The look-alike test.** 5,658 passages from 40 public sources: eCFR parts
with parallel rules (OSHA, FMCSA, DOT, CFPB, HIPAA, FMLA, FLSA, FTC), IRS
Publications 15 and 15-A, NIST SP 800-63B-4, twelve Kubernetes documentation
pages (CC BY 4.0) and four PostgreSQL documentation chapters (PostgreSQL
License). 440 questions, each with one answer passage and one named look-alike.
Claude Sonnet 5 drafted each question from its answer and the five passages
nearest to it by BM25; GPT-6 Sol kept a draft only if it picked the answer over
the look-alike blind; every kept question was read by eye and 86 were dropped;
the set was frozen before any tested model saw it. The score is the share of
questions whose first result is the answer. Every failed or missing reply
counts as a miss.

**Single embedding models**, answer ranked first, of 440 (30 September 2026):

| Model | First | Share |
| --- | --- | --- |
| Voyage 4 Large | 349 | 79.3% |
| Voyage 4 Lite | 336 | 76.4% |
| SIE Qwen3 Embedding 8B | 316 | 71.8% |
| Cohere Embed v4 | 309 | 70.2% |
| SIE Qwen3 Embedding 4B | 306 | 69.5% |
| OpenAI text-embedding-3-large | 295 | 67.0% |
| OpenAI text-embedding-3-small | 266 | 60.5% |

Head to head against OpenAI text-embedding-3-large, on the same 440 questions
(questions only SIE ranked first / only OpenAI ranked first; exact two-sided
McNemar; difference in points with a 95% interval clustered by the 40 sources):

| SIE model | SIE only / OpenAI only | McNemar p | Difference, 95% interval |
| --- | --- | --- | --- |
| Qwen3 Embedding 8B | 54 / 33 | 0.031 | +4.8 (-0.3 to +9.8) |
| Qwen3 Embedding 4B | 47 / 36 | 0.27 | +2.5 (-2.1 to +7.1) |

Against text-embedding-3-small: Qwen3 Embedding 8B 71 / 21 and Qwen3 Embedding
4B 71 / 31 (both p < 0.001). The pre-registered primary comparison was Qwen3
Embedding 4B against text-embedding-3-large; Qwen3 Embedding 8B was added by
amendment. The 8B row ran without its query instruction: the open-source
server's `qwen2_flash` adapter did not apply the configured default instruction
(fixed in sie #621 and #623), so its questions were encoded as
`Instruct: \nQuery:`. Voyage 4 Lite ranked the answer first on 59 questions SIE
Qwen3 Embedding 4B missed; the reverse happened on 29 (p = 0.002).

**Eight public MTEB retrieval tasks** (CQADupstackPhysics, CosQA, FiQA2018,
LegalBenchConsumerContractsQA, NFCorpus, SCIDOCS, SciFact, StackOverflowQA;
6,200 queries), share of queries whose first result is relevant, averaged over
tasks: Voyage 4 Large 54.7%, SIE Qwen3 Embedding 8B 54.8%, SIE Qwen3 Embedding
4B 54.6%, Cohere Embed v4 52.3%, OpenAI text-embedding-3-large 50.8%, Voyage 4
Lite 50.5%, OpenAI text-embedding-3-small 47.5%. Qwen3 Embedding 4B minus
text-embedding-3-large is +3.8 points (task-stratified bootstrap 95% interval
+2.7 to +4.9); in nDCG@10, 0.599 against 0.565. Qwen3 Embedding 8B's nDCG@10 is
0.608. Public tasks may overlap an open model's training data.

**Embedding plus reranking**, on a frozen 100-question sample stratified by
source (seed 20261007), answer ranked first, of 100:

| Pipeline | First |
| --- | --- |
| Voyage 4 Lite top 10, reranked by Voyage rerank-3 | 90 |
| SIE Qwen3 Embedding 8B (query instruction sent) top 10, reranked by SIE Qwen3 Reranker 4B | 82 |
| SIE Qwen3 Embedding 4B top 10, reranked by ZeroEntropy zerank-2 on SIE | 81 |
| Octen Embedding 8B top 10, reranked by zerank-2 on SIE | 81 |
| Voyage 4 Lite top 10, reranked by zerank-2 on SIE | 81 |
| SIE Qwen3 Embedding 4B top 10, reranked by SIE Qwen3 Reranker 4B | 80 |
| Octen Embedding 8B top 10, reranked by SIE Qwen3 Reranker 4B | 80 |
| Voyage 4 Lite top 10, reranked by SIE Qwen3 Reranker 4B | 79 |
| Octen Embedding 8B alone | 72 |

On the same sample, single models scored Voyage 4 Large 81, Voyage 4 Lite 76,
Cohere Embed v4 73, SIE Qwen3 Embedding 8B 72 and SIE Qwen3 Embedding 4B 65.
The Qwen3 Reranker rows use float32 scoring; an earlier run of SIE's 4B
pipeline on the released server, which normalised scores in bfloat16 and tied
many candidates, scored 75. A further screen of two open rerankers, Qwen3
Reranker 8B and mxbai-rerank-large-v2, also stayed below 88. No pipeline was
confirmed on fresh questions.

Evidence:
[the pilot and its gates](https://huggingface.co/datasets/superlinked/sie-task-evidence/tree/d971b789e2701e91a68dfe65525f77c035c3e234/evaluations/2026-10-07/lookalike-search-pilot-results-v1),
[the float32 reranker follow-up](https://huggingface.co/datasets/superlinked/sie-task-evidence/tree/168088b2582ef2b25d780b02978e5d89e93f3aca/evaluations/2026-10-08/lookalike-search-f32-rerank-followup-v1),
and the frozen inputs, split, gold and scorer at
[superlinked/sie-task-evidence@f960e26f](https://huggingface.co/datasets/superlinked/sie-task-evidence/tree/f960e26fa427b86286694891d449e3158c0d4b84/lookalike-search-pilot/2026-10-07).

## Limits

One gold passage per question; about 15 of the 440 questions probably have a
second valid answer, and they count against every model alike. The questions
were written and filtered by Anthropic and OpenAI models; a vendor-family effect
was not measured. The 100-question sample is a screen: its intervals are wide,
and the 340 remaining questions were never used.
