Run topk-embed-v1 at 50k tokens/s on a single GPU
topk-embed-v1 now runs in SIE, the open-source Superlinked Inference Engine. TopK announced the model family in its launch post on September 24. Both open models, topk-embed-v1-xsmall (0.8B) and topk-embed-v1-small (2B), ship in SIE v0.9.0 and embed text and page images behind one API.
On one L40S, SIE encodes 2B documents at 48,000 to 52,000 tokens/s, 1.5x the pipeline on the model’s Hugging Face card, and every test query returns the same top document. Both pipelines ran on the same hardware, one NVIDIA L40S, at the same bf16 precision. Against that pipeline:
What TopK built
TopK built these models for search over text and page images. Every text token, and every patch of a page, keeps its own vector, and MaxSim scores a query by matching each of its vectors to the closest document vector. In its announcement, TopK reports 65.22 nDCG@10 for the 2B model on native ViDoRe v3 image queries and 62.88 for the 0.8B, and runs a larger production model inside its own database.
1.5x the document throughput on one GPU
SIE runs the same checkpoints and holds 48,000 to 52,000 tokens/s at every document length we tried, against 31,000 to 34,000 for the Hugging Face pipeline. Getting there took a long list of GPU optimisations, and the models run on the same infrastructure as everything else in SIE: one API over shared GPUs that batch requests and scale with demand.
View chart data and sources
| Workload, 2B model | Hugging Face pipeline | SIE | SIE vs pipeline |
|---|---|---|---|
| 256 documents of 294 tokens | 31,400 tokens/s | 49,900 tokens/s | 1.59x |
| 64 documents of 1,230 tokens | 34,100 tokens/s | 52,000 tokens/s | 1.52x |
| 16 documents of 5,158 tokens | 33,200 tokens/s | 50,200 tokens/s | 1.51x |
| 8 documents of 8,192 tokens | 31,800 tokens/s | 48,200 tokens/s | 1.52x |
| 256 queries of about 15 tokens | 396 queries/s | 2,814 queries/s | 7.1x |
| 32 page images of 1,230 image tokens | 8.0 pages/s | 10.4 pages/s | 1.30x |
| One 10-token query, mean latency | 81 ms | 6 ms | 13x less |
The 0.8B model gains 1.4 to 1.7x on documents (up to 96,100 tokens/s), 12x on query throughput and 1.14x on pages, and one query takes 3 ms instead of 84 ms.
One NVIDIA L40S on Modal for each pipeline, both computing in bf16, the dtype in the checkpoint’s own config. Both pipelines got the same inputs and were called from Python after warm-up, with no server in between. Called this way, the model accepts documents up to the checkpoint’s limit of 8,192 tokens, so every length was encoded in full; a running SIE server caps 2B documents at 4,000 tokens. The Hugging Face pipeline is sentence-transformers 6 with TopK’s model code at the versions in its requirements.txt, at its better batch size of 8 or 32; SIE ran at its best batch setting. Token counts include the model’s Query: or Document: prefix. The inputs are synthetic, which does not change the speed: the model does the same work for any text of a given length.
Through SIE’s HTTP server, with clients on the same machine, the 2B model encodes 39,000 to 42,000 tokens/s with 16 to 32 clients, answers one query in 12 ms and encodes 10.5 pages/s with 8 clients. That is about 80% of the in-process rate.
Quality scores from TopK’s launch post and model cards. Speed measured by Superlinked, September 2026.
The query numbers carry one caveat: the pipeline only ran at batch sizes 8 and 32, and a larger batch would likely narrow the 7x gap on the 2B and the 12x gap on the 0.8B.
Same top document as the Hugging Face pipeline
We checked every output against the pipeline on an L4, both in bf16, with documents up to 7,590 tokens long and full pages. Every query returned the same top document, and MaxSim scores differed by at most 0.026.
Search millions of pages
Exact MaxSim over millions of pages is too slow for every query, so you search a compact encoding first and re-rank the shortlist with MaxSim, the way TopK’s database does. The :muvera profile returns one dense vector per item and is in v0.9.0. The :smve profile returns a sparse vector built with TopK’s SMVE method; on the 2B model its top 100 held 86% of each query’s exact MaxSim top 10 on BEIR SciFact, against 67% for MUVERA. It ships with the next SIE release.
Run it
Sign up to the cloud at superlinked.com/cloud.