Why did we open-source our inference engine? Read the post
← All Posts

Run topk-embed-v1 at 50k tokens/s on a single GPU

topk-embed-v1 now runs in SIE, the open-source Superlinked Inference Engine. TopK announced the model family in its launch post on September 24. Both open models, topk-embed-v1-xsmall (0.8B) and topk-embed-v1-small (2B), ship in SIE v0.9.0 and embed text and page images behind one API.

On one L40S, SIE encodes 2B documents at 48,000 to 52,000 tokens/s, 1.5x the pipeline on the model’s Hugging Face card, and every test query returns the same top document. Both pipelines ran on the same hardware, one NVIDIA L40S, at the same bf16 precision. Against that pipeline:

1.5x Document throughput
7x Query throughput
13x Lower single-query latency
1.3x Page throughput

What TopK built

TopK built these models for search over text and page images. Every text token, and every patch of a page, keeps its own vector, and MaxSim scores a query by matching each of its vectors to the closest document vector. In its announcement, TopK reports 65.22 nDCG@10 for the 2B model on native ViDoRe v3 image queries and 62.88 for the 0.8B, and runs a larger production model inside its own database.

1.5x the document throughput on one GPU

SIE runs the same checkpoints and holds 48,000 to 52,000 tokens/s at every document length we tried, against 31,000 to 34,000 for the Hugging Face pipeline. Getting there took a long list of GPU optimisations, and the models run on the same infrastructure as everything else in SIE: one API over shared GPUs that batch requests and scale with demand.

Document throughput, 2B model · tokens per second · HF is the Hugging Face pipeline · lengths in tokens
HF · 294
31.4k
SIE · 294
49.9k
HF · 1,230
34.1k
SIE · 1,230
52.0k
HF · 5,158
33.2k
SIE · 5,158
50.2k
HF · 8,192
31.8k
SIE · 8,192
48.2k

One L40S, both pipelines in bf16, called from Python after warm-up. Measured by Superlinked, September 2026.

View chart data and sources
Workload, 2B modelHugging Face pipelineSIESIE vs pipeline
256 documents of 294 tokens31,400 tokens/s49,900 tokens/s1.59x
64 documents of 1,230 tokens34,100 tokens/s52,000 tokens/s1.52x
16 documents of 5,158 tokens33,200 tokens/s50,200 tokens/s1.51x
8 documents of 8,192 tokens31,800 tokens/s48,200 tokens/s1.52x
256 queries of about 15 tokens396 queries/s2,814 queries/s7.1x
32 page images of 1,230 image tokens8.0 pages/s10.4 pages/s1.30x
One 10-token query, mean latency81 ms6 ms13x less

The 0.8B model gains 1.4 to 1.7x on documents (up to 96,100 tokens/s), 12x on query throughput and 1.14x on pages, and one query takes 3 ms instead of 84 ms.

One NVIDIA L40S on Modal for each pipeline, both computing in bf16, the dtype in the checkpoint’s own config. Both pipelines got the same inputs and were called from Python after warm-up, with no server in between. Called this way, the model accepts documents up to the checkpoint’s limit of 8,192 tokens, so every length was encoded in full; a running SIE server caps 2B documents at 4,000 tokens. The Hugging Face pipeline is sentence-transformers 6 with TopK’s model code at the versions in its requirements.txt, at its better batch size of 8 or 32; SIE ran at its best batch setting. Token counts include the model’s Query: or Document: prefix. The inputs are synthetic, which does not change the speed: the model does the same work for any text of a given length.

Through SIE’s HTTP server, with clients on the same machine, the 2B model encodes 39,000 to 42,000 tokens/s with 16 to 32 clients, answers one query in 12 ms and encodes 10.5 pages/s with 8 clients. That is about 80% of the in-process rate.

Quality scores from TopK’s launch post and model cards. Speed measured by Superlinked, September 2026.

The query numbers carry one caveat: the pipeline only ran at batch sizes 8 and 32, and a larger batch would likely narrow the 7x gap on the 2B and the 12x gap on the 0.8B.

Same top document as the Hugging Face pipeline

We checked every output against the pipeline on an L4, both in bf16, with documents up to 7,590 tokens long and full pages. Every query returned the same top document, and MaxSim scores differed by at most 0.026.

Search millions of pages

Exact MaxSim over millions of pages is too slow for every query, so you search a compact encoding first and re-rank the shortlist with MaxSim, the way TopK’s database does. The :muvera profile returns one dense vector per item and is in v0.9.0. The :smve profile returns a sparse vector built with TopK’s SMVE method; on the 2B model its top 100 held 86% of each query’s exact MaxSim top 10 on BEIR SciFact, against 67% for MUVERA. It ships with the next SIE release.

Run it

Sign up to the cloud at superlinked.com/cloud.

Open source inference for agents

Open-source inference for the models behind your agents. Run it yourself, or let us run it for you.

Contact us

Tell us about your use case and we'll get back to you shortly.

Apply for an inference grant

Free capacity on our hosted cluster for selected projects. Tell us what you run and we reply by email.