Why did we open-source our inference engine? Read the post

← Catalog

Qwen/Qwen3-Reranker-0.6B

Open comparison →

Primitive: /score · Score · Qwen3

The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. Building upon the dense foundational models of the Qwen3 series, it provides a comprehensive range of text embeddings and reranking models in various sizes (0.6B, 4B, and 8B).

Long context

Overview

Hardware: — drives latency, throughput & cost

Size596M params
Tasks /score
Licenseapache-2.0
Latency81 ms
Throughput1.6K tok/s
Cost$0.138 /1M tok

Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.

Scoring

Inputstext
Max sequence length32,768

Benchmarks

AskUbuntuDupQuestions

technology reranking en

Duplicate question detection from AskUbuntu

Corpus: 6,743 Queries: 360
Quality
ndcg at 10 0.6519
map at 10 0.4965
mrr at 10 0.7583
Performance L4 b1 c16
Corpus 1.8K tok/s
Corpus p50 92.9ms
Query 1.8K tok/s
Query p50 92.9ms
Reference →

CQADupstackPhysicsRetrieval

scientific retrieval en

Duplicate question retrieval from StackExchange Physics

Corpus: 38,314 Queries: 1,039
Quality
map at 10 0.4320
mrr at 10 0.4993
ndcg at 10 0.4955
Reference →

CosQA

technology retrieval en

Code search with natural language queries

Corpus: 6,267 Queries: 500
Quality
map at 10 0.2785
mrr at 10 0.2915
ndcg at 10 0.3598
Reference →

FiQA2018

finance retrieval en

Financial opinion mining and question answering

Corpus: 57,599 Queries: 648
Quality
map at 10 0.3683
mrr at 10 0.5374
ndcg at 10 0.4561
Reference →

LegalBenchConsumerContractsQA

legal retrieval en

Question answering on consumer contracts

Corpus: 153 Queries: 396
Quality
map at 10 0.7876
mrr at 10 0.7886
ndcg at 10 0.8320
Reference →

MMarcoReranking

general reranking zh

Multilingual MARCO passage reranking (Chinese)

Quality
ndcg at 10 0.0858
map at 10 0.0576
mrr at 10 0.8158
Performance L4 b1 c16
Corpus 18.7K tok/s
Corpus p50 69.8ms
Query 1.5K tok/s
Query p50 69.8ms
Reference →

NFCorpus

medical retrieval en

Biomedical literature search from NutritionFacts.org

Corpus: 3,593 Queries: 323
Quality
map at 10 0.2882
mrr at 10 0.5998
ndcg at 10 0.3904
Reference →

SCIDOCS

scientific retrieval en

Citation prediction, document classification, and recommendation for scientific papers

Corpus: 25,656 Queries: 1,000
Quality
map at 10 0.1308
mrr at 10 0.3745
ndcg at 10 0.2162
Reference →

SciFact

scientific retrieval en

Scientific claim verification using research literature

Corpus: 5,183 Queries: 300
Quality
map at 10 0.7242
mrr at 10 0.7373
ndcg at 10 0.7678
Reference →

StackOverflowQA

technology retrieval en

Programming question answering from Stack Overflow

Corpus: 19,931 Queries: 1,994
Quality
map at 10 0.9234
mrr at 10 0.9237
ndcg at 10 0.9340
Reference →

Open source inference for agents

Open-source inference for the models behind your agents. Run it yourself, or let us run it for you.

Github 3.3K

Contact us

Tell us about your use case and we'll get back to you shortly.

Apply for an inference grant

Free capacity on our hosted cluster for selected projects. Tell us what you run and we reply by email.