Why did we open-source our inference engine? Read the post

← Catalog

tencent/R3-rerank-0.6b

Open comparison →

Primitive: /score · Score · Qwen3

The latest agent skill reranking model at the 0.6B scale. R3-Reranker is the cross-encoder (rerank) stage of R3-Skill's two-stage retriever for query-conditional agent skill retrieval. It scores each (query, skill) pair jointly, paired with R3-Embedding-0.6B for recall.

Overview

Hardware: — drives latency, throughput & cost

Size596M params
Tasks /score
Licenseapache-2.0
Latency449 ms
Throughput2.1K tok/s
Cost$0.103 /1M tok

Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.

Scoring

Inputstext
Max sequence length4,096

Benchmarks

R3Skill (candidates: R3-embedding-0.6b, k=20)

technology retrieval en

Agent skill retrieval: match user requests to the SKILL.md that solves them (Tencent R3 release set)

Quality
ndcg at 10 0.8228
Performance L4 b1 c16
Corpus 25.4K tok/s
Corpus p50 469.9ms
Query 2.2K tok/s
Query p50 469.9ms
Reference →

R3Skill (candidates: R3-embedding-0.6b, k=20, limit 384)

technology retrieval en

Agent skill retrieval: match user requests to the SKILL.md that solves them (Tencent R3 release set)

Quality
ndcg at 10 0.8264
Performance L4 b1 c16
Corpus 25.1K tok/s
Corpus p50 429.0ms
Query 2.1K tok/s
Query p50 429.0ms
Performance RTX-PRO-6000-Blackwell-Server-Edition b1 c16
Corpus 118.0K tok/s
Corpus p50 105.6ms
Query 9.8K tok/s
Query p50 105.6ms
Reference →

Open source inference for agents

Open-source inference for the models behind your agents. Run it yourself, or let us run it for you.

Github 2.8K

Contact us

Tell us about your use case and we'll get back to you shortly.

Apply for an inference grant

Free capacity on our hosted cluster for selected projects. Tell us what you run and we reply by email.