Why did we open-source our inference engine? Read the post

← Catalog

tencent/R3-embedding-0.6b

Open comparison →

Primitive: /encode · Encode · Qwen3

The latest agent skill retrieval model at the 0.6B scale. R3-Embedding is the bi-encoder (recall) stage of R3-Skill's two-stage retriever for query-conditional agent skill retrieval.

Dense

Overview

Hardware: — drives latency, throughput & cost

Size596M params
Tasks /encode
Licenseapache-2.0
Latency197 ms
Throughput24.5K tok/s
Cost$0.0091 /1M tok

Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.

Embedding

Output typesDense
Dimensionsdense: 1,024
Max sequence length4,096
Inputstext

Benchmarks

R3Skill

technology retrieval en

Agent skill retrieval: match user requests to the SKILL.md that solves them (Tencent R3 release set)

Corpus: 2,050 Queries: 5,696
Quality
ndcg at 10 0.8158
Performance L4 b1 c16
Corpus 24.5K tok/s
Corpus p50 197.0ms
Query 22.4K tok/s
Query p50 113.7ms
Reference →

Open source inference for agents

Open-source inference for the models behind your agents. Run it yourself, or let us run it for you.

Github 2.8K

Contact us

Tell us about your use case and we'll get back to you shortly.

Apply for an inference grant

Free capacity on our hosted cluster for selected projects. Tell us what you run and we reply by email.