tencent/R3-embedding-0.6b
Primitive: /encode · Encode ·
Qwen3
The latest agent skill retrieval model at the 0.6B scale. R3-Embedding is the bi-encoder (recall) stage of R3-Skill's two-stage retriever for query-conditional agent skill retrieval.
View on Hugging Face → Fine-tuned from Qwen/Qwen3-Embedding-0.6B
Overview
Hardware: — drives latency, throughput & cost
| Size | 596M params |
|---|---|
| Tasks | /encode |
| License | apache-2.0 |
| Latency | 197 ms |
| Throughput | 24.5K tok/s |
| Cost | $0.0091 /1M tok |
Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.
Embedding
| Output types | Dense |
|---|---|
| Dimensions | dense: 1,024 |
| Max sequence length | 4,096 |
| Inputs | text |
Benchmarks
R3Skill
Agent skill retrieval: match user requests to the SKILL.md that solves them (Tencent R3 release set)