Qwen/Qwen3-Embedding-4B
Primitive: /encode · Encode ·
Qwen3
The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. Building upon the dense foundational models of the Qwen3 series, it provides a comprehensive range of text embeddings and reranking models in various sizes (0.6B, 4B, and 8B).
View on Hugging Face → Fine-tuned from Qwen/Qwen3-4B-Base
Overview
Hardware: — drives latency, throughput & cost
| Size | 4.0B params |
|---|---|
| Tasks | /encode |
| License | apache-2.0 |
| Latency | 464 ms |
| Throughput | 5.7K tok/s |
| Cost | $0.039 /1M tok |
Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.
Embedding
| Output types | Dense |
|---|---|
| Dimensions | dense: 2,560 |
| Max sequence length | 32,768 |
| Inputs | text |
Benchmarks
FiQA2018
Financial opinion mining and question answering
NFCorpus
Biomedical literature search from NutritionFacts.org
NanoFiQA2018Retrieval
Smaller subset of the FiQA financial QA dataset