Qwen/Qwen3-VL-Embedding-2B
Primitive: /encode · Encode ·
Qwen3-VL
The Qwen3-VL-Embedding and Qwen3-VL-Reranker model series are the latest additions to the Qwen family, built upon the recently open-sourced and powerful Qwen3-VL foundation model.
View on Hugging Face → Guide: Qwen3 embeddings and rerankers guide → Fine-tuned from Qwen/Qwen3-VL-2B-Instruct
Overview
Hardware: — drives latency, throughput & cost
| Size | 2.1B params |
|---|---|
| Tasks | /encode |
| License | apache-2.0 |
| Latency | 36 ms |
| Throughput | 494 tok/s |
| Cost | $0.450 /1M tok |
Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.
Embedding
| Output types | Dense |
|---|---|
| Dimensions | dense: 2,048 |
| Max sequence length | 32,768 |
| Inputs | text · image · video |
Benchmarks
CQADupstackPhysicsRetrieval
Duplicate question retrieval from StackExchange Physics
CosQA
Code search with natural language queries
FiQA2018
Financial opinion mining and question answering
Flickr30kI2TRetrieval
Image-to-text retrieval: retrieve captions from images
LegalBenchConsumerContractsQA
Question answering on consumer contracts
NFCorpus
Biomedical literature search from NutritionFacts.org
SCIDOCS
Citation prediction, document classification, and recommendation for scientific papers
SciFact
Scientific claim verification using research literature
StackOverflowQA
Programming question answering from Stack Overflow