nvidia/nemotron-colembed-vl-4b-v2
Primitive: /encode · Encode ·
Qwen3-VL
The nvidia/nemotron-colembed-vl-4b-v2 is a state-of-the-art late interaction embedding model that ranks No. 3 in the ViDoRe V3: a comprehensive evaluation of retrieval for enterprise use-case benchmark, (as of Jan 26, 2026) with a score of `61.42` on 8 public tasks.
Overview
Hardware: — drives latency, throughput & cost
| Size | 4.8B params |
|---|---|
| Tasks | /encode |
| License | cc-by-nc-4.0 |
| Languages | multilingual |
| Latency | — |
| Throughput | — |
| Cost | — /1M tok |
Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.
Embedding
| Output types | Multi-Vec |
|---|---|
| Dimensions | multivector: 2,560 |
| Max sequence length | 8,192 |
| Inputs | text · image |
Benchmarks
Vidore3ComputerScienceRetrieval
Visual document retrieval on computer science papers and slides
Vidore3FinanceEnRetrieval
Visual document retrieval on financial reports
Vidore3HrRetrieval
Visual document retrieval on HR-related documents
Vidore3PharmaceuticalsRetrieval
Visual document retrieval on pharmaceutical documents