Why did we open-source our inference engine? Read the post

← Catalog

nvidia/nemotron-colembed-vl-4b-v2

Open comparison →

Primitive: /encode · Encode · Qwen3-VL

The nvidia/nemotron-colembed-vl-4b-v2 is a state-of-the-art late interaction embedding model that ranks No. 3 in the ViDoRe V3: a comprehensive evaluation of retrieval for enterprise use-case benchmark, (as of Jan 26, 2026) with a score of `61.42` on 8 public tasks.

MultimodalMultilingualLong contextMulti-vector

Overview

Hardware: — drives latency, throughput & cost

Size4.8B params
Tasks /encode
Licensecc-by-nc-4.0
Languagesmultilingual
Latency
Throughput
Cost /1M tok

Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.

Embedding

Output typesMulti-Vec
Dimensionsmultivector: 2,560
Max sequence length8,192
Inputstext · image

Benchmarks

Vidore3ComputerScienceRetrieval

technology retrieval en

Visual document retrieval on computer science papers and slides

default_lang-eng
Quality
map at 10 0.6818
mrr at 10 0.9077
ndcg at 10 0.7959
muvera_lang-eng_limit-256
Quality
ndcg at 10 0.5491
map at 10 0.4479
mrr at 10 0.6088
Reference →

Vidore3FinanceEnRetrieval

finance retrieval en

Visual document retrieval on financial reports

default_lang-eng
Quality
map at 10 0.5592
mrr at 10 0.7910
ndcg at 10 0.6823
muvera_lang-eng_limit-256
Quality
ndcg at 10 0.3245
map at 10 0.2195
mrr at 10 0.3970
Reference →

Vidore3HrRetrieval

general retrieval en

Visual document retrieval on HR-related documents

Quality
ndcg at 10 0.3081
map at 10 0.1960
mrr at 10 0.3622
Reference →

Vidore3PharmaceuticalsRetrieval

medical retrieval en

Visual document retrieval on pharmaceutical documents

default_lang-eng
Quality
map at 10 0.5696
mrr at 10 0.7925
ndcg at 10 0.6809
muvera_lang-eng_limit-256
Quality
ndcg at 10 0.5023
map at 10 0.3996
mrr at 10 0.5714
Reference →

Open source inference for agents

Open-source inference for the models behind your agents. Run it yourself, or let us run it for you.

Github 2.8K

Contact us

Tell us about your use case and we'll get back to you shortly.

Apply for an inference grant

Free capacity on our hosted cluster for selected projects. Tell us what you run and we reply by email.