Why did we open-source our inference engine? Read the post

← Catalog

nvidia/llama-nemoretriever-colembed-3b-v1

Open comparison →

Primitive: /encode · Encode · llama_nemoretrievercolembed

The nvidia/llama-nemoretriever-colembed-3b-v1 is a late interaction embedding model fine-tuned for query-document retrieval. Users can input `queries`, which are text, or `documents` which are page images, to the model.

MultimodalMultilingualLong contextMulti-vector

Overview

Hardware: — drives latency, throughput & cost

Size4.4B params
Tasks /encode
Licenseother
Languagesmultilingual
Latency6.1 s
Throughput0.7 img/s
Cost /1M tok

Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.

Embedding

Output typesMulti-Vec
Dimensionsmultivector: 3,072
Max sequence length8,192
Inputstext · image

Benchmarks

Vidore3ComputerScienceRetrieval

technology retrieval en

Visual document retrieval on computer science papers and slides

default_lang-eng
Performance L4 b1 c4
Corpus 0.6 img/s
Corpus p50 6.0s
Query 381 tok/s
Query p50 182.9ms
muvera
Quality
ndcg at 10 0.5585
muvera_lang-eng_limit-256
Quality
ndcg at 10 0.7733
map at 10 0.6621
mrr at 10 0.8546
Reference →

Vidore3FinanceEnRetrieval

finance retrieval en

Visual document retrieval on financial reports

default_lang-eng
Performance L4 b1 c4
Corpus 0.6 img/s
Corpus p50 6.1s
Query 502 tok/s
Query p50 152.7ms
muvera
Quality
ndcg at 10 0.2794
muvera_lang-eng_limit-256
Quality
ndcg at 10 0.5632
map at 10 0.4576
mrr at 10 0.6416
Reference →

Vidore3HrRetrieval

general retrieval en

Visual document retrieval on HR-related documents

default
Quality
ndcg at 10 0.6513
map at 10 0.5053
mrr at 10 0.7844
Performance L4 b1 c16
Corpus 0.9 img/s
Corpus p50 17.9s
Query 689 tok/s
Query p50 740.7ms
muvera
Quality
ndcg at 10 0.3291
muvera_lang-eng_limit-256
Quality
ndcg at 10 0.5003
map at 10 0.3830
mrr at 10 0.5998
Reference →

Vidore3PharmaceuticalsRetrieval

medical retrieval en

Visual document retrieval on pharmaceutical documents

default_lang-eng
Performance L4 b1 c4
Corpus 0.7 img/s
Corpus p50 6.0s
Query 420 tok/s
Query p50 185.5ms
muvera
Quality
ndcg at 10 0.4697
muvera_lang-eng_limit-256
Quality
ndcg at 10 0.7338
map at 10 0.6232
mrr at 10 0.8296
Reference →

VidoreTabfquadRetrieval

general retrieval fr

Visual document retrieval on French tabular QA

Quality
ndcg at 10 0.8191
Reference →

Open source inference for agents

Open-source inference for the models behind your agents. Run it yourself, or let us run it for you.

Github 3.3K

Contact us

Tell us about your use case and we'll get back to you shortly.

Apply for an inference grant

Free capacity on our hosted cluster for selected projects. Tell us what you run and we reply by email.