Why did we open-source our inference engine? Read the post

← Catalog

TomoroAI/tomoro-colqwen3-embed-4b

Open comparison →

Primitive: /encode · Encode · Qwen3-VL

TomoroAI/tomoro-colqwen3-embed-4b is a state-of-the-art ColPali-style multimodal embedding model. It maps text queries, visual documents (images, PDFs) or short videos into aligned multi-vector embeddings.

MultimodalMultilingualLong contextMulti-vector

Overview

Hardware: — drives latency, throughput & cost

Size4.4B params
Tasks /encode
Licenseapache-2.0
Languagesmultilingual
Latency
Throughput
Cost /1M tok

Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.

Embedding

Output typesMulti-Vec
Dimensionsmultivector: 320
Max sequence length8,192
Inputstext · image

Benchmarks

Vidore3ComputerScienceRetrieval

technology retrieval en

Visual document retrieval on computer science papers and slides

default_lang-eng
Quality
map at 10 0.6491
mrr at 10 0.8847
ndcg at 10 0.7690
muvera_lang-eng_limit-256
Quality
ndcg at 10 0.8558
map at 10 0.7873
mrr at 10 0.9134
Reference →

Vidore3FinanceEnRetrieval

finance retrieval en

Visual document retrieval on financial reports

default_lang-eng
Quality
map at 10 0.5816
mrr at 10 0.8172
ndcg at 10 0.6970
muvera_lang-eng_limit-256
Quality
ndcg at 10 0.7161
map at 10 0.6000
mrr at 10 0.7717
Reference →

Vidore3HrRetrieval

general retrieval en

Visual document retrieval on HR-related documents

Quality
ndcg at 10 0.5851
map at 10 0.4758
mrr at 10 0.6364
Reference →

Vidore3PharmaceuticalsRetrieval

medical retrieval en

Visual document retrieval on pharmaceutical documents

default_lang-eng
Quality
map at 10 0.5754
mrr at 10 0.7978
ndcg at 10 0.6852
muvera_lang-eng_limit-256
Quality
ndcg at 10 0.7916
map at 10 0.6890
mrr at 10 0.8737
Reference →

Open source inference for agents

Open-source inference for the models behind your agents. Run it yourself, or let us run it for you.

Github 2.8K

Contact us

Tell us about your use case and we'll get back to you shortly.

Apply for an inference grant

Free capacity on our hosted cluster for selected projects. Tell us what you run and we reply by email.