TomoroAI/tomoro-colqwen3-embed-4b
Primitive: /encode · Encode ·
Qwen3-VL
TomoroAI/tomoro-colqwen3-embed-4b is a state-of-the-art ColPali-style multimodal embedding model. It maps text queries, visual documents (images, PDFs) or short videos into aligned multi-vector embeddings.
View on Hugging Face → Fine-tuned from Qwen/Qwen3-VL-4B-Instruct
Overview
Hardware: — drives latency, throughput & cost
| Size | 4.4B params |
|---|---|
| Tasks | /encode |
| License | apache-2.0 |
| Languages | multilingual |
| Latency | — |
| Throughput | — |
| Cost | — /1M tok |
Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.
Embedding
| Output types | Multi-Vec |
|---|---|
| Dimensions | multivector: 320 |
| Max sequence length | 8,192 |
| Inputs | text · image |
Benchmarks
Vidore3ComputerScienceRetrieval
Visual document retrieval on computer science papers and slides
Vidore3FinanceEnRetrieval
Visual document retrieval on financial reports
Vidore3HrRetrieval
Visual document retrieval on HR-related documents
Vidore3PharmaceuticalsRetrieval
Visual document retrieval on pharmaceutical documents