naver/v-splade-quality
Primitive: /encode · Encode ·
ModernVBERT
V-SPLADE is Naver's visual SPLADE model: sparse lexical embeddings over document images and text on a ModernVBERT backbone, bringing interpretable SPLADE-style retrieval to visual documents
Overview
Hardware: — drives latency, throughput & cost
| Size | 330M params |
|---|---|
| Tasks | /encode |
| License | apache-2.0 |
| Languages | en |
| Latency | — |
| Throughput | — |
| Cost | — /1M tok |
Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.
Embedding
| Output types | Sparse |
|---|---|
| Dimensions | sparse: 50,368 |
| Max sequence length | 7,999 |
| Inputs | text · image |
Benchmarks
Vidore3ComputerScienceRetrieval
Visual document retrieval on computer science papers and slides
Vidore3FinanceEnRetrieval
Visual document retrieval on financial reports
Vidore3HrRetrieval
Visual document retrieval on HR-related documents
Vidore3PharmaceuticalsRetrieval
Visual document retrieval on pharmaceutical documents