Why did we open-source our inference engine? Read the post

← Catalog

naver/v-splade-quality

Open comparison →

Primitive: /encode · Encode · ModernVBERT

V-SPLADE is Naver's visual SPLADE model: sparse lexical embeddings over document images and text on a ModernVBERT backbone, bringing interpretable SPLADE-style retrieval to visual documents

MultimodalSparse

Overview

Hardware: — drives latency, throughput & cost

Size330M params
Tasks /encode
Licenseapache-2.0
Languagesen
Latency
Throughput
Cost /1M tok

Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.

Embedding

Output typesSparse
Dimensionssparse: 50,368
Max sequence length7,999
Inputstext · image

Benchmarks

Vidore3ComputerScienceRetrieval

technology retrieval en

Visual document retrieval on computer science papers and slides

Quality
ndcg at 10 0.5839
map at 10 0.4565
mrr at 10 0.7488
Reference →

Vidore3FinanceEnRetrieval

finance retrieval en

Visual document retrieval on financial reports

Quality
ndcg at 10 0.4651
map at 10 0.3504
mrr at 10 0.5965
Reference →

Vidore3HrRetrieval

general retrieval en

Visual document retrieval on HR-related documents

Quality
ndcg at 10 0.4481
map at 10 0.3195
mrr at 10 0.5624
Reference →

Vidore3PharmaceuticalsRetrieval

medical retrieval en

Visual document retrieval on pharmaceutical documents

Quality
ndcg at 10 0.5324
map at 10 0.4196
mrr at 10 0.6386
Reference →

Open source inference for agents

Open-source inference for the models behind your agents. Run it yourself, or let us run it for you.

Github 2.8K

Contact us

Tell us about your use case and we'll get back to you shortly.

Apply for an inference grant

Free capacity on our hosted cluster for selected projects. Tell us what you run and we reply by email.