Why did we open-source our inference engine? Read the post

← Catalog

google/siglip2-base-patch16-224

Open comparison →

Primitive: /encode · Encode · SigLIP

SigLIP 2 extends the pretraining objective of SigLIP with prior, independently developed techniques into a unified recipe, for improved semantic understanding, localization, and dense features.

MultimodalDense

Overview

Hardware: — drives latency, throughput & cost

Size375M params
Tasks /encode
Licenseapache-2.0
Latency159 ms
Throughput699 tok/s
Cost$0.318 /1M tok

Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.

Embedding

Output typesDense
Dimensionsdense: 768
Max sequence length64
Inputstext · image

Benchmarks

Flickr30kI2TRetrieval

general retrieval en

Image-to-text retrieval: retrieve captions from images

Corpus: 31,783 Queries: 1,000
Quality
ndcg at 10 0.8156
map at 10 0.7254
mrr at 10 0.9302
Performance L4 b1 c8
Corpus 699 tok/s
Corpus p50 158.9ms
Query 15.1 mpix/s
Query p50 88.5ms
Reference →

Open source inference for agents

Open-source inference for the models behind your agents. Run it yourself, or let us run it for you.

Github 3.3K

Contact us

Tell us about your use case and we'll get back to you shortly.

Apply for an inference grant

Free capacity on our hosted cluster for selected projects. Tell us what you run and we reply by email.