Why did we open-source our inference engine? Read the post

google/siglip-so400m-patch14-224

Architecture
Parameters
400M
Tasks
Encode
Outputs
Dense
Dimensions
Dense: 1,152
Max Sequence Length
64 tokens
License

Benchmarks

Flickr30kI2TRetrieval

general retrieval en

Quality
ndcg at 10 0.8382
map at 10 0.7479
mrr at 10 0.9353
Performance L4-SPOT b1 c8
Corpus TPS 223
Corpus p50 395.0ms
Query TPS 11
Query p50 392.1ms
Performance L4 b1 c16
Corpus TPS 473
Corpus p50 484.7ms
Query TPS 22
Query p50 425.9ms

Self-hosted inference for search & document processing

Cut API costs by 50x, boost quality with 85+ SOTA models, and keep your data in your own cloud.

Github
1.5K

Contact us

Tell us about your use case and we'll get back to you shortly.