Why did we open-source our inference engine? Read the post

← Catalog

lightonai/mLateOn

Open comparison →

Primitive: /score · Score · ModernBERT

State-of-the-Art Multilingual ColBERT Retrieval Model by LightOn

MultilingualLong context

Overview

Hardware: — drives latency, throughput & cost

Size307M params
Tasks /encode · /score
Licenseapache-2.0
Languagesen, fr, de, it, es, pt, sv, no, ar, code
Latency
Throughput
Cost /1M tok

Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.

Scoring

Inputstext
Max sequence length8,192

Benchmarks

AskUbuntuDupQuestions

technology reranking en

Duplicate question detection from AskUbuntu

Corpus: 6,743 Queries: 360
Quality
ndcg at 10 0.6627
map at 10 0.5108
mrr at 10 0.7384
Reference →

CMedQAv1-reranking

medical reranking zh

Chinese medical question answering reranking (v1)

Corpus: 100,000 Queries: 2,000
Quality
ndcg at 10 0.7572
map at 10 0.6975
mrr at 10 0.7514
Reference →

CMedQAv2-reranking

medical reranking zh

Chinese medical question answering reranking (v2)

Corpus: 108,000 Queries: 4,000
Quality
ndcg at 10 0.7646
map at 10 0.7010
mrr at 10 0.7612
Reference →

MMarcoReranking

general reranking zh

Multilingual MARCO passage reranking (Chinese)

Quality
ndcg at 10 0.3284
map at 10 0.2652
mrr at 10 0.2677
Reference →

T2Reranking

general reranking zh

Chinese passage ranking benchmark

Quality
ndcg at 10 0.7460
map at 10 0.5776
mrr at 10 0.8002
Reference →

Open source inference for agents

Open-source inference for the models behind your agents. Run it yourself, or let us run it for you.

Github 2.8K

Contact us

Tell us about your use case and we'll get back to you shortly.

Apply for an inference grant

Free capacity on our hosted cluster for selected projects. Tell us what you run and we reply by email.