Why did we open-source our inference engine? Read the post

← Catalog

BAAI/bge-reranker-v2-m3

Open comparison →

Primitive: /score · Score · XLM-RoBERTa

More details please refer to our Github: FlagEmbedding.

MultilingualLong context

Overview

Hardware: — drives latency, throughput & cost

Size568M params
Tasks /score
Licenseapache-2.0
Languagesmultilingual
Latency92 ms
Throughput30.0K tok/s
Cost$0.0074 /1M tok

Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.

Scoring

Inputstext
Max sequence length8,192

Benchmarks

AskUbuntuDupQuestions

technology reranking en

Duplicate question detection from AskUbuntu

Corpus: 6,743 Queries: 360
Quality
ndcg at 10 0.6763
map at 10 0.5245
mrr at 10 0.7611
Performance L4 b1 c16
Query 6.5K tok/s
Query p50 40.6ms
Reference →

CMedQAv1Reranking

medical reranking zh

Chinese medical question answering reranking (v1)

Corpus: 100,000 Queries: 2,000
Quality
map at 10 0.8192
mrr at 10 0.8550
Reference →

CMedQAv2Reranking

medical reranking zh

Chinese medical question answering reranking (v2)

Corpus: 108,000 Queries: 4,000
Quality
map at 10 0.8287
mrr at 10 0.8608
Reference →

CQADupstackPhysicsRetrieval

scientific retrieval en

Duplicate question retrieval from StackExchange Physics

Corpus: 38,314 Queries: 1,039
Quality
map at 10 0.3618
mrr at 10 0.4262
ndcg at 10 0.4167
Reference →

CQADupstackPhysicsRetrieval (candidates: gte-multilingual-base, k=50)

scientific retrieval en

Duplicate question retrieval from StackExchange Physics

Performance L4 b1 c16
Query 26.5K tok/s
Query p50 89.8ms
Reference →

CosQA

technology retrieval en

Code search with natural language queries

Corpus: 6,267 Queries: 500
Quality
map at 10 0.2169
mrr at 10 0.2331
ndcg at 10 0.2897
Reference →

CosQA (candidates: gte-multilingual-base, k=50)

technology retrieval en

Code search with natural language queries

Performance L4 b1 c16
Query 13.9K tok/s
Query p50 66.1ms
Reference →

FiQA2018

finance retrieval en

Financial opinion mining and question answering

Corpus: 57,599 Queries: 648
Quality
map at 10 0.3675
mrr at 10 0.5479
ndcg at 10 0.4504
Reference →

FiQA2018 (candidates: gte-multilingual-base, k=50)

finance retrieval en

Financial opinion mining and question answering

Performance L4 b1 c16
Query 30.0K tok/s
Query p50 90.2ms
Reference →

LegalBenchConsumerContractsQA

legal retrieval en

Question answering on consumer contracts

Corpus: 153 Queries: 396
Quality
map at 10 0.7734
mrr at 10 0.7757
ndcg at 10 0.8156
Reference →

LegalBenchConsumerContractsQA (candidates: gte-multilingual-base, k=50)

legal retrieval en

Question answering on consumer contracts

Performance L4 b1 c16
Query 51.3K tok/s
Query p50 93.5ms
Reference →

MMarcoReranking

general reranking zh

Multilingual MARCO passage reranking (Chinese)

Quality
map at 10 0.3354
mrr at 10 0.3401
Performance L4 b1 c16
Reference →

NFCorpus

medical retrieval en

Biomedical literature search from NutritionFacts.org

Corpus: 3,593 Queries: 323
default_bm25-k-100
Quality
ndcg at 10 0.3285
map at 10 0.2339
mrr at 10 0.5301
default_bm25-k-100_qrels-k-50
Quality
ndcg at 10 0.3930
map at 10 0.2792
mrr at 10 0.5996
default_bm25-k-50_qrels-k-50
Quality
ndcg at 10 0.4243
map at 10 0.3065
mrr at 10 0.6220
default_bm25-k-50_qrels-k-10
Quality
ndcg at 10 0.4029
map at 10 0.2832
mrr at 10 0.6087
default_bm25-k-100_qrels-k-10
Quality
ndcg at 10 0.3759
map at 10 0.2621
mrr at 10 0.5876
default_bm25-k-50
Quality
ndcg at 10 0.3306
map at 10 0.2375
mrr at 10 0.5380
default_bm25-k-200_qrels-k-50
Quality
ndcg at 10 0.3674
map at 10 0.2602
mrr at 10 0.5765
default_candidates-k-50_candidates-model-Alibaba-NLP__gte-multilingual-base
Quality
map at 10 0.2510
mrr at 10 0.5638
ndcg at 10 0.3530
default_bm25-k-200
Quality
ndcg at 10 0.3280
map at 10 0.2331
mrr at 10 0.5264
default_bm25-k-200_qrels-k-10
Quality
ndcg at 10 0.3557
map at 10 0.2497
mrr at 10 0.5677
Reference →

NFCorpus (candidates: gte-multilingual-base, k=50)

medical retrieval en

Biomedical literature search from NutritionFacts.org

Performance L4 b1 c16
Query 35.3K tok/s
Query p50 109.6ms
Reference →

SCIDOCS

scientific retrieval en

Citation prediction, document classification, and recommendation for scientific papers

Corpus: 25,656 Queries: 1,000
Quality
map at 10 0.1076
mrr at 10 0.3207
ndcg at 10 0.1871
Reference →

SCIDOCS (candidates: gte-multilingual-base, k=50)

scientific retrieval en

Citation prediction, document classification, and recommendation for scientific papers

Performance L4 b1 c16
Query 29.4K tok/s
Query p50 98.7ms
Reference →

SciFact

scientific retrieval en

Scientific claim verification using research literature

Corpus: 5,183 Queries: 300
Quality
map at 10 0.7013
mrr at 10 0.7184
ndcg at 10 0.7456
Reference →

SciFact (candidates: gte-multilingual-base, k=50)

scientific retrieval en

Scientific claim verification using research literature

Performance L4 b1 c16
Query 33.4K tok/s
Query p50 113.2ms
Reference →

StackOverflowQA

technology retrieval en

Programming question answering from Stack Overflow

Corpus: 19,931 Queries: 1,994
Quality
map at 10 0.5545
mrr at 10 0.5631
ndcg at 10 0.6158
Reference →

StackOverflowQA (candidates: gte-multilingual-base, k=50)

technology retrieval en

Programming question answering from Stack Overflow

Performance L4 b1 c16
Query 53.1K tok/s
Query p50 111.1ms
Reference →

T2Reranking

general reranking zh

Chinese passage ranking benchmark

Quality
map at 10 0.5630
mrr at 10 0.7781
Reference →

Open source inference for agents

Open-source inference for the models behind your agents. Run it yourself, or let us run it for you.

Github 3.3K

Contact us

Tell us about your use case and we'll get back to you shortly.

Apply for an inference grant

Free capacity on our hosted cluster for selected projects. Tell us what you run and we reply by email.