Why did we open-source our inference engine? Read the post

← Catalog

knowledgator/gliclass-large-v3.0

Open comparison →

Primitive: /extract · Extract · DeBERTa

This is an efficient zero-shot classifier inspired by GLiNER work. It demonstrates the same performance as a cross-encoder while being more compute-efficient because classification is done at a single forward path.

Overview

Hardware: — drives latency, throughput & cost

Size439M params
Tasks /extract
Licenseapache-2.0
Latency94 ms
Throughput27.2K tok/s
Cost$0.0082 /1M tok

Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.

Extraction

Output kindsClass Labels
Inputstext
Max sequence length512

Benchmarks

AG News

news classification en

Topic classification of news articles into world, sports, business, and sci/tech categories

Corpus: 7,600 Queries: 7,600
Quality
accuracy 0.7429
Performance L4 b1 c16
Extract 27.2K tok/s
Extract p50 93.8ms
Reference →

medical_questions_pairs

medical classification en

Classify whether two medical questions ask the same thing (question-pair similarity)

Corpus: 3,048 Queries: 3,048
Quality
accuracy 0.4820
Reference →

Open source inference for agents

Open-source inference for the models behind your agents. Run it yourself, or let us run it for you.

Github 2.8K

Contact us

Tell us about your use case and we'll get back to you shortly.

Apply for an inference grant

Free capacity on our hosted cluster for selected projects. Tell us what you run and we reply by email.