Why did we open-source our inference engine? Read the post

← Catalog

fastino/gliguard-LLMGuardrails-300M

Open comparison →

Primitive: /extract · Extract · DeBERTa

GLiGuard is a compact, encoder-based guardrail model for LLM safety moderation built on the `GLiNER2` interface. Instead of generating moderation verdicts autoregressively, it treats safety as structured classification: you provide task names and candidate labels at inference time, and the model ...

Overview

Hardware: — drives latency, throughput & cost

Size208M params
Tasks /extract
Licenseapache-2.0
Languagesen
Latency55 ms
Throughput47.5K tok/s
Cost$0.0047 /1M tok

Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.

Extraction

Output kindsClass Labels
Inputstext
Max sequence length512

Benchmarks

toxic-chat

safety classification en

Toxicity classification of real user-AI conversations (ToxicChat)

Corpus: 2,853 Queries: 2,853
Quality
f1 0.5756
precision 0.4154
recall 0.9365
accuracy 0.9016
Performance L4 b1 c16
Extract 47.5K tok/s
Extract p50 55.2ms
Reference →

Open source inference for agents

Open-source inference for the models behind your agents. Run it yourself, or let us run it for you.

Github 2.8K

Contact us

Tell us about your use case and we'll get back to you shortly.

Apply for an inference grant

Free capacity on our hosted cluster for selected projects. Tell us what you run and we reply by email.