Why did we open-source our inference engine? Read the post

← Catalog

Qwen/Qwen3-4B-Instruct-2507

Open comparison →

Primitive: /generate · Generate · Qwen3

We introduce the updated version of the Qwen3-4B non-thinking mode, named Qwen3-4B-Instruct-2507, featuring the following key enhancements:

Long contextTool callingConstrained outputStreamingCodeSQL

Overview

Hardware: — drives latency, throughput & cost

Size4.0B params
Tasks /generate
Licenseapache-2.0
Latency576 ms
Throughput472 tok/s
Cost$1.78 /1M tok

Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.

Generation

CapabilitiesTool calling · Constrained output (JSON Schema, Regex) · Streaming · Code · SQL
Context length32,768
Max output tokens4,096

Benchmarks

CaseHOLD

legal generation en

Legal holding selection from US case law (CaseHOLD)

Quality
accuracy 0.6033
Performance RTX-PRO-6000 b1 c4
Throughput 441 tok/s
p50 latency 607.3ms
Reference →

GPQA Diamond

scientific generation en

Graduate-level, expert-validated science questions, diamond subset (GPQA)

Quality
accuracy 0.4444
Performance RTX-PRO-6000 b1 c4
Throughput 495 tok/s
p50 latency 1.2s
Reference →

MedQA

medical generation en

US medical licensing exam questions (MedQA / USMLE)

Quality
accuracy 0.5700
Performance RTX-PRO-6000 b1 c4
Throughput 475 tok/s
p50 latency 545.4ms
Reference →

MMLU-Pro

general generation en

Multi-discipline reasoning across 14 subjects (MMLU-Pro)

Quality
accuracy 0.5333
Performance RTX-PRO-6000 b1 c4
Throughput 468 tok/s
p50 latency 446.4ms
Reference →

Open source inference for agents

Open-source inference for the models behind your agents. Run it yourself, or let us run it for you.

Contact us

Tell us about your use case and we'll get back to you shortly.

Apply for an inference grant

Free capacity on our hosted cluster for selected projects. Tell us what you run and we reply by email.