Qwen/Qwen3-4B-Instruct-2507
Primitive: /generate · Generate ·
Qwen3
We introduce the updated version of the Qwen3-4B non-thinking mode, named Qwen3-4B-Instruct-2507, featuring the following key enhancements:
Overview
Hardware: — drives latency, throughput & cost
| Size | 4.0B params |
|---|---|
| Tasks | /generate |
| License | apache-2.0 |
| Latency | 576 ms |
| Throughput | 472 tok/s |
| Cost | $1.78 /1M tok |
Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.
Generation
| Capabilities | Tool calling · Constrained output (JSON Schema, Regex) · Streaming · Code · SQL |
|---|---|
| Context length | 32,768 |
| Max output tokens | 4,096 |
Benchmarks
CaseHOLD
Legal holding selection from US case law (CaseHOLD)
GPQA Diamond
Graduate-level, expert-validated science questions, diamond subset (GPQA)
MedQA
US medical licensing exam questions (MedQA / USMLE)
MMLU-Pro
Multi-discipline reasoning across 14 subjects (MMLU-Pro)