Qwen/Qwen3-0.6B
Primitive: /generate · Generate ·
Qwen3
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual ...
View on Hugging Face → Fine-tuned from Qwen/Qwen3-0.6B-Base
Overview
Hardware: — drives latency, throughput & cost
| Size | 752M params |
|---|---|
| Tasks | /generate |
| License | apache-2.0 |
| Latency | 413 ms |
| Throughput | 595 tok/s |
| Cost | $1.41 /1M tok |
Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.
Generation
| Capabilities | Streaming |
|---|---|
| Context length | 4,096 |
| Max output tokens | 1,024 |
Benchmarks
CaseHOLD
Legal holding selection from US case law (CaseHOLD)
GPQA Diamond
Graduate-level, expert-validated science questions, diamond subset (GPQA)
MedQA
US medical licensing exam questions (MedQA / USMLE)
MMLU-Pro
Multi-discipline reasoning across 14 subjects (MMLU-Pro)