Qwen/Qwen3.5-4B
Primitive: /generate · Generate ·
Qwen3 MoE
> This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.
View on Hugging Face → Fine-tuned from Qwen/Qwen3.5-4B-Base
Overview
Hardware: — drives latency, throughput & cost
| Size | 4.7B params |
|---|---|
| Tasks | /generate |
| License | apache-2.0 |
| Latency | 762 ms |
| Throughput | 353 tok/s |
| Cost | $2.38 /1M tok |
Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.
Generation
| Capabilities | Tool calling · Constrained output (JSON Schema, Regex) · Streaming |
|---|---|
| Context length | 8,192 |
| Max output tokens | 4,096 |
Benchmarks
CaseHOLD
Legal holding selection from US case law (CaseHOLD)
GPQA Diamond
Graduate-level, expert-validated science questions, diamond subset (GPQA)
MedQA
US medical licensing exam questions (MedQA / USMLE)
MMLU-Pro
Multi-discipline reasoning across 14 subjects (MMLU-Pro)