Qwen3.6 35B A3B · MTP
nvidia/qwen3.6-35b-a3b
nvidia/qwen3.6-35b-a3b-nvfp4 on HuggingFace ↗
Throughput
Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.
Recipe
- Profile
- qwen36-35b-a3b-mtp-eugr
- Engine
- vLLM
- Context
- 256k · fp8 KV
- Served as
- nvidia/Qwen3.6-35B-A3B-NVFP4
- Draft
- MTP · n3
Run on your Spark
single node
shell
spark inference up qwen36-35b-a3b-mtp-eugr
Why we run it
Golden fleet target — auto-scaffolded from recipe opencode-qwen36-250k.
Bench notes
PBM 4k @ 86.3 / 50k @ 78.8 / 100k @ 31.5 tok/s — perfbench-metrics — profile=qwen36-35b-a3b-mtp-eugr (post eugr vLLM 1ea84d74b, 2026-08-04); Editor's pick
Measurement history
Context ladder
Older bench-v2 / golden cells at each benched context window (single measurement).
| Context | KV | Throughput |
|---|---|---|
| @ 256k peak | — | 75.7t/s |
Benchmark runs
Recorded inference benchmark sessions for this model's profile.
| Date | Profile | Method | Avg | Session t/s | Range | Fill | Tool |
|---|---|---|---|---|---|---|---|
| 2026-08-04latest | qwen36-35b-a3b-mtp-eugr | bench-agent-v2v2.0 | 75.7t/s | 81.0 · 70.4 | 70.4–81.0 | ~50,000 | fail |
| bench-v2 avg 75.7 decode tok/s (2 sessions, ~50k ctx fill, tool_ok=False) | |||||||
| 2026-06-29latest | opencode-qwen36-250k | bench-agent-v2v2.0 | 58.9t/s | 57.9 · 59.9 | 57.9–59.9 | ~50,000 | ok |
| bench-v2 58.9 tok/s (~50k fill) | llama-benchy pp2048/tg128/d0 pp=5451 tg=75.4 tok/s | |||||||
| 2026-06-21 | qwen36-nvfp4 | bench-agent | 73.5t/s | 73.3 · 73.6 · 73.6 | 73.3–73.6 | @ 64k | — |
| agent bench avg 73.5 tok/s over 3 sessions × 3 turns (2304 tok in 31.3s) | |||||||