← Leaderboard

Qwen3.6 35B A3B · MTP

NVFP4 MTP · n3 MoE · 3B active / 35B total vLLM AgentsGeneralReasoning

Throughput

97 48 0 86.3 4k 78.8 50k 31.5 100k tok/s

Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.

Recipe

Profile
qwen36-35b-a3b-mtp-eugr
Engine
vLLM
Context
256k · fp8 KV
Served as
nvidia/Qwen3.6-35B-A3B-NVFP4
Draft
MTP · n3

Run on your Spark

single node shell
spark inference up qwen36-35b-a3b-mtp-eugr

Why we run it

Golden fleet target — auto-scaffolded from recipe opencode-qwen36-250k.

Bench notes

PBM 4k @ 86.3 / 50k @ 78.8 / 100k @ 31.5 tok/s — perfbench-metrics — profile=qwen36-35b-a3b-mtp-eugr (post eugr vLLM 1ea84d74b, 2026-08-04); Editor's pick

Measurement history

Context ladder

Older bench-v2 / golden cells at each benched context window (single measurement).

Context KV Throughput
@ 256k peak 75.7t/s

Benchmark runs

Recorded inference benchmark sessions for this model's profile.

Date Profile Method Avg Session t/s Range Fill Tool
2026-08-04latest qwen36-35b-a3b-mtp-eugr bench-agent-v2v2.0 75.7t/s 81.0 · 70.4 70.4–81.0 ~50,000 fail
bench-v2 avg 75.7 decode tok/s (2 sessions, ~50k ctx fill, tool_ok=False)
2026-06-29latest opencode-qwen36-250k bench-agent-v2v2.0 58.9t/s 57.9 · 59.9 57.9–59.9 ~50,000 ok
bench-v2 58.9 tok/s (~50k fill) | llama-benchy pp2048/tg128/d0 pp=5451 tg=75.4 tok/s
2026-06-21 qwen36-nvfp4 bench-agent 73.5t/s 73.3 · 73.6 · 73.6 73.3–73.6 @ 64k
agent bench avg 73.5 tok/s over 3 sessions × 3 turns (2304 tok in 31.3s)
Benchmarked 2026-08-04
SparkBench · GB10 · single node