← Leaderboard

Qwen3.6 35B

Q4 MoE · 3B active / 35B total llama.cpp AgentsGeneralReasoning

Throughput

52 26 0 46.8 4k 20.2 50k 7.9 100k tok/s

Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.

Recipe

Profile
qwen36-q4-llama
Engine
llama.cpp
Context
256k · q8_0 KV
Served as
qwen3.6-35b-a3b-q4

Run on your Spark

single node shell
spark inference up qwen36-q4-llama

Why we run it

Golden fleet target — auto-scaffolded from recipe qwen36-q4-llama.

Bench notes

PBM 4k @ 46.8 tok/s — perfbench-metrics — profile=qwen36-q4-llama

Measurement history

Context ladder

Older bench-v2 / golden cells at each benched context window (single measurement).

Context KV Throughput
@ 32k peak golden 33.6t/s

Benchmark runs

Recorded inference benchmark sessions for this model's profile.

Date Profile Method Avg Session t/s Range Fill Tool
2026-06-28latest qwen36-q4-llama bench-agent-v2v2.0 33.6t/s ~50,000 fail
golden 256k/q8_0 @ 33.6 tok/s — fill~50000 — bench-agent-v2 — tool_ok=False
2026-06-21 qwen36-q4-llama bench-agent 48.6t/s 48.7 · 48.6 · 48.6 48.6–48.7 @ 32k
agent bench avg 48.6 tok/s over 3 sessions × 3 turns (2304 tok in 47.4s)
Benchmarked 2026-07-10
SparkBench · GB10 · single node