← Leaderboard

Qwen3-Coder-30B-A3B-Instruct

Q4_K_M MoE · 3B active / 30B total llama.cpp AgentsCode

Throughput

47 24 0 42.1 4k 4.7 50k 2.2 100k tok/s

Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.

Recipe

Profile
unsloth-qwen3-coder-30b-a3b-instruct-llama
Engine
llama.cpp
Context
256k · q8_0 KV
Served as
qwen3-coder-30b-a3b-instruct

Run on your Spark

single node shell
spark inference up unsloth-qwen3-coder-30b-a3b-instruct-llama

Why we run it

DGX Spark forum pick for agentic coding. MoE coder fits 128GB; Q4 + Q5 GGUF for llama.cpp bake-off path.

Bench notes

PBM 4k @ 42.1 tok/s — perfbench-metrics — profile=unsloth-qwen3-coder-30b-a3b-instruct-llama

Measurement history

Context ladder

Older bench-v2 / golden cells at each benched context window (single measurement).

Context KV Throughput
@ 32k peak golden 13.0t/s

Benchmark runs

Recorded inference benchmark sessions for this model's profile.

Date Profile Method Avg Session t/s Range Fill Tool
2026-06-28latest unsloth-qwen3-coder-30b-a3b-instruct-llama bench-agent-v2v2.0 13.0t/s ~50,000 fail
golden 256k/q8_0 @ 13.0 tok/s — fill~50000 — bench-agent-v2 — tool_ok=False
2026-06-21 unsloth-qwen3-coder-30b-a3b-instruct-llama bench-agent 48.7t/s 46.0 · 45.7 · 54.3 45.7–54.3 @ 32k
agent bench avg 48.7 tok/s over 3 sessions × 3 turns (2256 tok in 46.6s)
Benchmarked 2026-07-10
SparkBench · GB10 · single node