Qwen3-Coder-30B-A3B-Instruct
unsloth/qwen3-coder-30b-a3b-instruct
unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF on HuggingFace ↗
Throughput
Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.
Recipe
- Profile
- unsloth-qwen3-coder-30b-a3b-instruct-llama
- Engine
- llama.cpp
- Context
- 256k · q8_0 KV
- Served as
- qwen3-coder-30b-a3b-instruct
Run on your Spark
single node
shell
spark inference up unsloth-qwen3-coder-30b-a3b-instruct-llama
Why we run it
DGX Spark forum pick for agentic coding. MoE coder fits 128GB; Q4 + Q5 GGUF for llama.cpp bake-off path.
Bench notes
PBM 4k @ 42.1 tok/s — perfbench-metrics — profile=unsloth-qwen3-coder-30b-a3b-instruct-llama
Measurement history
Context ladder
Older bench-v2 / golden cells at each benched context window (single measurement).
| Context | KV | Throughput |
|---|---|---|
| @ 32k peak golden | — | 13.0t/s |
Benchmark runs
Recorded inference benchmark sessions for this model's profile.
| Date | Profile | Method | Avg | Session t/s | Range | Fill | Tool |
|---|---|---|---|---|---|---|---|
| 2026-06-28latest | unsloth-qwen3-coder-30b-a3b-instruct-llama | bench-agent-v2v2.0 | 13.0t/s | — | — | ~50,000 | fail |
| golden 256k/q8_0 @ 13.0 tok/s — fill~50000 — bench-agent-v2 — tool_ok=False | |||||||
| 2026-06-21 | unsloth-qwen3-coder-30b-a3b-instruct-llama | bench-agent | 48.7t/s | 46.0 · 45.7 · 54.3 | 45.7–54.3 | @ 32k | — |
| agent bench avg 48.7 tok/s over 3 sessions × 3 turns (2256 tok in 46.6s) | |||||||