Qwen3-30B-A3B
nvidia/qwen3-30b-a3b
nvidia/Qwen3-30B-A3B-NVFP4 on HuggingFace ↗
Throughput
- @ 4k 74.2t/s
Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.
Recipe
- Profile
- nvidia-qwen3-30b-a3b-eugr
- Engine
- vLLM
- Context
- 40k · fp8 KV
- Served as
- qwen3-30b-a3b
Run on your Spark
single node
shell
spark inference up nvidia-qwen3-30b-a3b-eugr
Why we run it
Faster MoE sibling for interactive agent loops — lower latency and memory pressure when max intelligence isn't required.
Bench notes
PBM 4k @ 74.2 tok/s — perfbench-metrics — profile=nvidia-qwen3-30b-a3b-eugr
Measurement history
Context ladder
Older bench-v2 / golden cells at each benched context window (single measurement).
| Context | KV | Throughput |
|---|---|---|
| @ 40k peak golden | — | 72.3t/s |
Benchmark runs
Recorded inference benchmark sessions for this model's profile.
| Date | Profile | Method | Avg | Session t/s | Range | Fill | Tool |
|---|---|---|---|---|---|---|---|
| 2026-06-28latest | nvidia-qwen3-30b-a3b-eugr | bench-agent-v2v2.0 | 72.3t/s | — | — | ~18,432 | fail |
| golden 40k/fp8 @ 72.3 tok/s — fill~18432 — bench-agent-v2 — tool_ok=False | |||||||
| 2026-06-21 | nvidia-qwen3-30b-a3b-eugr | bench-agent | 74.6t/s | 74.2 · 74.7 · 75.0 | 74.2–75.0 | @ 40k | — |
| agent bench avg 74.6 tok/s over 3 sessions × 3 turns (2304 tok in 30.9s) | |||||||