Qwen3.8-27B MTP
radixark/qwen3.8-27b
RadixArk/Qwen3.8-27B-NVFP4 on HuggingFace ↗
Throughput
Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.
Recipe
- Profile
- radixark-qwen3-8-27b-mtp-eugr
- Engine
- vLLM
- Context
- 64k · fp8 KV
- Served as
- qwen3.8-27b-mtp
- Draft
- MTP · n3
Run on your Spark
single node
shell
spark inference up radixark-qwen3-8-27b-mtp-eugr
Why we run it
Qwen3.8-27B NVFP4 target; golden serve is native MTP k=3.
Bench notes
PBM 4k @ 25.1 / 50k @ 9.0 / 100k @ 6.6 tok/s — perfbench-metrics — profile=radixark-qwen3-8-27b-mtp-eugr