← Leaderboard

Qwen3.8-27B MTP

NVFP4 MTP · n3 27B vLLM AgentsGeneralReasoning

Throughput

28 14 0 25.1 4k 9.0 50k 6.6 100k tok/s

Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.

Recipe

Profile
radixark-qwen3-8-27b-mtp-eugr
Engine
vLLM
Context
64k · fp8 KV
Served as
qwen3.8-27b-mtp
Draft
MTP · n3

Run on your Spark

single node shell
spark inference up radixark-qwen3-8-27b-mtp-eugr

Why we run it

Qwen3.8-27B NVFP4 target; golden serve is native MTP k=3.

Bench notes

PBM 4k @ 25.1 / 50k @ 9.0 / 100k @ 6.6 tok/s — perfbench-metrics — profile=radixark-qwen3-8-27b-mtp-eugr

Benchmarked 2026-08-19
SparkBench · GB10 · single node