← Leaderboard

Qwen3.8-27B DFlash2

FP8 DFlash2 · n7 27B vLLM AgentsGeneralReasoning

Throughput

22 11 0 19.8 4k 6.7 50k 4.1 100k tok/s

Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.

Recipe

Profile
qwen-qwen3-8-27b-dflash2-eugr
Engine
vLLM
Context
64k · fp8 KV
Served as
qwen3.8-27b-dflash2
Draft
DFlash2 · n7

Run on your Spark

single node shell
spark inference up qwen-qwen3-8-27b-dflash2-eugr

Why we run it

Official Qwen3.8-27B FP8 (unquantized lm_head) for DFlash2 on eugr.

Bench notes

PBM 4k @ 19.8 / 50k @ 6.7 / 100k @ 4.1 tok/s — perfbench-metrics — profile=qwen-qwen3-8-27b-dflash2-eugr

Benchmarked 2026-08-19
SparkBench · GB10 · single node