← Leaderboard

Qwen3.8-27B DSpark

NVFP4 DSpark · n7 27B vLLM AgentsGeneralReasoning

Throughput

25 12 0 22.0 4k 10.1 50k 6.4 100k tok/s

Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.

Published evals

Vendor-card figures for qwen/qwen3.8-27b — not this pack. Not measured on this Spark.

Terminal-Bench 2.1
73.0% Qwen model card ↗ 2026-08-14 Terminal-Bench 2.1 with Terminus (vendor).
SWE-bench Pro
61.7% Qwen model card ↗ 2026-08-14 Claude Code harness; refined SWE-bench Pro set (vendor).
LiveCodeBench
90.3% Qwen model card ↗ 2026-08-14 LiveCodeBench v6 (vendor).
GPQA Diamond
89.2% Qwen model card ↗ 2026-08-14

Recipe

Profile
radixark-qwen3-8-27b-dspark-eugr
Engine
vLLM
Context
64k · fp8 KV
Served as
qwen3.8-27b-dspark
Draft
DSpark · n7

Run on your Spark

single node shell
spark inference up radixark-qwen3-8-27b-dspark-eugr

Why we run it

Public identity for NVFP4 + DSpark (Doopeworld drafter). Weights stay at radixark/qwen3.8-27b so MTP and DSpark do not collapse to one site row.

Bench notes

PBM 4k @ 22.0 / 50k @ 10.1 / 100k @ 6.4 tok/s — perfbench-metrics — profile=radixark-qwen3-8-27b-dspark-eugr

Benchmarked 2026-08-19
SparkBench · GB10 · single node