← Leaderboard

Qwen3.8-Flash-Next

NVFP4 MTP · n3 MoE · 6B active / 125B total flashnext AgentsMultimodalReasoning

Throughput

40 20 0 36.1 4k 21.6 50k 19.6 100k tok/s

Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.

Published evals

Vendor-card figures for qwen/qwen3.8-flash-next — not this pack. Not measured on this Spark.

Terminal-Bench 2.1
82.9% NVIDIA NVFP4 model card ↗ 2026-08-31 NVIDIA NVFP4 column; temp=1.0, top_p=0.95, reasoning_effort=xhigh. FP8 baseline on the same card is 83.3. Not the Mia pack, not measured on Sparky.
SWE-bench Pro
62.5% Qwen model card ↗ 2026-08-26 Claude Code harness; refined SWE-bench Pro set (vendor). DeepSWE 1.1 is 58.7 on the same card (not a site column).
LiveCodeBench
91.9% Qwen model card ↗ 2026-08-26 LiveCodeBench v6 (vendor).
GPQA Diamond
91.7% Qwen model card ↗ 2026-08-26

Recipe

Profile
mia-ailab-qwen3.8-flash-next-vllm
Engine
flashnext
Context
256k · fp8 KV
Served as
qwen3.8-flash-next
Draft
MTP · n3

Run on your Spark

single node shell
spark inference up mia-ailab-qwen3.8-flash-next-vllm

Why we run it

Qwen3.8-Flash-Next 125B-A6B VLM on one Spark via MiaAI NVFP4 (~99 GiB) plus a first-boot PLE mmap table (~27 GiB). Official NVIDIA NVFP4 ~135 GB does not fit. Served by spark engine flashnext (vllm/vllm-openai:qwen38-flash-next), not eugr. Companion A/B to golden 27B SGLang DFlash2. Site row is native PBM 36.1/21.6/19.6; recipe stays testing until bench v2.

Bench notes

PBM 4k @ 36.1 / 50k @ 21.6 / 100k @ 19.6 tok/s — perfbench-metrics — native 262k MTP3 vision on — profile=mia-ailab-qwen3.8-flash-next-vllm — YaRN 512k companion 36.3/22.0/18.1 plus 200k=31.3 / 300k=12.7

Benchmarked 2026-09-11
SparkBench · GB10 · single node