← Leaderboard

Qwen3.6-35B-A3B NVFP4 (Mia recipe)

unsloth/qwen3.6-35b-a3b-nvfp4 unsloth/Qwen3.6-35B-A3B-NVFP4 on HuggingFace ↗
NVFP4 MTP · n2 MoE · 3B active / 35B total vLLM AgentsMultimodal

Throughput

79 39 0 70.3 4k 24.3 50k 11.2 100k tok/s

Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.

Recipe

Profile
mia-unsloth-qwen3-6-35b-a3b-nvfp4-eugr
Engine
vLLM
Context
256k · fp8 KV
Served as
unsloth/Qwen3.6-35B-A3B-NVFP4
Draft
MTP · n2

Run on your Spark

single node shell
spark inference up mia-unsloth-qwen3-6-35b-a3b-nvfp4-eugr

Why we run it

Exact HF weights for MiaAI-Lab Unsloth-Qwen3.6-35b-NVFP4-DGX-Spark recipe comparison on eugr (mixed FP8 dense + NVFP4 MoE).

Bench notes

golden 256k/fp8 @ 23.8 tok/s — fill~50000 — bench-agent-v2 — tool_ok=False

Measurement history

Context ladder

Older bench-v2 / golden cells at each benched context window (single measurement).

Context KV Throughput
@ 256k peak golden fp8 23.8t/s
Benchmarked 2026-07-30
SparkBench · GB10 · single node