Qwen3.6-35B-A3B NVFP4 (Mia recipe)
Throughput
Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.
Published evals
Vendor-card figures for qwen/qwen3.6-35b-a3b — not this pack. Not measured on this Spark.
- SWE-bench Verified
- 73.4% Qwen model card ↗ 2026-04-21 Vendor figure; internal agent scaffold (bash + file-edit).
- Terminal-Bench 2.0
- 51.5% Qwen launch blog ↗ 2026-04-14 Vendor figure on the 35B-A3B launch post.
- SWE-bench Pro
- 49.5% Qwen launch blog ↗ 2026-04-14
Independent board on Artificial Analysis ↗ — we do not copy their scores.
Recipe
- Profile
- mia-unsloth-qwen3-6-35b-a3b-nvfp4-eugr
- Engine
- vLLM
- Context
- 256k · fp8 KV
- Served as
- unsloth/Qwen3.6-35B-A3B-NVFP4
- Draft
- MTP · n2
Run on your Spark
single node
shell
spark inference up mia-unsloth-qwen3-6-35b-a3b-nvfp4-eugr
Why we run it
Exact HF weights for MiaAI-Lab Unsloth-Qwen3.6-35b-NVFP4-DGX-Spark recipe comparison on eugr (mixed FP8 dense + NVFP4 MoE).
Bench notes
golden 256k/fp8 @ 23.8 tok/s — fill~50000 — bench-agent-v2 — tool_ok=False
Measurement history
Context ladder
Older bench-v2 / golden cells at each benched context window.
| Context | KV | Throughput |
|---|---|---|
| @ 256k peak golden | — | 23.8t/s |
| @ 256k peak golden | fp8 | 23.8t/s |
Benchmark runs
Recorded inference benchmark sessions for this model's profile (single run).
| Date | Profile | Method | Avg | Session t/s | Range | Fill | Tool |
|---|---|---|---|---|---|---|---|
| 2026-07-30latest | mia-unsloth-qwen3-6-35b-a3b-nvfp4-eugr | bench-agent-v2v2.0 | 23.8t/s | — | — | ~50,000 | fail |
| golden 256k/fp8 @ 23.8 tok/s — fill~50000 — bench-agent-v2 — tool_ok=False | |||||||