What actually runs on a DGX Spark.
SparkBench is a community lab for the NVIDIA GB10. We benchmark every model on real hardware, publish reproducible recipes for coding, agents, and reasoning, and ship the tooling so you can run it on your own box.
- Speed
- SWE
- TB
Speed vs published agents
PBM tok/s @ 4k on this GB10 vs a relative blend of published SWE-Verified (70.6–79.0) and Terminal-Bench (36.2–82.9). Hover a dot for the raw scores — the axis rescales as the board changes. Bright dots sit on the Pareto front: nothing else is both faster and better.
Leaderboard
| # | Model | For | Params | Curve | Throughput | SWE | TB | |
|---|---|---|---|---|---|---|---|---|
| 01 | Qwen3.6-35B-A3B | AgentsGeneralReasoning | 3B / 35B |
86.3t/s@ 4k
|
73.4% | 51.5% · 2.0 | AA ↗ HF ↗ | |
| 02 | Mellum2 12B MoE Opus Thinking | CodeReasoning | 2.5B / 12B | 4k |
83.8t/s@ 4k
|
— | — | HF ↗ |
| 03 | Qwen3-30B-A3B | AgentsGeneralReasoning | 3B / 30B | 4k |
74.2t/s@ 4k
|
— | — | HF ↗ |
| 04 | Ornith 1.5 35B-A3B · MTP | AgentsCodeReasoning | 3B / 35B |
70.8t/s@ 4k
|
79.0% | 67.8% · 2.1 | HF ↗ | |
| 05 | Qwen3.6-35B-A3B NVFP4 (Mia recipe) | AgentsMultimodal | 3B / 35B |
70.3t/s@ 4k
|
73.4% | 51.5% · 2.0 | AA ↗ HF ↗ | |
| 06 | qwen3-coder-next | AgentsCode | 3B / 80B |
58.9t/s@ 4k
|
70.6% | 36.2% · 2.0 | AA ↗ HF ↗ | |
| 07 | Qwen3.6 35B | Reasoning | 3B / 35B | — |
—
|
73.4% | 51.5% · 2.0 | AA ↗ HF ↗ |
| 08 | Qwen3.6 35B | AgentsGeneralReasoning | 3B / 35B |
46.8t/s@ 4k
|
73.4% | 51.5% · 2.0 | AA ↗ HF ↗ | |
| 09 | Ornith 1.0 35B | General | 3B / 35B |
46.5t/s@ 4k
|
75.6% | 64.2% · 2.1 | HF ↗ | |
| 10 | Qwen3-Coder-30B-A3B-Instruct | AgentsCode | 3B / 30B |
42.1t/s@ 4k
|
51.6% | — | HF ↗ | |
| 11 | qwen3-coder-next | AgentsCode | 3B / 80B |
38.7t/s@ 4k
|
70.6% | 36.2% · 2.0 | AA ↗ HF ↗ | |
| 12 | Qwen3.8-27B | AgentsGeneralReasoning | 27B |
38.1t/s@ 4k
|
— | 73.0% · 2.1 | AA ↗ HF ↗ | |
| 13 | Ornith 1.0 35B | General | 3B / 35B |
37.8t/s@ 4k
|
75.6% | 64.2% · 2.1 | HF ↗ | |
| 14 | Qwen3.8-Flash-Next | AgentsMultimodalReasoning | 6B / 125B |
36.1t/s@ 4k
|
— | 82.9% · 2.1 | AA ↗ HF ↗ | |
| 15 | qwen-agentworld-35b-a3b | General | 3B / 35B |
29.6t/s@ 4k
|
— | — | HF ↗ | |
| 16 | Laguna S 2.1 | AgentsCode | 8B / 118B |
29.1t/s@ 4k
|
— | 70.2% · 2.1 | HF ↗ | |
| 17 | qwen3.6-27b-aeon-ultimate-uncensored | General | 27B |
26.0t/s@ 4k
|
— | — | HF ↗ | |
| 18 | glm-4.7-flash-30b-a3b-q4_k_m-5gb | General | 3B / 30B | — |
—
|
59.2% | — | HF ↗ |
| 19 | DeepSeek V4 Flash 0731 (SparkInfer) | AgentsCodeReasoning | 13B / 180B |
23.4t/s@ 4k
|
79.0% | 82.7% · 2.1 | HF ↗ | |
| 20 | Gemma-4-26B-A4B-IT | GeneralMultimodalReasoning | 4B / 26B |
23.0t/s@ 4k
|
17.4% | 34.2% · 2.0 | AA ↗ HF ↗ | |
| 21 | Qwen3.8-27B DSpark | AgentsGeneralReasoning | 27B |
22.0t/s@ 4k
|
— | 73.0% · 2.1 | AA ↗ HF ↗ | |
| 22 | Laguna XS 2.1 | AgentsCode | 3B / 33B |
21.4t/s@ 4k
|
70.9% | 37.5% · 2.0 | HF ↗ | |
| 23 | Step-3.7-Flash | AgentsMultimodal | 11B / 198B |
20.5t/s@ 4k
|
76.5% | 59.6% · 2.1 | HF ↗ | |
| 24 | Gemma-4-12B-IT | GeneralMultimodal | 12B |
19.8t/s@ 4k
|
— | — | HF ↗ | |
| 25 | Qwen3.8-27B DFlash2 | AgentsGeneralReasoning | 27B |
19.8t/s@ 4k
|
— | 73.0% · 2.1 | AA ↗ HF ↗ | |
| 26 | DeepSeek V4 Flash (DwarfStar) | AgentsCodeReasoning | 13B / 180B | — |
—
|
79.0% | 82.7% · 2.1 | HF ↗ |
| 27 | Gemma 4 12B Coder (Fable5×Composer2.5) | Code | 12B |
19.1t/s@ 4k
|
— | — | HF ↗ | |
| 28 | Gemma 4 12B Opus Reasoning | Reasoning | 12B | 4k |
19.1t/s@ 4k
|
— | — | HF ↗ |
| 29 | Gemma 4 12B Agentic v2 (Fable5×Composer2.5) | AgentsCode | 12B | 4k |
18.9t/s@ 4k
|
— | — | HF ↗ |
| 30 | Qwen3.6-27B | AgentsGeneralReasoning | 27B |
17.3t/s@ 4k
|
77.2% | 59.3% · 2.0 | AA ↗ HF ↗ | |
| 31 | qwythos-9b-claude-mythos-5-1m | General | 9B |
12.6t/s@ 4k
|
— | — | HF ↗ | |
| 32 | Qwen3 6 27B | General | 27B |
12.1t/s@ 4k
|
77.2% | 59.3% · 2.0 | AA ↗ HF ↗ | |
| 33 | qwen3.6-27b-aeon-ultimate-uncensored-text-nvfp4-mtp-xs | General | 27B | — |
—
|
— | — | HF ↗ |
| 34 | thinkingcap-qwen3.6-27b | General | 27B | — |
—
|
— | — | private |
| 35 | Qwen3.6-27B | GeneralReasoning | 27B |
10.6t/s@ 4k
|
77.2% | 59.3% · 2.0 | AA ↗ HF ↗ | |
| 36 | Qwen3.6-27B (unsloth) | AgentsGeneralReasoning | 27B |
9.4t/s@ 4k
|
77.2% | 59.3% · 2.0 | AA ↗ HF ↗ | |
| 37 | Qwen3.6-27B | GeneralReasoning | 27B |
9.1t/s@ 4k
|
77.2% | 59.3% · 2.0 | AA ↗ HF ↗ | |
| 38 | thinkingcap-qwen3.6-27b | General | 27B | — |
—
|
— | — | private |
Recipes by task
General
Agents
Reasoning
Code
Run it on your own Spark
SparkBench is the operator tool behind this site. Portal, model inventory, three inference engines (vLLM, llama.cpp, ds4), reproducible bench v2, and an OpenAI gateway — one CLI on your GB10.
- One bootstrap command — clone, host env, portal, APIs, CLI.
- Three engines — eugr, llama.cpp, ds4. One GPU at a time.
- Real recipes — auto-scaffolded from weights; golden map in git.
- Your data, your box — nothing leaves the LAN unless you ship it here.
# One command — core stack (no GPU engine yet) curl -fsSL https://raw.githubusercontent.com/shawnmarck/sparkbench/main/scripts/bootstrap-sparkbench.sh | sudo bash # Then pick an engine + gateway sudo bash install/spark-install engine eugr sudo bash install/spark-install gateway # Switch, bench, serve spark inference list spark inference up qwen36-nvfp4 spark inference bench