Open-source · DGX Spark · GB10
What actually runs on a DGX Spark.
SparkBench is a community lab for the NVIDIA GB10. We benchmark every model on real hardware, publish reproducible recipes for coding, agents, and reasoning, and ship the tooling so you can run it on your own box.
32
Recipes benched
86.3 t/s @ 4k
Top @ 4k fill
3
Inference engines
1
GB10 box
Leaderboard
Rankings are by throughput (tok/s) on a single GB10
— not by intelligence, coding quality, or SWE benchmarks.
We can't wait to add those. Contributions welcome ↗
Measured at 4k context fill (PBM)
— decode tok/s after prefilling a fixed token context (perfbench-metrics).
Same agent-style workload as bench v2, but comparable across models at 4k / 50k / 100k fills.
| # | Model | For | Engine | Params | Curve | Throughput | |
|---|---|---|---|---|---|---|---|
| 01 | Qwen3.6 35B A3B · MTP | AgentsGeneralReasoning | vLLM | 3B / 35B |
86.3t/s@ 4k
|
HF ↗ | |
| 02 | Mellum2 12B MoE Opus Thinking | CodeReasoning | llama.cpp | 2.5B / 12B | — |
83.8t/s@ 4k
|
HF ↗ |
| 03 | Qwen3-30B-A3B | AgentsGeneralReasoning | vLLM | 3B / 30B | — |
74.2t/s@ 4k
|
HF ↗ |
| 04 | Qwen3.6-35B-A3B NVFP4 (Mia recipe) | AgentsMultimodal | vLLM | 3B / 35B |
70.3t/s@ 4k
|
HF ↗ | |
| 05 | qwen3-coder-next | AgentsCode | vLLM | 3B / 80B |
58.9t/s@ 4k
|
HF ↗ | |
| 06 | Qwen3.6 35B | Reasoning | vLLM | 3B / 35B | — |
—
|
HF ↗ |
| 07 | Qwen3.6 35B | AgentsGeneralReasoning | llama.cpp | 3B / 35B |
46.8t/s@ 4k
|
HF ↗ | |
| 08 | Ornith 1.0 35B | General | llama.cpp | 3B / 35B |
46.5t/s@ 4k
|
HF ↗ | |
| 09 | Qwen3-Coder-30B-A3B-Instruct | AgentsCode | llama.cpp | 3B / 30B |
42.1t/s@ 4k
|
HF ↗ | |
| 10 | qwen3-coder-next | AgentsCode | vLLM | 3B / 80B |
38.7t/s@ 4k
|
HF ↗ | |
| 11 | Ornith 1.0 35B | General | llama.cpp | 3B / 35B |
37.8t/s@ 4k
|
HF ↗ | |
| 12 | qwen-agentworld-35b-a3b | General | vLLM | 3B / 35B |
29.6t/s@ 4k
|
HF ↗ | |
| 13 | Laguna S 2.1 | AgentsCode | vLLM | 8B / 118B |
29.1t/s@ 4k
|
HF ↗ | |
| 14 | qwen3.6-27b-aeon-ultimate-uncensored | General | vLLM | 27B |
26.0t/s@ 4k
|
HF ↗ | |
| 15 | glm-4.7-flash-30b-a3b-q4_k_m-5gb | General | llama.cpp | 3B / 30B | — |
—
|
HF ↗ |
| 16 | Gemma-4-26B-A4B-IT | GeneralMultimodalReasoning | vLLM | 4B / 26B |
23.0t/s@ 4k
|
HF ↗ | |
| 17 | Laguna XS 2.1 | AgentsCode | vLLM | 3B / 33B |
21.4t/s@ 4k
|
HF ↗ | |
| 18 | Step-3.7-Flash | AgentsMultimodal | llama.cpp | 11B / 198B |
20.5t/s@ 4k
|
HF ↗ | |
| 19 | Gemma-4-12B-IT | GeneralMultimodal | llama.cpp | 12B |
19.8t/s@ 4k
|
HF ↗ | |
| 20 | DeepSeek V4 Flash (DwarfStar) | AgentsCodeReasoning | ds4 | 13B / 180B | — |
—
|
HF ↗ |
| 21 | Gemma 4 12B Coder (Fable5×Composer2.5) | Code | llama.cpp | 12B |
19.1t/s@ 4k
|
HF ↗ | |
| 22 | Gemma 4 12B Opus Reasoning | Reasoning | llama.cpp | 12B | — |
19.1t/s@ 4k
|
HF ↗ |
| 23 | Gemma 4 12B Agentic v2 (Fable5×Composer2.5) | AgentsCode | llama.cpp | 12B | — |
18.9t/s@ 4k
|
HF ↗ |
| 24 | Qwen3.6-27B | AgentsGeneralReasoning | vLLM | 27B |
17.3t/s@ 4k
|
HF ↗ | |
| 25 | qwythos-9b-claude-mythos-5-1m | General | vLLM | 9B |
12.6t/s@ 4k
|
HF ↗ | |
| 26 | Qwen3 6 27B | General | vLLM | 27B |
12.1t/s@ 4k
|
HF ↗ | |
| 27 | qwen3.6-27b-aeon-ultimate-uncensored-text-nvfp4-mtp-xs | General | vLLM | 27B | — |
—
|
HF ↗ |
| 28 | thinkingcap-qwen3.6-27b | General | vLLM | 27B | — |
—
|
private |
| 29 | Qwen3.6-27B | GeneralReasoning | vLLM | 27B |
10.6t/s@ 4k
|
HF ↗ | |
| 30 | Qwen3.6-27B (unsloth) | AgentsGeneralReasoning | vLLM | 27B |
9.4t/s@ 4k
|
HF ↗ | |
| 31 | Qwen3.6-27B | GeneralReasoning | llama.cpp | 27B |
9.1t/s@ 4k
|
HF ↗ | |
| 32 | thinkingcap-qwen3.6-27b | General | llama.cpp | 27B | — |
—
|
private |
Recipes by task
General
Agents
Reasoning
Code
Run it on your own Spark
SparkBench is the operator tool behind this site. Portal, model inventory, three inference engines (vLLM, llama.cpp, ds4), reproducible bench v2, and an OpenAI gateway — one CLI on your GB10.
- One bootstrap command — clone, host env, portal, APIs, CLI.
- Three engines — eugr, llama.cpp, ds4. One GPU at a time.
- Real recipes — auto-scaffolded from weights; golden map in git.
- Your data, your box — nothing leaves the LAN unless you ship it here.
install
shell
# One command — core stack (no GPU engine yet) curl -fsSL https://raw.githubusercontent.com/shawnmarck/sparkbench/main/scripts/bootstrap-sparkbench.sh | sudo bash # Then pick an engine + gateway sudo bash install/spark-install engine eugr sudo bash install/spark-install gateway # Switch, bench, serve spark inference list spark inference up qwen36-nvfp4 spark inference bench