Open-source · DGX Spark · GB10

What actually runs on a DGX Spark.

SparkBench is a community lab for the NVIDIA GB10. We benchmark every model on real hardware, publish reproducible recipes for coding, agents, and reasoning, and ship the tooling so you can run it on your own box.

SparkBench portal: Inference, Models, and Explore tabs
Portal — switch profiles, browse models, explore HuggingFace
32
Recipes benched
86.3 t/s @ 4k
Top @ 4k fill
3
Inference engines
1
GB10 box

Leaderboard

Sorted by PBM @ 4k · 32 recipes · Data updated 2026-08-04
Rankings are by throughput (tok/s) on a single GB10 — not by intelligence, coding quality, or SWE benchmarks. We can't wait to add those. Contributions welcome ↗
Measured at 4k context fill (PBM) — decode tok/s after prefilling a fixed token context (perfbench-metrics). Same agent-style workload as bench v2, but comparable across models at 4k / 50k / 100k fills.
# Model For Engine Params Curve Throughput
01 Qwen3.6 35B A3B · MTP
nvidia/qwen3.6-35b-a3b NVFP4 MTP · n3
AgentsGeneralReasoning vLLM 3B / 35B
86.3t/s@ 4k
02 Mellum2 12B MoE Opus Thinking
yuxinlu1/mellum2-12b-opus-thinking Q4
CodeReasoning llama.cpp 2.5B / 12B
83.8t/s@ 4k
03 Qwen3-30B-A3B
nvidia/qwen3-30b-a3b NVFP4
AgentsGeneralReasoning vLLM 3B / 30B
74.2t/s@ 4k
04 Qwen3.6-35B-A3B NVFP4 (Mia recipe)
unsloth/qwen3.6-35b-a3b-nvfp4 NVFP4 MTP · n2
AgentsMultimodal vLLM 3B / 35B
70.3t/s@ 4k
05 qwen3-coder-next
saricles/qwen3-coder-next NVFP4
AgentsCode vLLM 3B / 80B
58.9t/s@ 4k
06 Qwen3.6 35B
unsloth/qwen3.6-35b-a3b-nvfp4-fast Q4
Reasoning vLLM 3B / 35B
07 Qwen3.6 35B
unsloth/qwen3.6-35b-a3b Q4
AgentsGeneralReasoning llama.cpp 3B / 35B
46.8t/s@ 4k
08 Ornith 1.0 35B
s-batman/ornith-1.0-35b-nvfp4 NVFP4
General llama.cpp 3B / 35B
46.5t/s@ 4k
09 Qwen3-Coder-30B-A3B-Instruct
unsloth/qwen3-coder-30b-a3b-instruct Q4_K_M
AgentsCode llama.cpp 3B / 30B
42.1t/s@ 4k
10 qwen3-coder-next
qwen/qwen3-coder-next
AgentsCode vLLM 3B / 80B
38.7t/s@ 4k
11 Ornith 1.0 35B
deepreinforce-ai/ornith-1.0-35b
General llama.cpp 3B / 35B
37.8t/s@ 4k
12 qwen-agentworld-35b-a3b
qwen/qwen-agentworld-35b-a3b
General vLLM 3B / 35B
29.6t/s@ 4k
13 Laguna S 2.1
poolside/laguna-s-2.1 NVFP4 DFlash · n3
AgentsCode vLLM 8B / 118B
29.1t/s@ 4k
14 qwen3.6-27b-aeon-ultimate-uncensored
aeon-7/qwen3.6-27b-aeon-ultimate-uncensored NVFP4 DFlash · n10
General vLLM 27B
26.0t/s@ 4k
15 glm-4.7-flash-30b-a3b-q4_k_m-5gb
inference-snaps/glm-4.7-flash-30b-a3b-q4_k_m-5gb Q4
General llama.cpp 3B / 30B
16 Gemma-4-26B-A4B-IT
google/gemma-4-26b-a4b-it
GeneralMultimodalReasoning vLLM 4B / 26B
23.0t/s@ 4k
17 Laguna XS 2.1
poolside/laguna-xs-2.1 DFlash · n15
AgentsCode vLLM 3B / 33B
21.4t/s@ 4k
18 Step-3.7-Flash
stepfun-ai/step-3.7-flash IQ4_XS
AgentsMultimodal llama.cpp 11B / 198B
20.5t/s@ 4k
19 Gemma-4-12B-IT
google/gemma-4-12b-it Q4_K_M
GeneralMultimodal llama.cpp 12B
19.8t/s@ 4k
20 DeepSeek V4 Flash (DwarfStar)
antirez/deepseek-v4-flash
AgentsCodeReasoning ds4 13B / 180B
21 Gemma 4 12B Coder (Fable5×Composer2.5)
yuxinlu1/gemma-4-12b-coder-fable5-composer2.5-v1 Q4
Code llama.cpp 12B
19.1t/s@ 4k
22 Gemma 4 12B Opus Reasoning
yuxinlu1/gemma-4-12b-opus-reasoning Q4
Reasoning llama.cpp 12B
19.1t/s@ 4k
23 Gemma 4 12B Agentic v2 (Fable5×Composer2.5)
yuxinlu1/gemma-4-12b-agentic-v2 Q4
AgentsCode llama.cpp 12B
18.9t/s@ 4k
24 Qwen3.6-27B
qwen/qwen3.6-27b FP8 DFlash · n10
AgentsGeneralReasoning vLLM 27B
17.3t/s@ 4k
25 qwythos-9b-claude-mythos-5-1m
empero-ai/qwythos-9b-claude-mythos-5-1m
General vLLM 9B
12.6t/s@ 4k
26 Qwen3 6 27B
nvidia/qwen3.6-27b NVFP4
General vLLM 27B
12.1t/s@ 4k
27 qwen3.6-27b-aeon-ultimate-uncensored-text-nvfp4-mtp-xs
aeon-7/qwen3.6-27b-aeon-ultimate-uncensored-text-nvfp4-mtp-xs NVFP4 MTP
General vLLM 27B
28 thinkingcap-qwen3.6-27b
sakamakismile/thinkingcap-qwen3.6-27b
General vLLM 27B
29 Qwen3.6-27B
rdtand/qwen3.6-27b PrismaQuant
GeneralReasoning vLLM 27B
10.6t/s@ 4k
30 Qwen3.6-27B (unsloth)
unsloth/qwen3.6-27b NVFP4
AgentsGeneralReasoning vLLM 27B
9.4t/s@ 4k
31 Qwen3.6-27B
kaitchup/qwen3.6-27b
GeneralReasoning llama.cpp 27B
9.1t/s@ 4k
32 thinkingcap-qwen3.6-27b
protolabsai/thinkingcap-qwen3.6-27b Q5_K_M
General llama.cpp 27B

Recipes by task

Pick by what you're building

Run it on your own Spark

Same tool that generates this leaderboard
SparkBench portal: Inference, Models, and Explore tabs
Portal — switch profiles, browse models, explore HuggingFace

SparkBench is the operator tool behind this site. Portal, model inventory, three inference engines (vLLM, llama.cpp, ds4), reproducible bench v2, and an OpenAI gateway — one CLI on your GB10.

  • One bootstrap command — clone, host env, portal, APIs, CLI.
  • Three engines — eugr, llama.cpp, ds4. One GPU at a time.
  • Real recipes — auto-scaffolded from weights; golden map in git.
  • Your data, your box — nothing leaves the LAN unless you ship it here.
install shell
# One command — core stack (no GPU engine yet)
curl -fsSL https://raw.githubusercontent.com/shawnmarck/sparkbench/main/scripts/bootstrap-sparkbench.sh | sudo bash

# Then pick an engine + gateway
sudo bash install/spark-install engine eugr
sudo bash install/spark-install gateway

# Switch, bench, serve
spark inference list
spark inference up qwen36-nvfp4
spark inference bench