Open-source · DGX Spark · GB10

What actually runs on a DGX Spark.

SparkBench is a community lab for the NVIDIA GB10. We benchmark every model on real hardware, publish reproducible recipes for coding, agents, and reasoning, and ship the tooling so you can run it on your own box.

SparkBench portal: Inference, Models, and Explore tabs
Portal — switch profiles, browse models, explore HuggingFace

Speed vs published agents

11 recipes with a published SWE or TB score

PBM tok/s @ 4k on this GB10 vs a relative blend of published SWE-Verified (70.6–79.0) and Terminal-Bench (36.2–82.9). Hover a dot for the raw scores — the axis rescales as the board changes. Bright dots sit on the Pareto front: nothing else is both faster and better.

Leaderboard

Sorted by PBM @ 4k · 38 recipes · Data updated 2026-09-11
Default ranking is throughput (tok/s) on a single GB10 — SWE and Terminal-Bench columns are published vendor figures for the base model, not measured on this Spark. Click a column header to sort. How we cite ↗
Measured at 4k context fill (PBM) — decode tok/s after prefilling a fixed token context (perfbench-metrics). Full ladder is 4k + 50k (and 100k when the window fits), one recipe per model family — the fastest pack. Weaker quants and 4k-only benches live under All recipes.
# Model For Params Curve Throughput SWE TB
01 Qwen3.6-35B-A3B
nvidia/qwen3.6-35b-a3b NVFP4 MTP · n3
AgentsGeneralReasoning 3B / 35B
86.3t/s@ 4k
73.4% 51.5% · 2.0
02 Mellum2 12B MoE Opus Thinking
yuxinlu1/mellum2-12b-opus-thinking Q4
CodeReasoning 2.5B / 12B 4k
83.8t/s@ 4k
— —
03 Qwen3-30B-A3B
nvidia/qwen3-30b-a3b NVFP4
AgentsGeneralReasoning 3B / 30B 4k
74.2t/s@ 4k
— —
04 Ornith 1.5 35B-A3B · MTP
ornith-ai/ornith-1.5-35b-a3b NVFP4 MTP · n1 Editor's pick
AgentsCodeReasoning 3B / 35B
70.8t/s@ 4k
79.0% 67.8% · 2.1
05 Qwen3.6-35B-A3B NVFP4 (Mia recipe)
unsloth/qwen3.6-35b-a3b-nvfp4 NVFP4 MTP · n2
AgentsMultimodal 3B / 35B
70.3t/s@ 4k
73.4% 51.5% · 2.0
06 qwen3-coder-next
saricles/qwen3-coder-next NVFP4
AgentsCode 3B / 80B
58.9t/s@ 4k
70.6% 36.2% · 2.0
07 Qwen3.6 35B
unsloth/qwen3.6-35b-a3b-nvfp4-fast Q4
Reasoning 3B / 35B —
—
73.4% 51.5% · 2.0
08 Qwen3.6 35B
unsloth/qwen3.6-35b-a3b Q4
AgentsGeneralReasoning 3B / 35B
46.8t/s@ 4k
73.4% 51.5% · 2.0
09 Ornith 1.0 35B
s-batman/ornith-1.0-35b-nvfp4 NVFP4
General 3B / 35B
46.5t/s@ 4k
75.6% 64.2% · 2.1
10 Qwen3-Coder-30B-A3B-Instruct
unsloth/qwen3-coder-30b-a3b-instruct Q4_K_M
AgentsCode 3B / 30B
42.1t/s@ 4k
51.6% —
11 qwen3-coder-next
qwen/qwen3-coder-next
AgentsCode 3B / 80B
38.7t/s@ 4k
70.6% 36.2% · 2.0
12 Qwen3.8-27B
radixark/qwen3.8-27b NVFP4 DFlash2 · n8
AgentsGeneralReasoning 27B
38.1t/s@ 4k
— 73.0% · 2.1
13 Ornith 1.0 35B
deepreinforce-ai/ornith-1.0-35b
General 3B / 35B
37.8t/s@ 4k
75.6% 64.2% · 2.1
14 Qwen3.8-Flash-Next
mia-ailab/qwen3.8-flash-next NVFP4 MTP · n3
AgentsMultimodalReasoning 6B / 125B
36.1t/s@ 4k
— 82.9% · 2.1
15 qwen-agentworld-35b-a3b
qwen/qwen-agentworld-35b-a3b
General 3B / 35B
29.6t/s@ 4k
— —
16 Laguna S 2.1
poolside/laguna-s-2.1 NVFP4 DFlash · n3
AgentsCode 8B / 118B
29.1t/s@ 4k
— 70.2% · 2.1
17 qwen3.6-27b-aeon-ultimate-uncensored
aeon-7/qwen3.6-27b-aeon-ultimate-uncensored NVFP4 DFlash · n10
General 27B
26.0t/s@ 4k
— —
18 glm-4.7-flash-30b-a3b-q4_k_m-5gb
inference-snaps/glm-4.7-flash-30b-a3b-q4_k_m-5gb Q4
General 3B / 30B —
—
59.2% —
19 DeepSeek V4 Flash 0731 (SparkInfer)
0xsero/deepseek-v4-flash-0731-spark EXL3 DSpark
AgentsCodeReasoning 13B / 180B
23.4t/s@ 4k
79.0% 82.7% · 2.1
20 Gemma-4-26B-A4B-IT
google/gemma-4-26b-a4b-it
GeneralMultimodalReasoning 4B / 26B
23.0t/s@ 4k
17.4% 34.2% · 2.0
21 Qwen3.8-27B DSpark
radixark/qwen3.8-27b-dspark-nvfp4 NVFP4 DSpark · n7
AgentsGeneralReasoning 27B
22.0t/s@ 4k
— 73.0% · 2.1
22 Laguna XS 2.1
poolside/laguna-xs-2.1 DFlash · n15
AgentsCode 3B / 33B
21.4t/s@ 4k
70.9% 37.5% · 2.0
23 Step-3.7-Flash
stepfun-ai/step-3.7-flash IQ4_XS
AgentsMultimodal 11B / 198B
20.5t/s@ 4k
76.5% 59.6% · 2.1
24 Gemma-4-12B-IT
google/gemma-4-12b-it Q4_K_M
GeneralMultimodal 12B
19.8t/s@ 4k
— —
25 Qwen3.8-27B DFlash2
qwen/qwen3.8-27b FP8 DFlash2 · n7
AgentsGeneralReasoning 27B
19.8t/s@ 4k
— 73.0% · 2.1
26 DeepSeek V4 Flash (DwarfStar)
antirez/deepseek-v4-flash DSpark
AgentsCodeReasoning 13B / 180B —
—
79.0% 82.7% · 2.1
27 Gemma 4 12B Coder (Fable5×Composer2.5)
yuxinlu1/gemma-4-12b-coder-fable5-composer2.5-v1 Q4
Code 12B
19.1t/s@ 4k
— —
28 Gemma 4 12B Opus Reasoning
yuxinlu1/gemma-4-12b-opus-reasoning Q4
Reasoning 12B 4k
19.1t/s@ 4k
— —
29 Gemma 4 12B Agentic v2 (Fable5×Composer2.5)
yuxinlu1/gemma-4-12b-agentic-v2 Q4
AgentsCode 12B 4k
18.9t/s@ 4k
— —
30 Qwen3.6-27B
qwen/qwen3.6-27b FP8 DFlash · n10
AgentsGeneralReasoning 27B
17.3t/s@ 4k
77.2% 59.3% · 2.0
31 qwythos-9b-claude-mythos-5-1m
empero-ai/qwythos-9b-claude-mythos-5-1m
General 9B
12.6t/s@ 4k
— —
32 Qwen3 6 27B
nvidia/qwen3.6-27b NVFP4
General 27B
12.1t/s@ 4k
77.2% 59.3% · 2.0
33 qwen3.6-27b-aeon-ultimate-uncensored-text-nvfp4-mtp-xs
aeon-7/qwen3.6-27b-aeon-ultimate-uncensored-text-nvfp4-mtp-xs NVFP4 MTP
General 27B —
—
— —
34 thinkingcap-qwen3.6-27b
sakamakismile/thinkingcap-qwen3.6-27b
General 27B —
—
— —
35 Qwen3.6-27B
rdtand/qwen3.6-27b PrismaQuant
GeneralReasoning 27B
10.6t/s@ 4k
77.2% 59.3% · 2.0
36 Qwen3.6-27B (unsloth)
unsloth/qwen3.6-27b NVFP4
AgentsGeneralReasoning 27B
9.4t/s@ 4k
77.2% 59.3% · 2.0
37 Qwen3.6-27B
kaitchup/qwen3.6-27b
GeneralReasoning 27B
9.1t/s@ 4k
77.2% 59.3% · 2.0
38 thinkingcap-qwen3.6-27b
protolabsai/thinkingcap-qwen3.6-27b Q5_K_M
General 27B —
—
— —

Recipes by task

Pick by what you're building

Run it on your own Spark

Same tool that generates this leaderboard
SparkBench portal: Inference, Models, and Explore tabs
Portal — switch profiles, browse models, explore HuggingFace

SparkBench is the operator tool behind this site. Portal, model inventory, three inference engines (vLLM, llama.cpp, ds4), reproducible bench v2, and an OpenAI gateway — one CLI on your GB10.

  • One bootstrap command — clone, host env, portal, APIs, CLI.
  • Three engines — eugr, llama.cpp, ds4. One GPU at a time.
  • Real recipes — auto-scaffolded from weights; golden map in git.
  • Your data, your box — nothing leaves the LAN unless you ship it here.
install shell
# One command — core stack (no GPU engine yet)
curl -fsSL https://raw.githubusercontent.com/shawnmarck/sparkbench/main/scripts/bootstrap-sparkbench.sh | sudo bash

# Then pick an engine + gateway
sudo bash install/spark-install engine eugr
sudo bash install/spark-install gateway

# Switch, bench, serve
spark inference list
spark inference up qwen36-nvfp4
spark inference bench