← Leaderboard

DeepSeek V4 Flash 0731 (SparkInfer)

0xsero/deepseek-v4-flash-0731-spark 0xSero/deepseek-v4-flash-0731-spark on HuggingFace ↗
EXL3 DSpark MoE · 13B active / 180B total sparkinfer AgentsCodeReasoning

Throughput

29 15 0 23.4 4k 26.0 50k 8.4 100k tok/s

Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.

Published evals

Vendor-card figures for antirez/deepseek-v4-flash — not this pack. Not measured on this Spark.

SWE-bench Verified
79.0% DeepSeek V4 Flash model card ↗ 2026-04-26 Vendor SWE Verified (Resolved) on the Flash card. Applies to DeepSeek V4 Flash weights, not this GGUF pack specifically.
Terminal-Bench 2.1
82.7% DeepSeek V4 Flash changelog ↗ 2026-07-31 DeepSeek Harness minimal mode, max tier (vendor). Applies to DeepSeek V4 Flash weights, not this GGUF pack specifically.

Recipe

Profile
0xsero-deepseek-v4-flash-0731-sparkinfer
Engine
sparkinfer
Context
375k · auto KV
Served as
deepseek-v4-flash-0731
Draft
DSpark

Run on your Spark

single node shell
spark inference up 0xsero-deepseek-v4-flash-0731-sparkinfer

Why we run it

Mia/0xSero one-Spark EXL3 path — official FP4 still needs two Sparks.

Bench notes

bench-v2 avg 25.7 decode tok/s (2 sessions, ~50k ctx fill, tool_ok=True)

Benchmarked 2026-08-23
SparkBench · GB10 · single node