DeepSeek V4 Flash 0731 (SparkInfer)
0xsero/deepseek-v4-flash-0731-spark
0xSero/deepseek-v4-flash-0731-spark on HuggingFace ↗
Throughput
Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.
Published evals
Vendor-card figures for antirez/deepseek-v4-flash — not this pack. Not measured on this Spark.
- SWE-bench Verified
- 79.0% DeepSeek V4 Flash model card ↗ 2026-04-26 Vendor SWE Verified (Resolved) on the Flash card. Applies to DeepSeek V4 Flash weights, not this GGUF pack specifically.
- Terminal-Bench 2.1
- 82.7% DeepSeek V4 Flash changelog ↗ 2026-07-31 DeepSeek Harness minimal mode, max tier (vendor). Applies to DeepSeek V4 Flash weights, not this GGUF pack specifically.
Recipe
- Profile
- 0xsero-deepseek-v4-flash-0731-sparkinfer
- Engine
- sparkinfer
- Context
- 375k · auto KV
- Served as
- deepseek-v4-flash-0731
- Draft
- DSpark
Run on your Spark
single node
shell
spark inference up 0xsero-deepseek-v4-flash-0731-sparkinfer
Why we run it
Mia/0xSero one-Spark EXL3 path — official FP4 still needs two Sparks.
Bench notes
bench-v2 avg 25.7 decode tok/s (2 sessions, ~50k ctx fill, tool_ok=True)