← Leaderboard

Step-3.7-Flash

IQ4_XS MoE · 11B active / 198B total llama.cpp AgentsMultimodal

Throughput

23 11 0 20.5 4k 6.9 50k 2.5 100k tok/s

Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.

Recipe

Profile
stepfun-ai-step-3-7-flash-llama
Engine
llama.cpp
Context
256k · q8_0 KV
Served as
step-3.7-flash

Run on your Spark

single node shell
spark inference up stepfun-ai-step-3-7-flash-llama

Why we run it

StepFun frontier VLM MoE for Spark — IQ4_XS GGUF (~105 GB) + mmproj for llama.cpp. NVIDIA blog + StepFun benches validated on DGX Spark 128 GB.

Bench notes

PBM 4k @ 20.5 tok/s — perfbench-metrics — profile=stepfun-ai-step-3-7-flash-llama

Measurement history

Context ladder

Older bench-v2 / golden cells at each benched context window (single measurement).

Context KV Throughput
@ 32k peak golden 14.4t/s

Benchmark runs

Recorded inference benchmark sessions for this model's profile (single run).

Date Profile Method Avg Session t/s Range Fill Tool
2026-06-28latest stepfun-ai-step-3-7-flash-llama bench-agent-v2v2.0 14.4t/s ~50,000 ok
golden 256k/q8_0 @ 14.4 tok/s — fill~50000 — bench-agent-v2 — tool_ok=True
Benchmarked 2026-07-10
SparkBench · GB10 · single node