Step-3.7-Flash
stepfun-ai/step-3.7-flash
stepfun-ai/Step-3.7-Flash on HuggingFace ↗
Throughput
Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.
Recipe
- Profile
- stepfun-ai-step-3-7-flash-llama
- Engine
- llama.cpp
- Context
- 256k · q8_0 KV
- Served as
- step-3.7-flash
Run on your Spark
single node
shell
spark inference up stepfun-ai-step-3-7-flash-llama
Why we run it
StepFun frontier VLM MoE for Spark — IQ4_XS GGUF (~105 GB) + mmproj for llama.cpp. NVIDIA blog + StepFun benches validated on DGX Spark 128 GB.
Bench notes
PBM 4k @ 20.5 tok/s — perfbench-metrics — profile=stepfun-ai-step-3-7-flash-llama
Measurement history
Context ladder
Older bench-v2 / golden cells at each benched context window (single measurement).
| Context | KV | Throughput |
|---|---|---|
| @ 32k peak golden | — | 14.4t/s |
Benchmark runs
Recorded inference benchmark sessions for this model's profile (single run).
| Date | Profile | Method | Avg | Session t/s | Range | Fill | Tool |
|---|---|---|---|---|---|---|---|
| 2026-06-28latest | stepfun-ai-step-3-7-flash-llama | bench-agent-v2v2.0 | 14.4t/s | — | — | ~50,000 | ok |
| golden 256k/q8_0 @ 14.4 tok/s — fill~50000 — bench-agent-v2 — tool_ok=True | |||||||