glm-4.7-flash-30b-a3b-q4_k_m-5gb
inference-snaps/glm-4.7-flash-30b-a3b-q4_k_m-5gb
inference-snaps/glm-4.7-flash-30b-a3b-q4_k_m-5gb on HuggingFace ↗
Published evals
Vendor-card figures for this base model. Not measured on this Spark.
- SWE-bench Verified
- 59.2% GLM-4.7-Flash model card ↗ 2026-01-19 Vendor SWE-bench Verified. Applies to GLM-4.7-Flash weights, not this GGUF pack specifically.
Recipe
- Profile
- inference-snaps-glm-4-7-flash-30b-a3b-q4-k-m-5gb-llama
- Engine
- llama.cpp
- Context
- 32k · q8_0 KV
- Served as
- glm-4.7-flash-30b-a3b-q4
Run on your Spark
single node
shell
spark inference up inference-snaps-glm-4-7-flash-30b-a3b-q4-k-m-5gb-llama
Bench notes
bench-v2 avg 23.9 decode tok/s (2 sessions, ~14k ctx fill, tool_ok=True)
Measurement history
Context ladder
Older bench-v2 / golden cells at each benched context window (single measurement).
| Context | KV | Throughput |
|---|---|---|
| @ 32k peak | — | 23.9t/s |
Benchmark runs
Recorded inference benchmark sessions for this model's profile (single run).
| Date | Profile | Method | Avg | Session t/s | Range | Fill | Tool |
|---|---|---|---|---|---|---|---|
| 2026-07-11latest | inference-snaps-glm-4-7-flash-30b-a3b-q4-k-m-5gb-llama | bench-agent-v2v2.0 | 23.9t/s | 23.9 · 23.9 | — | ~14,745 | ok |
| bench-v2 avg 23.9 decode tok/s (2 sessions, ~14k ctx fill, tool_ok=True) | |||||||