← Leaderboard

glm-4.7-flash-30b-a3b-q4_k_m-5gb

inference-snaps/glm-4.7-flash-30b-a3b-q4_k_m-5gb inference-snaps/glm-4.7-flash-30b-a3b-q4_k_m-5gb on HuggingFace ↗
Q4 MoE · 3B active / 30B total llama.cpp General

Published evals

Vendor-card figures for this base model. Not measured on this Spark.

SWE-bench Verified
59.2% GLM-4.7-Flash model card ↗ 2026-01-19 Vendor SWE-bench Verified. Applies to GLM-4.7-Flash weights, not this GGUF pack specifically.

Recipe

Profile
inference-snaps-glm-4-7-flash-30b-a3b-q4-k-m-5gb-llama
Engine
llama.cpp
Context
32k · q8_0 KV
Served as
glm-4.7-flash-30b-a3b-q4

Run on your Spark

single node shell
spark inference up inference-snaps-glm-4-7-flash-30b-a3b-q4-k-m-5gb-llama

Bench notes

bench-v2 avg 23.9 decode tok/s (2 sessions, ~14k ctx fill, tool_ok=True)

Measurement history

Context ladder

Older bench-v2 / golden cells at each benched context window (single measurement).

Context KV Throughput
@ 32k peak — 23.9t/s

Benchmark runs

Recorded inference benchmark sessions for this model's profile (single run).

Date Profile Method Avg Session t/s Range Fill Tool
2026-07-11latest inference-snaps-glm-4-7-flash-30b-a3b-q4-k-m-5gb-llama bench-agent-v2v2.0 23.9t/s 23.9 · 23.9 — ~14,745 ok
bench-v2 avg 23.9 decode tok/s (2 sessions, ~14k ctx fill, tool_ok=True)
Benchmarked 2026-07-11
SparkBench · GB10 · single node