← Leaderboard

glm-4.7-flash-30b-a3b-q4_k_m-5gb

inference-snaps/glm-4.7-flash-30b-a3b-q4_k_m-5gb inference-snaps/glm-4.7-flash-30b-a3b-q4_k_m-5gb on HuggingFace ↗
Q4 MoE · 3B active / 30B total llama.cpp General

Recipe

Profile
inference-snaps-glm-4-7-flash-30b-a3b-q4-k-m-5gb-llama
Engine
llama.cpp
Context
32k · q8_0 KV
Served as
glm-4.7-flash-30b-a3b-q4

Run on your Spark

single node shell
spark inference up inference-snaps-glm-4-7-flash-30b-a3b-q4-k-m-5gb-llama

Bench notes

bench-v2 avg 23.9 decode tok/s (2 sessions, ~14k ctx fill, tool_ok=True)

Benchmarked 2026-07-11
SparkBench · GB10 · single node