← Leaderboard

Gemma 4 12B Opus Reasoning

Q4 12B llama.cpp Reasoning

Throughput

  • @ 4k 19.1t/s

Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.

Recipe

Profile
gemma4-12b-opus-reasoning-q4
Engine
llama.cpp
Context
32k · q8_0 KV
Served as
gemma4-12b-opus-reasoning-q4

Run on your Spark

single node shell
spark inference up gemma4-12b-opus-reasoning-q4

Why we run it

Golden fleet target — auto-scaffolded from recipe gemma4-12b-opus-reasoning-q4.

Bench notes

PBM 4k @ 19.1 tok/s — perfbench-metrics — profile=gemma4-12b-opus-reasoning-q4

Measurement history

Context ladder

Older bench-v2 / golden cells at each benched context window (single measurement).

Context KV Throughput
@ 32k peak golden 17.3t/s

Benchmark runs

Recorded inference benchmark sessions for this model's profile (single run).

Date Profile Method Avg Session t/s Range Fill Tool
2026-06-28latest gemma4-12b-opus-reasoning-q4 bench-agent-v2v2.0 17.3t/s ~14,745 ok
golden 32k/q8_0 @ 17.3 tok/s — fill~14745 — bench-agent-v2 — tool_ok=True
Benchmarked 2026-07-10
SparkBench · GB10 · single node