← Leaderboard

Gemma-4-12B-IT

Q4_K_M 12B llama.cpp GeneralMultimodal

Throughput

22 11 0 19.8 4k 10.0 50k 4.0 100k tok/s

Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.

Recipe

Profile
google-gemma-4-12b-it-llama
Engine
llama.cpp
Context
256k · q8_0 KV
Served as
gemma-4-12b-it

Run on your Spark

single node shell
spark inference up google-gemma-4-12b-it-llama

Why we run it

Best mid-size Gemma 4 for a single Spark: encoder-free unified multimodal (text/image/audio), 256K context, near-26B-MoE quality at ~12B dense. Apache 2.0, not gated. Prefer -it over base.

Bench notes

PBM 4k @ 19.8 tok/s — perfbench-metrics — profile=google-gemma-4-12b-it-llama

Measurement history

Context ladder

Older bench-v2 / golden cells at each benched context window (single measurement).

Context KV Throughput
@ 32k peak golden 14.1t/s

Benchmark runs

Recorded inference benchmark sessions for this model's profile.

Date Profile Method Avg Session t/s Range Fill Tool
2026-06-28latest google-gemma-4-12b-it-llama bench-agent-v2v2.0 14.1t/s ~50,000 ok
golden 256k/q8_0 @ 14.1 tok/s — fill~50000 — bench-agent-v2 — tool_ok=True
2026-06-21 google-gemma-4-12b-it-llama bench-agent 21.1t/s 21.1 · 21.1 · 20.9 20.9–21.1 @ 32k
agent bench avg 21.1 tok/s over 3 sessions × 3 turns (2304 tok in 109.4s)
Benchmarked 2026-07-10
SparkBench · GB10 · single node