← Leaderboard

Gemma-4-26B-A4B-IT

MoE · 4B active / 26B total vLLM GeneralMultimodalReasoning

Throughput

26 13 0 23.0 4k 22.0 50k 21.4 100k tok/s

Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.

Recipe

Profile
google-gemma-4-26b-a4b-it-eugr
Engine
vLLM
Context
256k · fp8 KV
Served as
gemma-4-26b-a4b-it

Run on your Spark

single node shell
spark inference up google-gemma-4-26b-a4b-it-eugr

Why we run it

Popular MoE pick (~3.8B active / 26B total) — same efficiency class as Qwen3.6 MoE. 256K context, text+image. Heavier than 12B but fast inference per token.

Bench notes

PBM 4k @ 23.0 tok/s — perfbench-metrics — profile=google-gemma-4-26b-a4b-it-eugr

Measurement history

Context ladder

Older bench-v2 / golden cells at each benched context window (single measurement).

Context KV Throughput
@ 8k peak golden 22.2t/s

Benchmark runs

Recorded inference benchmark sessions for this model's profile.

Date Profile Method Avg Session t/s Range Fill Tool
2026-06-28latest google-gemma-4-26b-a4b-it-eugr bench-agent-v2v2.0 22.2t/s ~50,000 fail
golden 256k/fp8 @ 22.2 tok/s — fill~50000 — bench-agent-v2 — tool_ok=False
2026-06-21 google-gemma-4-26b-a4b-it-eugr bench-agent 21.8t/s 21.2 · 21.0 · 23.0 21.0–23.0 @ 8k
agent bench avg 21.8 tok/s over 3 sessions × 3 turns (2023 tok in 93.1s)
Benchmarked 2026-07-10
SparkBench · GB10 · single node