← Leaderboard

Mellum2 12B MoE Opus Thinking

Q4 MoE · 2.5B active / 12B total llama.cpp CodeReasoning

Throughput

  • @ 4k 83.8t/s

Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.

Recipe

Profile
mellum2-12b-opus-q4
Engine
llama.cpp
Context
32k · q8_0 KV
Served as
mellum2-12b-opus-q4

Run on your Spark

single node shell
spark inference up mellum2-12b-opus-q4

Why we run it

Golden fleet target — auto-scaffolded from recipe mellum2-12b-opus-q4.

Bench notes

PBM 4k @ 83.8 tok/s — perfbench-metrics — profile=mellum2-12b-opus-q4

Measurement history

Context ladder

Older bench-v2 / golden cells at each benched context window.

Context KV Throughput
@ 32k peak golden q8_0 74.4t/s
@ 64k q8_0 27.0t/s
@ 96k q8_0 19.3t/s
@ 128k q8_0 14.6t/s

Benchmark runs

Recorded inference benchmark sessions for this model's profile (single run).

Date Profile Method Avg Session t/s Range Fill Tool
2026-06-28latest mellum2-12b-opus-q4 bench-agent-v2v2.0 74.4t/s ~14,745 ok
golden 32k/q8_0 @ 74.4 tok/s — fill~14745 — bench-agent-v2 — tool_ok=True
Benchmarked 2026-07-10
SparkBench · GB10 · single node