Mellum2 12B MoE Opus Thinking
yuxinlu1/mellum2-12b-opus-thinking
yuxinlu1/Mellum2-12B-A2.5B-Claude-4.6-4.8-Opus-Thinking-GGUF on HuggingFace ↗
Throughput
- @ 4k 83.8t/s
Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.
Recipe
- Profile
- mellum2-12b-opus-q4
- Engine
- llama.cpp
- Context
- 32k · q8_0 KV
- Served as
- mellum2-12b-opus-q4
Run on your Spark
single node
shell
spark inference up mellum2-12b-opus-q4
Why we run it
Golden fleet target — auto-scaffolded from recipe mellum2-12b-opus-q4.
Bench notes
PBM 4k @ 83.8 tok/s — perfbench-metrics — profile=mellum2-12b-opus-q4
Measurement history
Context ladder
Older bench-v2 / golden cells at each benched context window.
| Context | KV | Throughput |
|---|---|---|
| @ 32k peak golden | q8_0 | 74.4t/s |
| @ 64k | q8_0 | 27.0t/s |
| @ 96k | q8_0 | 19.3t/s |
| @ 128k | q8_0 | 14.6t/s |
Benchmark runs
Recorded inference benchmark sessions for this model's profile (single run).
| Date | Profile | Method | Avg | Session t/s | Range | Fill | Tool |
|---|---|---|---|---|---|---|---|
| 2026-06-28latest | mellum2-12b-opus-q4 | bench-agent-v2v2.0 | 74.4t/s | — | — | ~14,745 | ok |
| golden 32k/q8_0 @ 74.4 tok/s — fill~14745 — bench-agent-v2 — tool_ok=True | |||||||