qwythos-9b-claude-mythos-5-1m
empero-ai/qwythos-9b-claude-mythos-5-1m
empero-ai/qwythos-9b-claude-mythos-5-1m on HuggingFace ↗
Throughput
Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.
Recipe
- Profile
- empero-ai-qwythos-9b-claude-mythos-5-1m-eugr
- Engine
- vLLM
- Context
- 781k · fp8 KV
- Served as
- qwythos-9b-claude-mythos-5-1m
Run on your Spark
single node
shell
spark inference up empero-ai-qwythos-9b-claude-mythos-5-1m-eugr
Why we run it
Golden fleet target — auto-scaffolded from recipe empero-ai-qwythos-9b-claude-mythos-5-1m-eugr.
Bench notes
PBM 4k @ 12.6 tok/s — perfbench-metrics — profile=empero-ai-qwythos-9b-claude-mythos-5-1m-eugr
Measurement history
Context ladder
Older bench-v2 / golden cells at each benched context window.
| Context | KV | Throughput |
|---|---|---|
| @ 256k | fp8 | 12.2t/s |
| @ 320k | fp8 | 12.0t/s |
| @ 400k | fp8 | 12.0t/s |
| @ 500k | fp8 | 11.8t/s |
| @ 600k | fp8 | 11.6t/s |
| @ 781k peak golden | fp8 | 12.4t/s |
| @ 1M | fp8 | 10.8t/s |
Benchmark runs
Recorded inference benchmark sessions for this model's profile (single run).
| Date | Profile | Method | Avg | Session t/s | Range | Fill | Tool |
|---|---|---|---|---|---|---|---|
| 2026-06-28latest | empero-ai-qwythos-9b-claude-mythos-5-1m-eugr | bench-agent-v2v2.0 | 12.4t/s | — | — | ~50,000 | fail |
| golden 781k/fp8 @ 12.4 tok/s — fill~50000 — bench-agent-v2 — tool_ok=False | |||||||