qwen3.6-27b-aeon-ultimate-uncensored
aeon-7/qwen3.6-27b-aeon-ultimate-uncensored
aeon-7/qwen3.6-27b-aeon-ultimate-uncensored on HuggingFace ↗
Throughput
Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.
Recipe
- Profile
- aeon-qwen3-6-27b-dflash-n10
- Engine
- vLLM
- Context
- 32k · auto KV
- Served as
- qwen3.6-27b-dflash
- Draft
- DFlash · n10
Run on your Spark
single node
shell
spark inference up aeon-qwen3-6-27b-dflash-n10
Bench notes
PBM 4k @ 26.0 tok/s — perfbench-metrics — profile=aeon-qwen3-6-27b-dflash-n10
Measurement history
Context ladder
Older bench-v2 / golden cells at each benched context window (single measurement).
| Context | KV | Throughput |
|---|---|---|
| @ 32k peak | — | 24.5t/s |
Benchmark runs
Recorded inference benchmark sessions for this model's profile (single run).
| Date | Profile | Method | Avg | Session t/s | Range | Fill | Tool |
|---|---|---|---|---|---|---|---|
| 2026-07-10latest | aeon-qwen3-6-27b-dflash-n10 | bench-agent-v2v2.0 | 24.5t/s | 24.5 · 24.6 | 24.5–24.6 | ~14,745 | ok |
| bench-v2 avg 24.5 decode tok/s (2 sessions, ~14k ctx fill, tool_ok=True) | |||||||