← Leaderboard

qwen3.6-27b-aeon-ultimate-uncensored

aeon-7/qwen3.6-27b-aeon-ultimate-uncensored aeon-7/qwen3.6-27b-aeon-ultimate-uncensored on HuggingFace ↗
NVFP4 DFlash · n10 27B vLLM General

Throughput

29 15 0 26.0 4k 14.5 50k 11.4 100k tok/s

Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.

Recipe

Profile
aeon-qwen3-6-27b-dflash-n10
Engine
vLLM
Context
32k · auto KV
Served as
qwen3.6-27b-dflash
Draft
DFlash · n10

Run on your Spark

single node shell
spark inference up aeon-qwen3-6-27b-dflash-n10

Bench notes

PBM 4k @ 26.0 tok/s — perfbench-metrics — profile=aeon-qwen3-6-27b-dflash-n10

Measurement history

Context ladder

Older bench-v2 / golden cells at each benched context window (single measurement).

Context KV Throughput
@ 32k peak — 24.5t/s

Benchmark runs

Recorded inference benchmark sessions for this model's profile (single run).

Date Profile Method Avg Session t/s Range Fill Tool
2026-07-10latest aeon-qwen3-6-27b-dflash-n10 bench-agent-v2v2.0 24.5t/s 24.5 · 24.6 24.5–24.6 ~14,745 ok
bench-v2 avg 24.5 decode tok/s (2 sessions, ~14k ctx fill, tool_ok=True)
Benchmarked 2026-07-10
SparkBench · GB10 · single node