← Leaderboard

qwen-agentworld-35b-a3b

MoE · 3B active / 35B total vLLM General

Throughput

33 17 0 29.6 4k 22.3 50k 18.7 100k tok/s

Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.

Recipe

Profile
qwen-qwen-agentworld-35b-a3b-eugr
Engine
vLLM
Context
250k · fp8 KV
Served as
qwen-agentworld-35b-a3b

Run on your Spark

single node shell
spark inference up qwen-qwen-agentworld-35b-a3b-eugr

Why we run it

Golden fleet target — auto-scaffolded from recipe qwen-qwen-agentworld-35b-a3b-eugr.

Bench notes

PBM 4k @ 29.6 tok/s — perfbench-metrics — profile=qwen-qwen-agentworld-35b-a3b-eugr

Measurement history

Context ladder

Older bench-v2 / golden cells at each benched context window (single measurement).

Context KV Throughput
@ 250k peak golden 27.6t/s

Benchmark runs

Recorded inference benchmark sessions for this model's profile (single run).

Date Profile Method Avg Session t/s Range Fill Tool
2026-06-28latest qwen-qwen-agentworld-35b-a3b-eugr bench-agent-v2v2.0 27.6t/s ~50,000 fail
golden 250k/fp8 @ 27.6 tok/s — fill~50000 — bench-agent-v2 — tool_ok=False
Benchmarked 2026-07-10
SparkBench · GB10 · single node