← Leaderboard

Ornith 1.5 35B-A3B · MTP

Now testing NVFP4 MTP · n1 MoE · 3B active / 35B total vLLM AgentsCodeReasoning

This is the recipe currently serving on Sparky — the lab box we use for day-to-day testing.

Throughput

79 40 0 70.8 4k 36.5 50k 27.2 100k tok/s

Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.

Recipe

Profile
ornith-ai-ornith-1-5-35b-a3b-nvfp4-b12x-eugr
Engine
vLLM
Context
256k · fp8 KV
Served as
ornith-1.5-35b-a3b-nvfp4
Draft
MTP · n1

Run on your Spark

single node shell
spark inference up ornith-ai-ornith-1-5-35b-a3b-nvfp4-b12x-eugr

Why we run it

Agentic-coding MoE (3B active / 35B) on a single Spark: NVFP4, GB10-native b12x, in-checkpoint MTP n=1, language-model-only, chunked prefill. Native 256k context; YaRN presets stretch to 512k and 1M. Tools round-trip clean. Short concurrent decode ~94 tok/s at 1 stream and ~500 tok/s aggregate at 24 streams. Daily-driver alternative to Qwen3.6-35B when you want coding speed and 1M headroom more than 50k-fill retention.

Bench notes

PBM 4k @ 70.8 / 50k @ 36.5 / 100k @ 27.2 tok/s — perfbench-metrics — NVFP4 b12x MTP n=1, language-model-only + chunked prefill. MTP n=1/2/3 C1 coding 93.8/83.1/70.8 (keep n=1). C1 @ 24 seqs 93.8 vs 4 seqs 79.7. YaRN 512k and 1M boot (presets yarn_512k / yarn_1m). Concurrent 512-out C24 ~501 agg. tool_ok.

Benchmarked 2026-08-20
SparkBench · GB10 · single node