Ornith 1.5 35B-A3B · MTP
This is the recipe currently serving on Sparky — the lab box we use for day-to-day testing.
Throughput
Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.
Recipe
- Profile
- ornith-ai-ornith-1-5-35b-a3b-nvfp4-b12x-eugr
- Engine
- vLLM
- Context
- 256k · fp8 KV
- Served as
- ornith-1.5-35b-a3b-nvfp4
- Draft
- MTP · n1
Run on your Spark
spark inference up ornith-ai-ornith-1-5-35b-a3b-nvfp4-b12x-eugr
Why we run it
Agentic-coding MoE (3B active / 35B) on a single Spark: NVFP4, GB10-native b12x, in-checkpoint MTP n=1, language-model-only, chunked prefill. Native 256k context; YaRN presets stretch to 512k and 1M. Tools round-trip clean. Short concurrent decode ~94 tok/s at 1 stream and ~500 tok/s aggregate at 24 streams. Daily-driver alternative to Qwen3.6-35B when you want coding speed and 1M headroom more than 50k-fill retention.
Bench notes
PBM 4k @ 70.8 / 50k @ 36.5 / 100k @ 27.2 tok/s — perfbench-metrics — NVFP4 b12x MTP n=1, language-model-only + chunked prefill. MTP n=1/2/3 C1 coding 93.8/83.1/70.8 (keep n=1). C1 @ 24 seqs 93.8 vs 4 seqs 79.7. YaRN 512k and 1M boot (presets yarn_512k / yarn_1m). Concurrent 512-out C24 ~501 agg. tool_ok.