← Leaderboard

Ornith 1.0 35B

NVFP4 MoE · 3B active / 35B total llama.cpp General

Throughput

52 26 0 46.5 4k 19.2 50k 10.1 100k tok/s

Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.

Published evals

Vendor-card figures for deepreinforce-ai/ornith-1.0-35b — not this pack. Not measured on this Spark.

SWE-bench Verified
75.6% Ornith 1.0 model card ↗ 2026-06-25 OpenHands harness (vendor). Same figure as @ornith_ launch post.
Terminal-Bench 2.1
64.2% Ornith 1.0 launch ↗ 2026-06-25 Terminal-Bench 2.1 Terminus-2 (vendor).

Recipe

Profile
s-batman-ornith-1-0-35b-nvfp4-mtp-llama-2
Engine
llama.cpp
Context
586k · q8_0 KV
Served as
ornith-1.0-35b-nvfp4-mtp

Run on your Spark

single node shell
spark inference up s-batman-ornith-1-0-35b-nvfp4-mtp-llama-2

Why we run it

Community MTP GGUF variant — bench v2 golden workflow 2026-07-08.

Bench notes

PBM 4k @ 46.5 tok/s — perfbench-metrics — profile=s-batman-ornith-1-0-35b-nvfp4-mtp-llama-2

Measurement history

Context ladder

Older bench-v2 / golden cells at each benched context window.

Context KV Throughput
@ 192k peak golden 8.7t/s
@ 256k q8_0 6.9t/s

Benchmark runs

Recorded inference benchmark sessions for this model's profile (single run).

Date Profile Method Avg Session t/s Range Fill Tool
2026-07-08latest s-batman-ornith-1-0-35b-nvfp4-mtp-llama-2 bench-agent-v2v2.0 8.7t/s ~132,711 ok
golden 196k/q8_0 @ 8.7 tok/s — fill~133k — bench-agent-v2 — tool_ok=True
Benchmarked 2026-07-10
SparkBench · GB10 · single node