← Leaderboard

Laguna S 2.1

NVFP4 DFlash · n15 MoE · 8B active / 118B total vLLM AgentsCode

Throughput

29 14 0 25.6 4k 23.4 50k 11.0 100k tok/s

Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.

Recipe

Profile
poolside-laguna-s-2-1-dflash-eugr
Engine
vLLM
Context
256k · fp8 KV
Served as
laguna-s-2.1-dflash
Draft
DFlash · n15

Run on your Spark

single node shell
spark inference up poolside-laguna-s-2-1-dflash-eugr

Why we run it

Poolside Laguna S 2.1 NVFP4 (118B-A8B) + DFlash-NVFP4. PBM 4k/50k/100k: DFlash 25.6/23.4/11.0 tok/s; baseline (no DFlash) 18.9/11.7/9.4. Golden: poolside-laguna-s-2-1-dflash-eugr @ 262k fp8.

Bench notes

PBM 4k @ 25.6 tok/s (50k=23.4, 100k=11.0) — DFlash; baseline no-DFlash PBM 18.9/11.7/9.4 — profile=poolside-laguna-s-2-1-dflash-eugr

Benchmarked 2026-07-23
SparkBench · GB10 · single node