Laguna S 2.1
poolside/laguna-s-2.1
poolside/Laguna-S-2.1-NVFP4 on HuggingFace ↗
Throughput
Decode tok/s after a fixed context fill (PBM) — same agent-style workload as bench v2.
Recipe
- Profile
- poolside-laguna-s-2-1-dflash-eugr
- Engine
- vLLM
- Context
- 256k · fp8 KV
- Served as
- laguna-s-2.1-dflash
- Draft
- DFlash · n15
Run on your Spark
single node
shell
spark inference up poolside-laguna-s-2-1-dflash-eugr
Why we run it
Poolside Laguna S 2.1 NVFP4 (118B-A8B) + DFlash-NVFP4. PBM 4k/50k/100k: DFlash 25.6/23.4/11.0 tok/s; baseline (no DFlash) 18.9/11.7/9.4. Golden: poolside-laguna-s-2-1-dflash-eugr @ 262k fp8.
Bench notes
PBM 4k @ 25.6 tok/s (50k=23.4, 100k=11.0) — DFlash; baseline no-DFlash PBM 18.9/11.7/9.4 — profile=poolside-laguna-s-2-1-dflash-eugr