DFlash Just Got Faster: 4x Speed with 160 tok/s Locally

DFlash Just Got Faster: 4x Speed with 160 tok/s Locally

More

Summary

Fahd Mirza benchmarks DFlash with SGLang’s new SpecV2 overlapping scheduler on an NVIDIA H100 80GB GPU, demonstrating a 4.3x throughput improvement over baseline autoregressive inference and 1.5x over native multi-token prediction at 32 concurrent requests, reaching approximately 160 tokens per second on GSM-8K.

The video explains why this upgrade matters: standard speculative decoding staggers drafting and verification sequentially, leaving the GPU idle between phases. SGLang’s SpecV2 scheduler eliminates that idle time by running the DFlash drafter on block n+1 while the target model verifies block n in parallel — adding roughly 33% end-to-end throughput on top of DFlash’s existing advantage. DFlash’s core edge over traditional speculative decoding comes from KV injection: its draft model has direct access to the target model’s hidden states, making token predictions far more accurate than a blind separate drafter. This combination is now the default speculative decoding engine in SGLang, not an experimental branch.

Mirza walks through every relevant environment variable in detail — SPEC_V2, ENABLE_DFLASH_SPEC_V2, PLAN_STREAM, and context length override flags — and provides the full SGLang launch command with annotations. The 1.3–6.27 billion parameter DFlash model consumes just under 68GB VRAM combined with the target model on H100, and the drafter proposes 16 tokens per block in a single parallel forward pass. This functions as a practical setup reference for anyone deploying local LLM inference at scale.


📺 Source: Fahd Mirza · Published June 16, 2026
🏷️ Format: Benchmark Test

1 Item

Channels