Summary
Fahd Mirza demonstrates how to run DeepSeek’s DFlash speculative decoding method locally, pairing the open-source DeepSeek drafter model with Google’s Gemma 12B on a single Nvidia RTX A6000 GPU with 48GB VRAM. DFlash replaces the standard sequential single-token drafting approach with parallel block denoising: rather than guessing one token at a time based on previous guesses, the drafter reads Gemma 12B’s internal hidden states during a single forward pass and predicts an entire block of tokens simultaneously. Drafting cost stays flat regardless of block size, enabling significantly more accepted tokens per round.
The tutorial covers the full setup from scratch โ cloning the DeepSeek GitHub repo, creating a virtual environment, installing dependencies, downloading both the 12B target model and the lightweight DFlash drafter (totaling around 31GB VRAM), and running the provided eval.py benchmark script. The evaluation spans GSM8K (structured math) and MT-Bench (open chat), 20 samples each, measuring the accept length metric.
Results closely reproduce DeepSeek’s published paper numbers: an accept length of 5.44 on GSM8K versus the paper’s 5.45, and 3.01 on MT-Bench versus the paper’s 2.98 for Gemma 4 12B. Mirza contextualizes DFlash within DeepSeek’s broader inference toolkit, explaining its relationship to DSpark โ the more advanced system that bolts a Markov memory head and smart verification scheduling on top of DFlash’s parallel engine. This is a reproducible, hardware-specific walkthrough for anyone evaluating speculative decoding on local GPU setups.
๐บ Source: Fahd Mirza ยท Published July 03, 2026
๐ท๏ธ Format: Hands On Build







