Summary
Two Minute Papers host Károly Zsolnai-Fehér breaks down DeepSeek’s new DSpark research, which introduces three improvements to speculative decoding — a technique that uses a small “draft” model to predict multiple tokens at once, which a larger model then verifies in a single pass. The core idea has existed for years, but DSpark addresses its practical failure modes with targeted fixes.
The three contributions are: adding a small memory buffer to the draft model so consecutive token predictions stay coherent; an early-exit mechanism that identifies “doomed” tokens unlikely to survive verification and skips checking them; and a dynamic speculation depth that adjusts how many tokens to draft based on whether the current task (structured code or math) is predictable versus open-ended (creative writing). The result is a reported 60–85% speedup on DeepSeek Flash and Pro models measured against their prior MTP1 production baseline — a real gain Zsolnai-Fehér is careful to separate from a 661% throughput figure that appears only in edge cases where the old system was already resource-constrained.
The video is honest about limitations: DSpark is not a drop-in upgrade for closed APIs and requires a matching draft model, access to the target model’s token probabilities, and a compatible serving infrastructure. The gains are also workload-dependent, with structured tasks benefiting most. For teams running self-hosted DeepSeek inference, the research offers a practical path to significantly lower latency at no additional model quality cost.
📺 Source: Two Minute Papers · Published July 07, 2026
🏷️ Format: Deep Dive







