Another DeepSeek Moment Has Arrived

Another DeepSeek Moment Has Arrived

More

Summary

Two Minute Papers host Dr. Károly Zsolnai-Fehér covers a striking update to DeepSeek’s flash model — the smaller, faster member of DeepSeek’s model family — which has just received a revision that dramatically improves benchmark performance without changing the underlying model architecture or size. Several benchmark scores more than doubled compared to the previous version, and one metric improved by 7x, all from modifying only the post-training step.

The episode uses the update as a case study in the power of post-training. The base model provides raw knowledge and capabilities; post-training teaches the model strategy — when to use which ability, how to plan, how to verify its work, and how to recover from mistakes. The analogy offered is a skilled builder who has access to the right tools but finally learns how to sequence and check their work. The underlying brain is unchanged; the playbook is entirely new.

Beyond the benchmark numbers, Dr. Zsolnai-Fehér highlights that the updated flash model is open-weights, downloadable and ownable permanently with no session caps or weekly usage limits. It requires a capable machine to run locally but is also available via the Lambda API at costs significantly below frontier proprietary models. The video closes with a forward-looking projection: if this pace of post-training improvement continues, a free open model approaching current top-tier intelligence could feasibly run on a high-end consumer laptop within a year.


📺 Source: Two Minute Papers · Published August 03, 2026
🏷️ Format: News Analysis

1 Item

Channels

1 Item

Companies