DeepSeek’s New AI Is A Game Changer

DeepSeek’s New AI Is A Game Changer

More

Summary

Dr. Károly Zsolnai-Fehér of Two Minute Papers breaks down a new DeepSeek research paper introducing a visual reasoning technique that allows AI models to “point” at regions of an image while thinking — mirroring how humans use a finger to count objects rather than describing them in words. The core claim is striking: the approach uses roughly 90% fewer visual tokens than most frontier models while still matching or outperforming them on seven external benchmarks, with in-house benchmarks deliberately excluded to avoid the common practice of self-serving evaluation design.

Beyond counting tasks, the technique enables topological reasoning — for example, tracing a path through a maze — and produces interpretable visual thought processes that make it easier to diagnose model errors. DeepSeek achieved this through a policy distillation framework: multiple specialist teacher models (each strong at a specific visual task such as bounding boxes or maze tracing) train a single generalist student model that absorbs capabilities from all of them. The paper functions as a blueprint rather than a released model, meaning the technique could be applied to many existing systems, including open-source ones.

The video also honestly enumerates limitations: the visual pointing behavior requires a word cue to activate, fine-grained detail tasks like counting individual strands of hair remain problematic, and topological reasoning does not yet generalize robustly to out-of-distribution inputs. Dr. Zsolnai-Fehér frames the result as a meaningful step toward AI systems that are both cheaper to run and easier to understand.


📺 Source: Two Minute Papers · Published May 22, 2026
🏷️ Format: Deep Dive

1 Item

Channels

1 Item

Companies