Summary
This Two Minute Papers episode explores DeepSeek OCR, a research system that tackles the rising cost of long AI conversations by deliberately compressing what a model has to remember. Instead of feeding every word in as expensive text tokens, the approach renders text as an image and then compresses that image into a much smaller set of learned vision tokens, a bit like JPEG compression but built for language models.
The video walks through the headline numbers. At roughly 10x compression, the system recovers more than 96% of the original text, using about 90% less space. Pushed to nearly 20x, OCR accuracy drops to around 60%. It also explains the encoder design that keeps high-resolution documents affordable: local window attention first, then a 16x convolutional compressor, then global visual processing.
Beyond memory, the paper reports that the system can generate training data for LLMs and vision-language models at more than 200,000 pages per day on a single A100 GPU. The episode closes with honest caveats: the authors note that OCR alone does not prove true text compression, and later stress tests suggest the model may lean heavily on language priors, effectively guessing missing words from context. It is a clear, accessible look at an open system and its implications for long-context AI.
📺 Source: Two Minute Papers · Published October 11, 2026
🏷️ Format: Deep Dive







