Summary
Two Minute Papers host Dr. Károly Zsolnai-Fehér breaks down the newly disclosed architectural secret behind Google DeepMind’s Gemma 4, explaining how a 12-billion-parameter model running on a laptop achieves multimodal reasoning that trillion-parameter giants like DeepSeek cannot.
The key insight is that Gemma 4’s larger variant eliminates separate vision and audio encoders entirely. Rather than routing images through a dedicated vision transformer, the model slices them into patches and projects pixels directly as tokens into the main transformer. Audio gets the same treatment — sliced into 40-millisecond chunks and fed as tokens — forcing the system to learn perception and reasoning as a single unified process. This removes hundreds of millions of specialist parameters while blurring the boundary between seeing and thinking.
Zsolnai-Fehér argues this is among the most significant architectural ideas in recent open AI development. Gemma 4 has been downloaded over 300 million times and handles images, audio, and complex reasoning simultaneously at a fraction of the compute cost of closed competitors. He also notes that publishing the architectural details openly could benefit other projects like DeepSeek, which currently lack efficient multimodal capabilities despite having far more parameters — making Gemma 4 a potential blueprint for the broader open-source ecosystem.
📺 Source: Two Minute Papers · Published August 07, 2026
🏷️ Format: Deep Dive







