Summary
Matthew Berman breaks down OpenAI’s sweeping price cuts across the GPT-5.6 model family and argues the mechanism behind the cuts is more consequential than the cuts themselves. GPT-5.6 Luna, the smallest of the three models, received an 80% price reduction—now priced at $0.20 per million input tokens and $1.20 per million output tokens, undercutting frontier open-source models from China. GPT-5.6 Terra received a 20% cut to $2/$12 per million tokens, while Soul gained a 2.5x fast-mode speed increase at the same 2x price premium.
The video’s central argument is that these gains were achieved through recursive self-improvement: OpenAI used GPT-5.6 Soul, integrated with Codex, to continuously analyze production traffic, identify inefficiencies, redesign GPU kernels, and optimize speculative decoding—resulting in 20% lower serving costs and 15% better token generation efficiency. Berman draws a direct parallel to Andrej Karpathy’s auto-research open-source project and frames this as an automated AI research loop now running at frontier scale.
Using Artificial Analysis benchmark data, Berman compares Luna’s cost-per-completed-task (approximately $0.06) against competitors including GLM 5.2 Max, Kimi K3, and Claude Opus 5, showing Luna’s strong value-per-task position. The video closes with a pointed question about competitive dynamics: if the leading labs are entering recursive self-improvement cycles, how do smaller labs close the gap?
📺 Source: Matthew Berman · Published July 31, 2026
🏷️ Format: News Analysis







