Descriptions:
Wes Roth covers the launch of OpenAI’s GPT 5.6 family — Soul (flagship), Terra (mid-tier), and Luna (cost-efficient scale model) — with particular focus on a striking architectural detail: Soul autonomously post-trained Luna using a Codex prompt, representing a concrete example of AI models contributing to the training of subsequent models. Roth connects this to Andrej Karpathy’s open-source Auto Research project and broader ideas around recursive self-improvement, noting that OpenAI researchers described automated research pipelines as feeling close.
On benchmarks, the analysis is granular. GPT 5.6 Soul scores 53.6% on Agents Last Exam at a cost of $763, compared to Fable 5’s 40.5% at $2,300. On Artificial Analysis’s coding agent index, Soul hits 80 — 2.8 points above Fable 5 — while using less than half the output tokens at roughly one-third the cost. Roth walks through Terminal Bench 2.1 and Deep SWE results showing the same pattern: Soul leads on the intelligence-versus-cost curve. A new Ultra reasoning mode is also introduced, sitting above the existing Max tier.
The video also covers ChatGPT Work, OpenAI’s new desktop app competing with Claude’s similar offering, and demonstrates Soul’s improved design capabilities through a live Project Zomboid-style survival game built via computer use. For teams making model selection decisions, the cost-efficiency data presented here is directly actionable.
📺 Source: Wes Roth · Published July 09, 2026
🏷️ Format: Review







