A Model Explosion: GPT 5.6 Sol, Grok 4.5 and Meta Muse Rewrite the Rules

A Model Explosion: GPT 5.6 Sol, Grok 4.5 and Meta Muse Rewrite the Rules

More

Summary

AI Explained delivers one of the more analytically rigorous breakdowns of a frantic 24-hour model release window, covering GPT 5.6 Soul, Terror, and Luna from OpenAI; Grok 4.5 from xAI; and Meta’s Muse Spark 1.1 — three frontier releases that collectively reframe the value proposition around near-frontier capability at reduced cost.

The video goes deep on Agent’s Last Exam, a UC Berkeley-co-led benchmark spanning 55 industries with 300 experts crafting long-horizon tasks proven to have economic value. GPT 5.6 Soul scores 54% versus Claude Fable’s 45%, and the creator contextualizes this by noting that no single benchmark crossing triggered the shift from hand-coding to AI-first coding — raising the question of whether similar tipping points may be approaching in finance, operations, and other white-collar domains. Zapier’s Automation Bench, GDP-Val Elo ratings, and Artificial Analysis’s aggregate coding index are also covered, with consistent findings that Soul roughly matches or edges out Fable at about a third of the price.

Meta’s Muse Spark emerges as an unexpected wildcard: scoring within striking distance of Soul on Vibe Code Bench at approximately 35 times lower cost. The creator also runs a personal game-generation comparison between Soul Ultra and Fable, noting Soul’s speed advantage despite some interface regressions, and flags a separate Anthropic paper on AI consciousness as worth tracking independently.


📺 Source: AI Explained · Published July 10, 2026
🏷️ Format: News Analysis

1 Item

Channels

4 Items

Companies