Summary
Theo (t3.gg) reviews xAI’s newly launched Grok 4.7, a model Elon Musk had hyped for over a month as a potential frontier leader. Theo walks through SpaceX/xAI’s official blog claims and benchmark charts covering coding, electrical engineering, legal work, and clinical reasoning (Healthbench), while questioning why the company chose such an unusual, seemingly cherry-picked set of benchmarks for the release.
The video’s central argument is that public benchmarks increasingly fail to capture real-world model quality. Theo points to Cognition’s Terminal Bench and Frontier Code results showing Grok 4.7 sometimes scoring below its predecessor Grok 4.6, and compares Artificial Analysis’s intelligence index placement, where Grok 4.7 lands surprisingly low — even behind GLM 5.3 Max and Muse Spark 1.3. He also notes token-efficiency changes, with Grok 4.7 using more than double the tokens per task compared to Grok 4.6.
Despite the shaky benchmark story, Theo shares his own hands-on impressions, saying Grok 4.7 has outperformed Claude Fable in some real coding tasks and is one of his favorite models this year, offering a more nuanced, firsthand counterpoint to the mixed public reception on social media following the release.
📺 Source: Theo – t3․gg · Published September 22, 2026
🏷️ Format: News Analysis







