GPT 6 Astra, so good even OpenAI are worried

GPT 6 Astra, so good even OpenAI are worried

More

Summary

AI Explained delivers a thorough analysis of GPT-6 Astra, OpenAI’s latest frontier model, examining why it’s generating both excitement and internal concern at the company. The video systematically works through demanding benchmarks — Terminal Bench Science, Agents Last Exam, and Screen Spot Pro — where Astra achieves state-of-the-art results while reportedly using fewer tokens than Claude Fable 5.1, yielding lower effective costs despite identical listed API pricing.

The benchmark deep-dives go well beyond raw scores. Tasks covered include analyzing star brightness readings to detect repeating planetary transits, processing nearly 3GB of satellite imagery to track Greenland lake drainage across a 153-day season, and interpreting 237 MRI scans with precise injury measurement and labeling. These examples are intended to show that top benchmark performance now corresponds to genuine expert-level analytical capability across scientific domains. Cognition AI and Jane Street both provided endorsements for Astra’s coding performance, with Jane Street noting it outperforms Fable 5.1 on coding while Fable leads on trading intuition.

The second half of the video addresses concerns from within OpenAI about Astra’s “silent reasoning” — a capability that some researchers find unsettling. The channel also quotes Francois Chollet, creator of the ARC-AGI benchmark series, who now expects AGI-level benchmark saturation to arrive sooner than his earlier 2030 projection, citing faster-than-expected progress.


📺 Source: AI Explained · Published September 04, 2026
🏷️ Format: Deep Dive

1 Item

Channels

1 Item

Companies