Summary
The AI Search channel puts GPT-5.6 through a series of demanding creative and technical tests using OpenAI’s redesigned Codex desktop app. The model family consists of three tiers — Soul (largest, most expensive), Terra (mid-range), and Luna (smallest, fastest) — with most testing conducted on 5.6 Soul at maximum “Ultra” thinking mode.
Tests include building a real-time voice chat web app featuring an anime avatar with lip sync, using GPT Image for character generation and Google Gemini’s realtime voice model for audio — completed in a single prompt in 17 minutes with no follow-up corrections needed. A physics liquid simulation with adjustable gravity, viscosity, and vorticity controls plus webcam-based hand tracking also required no follow-up prompts (12 minutes). Additional tests cover multi-track music composition with instrument orchestration, 3D scene rendering from a reference image (where results remain imperfect), and a Manim mathematical butterfly animation using epicycles refined over several prompts totaling roughly 53 minutes of total compute time.
Throughout, results are compared against Claude Fable and open-source GLM 5.2. GPT-5.6 Soul generally leads on coding-heavy creative tasks but struggles with precise 3D spatial layout. The Codex desktop app’s multi-agent workflow, skill-saving, and scheduled automation features are highlighted as key differentiators over the web interface for complex multi-step projects.
📺 Source: AI Search · Published July 10, 2026
🏷️ Format: Review







