Summary
The How I AI channel puts OpenAI’s newly released GPT-5.6 model family — Soul, Terra, and Luna — through a custom evaluation framework, pitting the flagship Soul variant directly against Claude Fable 5 across tasks the creator uses daily. The video introduces the “How I AI Vibe Review Benchmark,” a structured test covering PRD generation, UI wireframing, prototype development, code debugging, and conversational quality, with GPT-5.5 serving as the LLM judge.
Soul comes in at $5 per million input tokens and $30 per million output tokens, compared to Fable’s $10/$50 — a meaningful cost difference that factors into the host’s overall recommendation. On the benchmark tasks, Soul is consistently praised for design personality and point of view: the host repeatedly notes that Soul produces more opinionated, visually distinctive prototypes compared to Fable’s cleaner but more generic outputs. One recurring tell identified is Soul’s apparent preference for forest green color schemes across generated designs.
The video also covers Terminal Bench 2.1 and cybersecurity evaluations cited in OpenAI’s release blog, and touches on the rollout dynamics of both models — including Anthropic’s limited Fable access under Claude subscriptions and uncertainty about how much Soul usage will be included. For developers and designers choosing between frontier models for creative and product work, this video offers a practical, subjective-but-structured perspective grounded in real workflow testing rather than pure benchmark numbers.
📺 Source: How I AI · Published July 09, 2026
🏷️ Format: Comparison







