Summary
Peter Yang runs GPT 5.5 and ChatGPT Images 2 through four practical side-by-side tests against Claude Opus 4.7 and Google’s Nano Banana 2, using OpenAI’s Codeex platform as the primary coding environment. The tests cover personal health advice (with DEXA scan data in context), front-end web design (a corgi cafe website), and two retro game recreations — Super Mario Bros. and F-Zero — built from scratch using image and text prompts.
The results split across models. Opus 4.7 wins on the Mario implementation and subtler front-end animation details, while GPT 5.5 produces the first fully functional F-Zero clone Yang has tested from any model, complete with competing AI bots and boost mechanics. On the health advice test, GPT 5.5 better respects data privacy guardrails while Opus delivers sharper personalized insight — a genuine tradeoff rather than a clear winner. For image generation, ChatGPT Images 2 outperforms Nano Banana 2 on character consistency and emotional expressiveness in anime-style birthday invite artwork.
Yang’s overall verdict is that GPT 5.5 has meaningfully closed the gap with Opus 4.7 in coding and creative tasks, particularly in complex game logic where earlier GPT versions consistently fell short. Developers routing workloads between OpenAI and Anthropic will find the specific task-level comparisons — including the F-Zero milestone — directly useful for deciding which model to call for front-end generation, agentic coding, and image-integrated design work.
📺 Source: Peter Yang · Published April 24, 2026
🏷️ Format: Comparison







