Summary
Wes Roth puts GPT-5.5 — internally codenamed “Spud” at OpenAI and confirmed by Greg Brockman as the start of a “new class of intelligence” — through a rigorous same-day stress test by using it to build a fully playable multiplayer real-time strategy game from scratch. The game pits four AI models against each other as autonomous players: Claude Sonnet, GPT-5.4 Mini, Groq 4.1 Fast, and Gemini 3 Flash Preview, competing across economic and military metrics with diplomacy mechanics actively being developed.
Roth walks through the multi-agent architecture that made the build possible: separate agents handled live browser-based testing, image generation via GPT Image 2.0 (including automatic background removal), and core coding — all coordinated through OpenRouter, which provides access to over 400 models. The entire prototype was completed in roughly one day at a cost of about $15, with the model also autonomously writing the game manual and generating asset prompts. The video demonstrates how GPT-5.5’s expanded 1-million-token context window and stronger instruction-following enable a qualitative shift in what a solo developer can accomplish in a single session.
Beyond the build demo, Roth reviews OpenAI’s disclosed metrics around the launch: 900M+ weekly ChatGPT users, 50M+ paying subscribers, 4M active Codex users, and a GDP-val benchmark score of approximately 85% — a level where domain experts rate the model’s output as equal to or better than humans with 12+ years of experience. GPT-5.5 is also the first OpenAI flagship model served on Nvidia’s GB200/GB300 NVL72 hardware, with Nvidia projecting up to a 35x reduction in per-token inference costs on that infrastructure.
📺 Source: Wes Roth · Published April 24, 2026
🏷️ Format: Hands On Build







