Summary
Fahd Mirza tests Claude Opus 5 head-to-head against Kimi K3 using a series of demanding single-file coding challenges evaluated live in the browser. The first test asks both models to generate a complete 3D WebGL landing page featuring a procedural multi-story building, real instanced mesh foliage, GSAP scroll-triggered camera choreography, and a seamless day-to-night lighting transition — all in a single self-contained file with no external dependencies.
Opus 5’s first attempt produces a partial result with rendering errors and missing animation. A second iteration, informed by a screenshot and error log, yields what Mirza calls a world-class result: a detailed low-poly architectural scene with foliage, soft lighting, glass railings, and working scroll animation. A follow-up prompt challenges both models to simulate eight culturally distinct international kebab skewers rotating and cooking simultaneously — complete with procedurally generated meat textures, flames, smoke, and reflections at 60fps — comparing outputs side by side.
Mirza references Anthropic’s benchmark positioning, which shows Opus 5 leading or matching top scores across agentic reasoning and real-world professional tasks, and notes its pricing at $5 to $25 per million tokens — the same as Opus 4.8, but marketed as a significant capability step up from Fable 5 at half that model’s cost. He closes with a respectful but direct call for Anthropic to address rate limiting and throttling, which he identifies as a meaningful practical barrier for heavy users building with the model daily.
📺 Source: Fahd Mirza · Published July 25, 2026
🏷️ Format: Comparison







