Summary
Fahd Mirza runs three frontier AI models through a demanding single-file coding gauntlet: simulate a full concrete plant network—trucks, batching, queuing, curing time constraints, and traffic delays—across five interactive browser tabs, with a self-verification tab that runs assertions against the model’s own output. The constraint that concrete cannot pause mid-pour and that evenly spaced departures don’t yield evenly spaced arrivals under traffic makes this a genuine reasoning challenge, not just a syntax test.
The three contestants are Kimi K3 (Moonshot AI’s new 2.8-trillion-parameter mixture-of-experts model with 896 experts, 16 active per token, 1M context, open weights releasing July 27th), Claude Fable 5 (Anthropic’s top closed model at $10 input / $50 output per million tokens, always-on thinking), and GLM 5.2 (ZhipuAI’s MIT-licensed open-weights model at roughly 744 billion parameters, approximately $3.15 output per million tokens). GLM 5.2 produced a working simulation that required a page reload to animate but ran correctly with truck cycle tracking and plant status displays. Fable 5 delivered a noticeably more polished interface with smoother animation and clearer multi-site visualization. Kimi K3 took longest to generate but has topped WebDev Arena’s front-end coding leaderboard, ranking first in six of seven domains above Fable 5.
Mirza notes all benchmark figures come from vendor launch posts and urges treating self-reported numbers with skepticism, while third-party leaderboard results for K3 look promising heading into its public weights release.
📺 Source: Fahd Mirza · Published July 17, 2026
🏷️ Format: Benchmark Test







