Summary
Fahd Mirza runs a structured three-round head-to-head between GLM 5.2 — THUDM’s fully open-weight 744B parameter MoE model (40B active per token, MIT licensed, available on Hugging Face) — and Anthropic’s Claude Opus 4.8 (closed, dense, $5/M input tokens), using the Hermes agent framework on an Ubuntu system with separate profiles for each model.
Round one asks both models to reconstruct a realistic wind tunnel aerodynamics dashboard from a photograph — complete with interactive sliders for angle of attack, Reynolds number, and Mach number, plus a self-selected bonus feature. GLM 5.2 produces live-updating charts with real data and an engineer’s notebook panel; Opus 4.8 generates a cleaner minimal layout but the sliders fail to update the graphs, a rare miss that Mirza acknowledges may not be representative.
Round two plants four distinct bugs in identical copies of a Fast API and SQLite full-stack app — a wrong field name, a 500 error, an incorrect HTTP method, and a silent filter clause failure — without telling either model what the bugs are. Opus 4.8 finds all four and identifies an additional error in an in-code comment that mischaracterized one of the bugs, verifying the underlying math independently. GLM 5.2 completes the task but is notably slower due to apparent rate limiting, which Mirza attributes to export control pressures on the Chinese lab. The video is a useful open-vs-closed reference for developers weighing cost, licensing, and raw capability.
📺 Source: Fahd Mirza · Published June 19, 2026
🏷️ Format: Comparison







