Summary
Fahd Mirza runs a hands-on three-way coding showdown between GLM-5.2, MiniMax M3, and Qwen 3.7 Max using the Hermes agent framework. Each model is tested against a deliberately buggy full-stack application — featuring broken API endpoints, incorrect group-standings logic, and mislabeled tournament data — via its respective API (ZhipuAI for GLM, MiniMax’s own platform, and Alibaba Cloud for Qwen). A planted red-herring comment is included as a trap; all three models avoided it and correctly identified every reported bug.
The real differentiation emerged in efficiency metrics extracted from each model’s SQLite activity log: Qwen 3.7 Max used the fewest tool calls and API round trips; MiniMax M3 cost less than half of GLM’s run despite burning more tokens re-reading files; and GLM 5.2 took the most steps and proved the most expensive despite reaching identical output quality. Mirza pushes back directly on GLM social-media hype, noting that Claude Opus 4.8 and GPT 5.5 at medium settings are both cheaper and more capable in his real-world experience.
A second test covering self-generating code is also included. The video provides a frank, data-backed assessment useful for developers choosing between frontier and near-frontier models for agentic coding pipelines where cost and token efficiency matter alongside raw correctness.
📺 Source: Fahd Mirza · Published June 22, 2026
🏷️ Format: Comparison







