Summary
Fahd Mirza runs three newly released frontier models — Alibaba’s Qwen 3.8 Max (2.4 trillion parameters), DeepSeek V4 Flash (13 billion active parameters per token, priced at $0.09/M input), and Moonshot’s Kimi K3 (2.8 trillion parameters, currently the only one with open weights) — through the same three challenges under identical conditions using the Hermes agent framework.
The first task is a real full-stack debugging problem: a FastAPI backend with Redis state management, running in Docker, with five planted bugs across the Dockerfile, compose file, backend, and frontend. All three models fixed every bug and self-verified. Qwen finished in roughly 6.5 minutes with particularly impressive behavior — it wrote its own verification script, caught a false-positive in its own test suite (a trailing newline issue), and confirmed the fix with byte-exact endpoint checks. Kimi K3 ran 17 live assertions and finished in around 8–9 minutes. DeepSeek took 23 minutes and lost time on unnecessary browser-style UI interaction the app never required. Verdict: Qwen edged Kimi on time and cleanup; DeepSeek was the slowest.
The second challenge — generating a 3D landing page as a single self-contained HTML file with scroll animation loaded from CDN, depicting a half-submerged oil rig — tests pure generation quality. The video provides a useful side-by-side evaluation for practitioners deciding which model to use for agentic coding tasks, with concrete timing, token counts, and qualitative observations about each model’s debugging strategy and verification rigor.
📺 Source: Fahd Mirza · Published August 04, 2026
🏷️ Format: Comparison







