Descriptions:
Fahd Mirza woke up at 4 AM in Sydney to put Kimi K3 through its paces the moment Moonshot AI’s API went live, and the result is one of the more technically honest assessments of the 2.8 trillion parameter model to emerge on launch day. The video leads with a live agentic test: K3, connected through the Hermes agent framework, is given a real seaport freight management application built with FastAPI and SQLite and told only to find and fix issues — no hints about what’s wrong or how many bugs exist.
The outcome is genuinely impressive. K3 identifies the planted draft-depth safety bug (a 9.8-meter vessel assigned to a 9.5-meter berth), verifies the fix from both directions, discovers two additional unplanted bugs, and autonomously kills a stale server process it detected running on port 8000 — without being asked. Mirza explicitly distinguishes this from benchmark gaming, walking through K3’s reasoning chain and confirming the fix in the running web interface.
Critically, Mirza also pushes back against launch-day hype: he quotes Moonshot’s own blog acknowledging noticeable gaps behind Claude 5.5 and GPT models in several evaluations. The actual benchmark picture is mixed — K3 leads on Program Bench and SWE Marathon, ties GPT-5.6 Soul on Terminal Bench, and trails on Deep SWE and Frontier SWE. The video rounds out with a standalone HTML water park simulation and a pure vision physics reasoning test. Model weights go public July 27th; the API is live today at platform.kimi.com.
📺 Source: Fahd Mirza · Published July 16, 2026
🏷️ Format: Hands On Build







