Summary
Cole Medin takes a critical look at Kimi K3, Moonshot AI’s recently released open-weight model that published benchmarks claim rivals or beats GPT-5.5 and Claude Opus 4.8 on agentic coding tasks. Rather than accepting those benchmarks at face value, Medin built a custom evaluation suite using his Archon repository, walking three models — Kimi K3, Kimi K2.7, and Claude Opus 4.8 — through dozens of real engineering tasks scored on a seven-dimension rubric (max 70 points), plus deliberately engineered “trap tasks” that target known failure modes in open-weight models.
The results tell a more nuanced story than the leaderboards suggest: Opus 4.8 averaged 64.3 out of 70 on real engineering tasks and hit an 8% failure rate on the trap suite. Kimi K3 performed competitively on the real tasks but jumped to a 36% failure rate on the designed traps — a gap that never surfaces in standard benchmarks. Medin spent millions of tokens on the evaluation and identifies specific reliability failure modes (such as partial task completion) that developers must engineer around when deploying open-weight models in production agentic pipelines.
On pricing, Kimi K3 runs at $3 per million input tokens and $15 per million output tokens via OpenRouter — roughly half the cost of Opus 4.8. The video also covers GLM and MiniMax as other open-weight alternatives with similar reliability patterns. Medin’s conclusion is that Kimi K3 is genuinely impressive but not a drop-in Opus replacement without explicit mitigation strategies for its failure modes.
📺 Source: Cole Medin · Published July 24, 2026
🏷️ Format: Benchmark Test







