Did Kimi K3 really beat Fable?

Did Kimi K3 really beat Fable?

More

Summary

Matthew Berman reviews Kimi K3, Moonshot AI’s newly released open-weights model, which has drawn comparisons to last year’s DeepSeek moment after topping several high-profile benchmarks. The model features 2.8 trillion parameters and a one-million token context window — making it the largest open-source model released to date — and is priced at $3 per million input tokens and $15 per million output tokens. On Arena AI’s front-end development benchmark, Kimi K3 scores 76% versus Fable 5 at 63%, marking the first time an open-weights model has ranked above all proprietary competitors on a major coding evaluation.

Berman walks through the economics carefully using DeepSuite, his preferred benchmark for not yet being contaminated by training data. DeepSuite shows Kimi K3 Max and GPT 5.6 Soul at approximately equal cost-per-task efficiency (~$4.70), which means that despite the lower headline price, Kimi K3 uses roughly twice as many tokens to complete the same task. For writing, an internal benchmark places Kimi K3 at 2840 ELO — jumping from rank 21 to first place and costing five times less than Claude Fable 5 at the same quality tier. The model also tops Vercel’s Next.js engineering benchmark with a 92% agent success rate.

Berman adds significant caveats throughout: benchmark saturation is a real concern, Anthropic has publicly accused Moonshot of distillation — using Anthropic model outputs as training signal — and US AI Czar David Sacks flagged Kimi K3’s Arena AI ranking as a national security data point, arguing that US regulatory friction slows American labs while Chinese developers operate under fewer constraints.


📺 Source: Matthew Berman · Published July 18, 2026
🏷️ Format: Review

1 Item

Channels

1 Item

Companies