Inkling by Thinking Machines: Benchmarks, Architecture & Real Tests

Inkling by Thinking Machines: Benchmarks, Architecture & Real Tests

More

Descriptions:

Thinking Machines Lab has released Inkling, a 1-trillion-parameter sparse mixture-of-experts model with 41 billion active parameters under an Apache 2.0 license — and Fahd Mirza puts it through its paces in this hands-on review. The video breaks down the model’s architecture: a 66-layer decoder-only transformer routing each token through 6 of 256 experts plus 2 shared experts, with a hierarchical path encoder for images and discrete tokenization for audio, all projected into a unified representation space before joint decoding.

On benchmarks, Inkling comfortably outperforms Nvidia’s Nemotron 3 Ultra across reasoning and agentic tasks and edges past Gemma K2.5 on most categories. It trails GLM 5.2 on agentic coding benchmarks like SWE-Bench Pro and Terminal Bench, and Kimi K2.6 leads on several multimodal tasks. Against closed models — Gemini 3.1 Pro, Claude 3B 5, and GPT 5.6 — it’s a mixed result, though it holds its own on MCP Atlas and IFBench.

Mirza tests the model with a demanding single-file HTML prompt requiring spatial reasoning and 3D visual simulation, finding strong results on tube geometry and depth ordering but weaker performance on the splash animation. A multilingual test across roughly 80 languages impresses, including low-resource languages. The video positions Inkling as the most credible American challenger yet in the open-weights race against Chinese labs like DeepSeek, Kimi, and GLM.


📺 Source: Fahd Mirza · Published July 15, 2026
🏷️ Format: Review

1 Item

Channels

1 Item

Companies