Summary
MiniMax has open-sourced M2.7, its 229-billion-parameter Mixture-of-Experts model, under a modified MIT license — and Fahd Mirza delivers one of the most thorough technical breakdowns available at launch. The architecture covers 256 experts per MoE layer with only 8 active per token (using a sigmoid scoring function with learned bias, distinct from standard softmax routing), a 196K context window enabled by a RoPE theta of 5 million, grouped query attention at a 6:1 ratio for memory efficiency, and three built-in multi-token prediction modules that implement speculative decoding at the architectural level for throughput gains.
The most striking part of the video is the explanation of MiniMax’s M2 Star iteration system. An earlier version of M2 built the entire training harness — hierarchical skills, persistent memory, guardrails, evaluation infrastructure — with one engineer in four days and zero human-written code. An internal M2.7 then ran over 100 autonomous rounds, analyzing failure trajectories, modifying its own scaffold, and deciding what to keep or revert. That autonomous self-improvement loop alone delivered a reported 30% performance gain. On SWE-bench Pro (real GitHub issue resolution), M2.7 scores 56.2 — within two points of Claude Sonnet 4.6, making it one of the strongest open-source models on software engineering tasks.
Deployment requires at minimum three H100 80GB GPUs with NVLink, and Mirza recommends SGLang over vLLM as the serving framework. Recommended inference settings are temperature 1.0, top-p 0.95, top-k 40, BFloat16 with FP8 quantization, and block size 128×128. A local inference demonstration is planned for a future video once hardware is provisioned.
📺 Source: Fahd Mirza · Published April 12, 2026
🏷️ Format: Deep Dive







