Summary
Fahd Mirza puts MiniCPM5 2B — the latest small model from ModelBench and Tsinghua University’s NLP lab — through a gauntlet of real-world tests on local hardware, pushing back against the lab’s own benchmark claims by running the model directly rather than trusting published numbers. The model is served via SGLang on a 48 GB VRAM GPU, though Mirza notes the weights themselves require only around 4 GB, making it accessible on consumer hardware.
Testing spans three categories: a tricky Oracle SQL partition-boundary bug (which the model eventually solves correctly after extended reasoning), a historical knowledge question about the 1689 Treaty of Nerchinsk, and a creative front-end coding challenge requiring a self-contained animated HTML file. Results are mixed but generally positive for the 2B size class — factual accuracy is solid, coding passes, but response verbosity and slow reasoning chains are flagged as practical drawbacks compared to competitors like Qwen.
Mirza also walks through the model’s three-stage training pipeline — base, mid-training, and post-training — with particular attention to the reinforcement learning plus on-policy distillation (OPD) phase, which the lab credits for an 11-point gain on reasoning and general benchmarks and nearly 7 points on agentic tasks. The video gives a grounded verdict on where MiniCPM5 2B fits relative to the crowded sub-4B model landscape.
📺 Source: Fahd Mirza · Published September 07, 2026
🏷️ Format: Review







