Summary
Sam Witteveen digs into VibeThinker 3B, a small language model from Waybo AI Lab — the AI research arm of the Chinese social network Weibo, operating out of Singapore. The model is built on Qwen 2.5 Coder 3B as a base and refined using a post-training recipe centered on reinforcement learning from verifiable rewards (RLVR), targeting structured reasoning domains like math and code rather than broad factual knowledge.
The central claim is that VibeThinker 3B matches or beats models roughly 300 times its size on specific tasks. On AMY and AMY 2026 math evaluations, it competes directly with Claude Opus 4.5, Gemini 3 Pro, Kimi 2.5, GLM 5.1, and DeepSeek V3. Coding benchmarks show similarly strong results against open-weight competitors. Where it falls short is on general knowledge evaluations like GPA Diamond, where raw parameter count still determines performance — the team makes no claim of general superiority.
Witteveen tests the model locally on a Dell workstation with an RTX Pro 6000 GPU, observing very long chains of thought even for simple tasks — a byproduct of training toward extended reasoning. The team’s test-time technique, Claim Level Reliability (CLR), generates multiple answer candidates and selects the most reliable one, providing a further accuracy boost that pushes results past some proprietary competitors on math and code tasks. Code and weights are publicly available.
📺 Source: Sam Witteveen · Published June 19, 2026
🏷️ Format: Benchmark Test







