Qwen3.6 (REAP 90pct GGUF): The Brain-Damaged Model

Qwen3.6 (REAP 90pct GGUF): The Brain-Damaged Model

More

Summary

Fahd Mirza takes a deep look at an aggressively pruned variant of Qwen 3.6 — a 35-billion-parameter mixture-of-experts model — compressed using a technique called REAP (Refined Expert Activation Pruning) at 90% intensity. The result is a model that drops from 35B to roughly 6B effective parameters, loading in about 6 seconds and consuming just under 4GB of VRAM on an NVIDIA RTX A6000 with 48GB capacity.

The core tension Mirza explores is the striking disconnect between the model’s official card (15/15 benchmark tests passed, all green) and its real-world capabilities: it cannot solve 17×23, name the capital of France, or complete a short story. He explains how REAP works — scoring each expert in a mixture-of-experts layer by multiplying the router’s trust weight against the expert’s output magnitude, keeping quiet-but-relied-upon specialists while cutting loud freeloaders — and argues the algorithm itself is principled. The problem is the dial: the original REAP paper only validates pruning up to 50%, and taking it to 90% pushes far beyond any claimed safe boundary.

The video functions as a sharp lesson in model card literacy. A 15/15 green scorecard does not guarantee a functioning model, especially when extreme compression techniques are applied well past their validated limits. Anyone evaluating open-source models or compressed variants will find Mirza’s breakdown of the gap between benchmark performance and actual usability directly applicable.


📺 Source: Fahd Mirza · Published June 21, 2026
🏷️ Format: Benchmark Test

1 Item

Channels