Summary
Fahd Mirza reviews Ridge, an architecture-aware quantization of Qwen3.8 27B produced by the Empero team — the same group behind a previously covered 9B distillation. The full Qwen3.8 27B model weighs 50GB; Ridge brings it down to 11.7GB at 3.7 bits per weight, making it runnable on a 16GB GPU card (or potentially 12GB with reduced context length).
The video’s key contribution is a clear explanation of why this quantization differs from generic approaches like IQ2. Qwen3.8 is a hybrid architecture mixing gated DeltaNet layers — which maintain internal state values sensitive to aggressive compression — with traditional full-attention layers every fourth block. Generic quantizers flatten everything uniformly, degrading the model’s ability to track earlier context. Ridge instead holds those sensitive state values at high precision and recovers file-size savings by compressing the less sensitive feed-forward layers in the middle of the stack, achieving better quality at the same file size.
Mirza serves the model using llama.cpp’s llama-server on an RTX 3060 with 48GB of VRAM, observing under 14GB consumption. He tests the model with a demanding creative coding prompt — a cinematic slow-motion tree-growing animation in a single HTML file — and finds the visual output genuinely impressive for a 3.7 BPW quant, correctly handling layered soil cross-sections, root/canopy mirroring, and atmospheric depth. A creative writing test in Mark Twain’s voice also fares well stylistically, though the model misses the underlying logic riddle. Antares 1B is also available in 3B and sub-1B variants.
📺 Source: Fahd Mirza · Published August 18, 2026
🏷️ Format: Hands On Build







