Summary
Fahd Mirza demonstrates Escha-W2, Asha Labs’ aggressively quantized version of Qwen3.8 27B, which compresses the model from 50GB at full precision down to just 10GB using 2-bit quantization — while matching or beating FP8 half-precision on major benchmarks. The key technical insight is non-uniform precision: attention layers, which are more sensitive to rounding errors, are kept at high precision, while the feed-forward layers (where there is substantial redundancy) are crushed to just 2 bits per weight. A fine-tuning pass is then applied over the entire compressed model to recover quality lost in compression.
The video walks through the full setup on Ubuntu with a 48GB VRAM GPU (the model can run on a 24GB consumer card with reduced KV cache). Asha Labs ships its own custom SGLang build with decode kernels tuned for this format, which the video installs and serves before connecting to a Hermes agent interface. The primary real-world test is a needle-in-a-haystack bug hunt inside Silo Trace, a full-stack factory feed-mill monitoring application with a Docker backend, SQL database, and React frontend. The 2-bit model finds and correctly fixes the bug — an inverted quality-check logic that was labeling in-spec batches as out-of-spec — though it takes roughly twice as long as the full-precision model.
A second test drops the model into a complex inheritance dispute requiring nuanced judgment and empathy, demonstrating that reasoning quality survives the compression. This is a practical guide for running near-full-quality large models on consumer or prosumer hardware.
📺 Source: Fahd Mirza · Published August 21, 2026
🏷️ Format: Hands On Build







