Qwen3.8-9B: Community Distillation of a Frontier Model: Run Locally

Qwen3.8-9B: Community Distillation of a Frontier Model: Run Locally

More

Summary

Fahd Mirza walks through installing and testing Qwen3.8-9B, a community-built 9-billion-parameter distillation of Alibaba’s 27-billion-parameter Qwen 3.8 model. The distillation was created by a team called Emperor, who generated 70,000 reasoning traces from the full 2.4-trillion-parameter teacher model and trained the smaller student model to replicate the teacher’s reasoning style — not just its answers. The result is a single-GPU model capable of frontier-style chain-of-thought reasoning.

The video covers the full installation on Ubuntu using vLLM for OpenAI-compatible serving, including a required fix for the flash linear attention dependency tied to the model’s gated delta net layers. Mirza then runs a demanding agentic coding benchmark: asking the model to build a live cryptocurrency tracker from scratch with WebSockets, Redis, Docker Compose, and an Nginx proxy — which the 9B model completes and serves successfully. A multilingual creative writing test across rare and endangered languages rounds out the evaluation.

Key numbers: MMLU improves from 0.55 to 0.75 over the base 9B model, and the model consumes approximately 44GB of VRAM at full context. The video doubles as a primer on how community distillation works and how hobbyists can replicate the process at home, making it a useful reference for anyone tracking lightweight reasoning models or the gap between official and community model releases.


📺 Source: Fahd Mirza · Published August 17, 2026
🏷️ Format: Hands On Build

1 Item

Channels

1 Item

Companies