Tess-4-27B + EAGLE-3: Local Reasoning, Nearly 2× Faster

Tess-4-27B + EAGLE-3: Local Reasoning, Nearly 2× Faster

More

Summary

Fahd Mirza demonstrates running Tess-4-27B — a reasoning-capable fine-tune of Qwen 3.6 27B by Miguel de Icaza — locally on an NVIDIA A100 80GB GPU, paired with EAGLE-3 speculative decoding via vLLM to achieve nearly 2× faster inference. The video opens with a clear explanation of how speculative decoding works: a small draft model guesses multiple tokens ahead while the large model verifies them in a single forward pass, delivering identical output quality with significantly fewer full model passes.

The main demo deploys the Tess-4-27B model through Hermes agent to debug a real full-stack Iron Ore price validator application — a practical agentic coding benchmark where the frontend was silently pointing to the wrong port. Mirza also covers Tess-4-27B’s design: trained on 64K long-context agentic traces, inheriting Qwen 3.6’s vision tower for multimodal capability, shipping under Apache 2.0, and distilling reasoning traces from a three-model teacher ensemble of Opus 4.8, GPT 5.5, and GLM 5.2. VRAM consumption sits at approximately 73GB on the A100.

The video closes with commentary on the broader HuggingFace ecosystem, noting a resurgence of independent model creators as major labs pull back on open-weight releases — while cautioning that not every community fine-tune is worth evaluating, pointing to a trending fake “Fable traces” model as an example of the noise problem on the platform.


📺 Source: Fahd Mirza · Published July 11, 2026
🏷️ Format: Hands On Build

1 Item

Channels