Gemma-4 26B A4B + vLLM: Best MoE Model of 2026: Running Locally

Gemma-4 26B A4B + vLLM: Best MoE Model of 2026: Running Locally

More

Summary

Fahd Mirza puts Google’s Gemma-4 26B A4B through its paces locally, starting with a clear explanation of what the model name actually means: 26 billion total parameters in a Mixture of Experts architecture, with only 4 billion active during any single inference pass. The design uses 128 experts plus one shared expert across 30 layers, activating just 8 of those experts per token — delivering the knowledge capacity of a 26B model at roughly the compute cost of a 4B model.

The setup runs on an NVIDIA H100 with 80GB of VRAM. Mirza walks through downloading the ~52GB model via Hugging Face CLI, serving it with vLLM at a 32K context window with automatic tool calling enabled and 90% GPU memory utilization. Total VRAM consumption settles around 75GB — higher than the raw model weight because vLLM allocates additional memory for KV cache, CUDA graphs, activation memory, and Torch compile cache, all of which Mirza explains clearly for viewers who have asked why consumption exceeds model size.

Capability testing spans three dimensions: a complex JavaScript coding challenge (a 2D snake-hunts-rat terrain simulation with distinct behavioral AI agents, scent trails, and a day/night cycle), multilingual structured output generation across dozens of languages simultaneously, and general knowledge questions requiring cross-regional comparison. The model handles all three impressively, positioning Gemma-4 26B A4B as a serious local deployment option for teams with H100-class hardware.


📺 Source: Fahd Mirza · Published April 04, 2026
🏷️ Format: Tutorial Demo

1 Item

Channels