Gemma 4 26B A4B vs Qwen3.5 35B A3B: MoE Models Battle Locally

Gemma 4 26B A4B vs Qwen3.5 35B A3B: MoE Models Battle Locally

More

Summary

Fahd Mirza runs a direct head-to-head comparison of two mixture-of-experts models — Gemma 4 26B A4B (activating 3.8 billion parameters per token) and Qwen 3.5 35B A3B (activating 3 billion parameters per token) — both served locally on an NVIDIA H100 80GB GPU using vLLM. The MoE architecture means both models execute at roughly the speed of a 4B dense model while drawing on the knowledge encoded across their much larger total parameter counts, a tradeoff the video frames as the core promise of the MoE design.

Testing spans a demanding single-file CRUD web application build (a pet hotel management system requiring dynamic state management, full booking logic, and polished UI in one HTML file), multilingual translation across multiple languages, and additional reasoning tasks. Gemma 4 is served with the Hermes tool-call parser and Gemma 4 reasoning parser; Qwen 3.5 uses the Qwen3 tool-call and reasoning parser. Both models produce functional applications in a single pass, with Qwen showing a narrow edge in instruction-following completeness and CRUD functionality coverage, while Gemma leads on multilingual output quality.

The comparison is a follow-up to a prior video in which dense Gemma 4 31B defeated Qwen 3.5 27B in a 3-2 decision, making this a parallel test of the MoE variants at comparable active parameter budgets. Mirza notes the subjectivity in evaluating generated UI and encourages viewer input on individual task verdicts, ultimately giving Qwen 3.5 35B A3B a slim overall win.


📺 Source: Fahd Mirza · Published April 07, 2026
🏷️ Format: Comparison

1 Item

Channels