Summary
Fahd Mirza pits two of the strongest open-weight dense models against each other in a live local benchmark: Google DeepMind’s Gemma 4 31B versus Alibaba’s Qwen 3.5 27B. Both models are Apache 2 licensed, multimodal, and run on a single NVIDIA H100 with 80GB of VRAM using vLLM, making the comparison directly applicable to developers selecting a model for self-hosted production or research workloads.
Tests cover four categories โ coding, reasoning, vision, and multilingual translation โ using identical prompts and parameters throughout. In the coding task, Gemma 4 produces a functional one-shot animated HTML simulation while Qwen 3.5’s output, despite generating nearly twice as much code, fails to animate on first run. Multilingual performance also favors Gemma 4. Qwen 3.5 leads on the AM 2026 math benchmark and GPQA, while official numbers across MMLU, GPQA Diamond, and LiveCodeBench show both models within roughly one percentage point of each other. Gemma 4 edges ahead on Codeforces ELO.
The video includes download size guidance (approximately 150GB total disk space), vLLM serve commands, KV cache configuration, and context length settings used for each test run โ enough detail to reproduce the comparison independently. Both models are also available in GGUF format for CPU inference via llama.cpp, broadening accessibility beyond high-end GPU users.
๐บ Source: Fahd Mirza ยท Published April 06, 2026
๐ท๏ธ Format: Comparison







