Gemma 4 Just Got a Massive Update (Tested Live Locally)

Gemma 4 Just Got a Massive Update (Tested Live Locally)

More

Descriptions:

Fahd Mirza puts Gemma 4’s latest patch update under the microscope in a live session running on an Nvidia H100 with vLLM on a fresh Ubuntu machine. Google shipped five fixes to the existing model weights this week — no retraining, no new checkpoint — and Mirza focuses on the two he can independently verify: the tool calling parser bug and the new vision token budget system.

The tool calling fix addresses a real bug where malformed output would break under long context. Mirza demonstrates this with a two-step weather query script: turn one correctly produces an empty content field and a populated tool_calls object with a structured get_weather call, while turn two uses the injected fake result to generate a coherent answer — proving the fix holds across a full exchange. The vision token budget feature, shown via a Hugging Face demo, lets users control how many tokens Gemma spends on image processing: at 70 tokens the model crops to 384×384 pixels using just 64 soft tokens, while bumping to 1,120 unlocks the fine detail needed for OCR.

Mirza notes that Flash Attention 4, chat template smoothing, and reduced laziness are listed as Google’s claims and left unverified — though FA4 turned out to be active on his H100 without any recompilation. With the model consuming just over 77GB of VRAM including KV cache, he wraps up by noting that the combination of Flash Attention 4 support, the tool calling fix, and controllable vision token budgets adds up to more than a routine patch — enough to wonder why Google didn’t just call it Gemini 4.1.


📺 Source: Fahd Mirza · Published July 17, 2026
🏷️ Format: Tutorial Demo

1 Item

Channels

1 Item

Companies