Gemini 3.8 Flash: Google is Back on AI Horse: Cheap and Fast

Gemini 3.8 Flash: Google is Back on AI Horse: Cheap and Fast

More

Summary

Fahd Mirza puts Google’s newly released Gemini 3.8 Flash through a rigorous set of real-world tests, framing the release as a pivotal moment for Google’s AI credibility after a prolonged period out of the spotlight. Three Flash model releases in six weeks, he argues, is either a sign of confidence or panic — and this video aims to find out which.

The centerpiece test is a live emergency dispatch triage system for New South Wales State Emergency Service — a full production stack comprising Postgres, FastAPI, Nginx, Docker Compose, and Redis. The system had a silent priority-ordering bug causing the largest flood emergency to be ranked below smaller incidents. Using the Hermes agent powered by Gemini 3.8 Flash, Mirza issues a single natural-language goal with no step-by-step guidance. The model reads every file in the stack, queries the live endpoint to observe the broken state, identifies the exact sort key issue, applies a single-line fix, and self-verifies with a Python assertion checking the full ordering tuple across all incidents — completing the task in two minutes and twelve seconds.

Mirza also runs vision tests involving OCR of a phone screenshot with social nuance, and multilingual comprehension. On published benchmarks, Gemini 3.8 Flash scores 73.7 on the Deep Sweep agentic leaderboard at $0.75 per million input tokens, outperforming Claude Fable 5 (70%, $21.63/task average) and Claude Opus 5 (74%, $11.84/task average) on a cost-adjusted basis. He concludes that Google has delivered a genuinely impressive workhorse model that the community is underrating.


📺 Source: Fahd Mirza · Published September 02, 2026
🏷️ Format: Review

1 Item

Channels

1 Item

Companies