Descriptions:
Fahd Mirza runs a hands-on evaluation of Google’s newly released Gemini 3.6 Flash, one of three Gemini Flash variants dropped simultaneously: 3.6 Flash as the general coding-capable workhorse upgrade, 3.5 Flash Light as a speed-optimized low-cost option that benchmarks punch above its weight class on SWE-bench Pro and OS World, and 3.5 Flash Cyber — a restricted security-focused variant locked to government partners under a product called Code Mender.
The primary test deploys 3.6 Flash through a Hermes agent wrapper against a FastAPI seaport freight-tracking application with a deliberately planted bug: a vessel draft validation flaw that silently approves ships too deep for assigned berths. The model successfully identifies and patches the bug without hints, but Mirza notes meaningful gaps — it neither restarts the backend server nor cleans up the stale database state autonomously, behaviors he’s observed in Gemini K3 and Fable on similar tasks. He also runs a multilingual poetry test prompting the model to generate original two-line verses in the style of each language’s most celebrated poet across all 83 languages simultaneously, then a multimodal survival-scenario challenge: analyzing an AI-generated Australian outback video and answering geographic and logistical questions to help a stranded tourist reach Perth.
Benchmark comparisons show 3.6 Flash beating 3.5 Flash on DeepSeek coding, OS World, and MLE-bench while using fewer tokens per task — though the 3.1 Pro gap on hard long-context reasoning tasks remains. Mirza describes the model as a solid incremental upgrade rather than a step-change competitive threat to models outside Google’s own lineup.
📺 Source: Fahd Mirza · Published July 22, 2026
🏷️ Format: Hands On Build







