Summary
Matthew Berman breaks down Google’s newly released Gemini 3.8 Flash, a lightweight model drawing attention for strong performance on coding and legal benchmarks at a fraction of frontier model pricing. On DeepSWE v1.1 — a long-horizon software engineering benchmark Berman considers the most accurate reflection of real-world utility — Gemini 3.8 Flash scores 73.7%, effectively matching Claude Opus 5 and edging out GPT-5.6 Soul at 72.7%. On the GDPVal knowledge-work benchmark, the picture is more mixed, with the model scoring 1545 versus Claude Opus 5’s 1824 and GPT-5.6 Soul’s 1710.
The model’s pricing is a standout feature: $0.75 per million input tokens and $3.75 per million output tokens at introductory rates, rising to $1.50/$7.50 after year-end but still cheaper than comparable mid-tier alternatives. On the Harvey legal agent benchmark, Gemini 3.8 Flash claims the top spot at 61.4%, making it particularly compelling for law firms and legal-adjacent enterprises. Google also released a restricted variant, Gemini 3.8 Flash Cyber, available only through the Fair Wind program for trusted security defenders, which scores 86.2% on the CyberGem benchmark — topping GPT-5.5 Cyber at 85.6%.
Berman rounds out the video with side-by-side visual generation comparisons against GPT-5.6 Soul and Claude Fable 5.1, arguing that enterprise teams should run their own internal benchmarks rather than defaulting to whichever frontier model is newest.
📺 Source: Matthew Berman · Published September 03, 2026
🏷️ Format: Review







