Summary
Fahd Mirza puts DeepSeek V4-Flash — freshly released from preview to general availability — through a demanding real-world test using the Hermes autonomous agent framework. Rather than standard benchmarks, the target is a deliberately broken full-stack application: a live NYC Subway arrivals board built on FastAPI, SQLite, Redis, and a plain HTML frontend, with five chained bugs planted across the stack. The most insidious is a silent Redis cache key mismatch that causes the app to return plausible-looking data while the caching layer does nothing — no error, no crash.
DeepSeek V4-Flash finds and fixes all five bugs, including the silent cache bug — something Mirza reports that both GPT-4 (‘Fable 5’) and an unnamed newer GPT model failed to fully resolve with the same prompt and codebase. The model also autonomously starts the backend and frontend services after patching. The changelog context is significant: V4-Flash has been post-trained from scratch (same architecture, fully rebuilt instincts), with SWE-bench scores jumping from single digits to the mid-50s and Terminal-bench climbing into the 80s.
A second test challenges the model to generate a 15-scene scroll-driven WebGL landing page in a single self-contained HTML file using raw WebGL and hand-written GLSL shaders — no libraries. The results are visually coherent if not perfectly scene-faithful. For developers evaluating coding agents and autonomous debugging tools, this video offers concrete, reproducible evidence of what the newly GA DeepSeek V4-Flash can handle in agentic settings.
📺 Source: Fahd Mirza · Published July 31, 2026
🏷️ Format: Benchmark Test







