Descriptions:
Fahd Mirza puts the newly released DeepSeek V4 Pro production build (0813) through two demanding real-world tests using the Hermes agent framework. The first challenge is a broken full-stack application — a FastAPI backend, Redis pub/sub worker, and SQLite store where jobs sit in a pending loop forever — handed to the model cold, with no hints. The model diagnoses and repairs the entire end-to-end pipeline in a single session, consuming fewer tokens than earlier preview builds required for simpler tasks.
The video contextualizes the result with freshly published agentic benchmarks: TerminalBench 2.1 climbed from 72 to 88 and CyberGym jumped from 53 to 83 compared to the previous preview release. Mirza overlays those numbers against competitors including Fable 5, Gemini K3, and Opus 4.8, noting that DeepSeek trades blows with models costing significantly more per token, while flagging that the self-reported nature of the benchmarks requires independent verification.
A second test asks the model to generate a single self-contained HTML/Canvas animation of a lamb roast on a spit — no libraries, no build step — to probe whether it can hold a complex interactive scene in memory and get physics and timing right simultaneously. The video also briefly mentions the Qwen 3.8 Max 2.4-trillion-parameter model appearing on Hugging Face, situating DeepSeek’s release within a broader wave of open-weight model activity.
📺 Source: Fahd Mirza · Published August 12, 2026
🏷️ Format: Hands On Build







