Summary
Jack Roberts walks through hands-on testing of the newly released DeepSeek model, positioning it against GPT-6 Astra (ChatGPT’s flagship) across a range of real-world tasks. The central argument is that DeepSeek delivers performance close to Astra at a fraction of the cost — reportedly 88 times cheaper — making deliberate model routing a practical money-saving strategy for builders.
The video benchmarks DeepSeek on three public leaderboards: Automation Bench (54.8%, up from 37.7% in V4), DeepSeek Suite V.1 software engineering tasks (74.2%, edging out Opus 5 and GPT 5.6), and Terminal Bench 2.1 (90.6%). Roberts then tests these claims against design work, business data analysis (50,000+ rows, 17 questions answered with 100% accuracy), and agentic operating system tasks.
A key practical tip is using the DeepSeek harness inside the CodeX framework so that Astra can delegate cost-intensive subtasks — like website text amendments and image generation via Higfield — to DeepSeek. Roberts also covers getting a direct Deepseek API key to avoid Open Router rate limits. The takeaway: DeepSeek excels at data analysis and iterating on existing design systems, while still falling short of Astra for generating polished designs from scratch.
📺 Source: Jack Roberts · Published September 13, 2026
🏷️ Format: Comparison







