Summary
Nate Herk put OpenAI Codex and Anthropic Claude Code head-to-head by giving both agents the exact same prompt: build a production-ready, originally branded Typeform alternative from scratch. The results were starkly different — one agent completed the task in 5 hours at a cost of around $800, while the other took 3 days and spent roughly $3,000. Both were given identical instructions to orchestrate specialized agents through research, build, and verify phases, with explicit instructions not to stop at a prototype.
The video walks through live demos of both finished apps — dubbed “Real Form” from one agent — examining UI quality, form-builder functionality, logic branching, theming, and publish flows. Herk uncovers real bugs in both outputs (broken page numbering, image preview failures, navigation dead-ends) and reflects on how much testing is still required even after an agent claims completion. He also notes the prompt itself was imperfect, suggesting a dedicated planning phase between research and build would likely improve outputs from both systems.
The takeaway is a nuanced split verdict: each agent has clear strengths in different scenarios, and the cost and time differences are significant enough to inform which tool to reach for depending on project scope. Anyone choosing between Codex and Claude Code for autonomous, long-running coding tasks will find this a useful practical reference.
📺 Source: Nate Herk | AI Automation · Published August 14, 2026
🏷️ Format: Comparison







