Descriptions:
Creator Nate Herk tests what GPT-5.6 Soul can do with a single prompt and no further human input: the model — running inside OpenAI’s Codex on the Ultra tier — autonomously orchestrates a complete narrated video production, from research and scriptwriting through voice synthesis, avatar rendering, and final edit assembly, without Nate recording, editing, or reviewing a single frame.
The video documents Soul’s full agentic workflow in detail. It split the script into sub-60-second segments for consistency, routed each through Nate’s authorized ElevenLabs voice clone, used Heygen’s API for avatar generation (including browser automation to force the correct Avatar V motion engine after the API couldn’t lock it), and assembled all visuals in Hyperframes with precise phrase-to-clip mapping. Separate quality-check agents then reviewed rendered frames for errors and cross-checked factual claims against OpenAI’s release notes before the video was considered done.
On benchmarks, Soul scored 91.9% on TerminalBench 2.1 (up from 85.6% for GPT-5.5) and 92.2% on BrowseComp. The full run spawned nine sub-agents and consumed approximately 450 million total tokens — which Nate estimates would cost around $300 via API. He notes that running on Codex Ultra’s maximum effort level drove significant over-delegation and token waste, and suggests the same output was likely achievable at roughly half the cost on a lower effort setting. The video is one of the most concrete first-hand tests of autonomous AI video production published to date.
📺 Source: Nate Herk | AI Automation · Published July 09, 2026
🏷️ Format: Showcase







