Descriptions:
HeyGen co-creator and tech lead James Russo delivers a conference talk at AI Engineer explaining why HTML, CSS, and JavaScript are the optimal medium for AI agents generating video content. The core argument: LLMs are trained on vast quantities of web data, making HTML their native language — forcing them to use custom DSLs, Lottie JSON, or Remotion’s React-based framework is analogous to asking Shakespeare to write poetry in Japanese.
The talk traces HeyGen’s year-long search through After Effects connectors, Lottie, Rive, and Remotion before landing on raw HTML as the winning approach. The breakthrough came around November of last year when Gemini Flash demonstrated a step-function improvement in generating coherent HTML-based video compositions without explicit format instruction. The result is Hyperframes, now open source, which solves the browser-to-video pipeline by freezing the browser clock and seeking deterministically frame-by-frame — enabling anything renderable in a browser (Three.js, WebGL, SVGs, Lottie animations, shaders) to appear in a deterministic MP4 output.
Russo notes independent alignment with conclusions from Andrej Karpathy and others who have identified HTML as the new markdown for LLM output. The session includes a live demonstration of an agent producing a polished product launch video in a single shot entirely from HTML, and positions Hyperframes as the missing composition layer between AI avatars and full-featured video production.
📺 Source: AI Engineer · Published July 21, 2026
🏷️ Format: Keynote Launch







