LTX CEO on Video Gen + Can we beat AI superforecasters?

LTX CEO on Video Gen + Can we beat AI superforecasters?

More

Descriptions:

In this episode of Cognitive Revolution, host Nathan Labenz sits down with Zeve, CEO of Lytrix and the team behind the LTX suite of video and music generation models. The conversation centers on how LTX approaches the deep integration of audio into video models — arguing that bolting audio on after pre-training produces inferior results, and that truly native audio requires joint token-level training from the start. Zeve explains how LTX currently handles video tokens and audio tokens with their own latent spaces, bridged through a cross-attention mechanism, and why that architecture produces substantially better lip sync fidelity.

The discussion then moves to an emerging challenge: robotics and other physical-world domains need their own token types — things like robot hand positions, finger pressure, or 3D structural vertex data — and LTX is building a post-training architecture that makes it easy for external builders to define and fine-tune new token streams on top of the existing backbone. Zeve frames this modularity as central to why the model should remain open, since different physical-data domains each require specialized fine-tuning.

The episode also touches on the competitive landscape of AI-generated music and video, the role of model taste in creative output (comparing Fable versus Opus for music video generation), and broader questions about cross-modal latent space merging. It is a technically rich conversation for anyone following the frontier of generative media models.


📺 Source: Cognitive Revolution “How AI Changes Everything” · Published July 08, 2026
🏷️ Format: Podcast