Summary
Machine Learning Street Talk hosts a sponsored deep-dive interview with Ming-Yu Liu, who leads research on Nvidia’s Cosmos 3 — a multimodal world model that unifies language understanding, video generation, audio synthesis, and physical action prediction under a single architecture. The conversation covers both the technical design of Cosmos 3 and its practical applications across autonomous driving and robotics.
Liu walks through the full architectural stack: starting from a pretrained language model, adding a vision encoder to form what he calls the “reasoning tower” (a vision-language model operating autoregressively), and then connecting a bidirectional diffusion “generation tower” that produces video, audio, and actions with cross-token coherence. The key design choice is treating actions as first-class tokens — enabling the model to generate sequences where each action influences the next visual observation, a prerequisite for closed-loop robot policy training.
The interview explores several practical dimensions: using Cosmos as a neural simulator for autonomous driving to replace or augment physical test fleets (“deploy 100 cars to 100 intersections without leaving the office”), the challenges of robot manipulation compared to navigation due to occlusion and deformation during contact, cross-embodiment generalization between humanoid variants, and the staged path from passive policy verification toward active reinforcement learning in simulation. Liu also engages with the question of what a “world model” actually is, arguing it is best understood as a collection of tools for achieving specific modeling goals — forward dynamics, inverse dynamics, and policy — rather than a single unified concept.
📺 Source: Machine Learning Street Talk · Published September 15, 2026
🏷️ Format: Interview







