Why World Models Could Change Robotics, 3D, and Creativity

Why World Models Could Change Robotics, 3D, and Creativity

More

Summary

In this a16z Latent Space interview, the founders of Accelerated Understanding — Anima Anandkumar and Justin Johnson — introduce Atlas, a new world model built around what they call “new view prediction,” a fundamental AI primitive distinct from the next-token prediction powering LLMs or the next-frame prediction underlying video models. Atlas takes a spatial context — any number of input images or a scene description — and generates what that world looks like from an arbitrary position in space and time.

The practical implications are demonstrated through bullet-time video generation: recreating Matrix-style frozen-time camera flyaround shots using just three iPhones instead of the hundreds of synchronized cameras previously required. Atlas simultaneously handles camera-conditioned generation, sparse 3D reconstruction from as few as one input frame, and robotic simulation. The founders draw a direct analogy to the NLP transition from task-specific models to large general-purpose language models, arguing that physical world understanding is undergoing the same unification.

The technical discussion covers why Atlas treats reconstruction as a high-context form of generation, enabling a principled continuum between the two. The team notes that unlike video models that implicitly assume next-frame consistency, Atlas explicitly reasons about 3D geometry, making it far more useful for robotics, architectural visualization, and film production. Atlas was publicly launched the day before the interview was recorded and received significant online attention for its bullet-time demonstrations.


📺 Source: a16z · Published September 04, 2026
🏷️ Format: Interview

1 Item

Channels

1 Item

People