The 100,000 Sandbox Problem — Akshat Bubna, Modal CTO

The 100,000 Sandbox Problem — Akshat Bubna, Modal CTO

More

Descriptions:

Latent Space hosts Modal CTO Akshat Bubna for an in-depth conversation on how Modal evolved from a general serverless runtime into a foundational layer for AI agent infrastructure, and what it means to scale sandboxed code execution to 100,000 concurrent environments. Bubna traces Modal’s origin — born from frustration with Kubernetes’s poor fit for bursty, GPU-heavy workloads — through its early GPU additions (predating ChatGPT) to its current position spanning 17 cloud providers and NeoCloud partners.

A central theme is the shift from developer experience (DX) to agent experience (AX): Modal redesigned its SDK so that agents can provision their own runtimes via decorator-based configuration rather than navigating hundreds of Kubernetes YAML files. Bubna discusses how the inference inflection point — where AI workloads now constantly alternate between GPU and CPU phases rather than being GPU-dominated — makes collocated, flexible compute increasingly important. He also covers Modal’s snapshotting primitives, which enabled projects like Ramp Inspect to deliver reactive background agents.

The episode addresses Modal’s strategic boundary decisions: focused on the model and agent lifecycle (data prep through inference and persistent agent deployment), deliberately capital-light by avoiding owned data centers, and differentiating on software rather than hardware. Engineers and architects building agent infrastructure will find specific detail on sandbox scaling, self-provisioning runtimes, and the economics of multi-cloud capacity pooling.


📺 Source: Latent Space · Published July 08, 2026
🏷️ Format: Podcast