Summary
In this a16z-hosted conversation, Simon Mo — co-founder of Infact and lead maintainer of vLLM, the open-source inference engine now running on half a million GPUs at any moment — sits down with a16z General Partner Matt Bornstein to trace how open-source AI evolved from academic curiosity to critical infrastructure. The discussion covers the engineering challenges unique to large language model serving: non-deterministic output lengths, complex batching and scheduling, and the GPU memory management problems that vLLM’s paged attention approach was designed to solve.
A substantial portion of the conversation focuses on shifting licensing norms. Where early open-weight releases used permissive Apache 2.0 terms, newer models from labs including Meta (Llama) and MiniMax are introducing commercial revenue thresholds and restrictions on derivative works as they attempt to sustain research costs — a tension the panel unpacks carefully, distinguishing between true open source and open weights. The licensing terms for Kimi and the controversy around Fireworks and Cursor building on top of Kimi’s model are cited as concrete examples.
Simon also highlights vLLM’s fast mode, which pushes open-weight models to 400–500 tokens per second, and argues that for many developer tasks the capability gap between open and closed frontier models has effectively closed today. The panel closes with a thought experiment: if GPU costs dropped 99%, would we return to a truly open AI ecosystem — and what role would moderation constraints play in that future?
📺 Source: a16z · Published August 06, 2026
🏷️ Format: Interview







