Summary
In this a16z show conversation, the hosts discuss the rapid rise of consumer personal AI agents and what would make an assistant worth paying for. They trace the arc from early consumer agents like Poke through OpenClaw and the recent wave of Instinct, Muse, Grockbot and ChatGPT’s new agent offerings, describing how quickly the space has moved in just a few weeks.
Guest David Pollen introduces AssistantBench, a consumer-facing benchmark that gives different AI assistants the same one-shot prompts, such as booking a flight to Chicago or finding a vegetarian restaurant within five blocks, and compares outcomes across 16 dimensions. The discussion explores where assistants succeed and fail in practice.
Other topics include proactivity as a source of defensibility and the trust risk of overstepping, why the general population may not care about being 10% more efficient, and the possible infrastructure needed for agent-to-agent interactions. The speakers also compare three approaches to bringing agents into group chats, including a silent note-taking agent that pings individuals separately, and consider whether Meta’s Muse charm is more about real-world data collection than hardware. It is the first in a planned series on personal agents.
📺 Source: a16z · Published September 29, 2026
🏷️ Format: Podcast







