Building Turbopuffer: Gergely Orosz (@pragmaticengineer ) × Simon Eskildsen (CEO)

Building Turbopuffer: Gergely Orosz (@pragmaticengineer ) × Simon Eskildsen (CEO)

More

Descriptions:

Gergely Orosz (The Pragmatic Engineer) sits down with Simon Eskildsen, founder and CEO of Turbopuffer, for an unusually technical founder interview about building a vector database on top of Amazon S3. Eskildsen traces his path from competitive programming (International Olympiad in Informatics) and Shopify to founding Turbopuffer after a concrete cost problem: while consulting for Readwise, he built a recommendation engine using vector embeddings that worked well but would have cost $30,000 per month — six times Readwise’s entire infrastructure budget — making it unshippable.

The core of the conversation covers Turbopuffer’s architectural bet: organizing and clustering vector files on S3 to achieve low query latency despite S3’s notoriously high P99 latency (~200ms per object fetch). Eskildsen explains why P99 latency matters more than median latency when a system fans out across many parallel requests, and walks through the summer of 2023 when he spent months finding an approach that could deliver acceptable query performance from object storage.

The interview is grounded in real production constraints and economics, making it a practical reference for engineers building RAG pipelines, semantic search systems, or any AI application where vector storage cost is a bottleneck. Eskildsen also reflects on how the collapse in embedding generation costs since 2023 has dramatically changed what is economically feasible to build — context that shapes the design space for anyone evaluating managed versus self-hosted vector database options today.


📺 Source: AI Engineer · Published August 03, 2026
🏷️ Format: Interview

1 Item

Channels