Summary
Claire of How I AI interviews Anker Goyel, CEO of BrainTrust — a startup specializing in AI evaluation and observability tooling — about how coding agents are transforming the way senior engineers approach complex infrastructure work.
Goyel’s central argument is that the real unlock of today’s capable models isn’t faster code generation but the ability to define rigorous success criteria and let an agent iterate creatively in the background. He describes using Codex to optimize slow queries running against billions of traces in BrainTrust’s product, where a user might search for 5,000 specific interactions across a 90-day window of data. Rather than treating this as a prompt-and-review task, Goyel structures it as an eval problem: write hard tests, set clear performance targets, and let the model find solutions autonomously. He also describes running multi-day remote experiments on EC2 to measure real S3 latency at 4,000 concurrent reads — work he notes would overwhelm a local machine.
On the limits of commercial agent tooling, Goyel is candid: off-the-shelf background agents work well for standard SaaS applications but fall short on complex, custom infrastructure. He maintains a personal concurrency limit of roughly four simultaneous foreground agent tasks and notes that both large and small engineering teams are increasingly building their own proprietary background agent systems. The episode is explicitly targeted at senior engineers, VPs of Engineering, and CTOs, and includes a segment demystifying how to build effective evals for AI-powered products without deep ML expertise.
📺 Source: How I AI · Published June 15, 2026
🏷️ Format: Interview







