Summary
Sumanyu Sharma, founder and CEO of Hamming AI, opens with an unusual credential: before working on voice agent reliability, he spent years monitoring thousands of hours of police radio at Citizen, the real-time crime alert app. That background shapes a talk that takes voice agent failure modes seriously — and quantifies them with unusual specificity.
Hamming currently monitors 10,000 deployed voice agents in production. The observed error rate is approximately 10% — agents skipping eligibility verification, applying unauthorized discounts, misrepresenting what a user said, or claiming to book appointments that were never created. With an estimated one trillion phone calls annually and the majority projected to be handled by voice agents within five years, even a 1% error rate implies 10 billion incidents per year. Sharma argues this makes voice agent reliability a safety problem, not just a product quality problem, and the blast radius of a centralized prompt change — affecting millions of users simultaneously — makes it qualitatively different from individual human error.
The talk presents a five-step improvement framework: identify failure categories, prioritize by frequency and severity, diagnose root causes, execute fixes, and measure for regressions. Sharma distinguishes between known problems trackable with standard evals (turn-taking latency, ASR errors) and emerging cross-conversation patterns only visible at scale — the latter being where Hamming focuses. He argues that teams consistently underinvest in production monitoring relative to pre-deployment evals, and that manual call listening, while not scalable, remains an irreplaceable starting point for building intuition.
📺 Source: AI Engineer · Published September 15, 2026
🏷️ Format: Deep Dive







