Summary
Midam Kim, ML engineer at ServiceNow and lifelong speech communication researcher, presents a linguistic framework for diagnosing voice AI failures at AI Engineer 2026. Starting from a personal anecdote — a customer service bot that repeatedly misidentified her name “Midam” as “Madam” despite multiple corrections — she builds a structured model of where voice agents diverge from the joint activity that defines human communication.
The framework organizes failure modes across two channels (listening and speaking) and four interdependent levels: sounds (STT phoneme recognition), words (language and domain understanding), interaction (turn-taking timing and interruption handling), and mental model (intent tracking across the full conversation). Kim argues these components cannot be optimized in isolation — improving ASR accuracy without addressing turn detection or context continuity will still produce frustrating interactions, because all four levels must stay aligned over time.
A key insight is that voice is uniquely ephemeral: unlike text chat, there is no visible history, so the only persistent artifact is the user’s mental model of what was communicated. When a bot fails to track that model — by asking for information already provided, cutting off a user mid-utterance, or applying TTS pronunciation rules that mangle an unfamiliar name — users escalate to human agents quickly. The talk gives engineers from STT, LLM, and TTS backgrounds a shared vocabulary for locating and reasoning about failure, grounded in decades of linguistic research on conversational norms.
📺 Source: AI Engineer · Published September 15, 2026
🏷️ Format: Keynote Launch







