Speech Recognition Is Not a Solved Problem — Pavan Muddireddy

Speech Recognition Is Not a Solved Problem — Pavan Muddireddy

More

Summary

Machine Learning Street Talk sits down with Pavan Muddireddy, audio research team lead at Mistral AI, for a deep technical conversation on why speech recognition remains an open problem despite years of progress. Muddireddy explains his day-to-day work at Mistral — a company spanning open-weight models, the Mistral.ai chat product, the Vibe Work agent platform, AI Studio, Forge fine-tuning, and MRA Compute — and how the convergence of pre-training and post-training techniques across modalities shapes his audio research.

The discussion gets into the mechanics of autoregressive speech generation, including a key architectural constraint: once a token is predicted and committed, it becomes frozen context that cannot be revised, unlike text editing. Muddireddy traces the evolution from hand-crafted mel spectrogram features toward end-to-end waveform inputs, explaining why removing encoder components at scale allows the model to find a better optimization point without imposed inductive bias — and makes scaling behavior more predictable.

A substantial portion covers noise robustness in ASR: how noise augmentation parallels flipping or scaling in vision models, why simulating varied acoustic conditions from limited data is essential, and how even frontier models still have meaningful accuracy gaps under real-world conditions. For practitioners working on voice AI, this episode offers rare firsthand perspective from a researcher actively running ablations on speech models at a frontier lab.


📺 Source: Machine Learning Street Talk · Published September 14, 2026
🏷️ Format: Podcast

1 Item

Channels

1 Item

Companies