Descriptions:
Machine Learning Street Talk sits down with researchers Ilia Shumailov and Alexander Panfilov to discuss their paper “Stealing Reasoning Traces from Proprietary LM APIs,” which reveals that encrypted reasoning blobs returned by frontier models — including GPT and Claude — can be decoded using smaller models within the same provider family. The paper reportedly garnered 3 million views within 40 hours of release.
The researchers explain how this architectural vulnerability exists across all three major providers they tested: Anthropic, OpenAI, and Google. Because the companies share similar architectural approaches, they share the same fundamental weakness: reasoning blobs can be replayed in arbitrary contexts, enabling attacks including jailbreaks, prompt injections, extraction of secrets from other users’ sessions, and unauthorized training on decoded reasoning traces. The conversation unpacks why models return these blobs to clients at all — largely due to stateless architecture and cost considerations — and why that design decision created the exposure.
The discussion covers responsible disclosure (all three labs acknowledged the report without legal pushback), potential mitigation paths at the architectural, system, and model levels, and why fixing the architectural layer alone is insufficient. For developers and security researchers working with reasoning model APIs from any major provider, this interview provides essential context on an unresolved structural risk in how intermediate computation is currently handled.
📺 Source: Machine Learning Street Talk · Published August 22, 2026
🏷️ Format: Interview







