Summary
Sebastian Fox, a medical doctor and co-founder of Composo (an AI evaluation systems company for high-stakes domains), opens with a case study that sets the stakes immediately: an AI ambient scribe that wrote up a routine headache and missed a single line — jaw pain on chewing — that would have flagged giant cell arteritis, a same-day steroid emergency that can cause blindness within days. The missed detail was not hallucinated; it simply never made it into the note. Nothing in the note was technically wrong. That, Fox argues, is exactly what makes these failures so dangerous.
Drawing on the largest real-world study of AI clinical notes, Fox presents production error rates that are difficult to dismiss: roughly 1 in 20 notes carries an error serious enough to cause significant patient harm, nearly 1 in 5 contains an important omission, and more than 1 in 10 contains a hallucination. Ambient scribes are already deployed in approximately a third of US medical practices, physician AI use doubled in the past year, and almost none of it is tracked through adverse event reporting systems.
The most technically substantive part of the talk explains why even sophisticated LLM-judge evaluation pipelines — with detailed faithfulness rubrics, worked examples, and deterministic NLP concept-counting — still wave through serious errors. Fox demonstrates this by running the same notes that contained errors through a strong judge system and finding that one in five clean passes still contained a buried serious error, usually an omission. The root cause: verification is only cheaper than generation for the easy cases (spot-the-difference between transcript and note), not for the hard ones (determining which differences actually matter clinically). The talk frames this as a general problem for any high-stakes AI deployment, not just healthcare.
📺 Source: AI Engineer · Published August 22, 2026
🏷️ Format: Deep Dive







