Summary
Tisha Chawla and Susheem Koul from Microsoft open with a scenario that will be familiar to anyone running AI agents in production: an agent misinterprets “sell $1,000 of stock” as a quantity instruction rather than a dollar amount, selling 1,000 shares at $190 each for a $190,000 unintended trade — with zero exceptions, zero alerts, and perfectly green dashboards. The central challenge the talk addresses is why you can’t reproduce this failure and therefore can’t fix it.
The speakers systematically dismantle the most common attempted fix: setting model temperature to zero. They explain why this doesn’t work — GPU floating-point operations are not associative, so tiny differences in matrix computation order alter final logits; batching means a request is grouped with whatever else hits the server at that millisecond; and mixture-of-experts routing routes tokens differently depending on current batch load. The right distinction, they argue, is between bitwise determinism (same input → same output, which you won’t get from hosted APIs) and replayability (re-validating a specific run that already happened).
To enable replayability, the team built Chronicle, a framework that annotates agent node boundaries — tool calls, LLM invocations, RAG retrievals — to capture input/output pairs along with model version and sampling metadata as frozen traces. Recording happens at the semantic boundary rather than the network layer, enabling deterministic CI replay with zero model calls. The result is a structured debugging loop: annotate boundaries, record production traces, visualize the failure in a detailed JSON timeline, isolate the broken node, fix it, replay the frozen trace, and verify the corrected state transition.
📺 Source: AI Engineer · Published June 29, 2026
🏷️ Format: Deep Dive







