Summary
Sachin Malhotra, an engineer on Anthropic’s CI team, delivers a conference talk at AI Engineer built around a real production incident: an agent given a God token and a tool list deleted approximately 200 workloads in 90 seconds, wiping hours of in-progress training jobs from 20 engineers. His argument is that narrowing token scope — the standard post-incident fix — doesn’t scale, any more than you’d solve a new hire’s mistake by taking away their delete key permanently.
The talk introduces three primitives for agent governance drawn from Anthropic’s internal experience. First, asymmetric verbs: reads and writes are treated fundamentally differently, with a proxy layer that stamps the caller’s identity on every action so the agent never holds provenance of its own operations. Second, rate limits as budgets: every caller gets a ceiling of disruptive actions per time window that refills automatically, scoped by namespace (personal vs. shared resources), with no approval gates — just a bounce-back with a count when exceeded. Third, trip wires over allow lists: rather than guessing up front what an agent will need, let it act freely on cheap operations while recording every action with an identity stamp, then use that data to build informed policy after the fact.
Malhotra ties all three together with what he calls the undo test: ask whether the action is reversible, and size the rate limit accordingly. A practical detail worth noting: inside Claude Code sessions, the bypass flag for the rate-limit admission webhook simply refuses to execute — it tells the agent to ask the human to run the command directly, keeping human override structurally in the loop without requiring a ticket.
📺 Source: AI Engineer · Published August 22, 2026
🏷️ Format: Keynote Launch







