Summary
Tisha Chawla and Susheem Koul from Microsoft take the stage at AI Engineer to address what they call the most expensive question in AI today: when an agent workflow generates a surprise bill, who actually spent the tokens? The talk frames the problem historically — SaaS controlled costs via seat limits, cloud via autoscaling policies — and argues that the agentic era lacks an equivalent control plane at the model-call boundary, not just the gateway level. Real examples cited include Uber’s AI budget exhausted within four months of deployment and multiple companies burning hundreds of millions of dollars from runaway agent loops.
The solution they present is a three-layer architecture they call TokenOps. The agent runtime is instrumented via a bridge layer containing three components: attribution (tagging every agent run to user dimensions), boundary annotation (a decorator applied to any method regardless of framework — LangChain or otherwise — that logs input/output to a ledger and acts as a bidirectional channel for the control plane), and a governor node that receives and executes steering actions pushed down from the control plane. The key design insight is that the system intervenes in-loop rather than halting: if a RAG retrieval tool is returning 20 chunks but the LLM only uses the first five, the control plane can push a cap down through boundary annotation without stopping the run.
The talk includes a live demo showing the full flow: attribution tagging, boundary-annotated agent runs, and real-time control plane policy adjustment. The goal throughout is a shift from token maxing — spending freely for exploration — to value maxing, where every token can be traced to a specific run, agent, and business outcome.
📺 Source: AI Engineer · Published August 22, 2026
🏷️ Format: Keynote Launch







