Summary
Sohail Shaikh and Ankush Rastogi of Prosodica present research at the AI Engineer conference on a performance failure pattern they call the “100-tool agent trap”: the common practice of loading every available tool definition into an agent’s context on every request. The talk demonstrates with benchmark data that this approach, while workable at small scale, collapses badly as tool catalogs grow.
The numbers are striking. Using the Berkeley Function Calling Leaderboard, SkillsBench scenarios, and synthetic tool pools, the team measured tool selection accuracy at various catalog sizes. A standard “fat agent” achieves roughly 78% accuracy at 10 tools โ acceptable, but not great. At 100 tools, accuracy drops below 40%. At 741 tools, it falls to 13.6%, meaning the agent selects the correct tool less than once in eight attempts. The root cause is the “lost in the middle” problem: LLMs pay stronger attention to content at the beginning and end of long contexts, and when hundreds of tool schemas are packed into the middle, the model stops reliably attending to them. A 741-tool schema also requires approximately 127,000 tokens before the user’s query is even considered โ at 100,000 requests per day, that compounds into billions of wasted tokens.
The proposed solution is semantic routing with just-in-time context injection โ essentially RAG applied to tool selection. Tool descriptions are embedded and stored in a vector database; at query time, only the top-K most semantically relevant tools (K=5 is the recommended default) are retrieved and injected into the prompt. This keeps model accuracy above 83% at all tested catalog sizes and reduces input token volume by approximately 99%. The presenters cite Anthropic’s own MCP documentation as reporting a 98.7% token reduction from the same architectural pattern, validating the approach independently.
๐บ Source: AI Engineer ยท Published June 28, 2026
๐ท๏ธ Format: Deep Dive







