Summary
At the AI Engineer conference, Archana Kamath (VP of Engineering, Inference Engine & AI Infrastructure at DigitalOcean) and engineer Tyler Gillam make the case that chasing benchmark leaderboards is the wrong way to pick an AI model — and that the right model depends entirely on the specific request, system prompt, cost tolerance, latency requirements, and end-user preference.
The talk introduces DigitalOcean’s Inference Light Router, a purpose-built open-source routing model that sits inside their managed inference engine. Unlike black-box auto-routing approaches, the system lets developers define explicit preferences per task — for example, always prefer GLM 5.2 for code generation with GPT 5.2 as a failover, or route bug-fixing tasks to whichever model in the pool has been fastest in the last 30 minutes. The routing decision itself adds under 200 milliseconds of latency and costs customers nothing extra. In DigitalOcean’s own evaluations, the specialized routing model outperforms GPT-5-series models on routing tasks at a fraction of the latency.
Tyler’s live demo walks through the DigitalOcean cloud console, showing preset router templates (software engineering, general writing, knowledge bases), custom task configurations, and side-by-side playground comparisons pitting Claude Opus directly against the router. The router matched 90% correctness versus Opus’s 95% — likely within judge margin of error — while consuming significantly fewer tokens and returning results faster. The session also covers the built-in evaluation framework for measuring and iterating on router performance over time.
📺 Source: AI Engineer · Published August 22, 2026
🏷️ Format: Keynote Launch







