The State of Model Routing — NVIDIA, Cognition, OpenRouter

The State of Model Routing — NVIDIA, Cognition, OpenRouter

More

Summary

At an AI Engineer conference focused on local and on-device AI deployment, a panel featuring Walden from Cognition (makers of Devon AI), Carter and Dane from NVIDIA, and a representative from OpenRouter explores the fast-evolving state of model routing — the practice of dynamically selecting which AI model handles a given task to balance cost, latency, and output quality.

Cognition’s Walden opens by discussing the company’s newly released Fusion router, which the team claims achieves better-than-frontier-model performance on certain task classes. The counterintuitive mechanism: smarter orchestrator models are better at delegation, so routing shouldn’t simply push users to cheaper, dumber models. Instead, Fusion keeps a frontier model in the orchestrator role while routing execution to smaller models — maintaining a running KV cache to avoid the cost of re-providing context, which Walden notes makes cached tokens roughly ten times cheaper. NVIDIA’s team discusses Flex Run, a technology that distills frontier models into smaller weight footprints and dynamically activates different weight classes based on task novelty and in-distribution probability — a technique particularly effective with open models where training data provenance is known.

The panel also covers emerging research into training models specifically for collaborative multi-model environments — both as orchestrators deciding what to delegate, and as “sidekick” executors optimized to follow another model’s instructions. The consensus view is that model routing is still early, with significant opportunity for startups to define the tooling layer, and that co-designing models with their routing harness will likely be a major architectural trend in production AI systems.


📺 Source: AI Engineer · Published August 06, 2026
🏷️ Format: Deep Dive

1 Item

Channels

3 Items

Companies