Beyond Static Intelligence: Evaluating Continual Learning — Parth Asawa, UC Berkeley

Beyond Static Intelligence: Evaluating Continual Learning — Parth Asawa, UC Berkeley

More

Summary

Parth Asawa, a PhD student at UC Berkeley, delivers a conference talk arguing that the AI field is systematically failing to measure one of the most important properties of modern agents: the ability to learn over time. Today’s standard benchmarks evaluate models as if they have no memory between tasks — a design choice, Asawa argues, that made sense for frozen checkpoints but is increasingly out of step with how agents are actually deployed.

The talk introduces a new evaluation framework built around three metrics measured on Pareto frontiers: reward (absolute task performance), gain (the improvement attributable specifically to prior experience rather than base model strength), and cost (the compute expenditure of the learning process). The gain metric is the methodological centerpiece — it requires running any system twice, once stateful and once with memory wiped between instances, so that the delta isolates what learning actually contributed. Without this separation, a stronger base model can look like a better learner when it is simply better to begin with.

Asawa grounds the framework in a concrete database exploration task where agents must learn schemas, table relationships, and data idiosyncrasies over repeated interactions — mirroring what human data engineers do naturally. The broader argument is that optimizing for point-in-time capability scores will produce models that are locally impressive but globally stagnant, and that the field needs richer longitudinal benchmarks before continual learning can be meaningfully advanced.


📺 Source: AI Engineer · Published August 12, 2026
🏷️ Format: Keynote Launch

1 Item

Channels