The Future of Evals: From LLM as a Judge to Agent as a Judge — Aparna Dhinakaran, Arize AI

The Future of Evals: From LLM as a Judge to Agent as a Judge — Aparna Dhinakaran, Arize AI

More

Descriptions:

Aparna Dhinakaran, co-founder of Arize AI, opens the evals track at the AI Engineer conference with a talk charting how evaluation methodology must evolve to keep pace with increasingly complex agentic systems. Drawing on Arize’s position running over 100 million evals per month across teams that average 12 distinct eval jobs (with top teams running more than 3,800 evaluators), she argues that the classical LLM-as-a-judge paradigm — fixed rubric, deterministic scoring — fundamentally breaks down when agents produce different trajectories on every run.

The core insight: when Arize’s own in-product agent Alex gained capabilities like dynamic UI generation, long-horizon memory, and cross-trace search, the team discovered it was failing in ways no static rubric could catch — context forgetting, tool repetition loops, inefficient trajectories. This led to what Dhinakaran calls the natural next step: evaluating agents with agents.

The talk announces the launch of Signal, Arize’s new agent-as-a-judge product. Signal is a long-running background agent that reads production traces, discovers emergent failure patterns, and — crucially — can generate a pull request proposing a fix. Dhinakaran frames agent-as-a-judge not as a replacement for deterministic evals or LLM judges, but as a third layer for problems those tools cannot address. The talk sets up a full evals track at AI Engineer featuring speakers from Term Bench, Uber, and Snorkel.


📺 Source: AI Engineer · Published July 24, 2026
🏷️ Format: Keynote Launch

1 Item

Channels

1 Item

Companies