How Evals and Prompts Shape Agent Behavior — Preetika Bhateja & Daniel Bump, YouTube Ads

How Evals and Prompts Shape Agent Behavior — Preetika Bhateja & Daniel Bump, YouTube Ads

More

Descriptions:

Preetika Bhateja and Daniel Bump, engineers on the YouTube Ads image and video models team at Google, share practical lessons from building production-grade evals for generative AI agents at the AI Engineer conference’s evals track. Their core message: getting evals right is a sequenced process, and jumping to scalable automated rating too early can cause more confusion than clarity.

The talk introduces a staged approach — starting with intuition-based ‘vibing’ (manually inspecting outputs) to understand failure patterns before committing to a formal eval framework. This early-stage, non-scalable approach lets teams make radical architectural changes quickly without an eval system constraining iteration. Once the agent’s foundation is stable, they recommend layering in a critique agent with a remediation loop before scaling to full eval infrastructure.

For teams working with human raters, Bhateja and Bump emphasize the value of collecting explanations alongside pass/fail judgments, particularly for multi-output evals where an ad might pass on brand safety but fail on accuracy. When transitioning to LLM-as-judge, they describe monitoring human-LLM agreement rates as a calibration mechanism, supplemented by spot-checking the reasoning chains behind automated verdicts. The talk also covers agent trace analysis as a complement to outcome-based evals, helping teams understand not just whether an output was correct but which path through the code produced it — a distinction that becomes critical as agent behavior grows non-deterministic at scale.


📺 Source: AI Engineer · Published July 24, 2026
🏷️ Format: Deep Dive

1 Item

Channels