Summary
Alfonso Graziano, tech lead at NearForm and O’Reilly author of “Learning AI Native Software Engineering,” presents a structured methodology for using AI coding agents to iteratively build and improve other AI agents. The core idea: since AI agents are just software, the same AI-powered development workflows that accelerate normal coding can be applied to agent optimization itself.
The talk introduces the concept of a golden dataset — a curated set of input/expected-output pairs developed with subject-matter experts — which functions as a test suite for non-deterministic systems. A companion scorer uses an LLM to evaluate agent outputs against this dataset, producing an accuracy baseline. From there, a coding agent (demonstrated using the Mastra framework) runs an automated optimization loop: generate a hypothesis, modify the target agent, run evals, compare metrics, and roll back or continue. Graziano shows a live run achieving incremental gains of 5% and 12% across iterations, with regressions triggering automatic rollbacks.
The methodology directly addresses the two most common agent failure modes — poor eval performance and unexpected behavior on live data — without requiring manual prompt rewriting or expensive model swaps. The session is structured around repeatable steps that teams can adopt immediately, and includes practical guardrails like preventing the coding agent from gaming scorers rather than genuinely improving the target system.
📺 Source: AI Engineer · Published June 28, 2026
🏷️ Format: Deep Dive







