Summary
Benedikt Sanftl (CEO) and Burak (CTO) of Mutagent present the concept they call the “Agentic AI Engineer” — an autonomous agent that manages the full development lifecycle of other AI agents, from specification through build, evaluation, deployment, production monitoring, diagnosis, and optimization. The core argument: as organizations scale to tens or hundreds of AI features, the human-in-the-loop development cycle becomes the bottleneck, and the only way to break it is to have agents run the loop itself.
The framework defines two paths: cold start (building an agent from scratch) and iteration on an existing deployed agent. Both paths cycle through the same stages — spec, build, eval, ship, monitor, diagnose, optimize — with agents automating the transitions between stages. A key insight on evaluation design: score-based LLM-as-judge approaches are insufficient because a numeric score doesn’t tell engineers what to fix. Binary criteria evals are preferred because a failed criterion directly implies a corrective action. Trajectory evaluation is also emphasized — checking every tool call in an agent’s session, not just the final output, since a wrong intermediate tool result propagates through to a wrong answer even if the final step looks correct.
The talk includes a live demonstration of Mutagent’s platform orchestrating agent improvements: automatically generating eval datasets from production traces, running trajectory analysis across 200 samples in parallel, proposing targeted prompt mutations for identified failure modes, and iterating without requiring a human to review individual traces. Sanftl positions this as the natural next step after current vibe-coding and manual agent iteration workflows.
📺 Source: AI Engineer · Published June 29, 2026
🏷️ Format: Keynote Launch







