Descriptions:
At the AI Engineer conference, Ben Hylak of Raindrop delivered a candid, practitioner-focused talk on what actually moves the needle when building production AI agents. Rather than surveying trendy frameworks, Hylak oriented the session around a core distinction: are teams optimizing for benchmark scores or for raising the floor of reliable real-world performance? He argued that most publicly circulated eval discourse remains stuck in the chatbot era — fact-checking style Q&A datasets — and fails to transfer to agents operating across finance, healthcare, and defense.
Hylak challenged the assumption that continual learning is widely deployed, noting that in practice very few production teams run it. He also drew a sharp line between what AI labs need from evaluations (general-purpose robustness) versus what product companies need (domain-specific, company-specific reliability). The talk touched on the spectrum of human oversight required across agent types, from autocomplete-style coding assistants where errors are trivially reversible to high-stakes autonomous agents like AI diagnostics where the responsibility calculus is entirely different.
The session included live Q&A soliciting failure modes from the audience, making it a useful window into where practitioners are actually struggling rather than what conference slides typically advertise. Engineers building agents in regulated or high-stakes domains will find the floor-versus-benchmark framing and the discussion of domain-specific eval design particularly relevant.
📺 Source: AI Engineer · Published August 12, 2026
🏷️ Format: Keynote Launch







