Summary
Vivek Trivedy, who leads applied research at LangChain, presented a systematic framework for continuously improving AI agents by treating the problem as a data mining challenge rather than a prompt-engineering exercise. The core loop he advocates: ship the agent, collect massive volumes of trace data, mine those traces to surface good and bad interactions, then run data-driven experiments to validate whether changes to prompts, tools, or orchestration actually improve outcomes.
A standout section covers LangChain’s work with Harvey on the legal benchmark, where the team showed that an open-source model could match Claude Opus-level trace judging capability at roughly two orders of magnitude lower cost — achieved through careful harness engineering informed by reading Opus reasoning patterns in the traces. Trivedy also addresses the fine-tuning inflection point: once prompt engineering hits diminishing returns, fine-tuning base models on narrow vertical tasks can push performance past frontier levels while allowing teams to shift from per-token API costs to fixed hardware costs.
The talk concludes with a look at LangChain’s LangSplat Engine, which automates the trace-reading and eval-generation loop for teams handling large volumes of agent telemetry. Engineers building agents at scale — particularly those wrestling with how to systematically learn from production failures rather than relying on ad hoc fixes — will find the methodology practically applicable.
📺 Source: AI Engineer · Published August 12, 2026
🏷️ Format: Keynote Launch







