How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel

How to Build Better AI Evals with Claude Code in 5 Steps | Shreya & Hamel

More

Summary

Peter Yang hosts Hamel Husain and Shreya Shankar — the instructors behind one of the most widely taken AI evaluation courses online, with over 4,500 students — for a practical walkthrough of how to build better AI evals using Claude Code. The session opens with a framework distinction that shapes the entire demo: “bottom-up” evals are data-driven and must come from the human analyst who has reviewed real traces; “top-down” evals encode subjective taste and judgment criteria. The guests explain that Claude Code is strong at helping you explore data and generate evaluation interfaces, but cannot surface bottom-up failure modes on its own — that analysis is irreducibly human.

The demo uses a real podcast post-production skill that generates newsletter takeaways, with Hamel and Shreya live-critiquing the host’s existing eval suite. They walk through how to use Claude Code as an agentic assistant to scan traces, classify error types, and draft structured LLM judge criteria — while flagging common mistakes like over-relying on character-length checks instead of semantic quality. A cited study of 400,000 Claude Code sessions is used to illustrate why domain expertise doubles verified task success rates compared to novice usage.

The second half of the video features a candid comparison of daily-driver AI coding tools among the three practitioners. Hamel makes a detailed case for Codex based on its mobile command-center interface, cross-thread orchestration (letting one Codex session steer others), and favorable subscription economics. Shreya and Peter offer counterpoints around output clarity and model behavior, giving viewers an unfiltered view of how experienced AI builders currently choose their tools.


📺 Source: Peter Yang · Published August 23, 2026
🏷️ Format: Tutorial Demo

1 Item

Channels

1 Item

Companies

1 Item

People