Your coding agent doesn’t always follow your rules — Talha Sheikh, Checkout.com

Your coding agent doesn’t always follow your rules — Talha Sheikh, Checkout.com

More

Summary

Talha Sheikh, an engineer at Checkout.com, argues that deterministic verification of AI coding agent output is a permanent architectural requirement — not a workaround for today’s imperfect models. The talk grows out of a familiar frustration: giving Claude Code a feature request, watching sub-agents run to apparent completion, and then discovering the result doesn’t work.

Sheikh built a tool called Vector V1 that hooks into Claude’s session lifecycle using Claude hooks. When a session ends, Vector automatically runs a declarative set of test cases against the output. Failures are fed back to the agent, which retries until all checks pass or a human needs to intervene — creating a closed verification loop without manual review of every task. The config-driven approach means test cases can be defined once and reused across sessions.

The talk’s central argument is that model capability and model reliability are distinct properties that don’t converge. Even as Claude Opus and future frontier models become more capable, giving an agent better instructions or more context is not equivalent to verifying its output — they address different failure modes. Sheikh extends this into a cost optimization argument: a sufficiently tight verification harness allows cheaper models (Claude Haiku, open-source alternatives) to handle constrained tasks reliably, reducing inference costs significantly. He closes by framing verification not as a proprietary product but as an industry pattern — observing that Anthropic, Meta, and most engineering teams are all independently building their own enforcement layers — and suggesting the field needs a shared, language-agnostic specification.


📺 Source: AI Engineer · Published July 08, 2026
🏷️ Format: Hands On Build

1 Item

Channels

1 Item

Companies