Summary
Matthew Berman walks through a model routing strategy that reduces AI development costs by roughly 68% without measurable quality loss. The core principle separates planning — where frontier model reasoning is genuinely required — from execution, where a well-specified plan can be implemented effectively by a cheaper coding model. In Berman’s workflow, Fable handles research, specification writing, and optional pull request review, while GPT-5.5 or a similarly cost-efficient model handles the actual code generation.
Berman backs the approach with explicit token-cost arithmetic: planning a feature with Fable at $10/M input and $50/M output tokens costs approximately $2, while coding with Fable adds another $7.50 — totaling $9.50. Offloading code generation to a model priced at $2/M input and $6/M output drops that phase to $1.02, bringing the combined total to $3.02. The math holds because a complete, well-structured spec removes the need for the execution model to perform high-level architectural reasoning.
The video covers both manual implementation — switching between Claude and OpenAI subscriptions — and automated model routing tools that handle the handoff without user intervention. Berman also demonstrates the spec format Fable produces (nearly 600 lines for a single feature), showing why a detailed spec enables cheaper models to perform at near-frontier quality during the execution phase.
📺 Source: Matthew Berman · Published July 07, 2026
🏷️ Format: Tutorial Demo







