You NEED to do this (HUGE AI SAVINGS)

You NEED to do this (HUGE AI SAVINGS)

More

Summary

Matthew Berman argues that the most effective AI users have moved beyond prompt engineering and are now focused on token-level optimization — selecting the right model for each sub-task rather than routing everything through a single expensive frontier model. The video opens by explaining why raw per-token pricing is a misleading metric without accounting for how many tokens a given model consumes to complete a task, a concept Berman calls ‘intelligence density.’

Using Artificial Analysis benchmark data, Berman compares GPT-5.6 Soul ($5/$30 per million input/output tokens), Claude Fable (higher still), and Kimmy K3 ($3/$15 per million tokens). Despite Kimmy K3’s lower headline price, the data shows it requires roughly twice as many tokens to complete the same tasks as GPT-5.6, making real-world cost-per-task nearly identical at around $95–$104. Claude Fable comes in at $2.75 per task completed versus GPT-5.6’s $1.04, making it substantially more expensive on an output-adjusted basis.

The actionable recommendation is a three-model workflow: use a frontier model like Claude Fable for high-level planning (where reasoning quality matters most and token volume is low), a cheaper fast model like Grok 4.5 or Cursor’s Composer for code execution (where output volume is highest and quality requirements are lower), and a second frontier model like GPT-5.6 for final review against the original spec. Berman also references a Graphile study on how different AI coding tools cluster their error types, recommending it for teams shipping AI-generated code.


📺 Source: Matthew Berman · Published July 23, 2026
🏷️ Format: Workflow Case Study

1 Item

Channels

2 Items

Companies

1 Item

People