Summary
Sharbel A. breaks down the hidden mechanics behind Claude Code’s token consumption and delivers seven concrete optimizations to reduce session costs. The foundational insight is that LLMs have no persistent memory — every new message re-sends the entire conversation history, causing costs to compound exponentially. A 3,000-token file read at turn four of a 40-turn session effectively costs 3,000 tokens 37 additional times, making long threads far more expensive than they appear.
The most impactful fix is free: use /clear between unrelated tasks. With 96% of one user’s spend coming from rereading history, resetting the context base beats any other single optimization. A subtler and costlier trap is switching models mid-session to save money — because the model is part of the cache key, dropping from Opus to Sonnet on a 200,000-token context can turn a 10-cent turn into a one-dollar turn. The same cache invalidation applies to changing effort levels, toggling fast mode, connecting or disconnecting MCP servers, and even upgrading Claude Code before resuming a long session.
The video also covers MCP tool overhead in detail — GitHub alone loads 26,000 tokens per session, Slack 21,000 — and explains Anthropic’s newer deferred tool-loading behavior, which reduces that overhead by 85% automatically. Sub-agent delegation is addressed honestly: it moves token costs rather than eliminating them, with Anthropic’s own numbers showing sub-agents burning roughly 9,800 tokens to save 5,700 in the main context. A free audit prompt is included that users can paste into Claude Code to identify their personal biggest token drains before applying any of the fixes.
📺 Source: Sharbel A. · Published August 19, 2026
🏷️ Format: Tutorial Demo







