Summary
Kyle Jaejun Lee, a software engineer at KRAFTON, describes what it actually takes to run a fleet of AI coding agents across three machines — a MacBook and two headless Linux boxes — as a daily production workflow. The talk, presented at AI Engineer, is structured around five concrete failure modes encountered when scaling beyond a single machine, and the specific engineering decisions that resolved each one.
Lee’s central insight is that flat agent setups quickly turn the human operator into the bottleneck: simultaneously acting as scheduler, memory, and reviewer across too many live contexts. His solution is a strict hierarchy — CEO, VP, manager, and worker agents — where each layer sees only its own scoped context and results flow upward to a single review inbox. State lives in files on disk rather than inside model context windows, which means agents survive context resets and machine crashes alike.
The five failures Lee documents include orchestrators doing work instead of delegating (fixed with CLI harnesses that make delegation the only available path), tmux pane overflow from runaway worker spawning, out-of-memory crashes from stacked Claude Code and MCP processes, credential collision between workspaces, and laptop power loss killing in-flight jobs. Cross-machine state transfer is handled via Git commits pushed and pulled over SSH with tmux send-keys. A web-based review gateway now serves as the single approval control point, and the system can cold-boot back to full operation with a single command.
📺 Source: AI Engineer · Published July 08, 2026
🏷️ Format: Workflow Case Study







