Summary
Dan Bjornn, senior data scientist at Lease End — a company connecting auto-lease holders with financing options — delivers a candid post-mortem on his team’s LLM fine-tuning journey. Built in late 2024, their customer messaging system used a fine-tuned model to classify inbound texts into six intent categories (call now, schedule later, opt out, and others), processing thousands of messages daily in real time. The system generated $12 million in revenue at a 50x ROI — but quietly accumulated technical debt the metrics didn’t surface.
The talk details two failure modes that reached production: the “confused confirmer,” where the model triggered an immediate call after a scheduling confirmation message, and the “overeager puppy,” where a customer saying “good morning” caused the system to call them instantly. Every fix risked regressions elsewhere, turning maintenance into a week-long whack-a-mole cycle of data gathering, LLM-as-judge labeling, manual validation, fine-tuning, and regression checking.
Bjornn introduces the concept of the “calcification tax”: the longer a fine-tuned model runs in production, the more rigid and fragile the surrounding system becomes. Teams get locked into specific model versions, lose cross-provider flexibility, and face triage decisions about which bugs are painful enough to justify a full retrain. The talk ultimately makes the case for prompt engineering and structured outputs as a more maintainable alternative to fine-tuning for narrowly scoped classification tasks — particularly when the cost of a wrong classification directly impacts customer experience.
📺 Source: AI Engineer · Published August 20, 2026
🏷️ Format: Workflow Case Study







