Stop Overpaying for Intelligence | DevDay 2026

Stop Overpaying for Intelligence | DevDay 2026

More

Summary

Two members of OpenAI’s AI deployment engineering team, Mandeep and Sapto, explain how developers can cut AI costs without losing quality. Their central point is to measure cost per task rather than cost per token. A cheaper model can end up costing more if it uses more tokens or fails and needs a human to step in, as Mandeep illustrates with a frustrating airline chatbot experience.

They cite customer examples. Perplexity reported that GPT-6 Astra produced 9% more accurate results at half the cost on its research benchmark, despite higher per-token pricing. Notion reported better accuracy at half the cost per task after migrating from 5.5 to 5.6. For right-sizing a model, they advise defining the task and the expected accuracy first, then finding the model, harness, and configuration that meet it.

The panel then covers four cost levers beyond model choice: prompt caching for repeated context, programmatic tool calling so a model runs a small program instead of reading every intermediate result, reasoning effort settings, and batch or flex API processing. A legal software company, Clio, is cited as a real-world example of programmatic tool calling.


📺 Source: OpenAI · Published October 07, 2026
🏷️ Format: Deep Dive

1 Item

Channels

1 Item

Companies