The cheapest token does not always make for the cheapest task

OpenAI advises measuring useful work per dollar rather than the price per token. A method to evaluate the actual cost-effectiveness of agents.

OpenAI is releasing a guide for enterprise leaders on managing AI spend, at a time when usage is shifting from simple conversation to agents that execute hundreds of steps. Its central argument shifts the focus: the price per token is a poor decision metric.

The reasoning is straightforward. A cheap model may fail, retry, or produce work that needs to be redone; a more capable model costs more per token but sometimes achieves an acceptable result faster, in fewer attempts, and with less review. The right measure, according to the publisher, is therefore not the unit cost but the useful work per dollar: completed tasks, time saved, and improved decisions.

From this comes a way to frame the investment. Treat AI as a portfolio, with broad access for everyday productivity, specialized workflows for repetitive work, and a small number of strategic bets backed by the company's own context. Funding follows maturity: exploration tests whether the model can handle the task, validation proves it on representative cases against a defined quality threshold, and production covers integrations, controls, and change management. Common building blocks—identity, trusted connectors, evaluations, observability, routing between models—deserve to be funded centrally, so that each new workflow becomes simpler and safer to launch.