Tokens tell CIOs what an AI agent consumed. They do not necessarily tell them what it cost to get useful work done.
As enterprises put agents to work across multistep processes, the total can also include tool calls, retries, failed attempts and human intervention. That pushes CIOs toward a harder question: What should they actually measure?
Most enterprises are still measuring token consumption and model calls, according to Will Sommer, senior director analyst at Gartner. But those consumption metrics do not necessarily capture everything an agent does before completing a task -- or whether the work ultimately produces a successful result.
A more useful way to measure cost may be a completed, quality-adjusted outcome rather than a token or model call.
Agentic systems can also be priced in different ways, including by tokens, compute usage or outcomes, according to Sommer. That can make it difficult for CIOs to compare competing systems based on price alone.
What counts as a successful outcome depends on the work an agent is performing, Sommer said.
"Some outcomes will be obvious, for example a customer service case satisfactorily closed," Sommer explained. "Others will be more complex. Our recommendation where possible is to identify an observable, countable, quality-adjusted unit that represents completed, valuable work."
Sommer said outcomes can range from the task level to the business level. As the outcome becomes more complex, the measure of success may need to become broader. An agent designed to prevent fraudulent insurance claims, for example, could be measured based on total costs avoided relative to a pre-established baseline.
Define the finish line first
The first step is defining when the work is actually finished.
"It's important to establish the finish line first, and then measure cost, quality and employee involvement against that standard," said Chris Bennett, vice president of the Global AI Practice at Unisys. "Without this clarity, two teams might use the same metric to describe very different types of work."
For example, an enterprise needs to decide whether an invoice counts as processed when an agent extracts the information or only after the information has been checked and entered correctly, Bennett said. Similarly, a support request could be considered complete when an agent drafts a response or only when the customer's issue has been resolved.
Defining that finish line is not necessarily an IT decision. Business leaders remain accountable for the outcomes and performance of agents operating within their processes, Sommer said, while CIOs and AI leaders are responsible for establishing the infrastructure, standards and controls needed to measure those outcomes.
What tokens don't capture
Defining the outcome is only part of the calculation. Enterprises also need to account for everything the agent does before it gets there.
The cost of the relevant outcome must include all of the failed attempts to achieve that outcome.
Will Sommersenior director analyst, Gartner
"The cost of the relevant outcome must include all of the failed attempts to achieve that outcome," Sommer said.
Bennett said enterprises should track each attempt and the time spent completing the task: "If an agent produces a usable answer on the fourth try, with assistance from an employee for checking and correction, then that reflects the true cost of the result."
Comparing those costs with the old process, including how much time or effort the agent actually saves, can help an enterprise assess whether the new process is worthwhile.
"If people are still doing most of the work behind the scenes, a successful final output doesn't provide much insight into the overall economics," Bennett said.
Those costs can become harder to track as agents make autonomous decisions about what to do next.
"Cost comes from how many downstream tool calls a request triggers, not the size of the prompt that started it," said Jyotika Singh, security researcher at Forcepoint X-Labs. Prompt size still affects token charges. The downstream activity Singh described can create additional costs that are harder to see from the initiating prompt alone.
Singh said enterprises can also miss internal reasoning tokens that don't appear in the visible prompt or output, as well as context repeatedly processed during a long-running session.
"The costs most likely to be missed are the ones generated by the agent's own follow-on decisions, not the ones tied to what a user typed," Singh said.
Measure agents on equal terms
Uber offers one example of measuring agent costs against the work being completed.
The company evaluates models used in its agentic software development workloads based in part on cost per completed task, along with output quality and model reliability, according to an August Uber Engineering post.
For its AI code review agent, the company measures cost per review alongside quality measures including precision and recall. Uber said it establishes target outcome metrics and evaluation benchmarks for each new managed agent.
Many organizations still lack the AI FinOps systems needed to measure cost and value per outcome effectively, Sommer said. Comparing the economics of competing agentic systems presents another challenge for CIOs.
Vendors may charge based on tokens, compute usage, completed outcomes or other measures. Token counts can also be difficult to compare across providers because similar levels of consumption may produce different volumes and quality of output.
Rather than comparing vendor pricing units alone, Sommer said CIOs need to standardize comparisons around defined tasks, workflows and business outcomes.
How to compare AI agents on equal terms
Pricing units vary across vendors, so CIOs need a common benchmark tied to completed, valuable work.
• Define the task and what counts as a successful result.
• Test each system on the same real work and data.
• Track how many attempts and tool calls the task requires.
• Include employee review, correction and intervention.
• Compare total cost alongside quality and reliability.
Bennett recommended a similar approach. Enterprises comparing agentic systems should select a set of real tasks and test each system using the same data and definition of a successful result, he said. They can then compare how often each system produced the right result, how many attempts it required and how much employee review was necessary.
"That takes more work than comparing pricing sheets, but it gives organizations a much better sense of what they are paying for," Bennett said.
Liz Hughes is an award-winning editor and writer covering AI and emerging technology and the former editor of AI Business and IoT World Today.