Back to the journal
Industry / Analysis

Accenture’s AI Spending Report Exposes a Measurement Gap

A new Accenture report finds companies struggling to link AI spending to financial results. Cutting the bill alone can reward the wrong outcome.

A balance scale compares an AI processor and receipt paper with a stack of completed task cards.

Accenture says less than a fifth of enterprise spending on AI tokens can be linked to a quantified financial outcome. Its report, published September 10, draws on a survey of 750 senior executives across 17 countries. Tokens are the units of data a model processes, and a common basis for charging for its use. A token bill tells a company what it consumed. It does not, by itself, tell the company what it gained.

The finding deserves attention without being mistaken for a verdict that the remaining spending was wasted. Missing evidence of value and evidence of no value are different things. My view is that finance teams should make useful completed work the center of the spending argument. Otherwise, a campaign to cut the most visible bill can reward a system that costs less to run while accomplishing less for the business.

What the survey can establish

The sample covers enterprises with annual revenue above $1 billion. The research includes 15 interviews conducted in June and July and a separate simulation, which should not be mistaken for additional respondents. The headline measure combines two survey answers: how much consumption connects to an outcome, and how much of that can be expressed financially. It is an estimate built from executives’ answers, not an independent transaction audit. The report cannot tell a particular business how much spending to cut.

There is already a useful vocabulary for this distinction. The FinOps Foundation’s guidance on technology finances separates the cost of consuming a resource from the cost of delivering a business result. A token is a resource; a resolved customer case is a result. Its guidance also recognizes that revenue can be difficult to attribute and that service levels or risk reduction may be useful measures. Counting dollars returned immediately is not the only disciplined way to judge an investment.

A smaller bill can buy worse results

Consider an invented comparison of two AI workflows handling comparable tasks over the same period. The first spends $100 on model use and produces 80 results that meet an agreed quality standard: $1.25 per accepted result. The second spends $60 but produces only 30 acceptable results: $2 each. Its bill is smaller, but each useful result is more expensive. These figures illustrate the arithmetic; they are not observed company results or quoted model prices.

Change the assumption and the conclusion changes. If the second workflow still delivers 80 acceptable results for $60, its model cost falls to $0.75 each. That would be an improvement on this measure. The point is not to defend expensive models. It is to require evidence about what survives a switch to a cheaper one. Neither calculation includes staff time or infrastructure, so neither establishes total profitability.

Even the word accepted needs an owner. Imagine a support system that counts a case as finished when it sends a response, while the customer has to contact the company again to get the problem fixed. The apparent productivity gain depends on where the company stops counting. In this hypothetical, measuring repeated contacts alongside resolution would expose work that a fast response counter misses. A team should not be able to improve its reported economics merely by declaring the task finished earlier.

Who gets charged, and who benefits?

Accenture recommends charging consuming teams for their AI use. That creates a further question: does the team receiving the bill also receive the benefit? Suppose a service team funds a better answer that saves an operations team from handling a later escalation. A budget rule focused solely on the service team’s spending could discourage the improvement. The relevant comparison would follow the whole process. This is a possible incentive conflict, not an observed finding from the survey.

The strongest objection to demanding measured outcomes is that exploration has value before a repeatable process exists. An experiment might teach a team which approach fails, improve an uncertain design, or make a later project possible. Insisting on immediate financial returns could eliminate useful learning. I would reserve a bounded learning budget, with a question to answer and a date to assess what was learned. That is more honest than assigning imaginary savings to justify every experiment.

For recurring production work, the standard should become firmer. A cheaper system that preserves useful output deserves credit. A more expensive one can also earn its place if the additional result warrants the cost. Evidence of reliable quality, fewer repeat attempts and lower total process costs would support either choice. The management failure is letting the invoice stand in for that comparison. An AI budget should buy work the organization values, and its accounting should be able to explain what that work was.

Explore More Stories ↗