AI Spend and AI Return: Why One Is Easy to Measure and One Is Not

An analysis for enterprise leaders of the measurement asymmetry in AI investment: why spend is precise and return is contested, what the 2026 data shows, and which decisions have to be made before deployment for the question to be answerable at all.

October 27, 202610 min read
ai roi measurement enterprisehow to measure ai roiai investment return

Almost every enterprise can state its AI spend to the invoice. Licences, compute, platform fees, integration work, headcount allocated to the programme. The number is precise, current and uncontested.

Far fewer can state the return with the same confidence. The answers tend to arrive as estimates, as hours saved per person extrapolated from a survey, or as a business metric that moved during a period when several other things also moved.

The asymmetry is structural rather than a failure of rigour. Spend is a transaction, recorded automatically by the systems that process it. Return is a counterfactual: what would have happened without this. Counterfactuals require a baseline, and baselines have to be captured before the intervention.

That timing is the whole problem. By the time anyone asks what the return was, the moment to make the question answerable has passed.

Key takeaways

What the data shows

Three independent sources point the same direction.

Domino Data Lab's Fifth Annual Enterprise AI Report, surveying 639 senior enterprise AI leaders at director level and above in organizations with revenues above $100 million, found that the share reporting returns growing no faster than investment has held at 57% since 2025. Over the same period, the share reporting improved production capability rose from 88% to 93%.

The regional distribution is wide. 51.1% of US respondents reported AI costs outpacing return, against 66.9% in the UK and 67% in continental Europe.

MIT NANDA research found that 95% of organizations running AI pilots saw no measurable profit and loss impact, with 5% of integrated pilots generating meaningful financial value.

Glean's Work AI Index found 87% of digital workers using AI while 13% of organizations perform significantly better as a result.

Three methodologies, one pattern. The input side improved measurably over the period. The return side did not move.

Why return is hard to measure

The baseline was never captured

This is the dominant cause and it is entirely preventable.

Demonstrating that a deployment improved something requires knowing the value of the thing beforehand. Most organizations deploy against processes they have never measured at the level the deployment affects. Cycle time may be known at an aggregate level. Touch time per case, rework rate, exception handling effort and cost per transaction usually are not.

Reconstructing a baseline afterward is possible and always disputable, because the reconstruction is performed by the people with an interest in the result.

Individual savings do not aggregate

The most common return measure is hours saved per person, usually derived from a survey or an extrapolation.

Those hours are real and they distribute across a day in fragments. Nothing collects them, and the process that determines team throughput is unchanged unless someone deliberately reclaims and redirects the capacity.

A programme reporting thousands of hours saved with no corresponding change in headcount, cycle time or output has measured an input rather than a return.

Attribution is genuinely difficult

Several things change at once. A process is redesigned, a system is upgraded, a team reorganizes, volume shifts, and an AI capability is deployed. Isolating the contribution of one is a real analytical problem.

It is worth noting that attribution is the smaller difficulty. A programme with a clean baseline and a contested attribution can at least argue about the size of the effect. A programme with no baseline cannot establish that an effect occurred.

The measured metric describes the tool

Latency, usage, adoption rate, output quality scores. These describe how the system performs. They are necessary and they are not returns.

The asymmetry is self-reinforcing: tool metrics are available without additional work, business metrics require measurement that was not done, so programmes report what they have.

Spend and return

Why one side is precise

SpendReturn
What it isA transactionA counterfactual
Recorded byFinance systems, automaticallyNothing, unless designed
Available whenImmediatelyOnly if a baseline existed beforehand
Contested?RarelyUsually
Improves withNothing requiredA measurement decision made before deployment

The two columns are asymmetric by construction. Closing the gap is a decision made at scoping rather than an analysis performed afterward.

What has to be decided before deployment

Five decisions. All of them are cheap before and expensive or impossible after.

Name the business metric. Not the task the system performs, the outcome that should change. Cycle time on a defined queue, cost per case, rework rate, first-time resolution, error rate. One primary metric, with at most two supporting ones.

Capture its current value. Including its variance over several cycles, since a single point cannot distinguish an effect from normal fluctuation.

Name the decision that will change. Who will do something differently, and what. A deployment that changes no decision cannot move a business metric regardless of how well it performs.

Decide the capacity reclaim. If the expected return is time, name what the freed hours are for before they are freed. Capacity that is not assigned gets absorbed and the programme cannot demonstrate what it delivered.

Define the stop condition. What result would cause the deployment to be reversed or reduced. A programme with no stop condition has no measurement that matters, since no outcome changes the decision.

The reclaim problem

Worth separating because it is the most common reason a real saving produces no measurable return.

A deployment saves four hours per week across a team of twenty. That is eighty hours, which is a meaningful number.

If nobody decided in advance what those hours are for, they distribute back into the same work. The queue clears slightly faster, people spend marginally longer on each case, and the eighty hours appear nowhere. No headcount changes, no throughput metric moves, no cost line falls.

The saving was real and the return was zero, because time is only a return when it is converted into something else: capacity redirected to a named activity, headcount not added that would otherwise have been, or throughput increased against a constrained resource.

Naming the conversion before deployment is a management decision that takes an hour and determines whether the programme can report anything.

A workable measurement structure

LayerWhat it answersSufficient alone?
SpendWhat did this costYes, available automatically
UsageAre people using itNo, measures access
Workflow changeDid a step change, and which oneNo, measures activity
Business metricDid the outcome move against baselineYes, if the baseline exists
Capacity conversionWhere did the freed time goRequired when the return is time
AttributionHow much of the movement is thisImproves confidence, not required to establish an effect

Programmes that can answer the ROI question have the fourth and fifth rows. Programmes that cannot usually have the first three and a reconstruction argument.

Where Horizon fits

Horizon is an AI-powered continuous discovery platform. Its relevance to this problem is that it produces the baseline as a byproduct of establishing what should change.

Discovery Cycles quantify where effort actually sits in a process, per area and per step, from the people doing the work. The Insights Dashboard ranks findings by impact and effort with traceability to source. The Initiatives Dashboard converts priorities into business cases with expected impact and owners attached, which means the metric and its current value are named before anything is built.

Despegar, the leading travel technology company in Latin America with operations in more than 20 countries and over 3,900 employees, ran discovery ahead of a CRM migration under exactly these conditions. Commercial operations were manual and poorly integrated, information was dispersed across tools, teams duplicated tasks and decisions were delayed by inconsistent data. None of that had been quantified.

Horizon gathered input from more than 170 collaborators across Media Sales and HR, analyzing over 124 hours of conversations. In under four weeks it produced insights across campaign activation, reporting, billing and collections, and a dashboard of 45 initiatives prioritized by impact and effort, with recommendations to guide the migration and define roles and standards across Product, Commercial, Finance and HR. Average satisfaction was 9.1 out of 10, against a manual comparison of 12 months.

The relevant property for this analysis is the form of the output. Initiatives arrived prioritized by impact and effort, which means each carried an expected effect and a scope before any of them was funded. That is the artifact a return measurement needs and the one most programmes lack.

As one participant described it, interviewing 200 people in two weeks would have been impossible manually, and the findings would not have come from a form.

That is one engagement under specific conditions rather than a projection for any organization.

Measurement checklist

Before funding any AI deployment:

  1. Can you name one primary business metric this should move?
  2. Do you know its current value and its variance over several cycles?
  3. Can you name the decision that will be made differently, and by whom?
  4. If the expected return is time, have you named what the freed hours are for?
  5. Is there a named owner accountable for the outcome, distinct from the build?
  6. What result would cause you to stop or reverse this?
  7. What else is changing in the same period that could move the same metric?
  8. Six months from now, what evidence would settle whether this worked?

Question 8 is the shortest version of the whole list. If the answer requires a reconstruction, the baseline work was not done.

FAQ

How do you measure ROI on enterprise AI?

Name one primary business metric before deployment, capture its current value and variance, name the decision that will change, and define what happens to any freed capacity. After deployment, compare against the baseline and account for other changes in the same period. Without a baseline captured beforehand, the measurement becomes a reconstruction, which is disputable by design.

What percentage of enterprises see positive AI ROI?

The 2026 Domino Data Lab survey of 639 senior enterprise AI leaders found 57% reporting returns growing at the same pace as investment or slower, unchanged from 2025. MIT NANDA research found 95% of organizations running pilots saw no measurable profit and loss impact. The figures vary by methodology and point consistently in the same direction.

Why is AI spend easy to measure and AI return hard?

Spend is a transaction recorded automatically by finance systems. Return is a counterfactual, meaning what would have happened otherwise, which requires a baseline captured before the intervention. Nothing captures that baseline unless someone decides to, and the decision has to be made at scoping.

Are hours saved a valid measure of AI return?

Only when converted. Individual hours saved distribute across a day in fragments and disappear unless someone reclaims and redirects them. A programme reporting thousands of hours saved with no change in headcount, throughput or cycle time has measured an input. Naming the conversion before deployment is what makes the hours countable.

How do you handle attribution when several things change at once?

Record what else changed in the same period and its plausible effect, and where possible stagger deployments across comparable units so one serves as a reference. Attribution is a genuine difficulty and a smaller one than the baseline problem: a contested attribution still establishes that an effect occurred, while a missing baseline does not.

Decide the measurement before the deployment

The ROI question is asked after the money is spent and answerable only if someone prepared for it beforehand.

Naming one metric, capturing its current value and deciding what happens to freed capacity takes a few hours at scoping. Skipping it makes every subsequent conversation about value an argument rather than a calculation.

See it. Fix it. Lead it.

See Horizon in action.

Ready to transform?

See Horizon in Action

Discover how AI-powered organizational discovery can uncover hidden opportunities in days, not months.

Get Started

Related Resources