The short answer
Enterprises have become considerably better at putting AI into production. They have not become better at getting returns from it.
(cite index="46-1">The Fifth Annual Domino Enterprise AI Report found that the share of enterprises whose ROI fails to outpace their investment has held at 57% since 2025, even as production capability continues to climb. The survey of 639 senior enterprise AI leaders found 93% report improved production capability in 2026, up from 88% in 2025, making this a two-year plateau.</cite>
Two numbers moving in opposite directions is the finding. Production capability is a solved-enough problem. Value realization is not, and it has not improved in a year during which deployment capability improved measurably.
(cite index="46-1">Independent research from MIT NANDA reached a parallel conclusion, finding that 95% of organizations running AI pilots saw no measurable profit and loss impact, with only 5% of integrated pilots generating meaningful financial value.</cite>
The gap has a name in the Domino research: (cite index="50-1">a last-mile gap between AI models in production and the business users who need to unlock their value.</cite>
That framing is useful because it locates the problem. It is not a modeling problem and no longer primarily a deployment problem. It is a problem of whether AI reaches the workflow where a decision gets made.
Key takeaways
- 93% of enterprises report improved AI production capability in 2026; 57% report returns growing no faster than investment, unchanged since 2025.
- The plateau has now held for two consecutive years, which rules out the explanation that value simply lags deployment.
- (cite index="51-1">One in ten organizations have no direct access to AI insights for business users at all.</cite>
- (cite index="50-1">Agentic AI is scaling ahead of governance: 43% have agentic AI in governed production, while 41% are piloting or scaling it without the governance to manage it.</cite>
- The gap between regions is substantial, with European and UK respondents reporting worse ROI outcomes than US ones.
- Measuring success at deployment rather than at decision is the common thread.
What the data shows
| Finding | 2025 | 2026 |
|---|---|---|
| Report improved production capability | 88% | 93% |
| ROI growing no faster than investment | 57% | 57% |
(cite index="46-1">The survey covered 639 senior enterprise AI leaders at Director level and above, at organizations with revenues of $100 million or greater, across North America, the United Kingdom, and continental Europe, in financial services and insurance, life sciences, and public sector organizations. It was conducted independently by BARC Research on behalf of Domino Data Lab in April 2026.</cite>
The regional distribution is worth noting, because it complicates the simple reading. (cite index="52-1">51.1% of US respondents report AI costs outpacing their ROI, against 66.9% of UK respondents and 67% of European ones.</cite>
That spread suggests the constraint is not purely technical. If it were, the numbers would converge, since the models and platforms are the same everywhere. Regulatory environment, operating model, and how AI is integrated into workflows differ considerably more than the technology does.
The measurement problem underneath the numbers
(cite index="51-1">Domino's framing is direct: AI success is measured at the wrong point. Most organizations measure deployment and model performance, but real value only happens when AI drives decisions in business workflows.</cite>
That distinction produces a specific and common failure pattern.
An organization deploys a model. It performs well against its evaluation metrics. It goes into production successfully. The program reports the deployment as a milestone, which it is. What nobody measures is whether a decision anywhere in the business changed as a result.
(cite index="51-1">The report notes that AI is in production but not in use, with one in ten organizations having no direct access to AI insights for business users.</cite>
A model that no business user can reach cannot change a decision, regardless of how well it performs.
Last mile
Where the value is lost between build and outcome
| Stage | Usually measured? | Where value leaks |
|---|---|---|
| Model performance | Yes, thoroughly | Rarely the constraint |
| Deployment to production | Yes, celebrated | Rarely the constraint |
| Reaching business users | Sometimes | One in ten have no direct access |
| Fitting into the decision moment | Rarely | Output arrives outside the workflow where the decision happens |
| Changing behavior | Rarely | Users receive a tool, not a new way of working |
| Business outcome | Sometimes, late | The metric nobody baselined |
Measurement density is highest at the top of the table and the value is lost at the bottom.
The governance divergence
The 2026 data adds a second finding that will shape the next two years.
(cite index="50-1">Expanding agentic AI use ranks as the top organizational priority for enterprise AI leaders in 2026 at 38.5%, tied with upskilling business users, ahead of every other investment type including the governance infrastructure meant to manage agents. 43% of organizations have agentic AI running in governed production, while 41% are piloting or scaling it without the governance to manage it, with those actively scaling outnumbering those merely piloting by more than two to one.</cite>
Roughly as many organizations are running agentic AI without governance as with it, and the ungoverned group is weighted toward scaling rather than experimenting.
(cite index="50-1">Governance maturity is the clearest dividing line between the two groups: among organizations whose governance is fully keeping pace with their AI activity, 67.5% have agentic AI running in governed production.</cite>
This connects directly to the ROI plateau. Gartner's prediction that over 40% of agentic AI projects will be canceled by the end of 2027 cites escalating costs, unclear business value, and inadequate risk controls. The Domino data shows the third condition present in roughly two-fifths of organizations that are already scaling.
What the plateau is not
Three explanations get offered and none survives the two-year duration.
"Value lags deployment." Plausible for one year. The plateau has now held across two, during which production capability rose measurably. If the lag explanation held, the 2026 ROI number should have improved against 2025 deployment gains.
"The models are not good enough yet." Model capability improved substantially over the same period. If capability were the constraint, the two lines would move together.
"It is too early to measure." The organizations in the sample have revenues above $100 million and AI leaders at Director level and above. These are not first-year programs.
The consistent alternative explanation is that the value was never designed for. Deployment was the objective, the workflow change was not scoped, and the outcome metric had no baseline.
What the organizations closing the gap do differently
The data points toward a set of practices rather than a technology choice.
They define the decision the AI is meant to change. Not the task it performs, but the decision that will be made differently and by whom.
They measure the business metric, not the model metric. Cycle time, cost per case, rework rate, quality. Model accuracy is a prerequisite, not an outcome.
They design the workflow before the deployment. Where the output arrives, who reviews it, what happens next, and what happens when it is wrong.
They baseline before they build. Improvement cannot be demonstrated against a metric that did not exist beforehand.
They put governance in the delivery path rather than after it. The governed group is the group that is scaling agentic AI successfully.
They know how the work actually happens. This is the least discussed and arguably the most determinative. A use case selected without operational evidence tends to target a step that was not the constraint.
Where Horizon fits
Horizon is an AI-powered continuous discovery platform. Its relevance to this data is at the point where use cases are selected, which is upstream of everything the last-mile gap describes.
The pattern in the research is that value is lost between a model that works and a decision that changes. A substantial part of that loss originates earlier, when a team chose a use case without knowing where the friction actually sat, which step was the constraint, or what proportion of volume followed the path the AI was scoped for.
Horizon runs AI-led discovery conversations across an organization, structures the evidence against processes and systems, ranks opportunities by impact, and converts them into initiatives with owners and business cases. The output includes the baseline the ROI question later depends on.
In the published Mercado Libre case, Horizon ran discovery across 2,000 employees in Finance and adjacent functions in five countries in four days, against a manual baseline of 11 to 20 weeks. Despegar produced 45 prioritized initiatives for a CRM migration in four weeks, against a manual version that had taken 12 months. Across deployments, 640 initiatives have been prioritized with an average identified value above $100K each.
Those are results from specific engagements rather than a projection for any organization.
What to check in your own program
- For each AI use case in production, can you name the decision it was meant to change?
- Did that decision change?
- Was there a baseline for the business metric before deployment?
- Can business users reach the output directly, in the workflow where they decide?
- Is anyone accountable for the outcome, as distinct from the deployment?
- What proportion of your AI reporting is model metrics versus business metrics?
- If you are scaling agentic AI, is governance keeping pace, or behind?
- Was the use case selected from operational evidence or from a workshop?
- What would cause you to stop a use case that is running?
FAQ
What is the last-mile gap in enterprise AI?
It is the gap between AI models running in production and the business users who need the output to change a decision. Domino's Fifth Annual Enterprise AI Report identifies it as the reason production capability keeps improving while returns do not: models are deployed successfully but do not consistently reach the workflow where decisions happen.
What percentage of enterprises get positive ROI from AI?
In Domino's 2026 survey of 639 senior enterprise AI leaders, 57% reported that returns are growing at the same pace as investment or slower, unchanged from 2025. MIT NANDA research reached a parallel conclusion, finding 95% of organizations running AI pilots saw no measurable profit and loss impact, with 5% of integrated pilots generating meaningful financial value.
Why has enterprise AI ROI not improved despite better deployment?
The data suggests success is being measured at the wrong point. Organizations measure model performance and deployment milestones, both of which have improved, rather than whether a business decision changed. Additional factors include AI output that does not reach business users in their workflow and use cases selected without evidence about where operational friction actually sits.
How many enterprises are running agentic AI without governance?
In the 2026 Domino data, 43% have agentic AI running in governed production while 41% are piloting or scaling without the governance to manage it. Within that second group, organizations actively scaling outnumber those merely piloting by more than two to one.
Is the AI ROI problem worse in Europe than the US?
By this data, yes. 51.1% of US respondents reported AI costs outpacing ROI, against 66.9% in the UK and 67% in continental Europe. The gap suggests the constraint is not purely technical, since the underlying models and platforms are broadly the same across regions.
What should enterprises measure instead of deployment?
The business metric the use case was meant to move, baselined before the build: cycle time, cost per case, rework rate, quality, or risk exposure. Alongside that, whether business users can reach the output in the workflow where the decision is made, and whether behavior actually changed after rollout.
The milestone moved
Getting a model into production used to be the achievement. The 2026 data says it no longer distinguishes anyone, since 93% report they can do it.
What distinguishes organizations now is whether anything downstream changed, and that is decided long before deployment, at the moment someone chose which problem the AI was going to solve.
See it. Fix it. Lead it.