Individual adoption of AI inside enterprises is close to universal. Organizational performance improvement is close to rare.
Glean's Work AI Index found that 87% of digital workers already use AI in their work, while only 13% of organizations perform significantly better as a result. That is a spread of 74 points between the input and the outcome.
The instinct on seeing that number is to question the tooling or the training. Both are usually fine. The gap opens somewhere else, and the location is consistent enough to describe.
Individual AI usage produces individual time savings. Those savings accrue to the person, in minutes distributed across a day, on tasks that were not the constraint on the work. Nothing collects them, nothing redirects them, and the process that determines the team's throughput is unchanged.
Performance improves when a workflow changes. Usage alone does not change workflows.
Key takeaways
- 87% of digital workers use AI; 13% of organizations perform significantly better because of it.
- Individual time savings distribute across a day in fragments that no process captures.
- The constraint on most enterprise work is a handoff, an approval or a dependency, none of which individual AI usage touches.
- Organizations measure usage because usage is easy to measure, which reinforces the gap.
- Closing it requires knowing where the process actually loses time, which most organizations cannot state.
- The organizations that close it change a workflow rather than distributing a tool.
Why individual savings do not aggregate
The savings are fragmented
An employee saves eight minutes drafting an email, four minutes summarizing a document, twelve minutes writing a first version of an analysis. Twenty-four minutes across a day.
Those minutes do not reappear anywhere measurable. They are absorbed into the same day, spent on the same queue, and the case that was waiting for an approval is still waiting.
Aggregating fragmented individual savings into organizational throughput requires a mechanism, and almost no organization has built one.
The saved task was rarely the bottleneck
This is the structural reason and it holds across functions.
Enterprise work is a sequence of steps performed by different people and systems. Throughput is governed by the slowest link, which in most processes is a wait state: an approval pending, an input missing, a dependency unresolved, a decision unmade.
Individual AI usage accelerates the steps where a person is actively working, which is the touch time. It does not touch the wait time. In processes where wait time dominates, the accelerated steps simply arrive at the queue sooner.
Nobody reclaims the capacity
Even where savings are real and concentrated, they need to be reclaimed deliberately. Capacity that is not reassigned gets absorbed, and the absorption is invisible because nothing was measuring the time in the first place.
Organizations that reclaim capacity do so by naming what the freed hours are for before the tool is deployed. That is a management decision rather than a technology one.
What the broader data shows
The pattern is not specific to one study.
Domino Data Lab's Fifth Annual Enterprise AI Report, surveying 639 senior enterprise AI leaders, found that the share of organizations whose returns fail to outpace investment has held at 57% since 2025, while 93% report improved production capability, up from 88%. Production improved and returns did not.
MIT NANDA research found that 95% of organizations running AI pilots saw no measurable profit and loss impact, with only 5% of integrated pilots generating meaningful financial value.
Three datasets, three methodologies, one direction. The input side of enterprise AI is working. The output side is not, and it has not moved in a year during which the input side improved measurably.
Gap anatomy
Where the value stops
| Stage | Status in most enterprises | Where value leaks |
|---|---|---|
| Access | Solved, 87% usage | Rarely the constraint |
| Individual task speed | Improving | Savings fragment across the day |
| Workflow change | Rarely attempted | The step that was accelerated was not the constraint |
| Capacity reclaim | Rarely planned | Freed hours absorbed without measurement |
| Process throughput | Unchanged | Wait time untouched |
| Business outcome | 13% significantly better | The metric nobody baselined |
Measurement density is highest in the top two rows. The value is lost in the middle three.
Why organizations keep measuring usage
Usage metrics are available. Licenses assigned, active users, repeat usage, prompt frequency. They come out of the platform without any additional work.
Workflow metrics require knowing what the workflow was before, which requires having measured it, which most organizations have not done for the processes AI is being applied to.
That produces a reporting asymmetry with a predictable consequence. Programs report the numbers they have, those numbers improve, and the question of whether anything downstream changed does not get asked because nobody can answer it.
A useful diagnostic: look at the last AI program review and count how many metrics described the tool versus how many described the operation.
What closing the gap requires
Four conditions, and the first is the one that blocks the rest.
Know where the process loses time
The distinction that matters is between touch time and wait time. Accelerating touch time in a process governed by wait time produces the 87/13 pattern reliably.
Establishing that split requires knowing what the people running the process actually spend their time on, which is not in any system because waiting, chasing and reconciling produce no transaction records of their own.
Change the workflow, not the task
A tool distributed to individuals changes tasks. Performance changes when a step is removed, a handoff is restructured, an approval is reassigned, or a queue is eliminated. Those are process decisions with named owners.
Plan the capacity reclaim in advance
Decide what the freed hours are for before deployment. Capacity that is not assigned gets absorbed, and the program then cannot demonstrate what it delivered.
Baseline before you build
Improvement cannot be demonstrated against a metric that did not exist beforehand. This is the condition most frequently skipped and the one that makes every subsequent ROI conversation unresolvable.
Where Horizon fits
Horizon is an AI-powered continuous discovery platform. Its relevance to this data is at the first condition, which is knowing where the process actually loses time.
Discovery Cycles run AI-led interviews across the roles that operate a process, asynchronously and at census scale. The Insights Dashboard quantifies the time impact per area and ranks findings by impact and effort, with traceability to the input behind each one. The Initiatives Dashboard converts priorities into business cases with owners attached, which is the artifact a workflow change requires.
Grupo HZ, a multi-company industrial holding operating across Argentina, Brazil and Chile with a shared services centre covering Administration, Finance, Procurement and Treasury, illustrates the gap at close range.
Before the engagement, spreadsheets, email and manual validations were the backbone of critical processes, with no integrated workflows and no visibility into where time was being lost or which processes had the highest automation potential. Any individual in that organization could have used AI to draft faster. The throughput constraint sat in manual validation between systems, which no individual tool reaches.
Horizon ran 55 structured interviews across six departments and quantified the time impact per area: Procurement at 48.3 hours per week, Accounts Payable at 34, Treasury at 21.55, Credit and Collections at 10.5, Sales and Commercial in Chile at 6.5, and Planning and Costing at 2.9. Roughly 123 hours per week, with 50 to 90% automation potential identified and specific solution types mapped to specific processes.
The engagement produced 33 actionable findings, saved 159 discovery hours against the manual approach, reallocated three FTEs to strategic work without adding headcount, and returned 182% ROI. The stated shift was from a framing of needing more people to a framing of needing better processes.
The relevant detail for this analysis is that three FTEs were reclaimed. That is capacity redirected rather than absorbed, and it happened because the hours had been located and quantified first.
That is one engagement under specific conditions rather than a projection for any organization.
What to check in your own program
- Can you state the split between touch time and wait time in the processes AI is being applied to?
- What proportion of your AI reporting describes the tool versus the operation?
- For each deployment, can you name the workflow step that changed?
- Was a baseline captured for the business metric before deployment?
- Was the capacity reclaim planned, and did anyone receive the freed hours?
- Do you know which of your processes are governed by wait time?
- If usage dropped to zero tomorrow, which business metric would move?
Question 7 is blunt and informative. If the answer is none, the program is delivering individual convenience rather than organizational performance, which may still be worth something and should be described accurately.
FAQ
What is the AI productivity gap?
The distance between individual AI usage inside organizations and measurable organizational performance improvement. Glean's Work AI Index found 87% of digital workers using AI while only 13% of organizations perform significantly better as a result.
Why does AI usage not improve company performance?
Because individual usage accelerates tasks rather than changing workflows. The savings fragment across a day, and the accelerated steps are frequently not the constraint on throughput. In processes where most elapsed time is spent waiting for approvals, inputs or dependencies, faster individual work arrives at the same queue sooner.
What percentage of companies get value from AI?
The figures vary by study and point the same direction. Glean's index puts significantly better organizational performance at 13%. Domino Data Lab's 2026 survey of 639 enterprise AI leaders found 57% reporting returns growing no faster than investment, unchanged since 2025. MIT NANDA research found 95% of pilots produced no measurable profit and loss impact.
How do you measure AI value beyond usage?
Measure the business metric the deployment was meant to move, baselined before the build: cycle time, cost per case, rework rate, quality or risk exposure. Alongside that, track whether a specific workflow step changed and whether freed capacity was reassigned to something named.
What is the difference between touch time and wait time?
Touch time is time someone spends actively working on a case. Wait time is time the case spends in a queue waiting for an approval, an input or a dependency. Individual AI tools reduce touch time. In processes governed by wait time, that produces no change in throughput.
Usage was never the goal
Reaching 87% adoption was a real achievement and it is now table stakes. It no longer distinguishes anyone, which means it can no longer be the headline of an AI program review.
What distinguishes organizations is whether a workflow changed, and that is decided by knowing where the process loses time before any tool is distributed.
See it. Fix it. Win it.