There is a pattern that shows up inside most mature automation estates and almost never inside the business case that created them.
A tool is deployed to catch errors, reconcile records or complete a step automatically. It works, within a condition: the configuration around it has to stay current. Reference data has to match, thresholds have to reflect the current policy, the interface it reads has to keep its shape.
Keeping that condition true falls to a person. So the arrangement inverts: for the tool to catch mistakes, someone first has to make sure the tool itself is correct.
When that upkeep slips, the failure is usually quiet. The automation runs, produces nothing alarming, and the problem surfaces downstream days later when a number does not reconcile.
Tool evaluations measure what a system does when it is working correctly. The variable that determines its real cost is what it requires to stay correct, and who carries that. That person rarely appears anywhere in the approval.
Key takeaways
- Maintenance load is part of the cost of an automation and it is usually estimated at zero.
- Four categories carry it: configuration upkeep, interface drift, exception absorption and verification.
- The person carrying it is frequently not the person who approved the deployment.
- Silent failure is the expensive mode, because the automation keeps running and nothing signals the degradation.
- Maintenance load scales with volume and with the number of systems the automation touches.
- Estimating it before deployment changes which candidates look attractive.
The four categories of maintenance load
1. Configuration upkeep
Reference data, thresholds, mappings, routing rules, exception lists. Each of these was correct at deployment and each one has a half-life.
A supplier is added and the mapping does not include it. A policy threshold changes and the rule still carries the old value. A cost centre is renamed. None of these breaks the automation. They cause it to apply yesterday's rules to today's work.
Someone has to notice and update, and in most deployments that responsibility is assumed rather than assigned.
2. Interface drift
Automations bind to the shape of what they read: a screen layout, a file format, an API response, a report structure. When the source changes, the binding fails.
The loud version of this is an immediate error, which is manageable. The quiet version is a field that moved and now reads a different value, which parses correctly and means something else.
3. Exception absorption
The automation handles the cases it was scoped for. Everything else routes to a person.
If the exception rate was understated at scoping, which is common, the absorbed volume is higher than planned. The team ends up operating the automation's overflow as a standing part of their week, and that work is attributed to the process rather than to the automation.
4. Verification
Someone checks that the automation is doing what it should, either by sampling output or by reconciling results against a second source.
This is the most invisible category, because it looks like ordinary diligence. It exists only because the automation is there, and it continues for as long as the automation does.
Why the business case omits it
Three reasons, and none involves anyone acting carelessly.
The savings are specific and the costs are diffuse. A business case can state that an automation removes ninety minutes per case across a known volume. The maintenance load is spread across several people, measured in interruptions rather than in blocks, and nobody has a figure for it.
The person who approves is not the person who maintains. The approval sits with a function lead or a platform owner. The upkeep lands on an analyst who was not in the meeting and whose time was not costed.
The load appears after the evaluation window. Configuration drift, interface changes and exception growth accumulate over quarters. A pilot measured over six weeks captures almost none of it, which makes early results systematically optimistic.
Why silent failure is the expensive mode
An automation that stops working loudly gets fixed. Someone notices, raises it, and the cost is bounded by the time to repair.
An automation that keeps running on stale configuration produces output that looks normal. There is no error, no alert and no gap in the record. The degradation is detected downstream, by a person noticing a number that does not reconcile, which can be days or weeks later.
Two consequences follow. The correction cost is larger, because the affected period has to be identified and reprocessed. And trust erodes in a way that creates permanent work: once a team has been surprised once, they add a verification step, and that step stays.
That is how an automation intended to remove a manual check ends up with a manual check in front of it.
Maintenance surface
Where the ongoing load comes from
| Category | Trigger | Failure mode | Scales with |
|---|---|---|---|
| Configuration upkeep | Policy, reference data or org change | Silent, applies outdated rules | Rate of business change |
| Interface drift | Source system or format change | Loud if structure breaks, silent if a field shifts | Number of systems touched |
| Exception absorption | Cases outside scope | Visible, lands on a person | Volume and exception rate |
| Verification | Any loss of trust | Permanent once established | Volume |
Only the third category is usually visible in a business case, and it is frequently understated.
How to estimate it before deployment
Six questions, answerable at scoping, that produce a usable estimate.
1. What has to stay current for this to remain correct?
List it explicitly: reference data, thresholds, mappings, routing rules. Each entry is a maintenance commitment with an owner.
2. How often does each of those change?
Reference data in a stable function changes rarely. Pricing, supplier and policy data in an active business changes continuously. The rate determines the load.
3. How many interfaces does it depend on, and who controls them?
An automation reading three systems owned by three teams has three sources of drift and no control over any of them. Dependencies owned elsewhere are the ones that change without notice.
4. What proportion of cases will route to a person?
This is the exception rate, and understating it is the most common estimation error. If nobody can state the current rate, that is the finding: the figure in the business case is a guess.
5. How would you know if it stopped being correct?
If the answer involves someone noticing downstream, the design has a silent failure mode and the detection cost belongs in the estimate.
6. Who holds the maintenance, and is their time costed?
A named person, with the hours entered into the business case. An unnamed owner means the load is real and unassigned, which is the condition that produces degradation.
Reading the answers
| Pattern | What it implies |
|---|---|
| Few configuration items, slow change rate | Low upkeep, favourable candidate |
| Several interfaces owned by other teams | Recurring drift, budget for it |
| Exception rate unknown | The business case rests on an estimate nobody can defend |
| Detection depends on downstream notice | Silent failure mode, add verification cost |
| No named maintenance owner | The load exists and will land somewhere unplanned |
| High business change rate in the affected area | Upkeep scales with the pace of the business |
Two or more rows in the lower half suggest the candidate is less attractive than the headline saving implies, or that the process should be simplified before it is automated.
The alternative that frequently wins
When the maintenance estimate is high, a second option becomes competitive: remove or simplify the step rather than automating it.
A step that exists because two systems disagree can be addressed by integrating them, which removes both the manual work and the automation that would have compensated for it. A step that exists because of a superseded policy can be eliminated. A process with high variation can be standardized first, which reduces the configuration surface before anything is built.
These options carry no ongoing maintenance load, which is what makes them durable. They appear in the comparison only when the automation option is costed honestly.
Where Horizon fits
Horizon is an AI-powered continuous discovery platform. Its relevance here is that the maintenance load of an existing estate, and the exception rate of a candidate process, are both established by asking the people who carry them.
Discovery Cycles run AI-led interviews across the roles that operate a process, surfacing the work people perform to keep systems correct, which appears in no system record because it is compensation rather than transaction. The Insights Dashboard quantifies the effort per finding and ranks by impact and effort, and the Initiatives Dashboard converts priorities into business cases with dispositions attached, including the ones that remove a step rather than automating it.
PedidosYa, the leading food delivery and quick commerce platform in Latin America, connecting users, businesses and riders across 15 local markets, provides a clear example of maintenance load operating in production.
Horizon ran 14 or more asynchronous AI interviews across Rider Payments and Partner Payments over three months without blocking a single calendar slot, surfacing 31 actionable findings across five high-impact process areas.
Two of them are maintenance load in its purest form. Rider wallet adjustments and cash tool failures required 3 to 4 hours per week of manual rework affecting roughly 2,400 riders weekly, caused by recurring retries in a 40-minute window with no automated resolution. Someone was rerunning a failed automated step by hand, every week, as a standing part of the operation.
Payment report downloads from one provider consumed around 40 hours per month, roughly 30 for the reports themselves and about 10 additional hours caused by download failures affecting tip processing. The automation existed, it failed predictably, and a person absorbed the failure.
Neither of these appeared in any system as a cost, because the work of compensating for an automation produces no transaction of its own. Both were quantified only once the people performing them were asked.
The engagement also identified partner billing control at 4 to 5 hours per week cross-referencing three data sources for 25,000 partners, and manual reconciliation in one market consuming roughly two hours per week against about 40 minutes elsewhere, driven by a visualization error forcing line-by-line comparison.
PedidosYa has since extended discovery to Tax, Collections and cross-market benchmarking, running it as a continuous programme.
That is one engagement under specific conditions rather than a projection for any organization.
Pre-deployment checklist
- What has to stay current for this automation to remain correct?
- How often does each of those items change?
- How many interfaces does it depend on, and who controls each?
- What proportion of cases will route to a person, and where does that number come from?
- How would you detect that it stopped being correct?
- Who holds the maintenance, and are their hours in the business case?
- What happens to the affected period if a silent failure runs for two weeks?
- Would a verification step be added after the first surprise, and is that costed?
- Have you compared this against removing, integrating or standardizing the step?
- What is the five-year cost including upkeep, rather than the first-year saving?
Question 6 is the one that most often has no answer. An unnamed maintenance owner is a load that exists and lands somewhere nobody planned.
Common mistakes
Costing the build and not the upkeep. Produces a business case that is accurate for the first quarter.
Understating the exception rate. The absorbed volume lands on a person and gets attributed to the process.
Treating a pilot result as the steady state. Configuration drift, interface changes and exception growth accumulate after the evaluation window.
Assuming the maintainer is the approver. The load usually lands on someone who was not in the meeting.
Ignoring the verification that forms after the first incident. It becomes permanent and it is pure addition.
Skipping the comparison against removal. An honest maintenance estimate frequently makes eliminating the step the better option.
FAQ
What is automation maintenance cost?
The ongoing effort required to keep an automation correct after deployment: updating configuration and reference data, adapting to changes in the systems it reads, handling cases outside its scope, and verifying that its output is still right. It is distinct from build cost and it continues for as long as the automation does.
Why is maintenance cost missing from automation business cases?
Because the savings are specific and the costs are diffuse. A business case can state the time removed per case. The maintenance load is spread across people in interruptions rather than blocks, lands on someone who was not in the approval, and accumulates after the evaluation window closes.
What is a silent automation failure?
An automation that keeps running on stale configuration or a changed input and produces output that looks normal. Nothing errors and nothing alerts. The problem surfaces downstream when a number does not reconcile, which makes both the correction and the trust cost larger than a visible failure.
How do you estimate automation maintenance before deployment?
List what has to stay current and how often each item changes, count the interfaces and who controls them, establish the real exception rate, define how a failure would be detected, and name the person who will carry the upkeep with their hours in the business case.
Should high maintenance cost stop an automation?
It should put removal, integration and standardization into the comparison. A step that exists because two systems disagree can be addressed by connecting them, which removes both the manual work and the automation that would have compensated for it. Those options carry no ongoing load, which is what makes them durable.
How do you find maintenance work that is already happening?
Ask the people who operate the process what they rerun, recheck or correct on a recurring basis. Compensating for an automation produces no transaction, so it appears in no system record and is only visible to the person performing it.
The question is what it costs to keep it right
Automation evaluations are built around what a system does when it is working. That is the easier property to measure and it is not the one that determines the total cost.
What a system requires to stay correct, and who absorbs that requirement, is the variable that decides whether the saving holds in year three.
See it. Fix it. Scale it.