Procure-to-pay discovery usually starts with open questions. Walk me through your process. Where does it slow down. What frustrates you.
Those produce a large volume of material and a weak signal. People describe what is most annoying rather than what is most expensive, and the two frequently differ. A step that irritates everyone weekly may consume fewer hours than one nobody mentions because it has always been that way.
A more productive order inverts it. Start from the small number of metrics that explain most of the value at stake, then work backward to the questions that would move them.
Across P2P engagements the metrics that matter converge on a short list. Once they are named, the discovery questions follow, the findings are comparable, and the business case writes itself because the baseline was established before anything was investigated.
Key takeaways
- A small number of metrics explain most of the value at stake in procure-to-pay.
- Open-ended discovery surfaces annoyance, which correlates weakly with cost.
- Cycle time matters less than the split between touch time and wait time within it.
- Exception rate and first-time match rate determine how much of the process the standard flow actually covers.
- Maverick spend and duplicate payment are symptoms of upstream process gaps rather than compliance failures.
- Starting from metrics makes findings comparable across business units and markets.
The metrics that explain the value
1. End-to-end cycle time, split into touch and wait
Requisition to payment, measured end to end, and then decomposed.
The decomposition is the useful part. If most elapsed time is a case waiting for an approval, an input or a clarification, then adding capacity or automating processing steps changes little. The constraint sits in queue structure and decision rights.
Working backward: where does a requisition wait, who is it waiting for, and what triggers the wait to end.
2. First-time match rate
The proportion of invoices that match purchase order and receipt without intervention.
This single number determines how much of the accounts payable operation is exception handling. A rate in the high nineties means AP is processing. A rate in the seventies means AP is investigating, and the effort per invoice is several times higher for the failures.
Working backward: what causes a match to fail, in what proportions, and where does the error originate.
3. Exception rate and distribution by cause
Not a single number, a distribution. Price variance, quantity variance, missing receipt, missing purchase order, wrong cost centre, duplicate submission, unapproved supplier.
Each cause has a different upstream owner and a different fix. Treating them as one exception rate produces an initiative to improve exception handling, which is a symptom-level response.
Working backward: for the two largest causes, which upstream step produces them.
4. Touch count per transaction
How many people handle a typical requisition and a typical invoice from start to finish.
This is a useful proxy for coordination cost and it is easy to collect. Each additional touch introduces a handoff, a queue and an opportunity for information to be lost.
Working backward: which touches add a decision and which add only a transfer.
5. Supplier onboarding cycle time
Request to first eligible transaction.
Long onboarding pushes the business toward workarounds, which then present as compliance problems. Onboarding time is frequently the upstream cause of maverick spend, and addressing the compliance symptom without the cycle time produces friction rather than compliance.
Working backward: which steps in onboarding are sequential that could be parallel, and which approvals could be risk-tiered.
6. Rate of spend outside the process
Purchases made without a requisition, outside contracted suppliers, or approved retroactively.
The instinct is to treat this as behaviour. It is more usefully read as a signal about the process: people route around a path that does not meet their need for speed, coverage or usability.
Working backward: for the largest category of off-process spend, what did the official path fail to provide.
7. Duplicate and error payment rate
Payments issued twice, at the wrong amount, or to the wrong party.
Low rates are still informative, because the detection effort is the cost rather than the payments themselves. If someone is checking every payment against a second source to keep the rate low, the checking is the recoverable item.
Working backward: what controls exist, what they catch, and how much effort they consume per item caught.
Why starting from metrics changes the discovery
Three effects, and they compound.
Findings become comparable. Two business units describing their process in their own terms produce accounts that cannot be ranked. Two business units reporting first-time match rate and exception distribution produce a comparison with a target.
Cost gets separated from annoyance. The step people complain about and the step consuming the most hours are frequently different. Metrics-first discovery surfaces both and distinguishes them.
The baseline exists before the intervention. The most common reason a P2P improvement cannot demonstrate its value is that nobody measured before. Starting from metrics resolves that as a side effect.
P2p metric to question
Working backward from each metric
| Metric | If it is poor, ask | Likely upstream owner |
|---|---|---|
| Wait time share of cycle | Who is each case waiting for, and why | Approval design, decision rights |
| First-time match rate | What causes the match to fail, in what proportion | Purchasing, receiving, supplier master |
| Exception distribution | Which upstream step produces the top two causes | Requisition quality, receipt discipline |
| Touch count | Which touches decide and which only transfer | Process design |
| Onboarding cycle time | Which steps are sequential that could be parallel | Compliance, supplier management |
| Off-process spend | What the official path failed to provide | Catalogue coverage, onboarding speed |
| Payment error rate | What the detection controls cost per item caught | Controls design |
Each metric names a question and an owner. A finding without both tends to become a recommendation nobody acts on.
Where the recoverable effort concentrates
Across procurement and accounts payable operations, the effort consistently pools in four places.
Exception investigation. Determining why an invoice did not match, which requires checking several systems and frequently contacting a supplier or a requester. The per-case effort is several times that of a clean transaction.
Manual validation between systems. Verifying that pricing, terms or supplier details agree across the purchasing system, the contract record and the ERP. This is compensation for an integration gap and it scales with volume.
Chasing approvals. Following up on requisitions and invoices sitting in queues, which is pure coordination and produces nothing except movement.
Report assembly. Building spend, accrual and supplier performance views by hand from several sources each cycle.
None of these four appear as a process step in any documentation. All of them consume hours that scale with transaction volume.
The maverick spend reframe
Worth separating out, because it is the P2P topic most often addressed at the wrong level.
Spend outside the process is usually treated as a compliance problem, addressed with policy reminders, approval tightening and reporting. That approach produces a temporary improvement and a durable increase in friction.
Read as a process signal instead, off-process spend indicates that the official path failed to provide something: speed, catalogue coverage, a supplier that could actually be onboarded in time, or a usable interface.
The diagnostic question is what the requester needed that the process did not supply. In most operations the answer concentrates in two or three categories, and addressing those reduces off-process spend more durably than any control does.
Where Horizon fits
Horizon is an AI-powered continuous discovery platform. In P2P its role is producing the effort baseline and the exception distribution, which are the two inputs the metrics above require and the two that no system records.
Discovery Cycles run AI-led interviews across procurement, accounts payable, treasury and the business units that raise requisitions, asynchronously and without blocking calendars. The Insights Dashboard quantifies the effort per finding and ranks by impact and effort, with traceability to the input behind each one. The Process Library structures the resulting documentation, which makes comparison across business units possible, and the Initiatives Dashboard converts priorities into business cases with owners.
Grupo HZ, a multi-company industrial holding with more than 65 years in packaging and paperboard, operating across Argentina, Brazil and Chile with a shared services centre covering Administration, Finance, Procurement and Treasury, provides a clear example of what the baseline looks like when it exists.
Before the engagement, spreadsheets, email and manual validations were the backbone of critical processes, with no integrated workflows and no visibility into where time was being lost. Pricing errors in the ERP went undetected without cross-checks, creating downstream issues in billing and auditing, which is the manual validation pattern described above.
Horizon ran 55 structured interviews across six departments asynchronously and quantified the time impact down to the hour, per area. Procurement carried 48.3 hours per week, the largest of any area by a wide margin. Accounts Payable carried 34 hours per week. Treasury carried 21.55. Credit and Collections 10.5, Sales and Commercial in Chile 6.5, and Planning and Costing 2.9.
Together, Procurement and Accounts Payable accounted for roughly two thirds of the 123 hours per week identified across the whole shared services operation. The engagement produced 33 actionable findings with 50 to 90% automation potential across all areas and specific solution types mapped to specific processes, saved 159 discovery hours against the manual approach, reallocated three FTEs to strategic work without adding headcount, and returned 182% ROI.
The distribution is the finding that matters. A prioritization based on transaction volume would not have produced that ranking, and the programme would have started somewhere other than where the hours actually were.
That is one engagement under specific conditions rather than a projection for any organization.
P2P improvement checklist
- Do you know end-to-end cycle time split into touch time and wait time?
- Do you know your first-time match rate?
- Do you have an exception distribution by cause, not only a rate?
- For the two largest exception causes, do you know which upstream step produces them?
- Do you know the touch count for a typical requisition and a typical invoice?
- Do you know supplier onboarding cycle time from request to first transaction?
- Do you know the rate of spend outside the process, and what the official path failed to provide?
- Do you know how much effort goes into exception investigation per case?
- Do you know how many hours per week go into manual validation between systems?
- Can you compare these metrics across business units and markets?
- Is there a baseline that would let you demonstrate improvement afterward?
Common mistakes
Starting with open questions. Produces annoyance-weighted material rather than cost-weighted material.
Treating exception rate as one number. Each cause has a different upstream owner and a different fix.
Optimizing touch time when wait time dominates. Common in approval-heavy procurement and consistently disappointing.
Addressing off-process spend as behaviour. Controls produce friction. Fixing what the official path failed to provide produces compliance.
Ignoring the cost of controls. A low error rate maintained by manual checking has moved the cost rather than removed it.
Improving without a baseline. Makes the result unverifiable and the next investment harder to justify.
FAQ
What are the key metrics for procure-to-pay?
End-to-end cycle time split into touch and wait, first-time match rate, exception distribution by cause, touch count per transaction, supplier onboarding cycle time, rate of spend outside the process, and duplicate or error payment rate. Together these explain most of the value at stake, and each one names a specific discovery question.
How do you improve the procure-to-pay process?
Start from the metrics rather than from open discovery, work backward from each poor metric to the upstream step that produces it, and quantify the effort consumed by exception investigation and manual validation. Those two typically carry the recoverable hours and appear as steps in no documentation.
What causes low first-time match rates?
The most common causes are price and quantity variances between purchase order and invoice, missing or late goods receipts, missing purchase orders where spend happened outside the process, and supplier master data that differs across systems. Each has a different upstream owner, which is why a single exception rate is less useful than a distribution.
Is maverick spend a compliance problem?
More usefully read as a process signal. People route around the official path when it fails to provide speed, catalogue coverage, a supplier that can be onboarded in time, or a usable interface. Tightening controls without addressing what the path failed to supply produces friction and a temporary improvement.
Should you automate accounts payable exceptions?
Only after establishing the distribution by cause, and only for causes that are stable and high-volume. Several common exception causes are better addressed upstream, by fixing requisition quality, receipt discipline or supplier master data, which removes the exception rather than automating its handling.
How do you build a business case for P2P improvement?
From the effort baseline. Quantify hours per week consumed by exception investigation, manual validation, approval chasing and report assembly, per area. That number, combined with the exception distribution, supports both the size of the opportunity and the sequencing of which cause to address first.
Ask the metric, then ask the question
Open discovery in procure-to-pay produces a lot of material and a weak ranking, because people describe what bothers them rather than what costs the most.
Naming the metrics first makes the findings comparable, separates cost from annoyance, and leaves a baseline behind that the next conversation about value will need.
See it. Fix it. Scale it.