Large transformation portfolios rarely suffer from a shortage of ideas. They suffer from several hundred of them, arriving from different functions, described at different levels of detail, with substantial overlap that nobody has the vantage point to see.
At that scale, prioritization frequently degrades into negotiation. Seniority, urgency and presentation quality shape the ranking more than evidence does, because there is no shared basis for comparing a procurement bottleneck against a reporting gap.
The fix has three parts, and the first one is usually skipped. Deduplicate before scoring, because scoring an overlapping portfolio ranks the same initiative several times under different names. Score on consistent criteria. Then sequence by dependency rather than by score, because the highest-scoring initiative is often blocked by a lower-scoring one.
Key takeaways
- Deduplication comes before scoring. Large portfolios routinely contain the same initiative described several times by different functions.
- Consistent criteria matter more than sophisticated ones. The value is in comparing on the same basis.
- Evidence strength deserves its own criterion, since an initiative backed by a workshop and one backed by operational evidence are not equivalent candidates.
- The ranked list is an input to the roadmap rather than the roadmap itself.
- Sequencing is governed by dependency, capacity and risk balance, none of which appear in a score.
- A prioritization process that never changes the ranking was a documentation exercise.
Step 1: Deduplicate before you score
In a portfolio assembled from multiple functions, overlap is the default rather than the exception. Three teams identify the same reconciliation problem from different angles and file three initiatives with different names, different sponsors and different proposed solutions.
Scoring that portfolio produces three entries near each other in the ranking, which reads as strong signal for a theme and is actually one problem counted three times.
Deduplication has two levels.
Same problem, different framing. A request for a report, a request for an integration and a request for a data owner may all originate in one data ownership gap. Merging them requires knowing the cause, which is why deduplication depends on evidence rather than on reading titles.
Same cause, different symptoms. Harder and more valuable. Several unrelated-looking initiatives across different functions can share one upstream mechanism. Fixing the mechanism retires all of them.
The output of this step is a smaller portfolio where each entry is a distinct problem with a distinct cause. That alone frequently changes the ranking more than the scoring does.
Step 2: Apply pass or fail gates
Not every initiative is ready to be scored. Scoring one that fails a basic condition produces false precision.
| Gate | Pass condition |
|---|---|
| Named problem | The initiative describes a problem, not only a proposed solution |
| Evidence exists | Someone can point to what produced this beyond an opinion |
| Business owner | A named person will sponsor the change and the outcome |
| Measurable outcome | There is a metric that would move, with a current value |
| Feasible path | There is a plausible route from decision to implementation |
An initiative failing a gate is not rejected. It returns to the intake with the specific gap named.
Step 3: Score on six criteria
Score each surviving initiative from 1 to 5. For effort, score in reverse, where 5 means a focused path and 1 means a heavy, dependency-laden one.
| Criterion | Weight | Score 1 means | Score 5 means |
|---|---|---|---|
| Business value | 25% | Marginal or hard to measure | Material effect on cost, cycle time, quality, risk or revenue |
| Evidence strength | 20% | Based on anecdote or a single workshop | Backed by operational evidence from the people doing the work |
| Feasibility | 15% | Path unclear, dependencies unresolved | Route to implementation is known |
| Adoption readiness | 15% | No owner, no user group, no change path | Owner, users, incentives and manager support are clear |
| Effort | 15% | Multi-year, high-dependency build | Focused implementation that can prove value quickly |
| Strategic fit | 10% | Interesting but peripheral | Directly supports a named enterprise priority |
Weighted score equals the sum across criteria of the score divided by 5, multiplied by the weight. That produces a total out of 100.
Two notes on using this.
Evidence strength is the criterion most often omitted and the most diagnostic. An initiative proposed confidently in a leadership meeting and one derived from what the people running the process described are not equivalent candidates, and a scoring model without this criterion treats them as such.
Do not over-optimize the decimals. The value of the model is the conversation it forces about which assumptions are weak and what would have to change for an initiative to move up. If the scoring meeting never changes the ranking, the model was a formality.
Step 4: Sequence, which is a different exercise
A ranked list is not a roadmap. Four considerations govern sequencing and none of them appear in the score.
Dependency
The highest-scoring initiative is frequently blocked by a lower-scoring one. A process redesign that depends on a system integration cannot start before it, regardless of relative scores. Map dependencies before committing a sequence.
Capacity
Initiatives compete for the same scarce people. Three high-scoring initiatives that all require the same integration team run sequentially whether or not the plan says otherwise.
Portfolio balance
A roadmap consisting only of quick wins produces visible activity and leaves the structural problems untouched. A roadmap consisting only of structural bets overloads the organization before anything demonstrates value. Most portfolios need a mix.
Evidence maturity
Some high-value initiatives are not ready to fund and are ready to investigate. Placing them in a "prove next" category preserves them without committing delivery capacity to an unresolved question.
Portfolio disposition
Four categories, not one ranked list
| Category | Value | Readiness | Action |
|---|---|---|---|
| Do now | High | High | Fund with owner, metric, timeline and governance path |
| Prove next | High | Lower | Run discovery, data validation or a small proof first |
| Redesign | Potentially high | Weak or unclear | Reframe the problem, clarify ownership, gather better evidence |
| Defer | Low or peripheral | Any | Hold unless strategy changes or new evidence appears |
Sorting into categories is what prevents the two common failures: funding only the easy initiatives, or funding only the ambitious ones.
The governance that keeps it working
A prioritization exercise run once produces a roadmap that decays. Three mechanisms keep it current.
A standing review cadence. Monthly or quarterly, with the portfolio reviewed as a portfolio rather than initiative by initiative.
A re-scoring trigger. New evidence, a completed initiative that changes a dependency, a strategy change, or an initiative that has been in "prove next" beyond a defined period.
A stop mechanism. Portfolios grow because initiatives enter and rarely leave. Without an explicit path to stop something, the portfolio expands until capacity is spread across too many things to finish any of them.
That third one is the least common and the most consequential. A governance process that can only add is not a portfolio process.
Where Horizon fits
Horizon is an AI-powered continuous discovery platform. It addresses the two steps that determine whether the rest of the model produces anything: deduplication, which requires knowing causes, and evidence strength, which requires knowing what the people doing the work actually described.
Discovery Cycles run AI-led interviews across functions and geographies simultaneously, which is what makes cross-team pattern detection possible. The Insights Dashboard groups findings by process, system and cause rather than by who reported them, so several separately filed problems sharing one mechanism appear as one finding. Each item is traceable to the source input, which is what makes evidence strength scoreable rather than asserted. The Initiatives Dashboard converts priorities into business cases with owners, effort and expected impact attached.
Mercado Libre had the problem this model is built for. Operating in 18 countries with more than 84,000 employees, it was carrying over 600 initiatives, many overlapping or misaligned, across fragmented systems including SAP, ARIBA, PIC, CLM, Jira, BigQuery, Tableau, Apache and Bypass. Its manual discovery model meant months of interviews covering under 1% of the workforce and requiring more than 100 coordination hours per initiative. Business objectives included ensuring that every initiative delivered ROI above 150%.
Horizon interviewed 2,000 employees in four days across Finance and related functions, reaching 100% of the targeted areas in five countries. Transcripts were fused with system data from SAP, ARIBA, PIC, BigQuery and Tableau to expose value gaps. Cross-country benchmarking revealed hidden duplications and best practices that were invisible from inside any single market.
The output was 24 initiatives co-created with ROI, effort and a step-by-step roadmap attached, all clearing the 150% ROI threshold, alongside auto-generated board decks, RFPs, process maps and Jira epics. The engagement produced $2.3M in projected annual savings from 12,000 to 13,000 hours automated, 76 FTE equivalents avoided, and first implementations live in under a week.
The reduction from more than 600 initiatives to 24 funded ones is the deduplication step operating at scale. Most of that reduction came from finding which initiatives were describing the same underlying problem, which is not possible to establish by reading titles.
That is one engagement under specific conditions rather than a projection for any organization.
Prioritization checklist
- Have you deduplicated the portfolio before scoring it?
- For overlapping initiatives, do you know whether they share a problem or a cause?
- Does every scored initiative describe a problem rather than only a solution?
- Does every scored initiative have a named business owner?
- Does every scored initiative have a metric with a current value?
- Is evidence strength one of your scoring criteria?
- Have you mapped dependencies before committing a sequence?
- Does the roadmap balance quick wins against structural work?
- Is there a category for initiatives that need proof rather than funding?
- Is there an explicit mechanism to stop an initiative?
- Did the scoring conversation change the ranking?
Question 11 is the test of whether the exercise was real.
Common mistakes
Scoring before deduplicating. Produces a ranking that counts the same problem several times and reads as consensus.
Omitting evidence strength. Treats an opinion and an observation as equivalent inputs.
Treating the ranked list as the roadmap. Ignores dependency, capacity and balance, all of which govern what can actually run.
Funding only quick wins. Produces visible activity and leaves the structural constraints in place.
No stop mechanism. The portfolio grows until capacity is spread too thin to complete anything.
Running it once. Evidence changes, dependencies resolve, priorities move. A ranking from two quarters ago describes a different organization.
FAQ
How do you prioritize transformation initiatives?
Deduplicate the portfolio first, since large portfolios routinely contain the same problem described several times. Apply pass or fail gates so that only initiatives with a named problem, evidence, an owner, a metric and a feasible path get scored. Score on business value, evidence strength, feasibility, adoption readiness, effort and strategic fit. Then sequence by dependency and capacity rather than by score alone.
What criteria should you use to rank transformation initiatives?
Six work well: business value, evidence strength, feasibility, adoption readiness, effort and strategic fit. Evidence strength is the one most often omitted and the most useful, because it distinguishes an initiative backed by operational evidence from one backed by a single workshop.
Why do transformation portfolios end up with too many initiatives?
Because initiatives enter through many channels and rarely leave. Without deduplication, the same problem appears several times under different sponsors. Without an explicit stop mechanism, initiatives that lose relevance stay in the portfolio consuming attention and planning capacity.
Is the highest-scoring initiative the one to start first?
Usually not. Sequencing is governed by dependency, capacity, portfolio balance and evidence maturity. A high-scoring initiative blocked by a lower-scoring integration cannot start first, and a roadmap of only high scorers frequently overloads the same scarce delivery teams.
How often should a transformation portfolio be re-prioritized?
Review as a portfolio monthly or quarterly, and re-score when new evidence arrives, when a completed initiative changes a dependency, when strategy shifts, or when an initiative has sat in a proof category beyond a defined period. Re-scoring on a fixed calendar without a trigger tends to produce the same ranking.
The ranking is the easy part
Most organizations can produce a ranked list. Fewer can produce one where each entry is a distinct problem, the evidence behind each is visible, and the sequence reflects what can actually run given dependencies and capacity.
The work that makes prioritization useful happens before the scoring meeting, in establishing what the initiatives actually are.
See it. Fix it. Lead it.