Atlas / GOVERN & LEAD / Strategy / Use-Case Prioritization
DEEP-DIVE · STRATEGY

Use-Case Discovery and Prioritization: From Wishlist to Portfolio

Most AI programs die of abundance, not scarcity: a hundred pilots, no portfolio logic, nothing in production. This is a scoring method that survives a steering committee, and a case for managing a portfolio rather than ranking a list.

TL;DR
  • The failure mode of enterprise AI is rarely too few ideas; it is too many, with no logic for choosing among them. Prioritization, not model quality, is where the value is won or lost.
  • Score candidates on value, feasibility and risk with coarse bands, not false decimals, then plot them on a value-feasibility matrix. Feasibility, dominated by data readiness and change management, is the axis teams flatter.
  • Prioritization is portfolio management, not a ranked list: balance quick wins, strategic bets and platform investments, sequence by dependency, and revisit quarterly. A one-time workshop is not a process.

The thousand-pilots trap

The AI programs I am asked to rescue rarely suffer from a shortage of ideas. They suffer from the opposite. An enthusiastic organization, given a capable technology and a mandate to experiment, generates use cases the way a healthy tree generates leaves. Every function has three, every workshop produces a dozen, and the intake spreadsheet crosses two hundred rows before anyone asks the harder question. The energy is real and it is worth protecting. But energy without a selection mechanism does not compound. It disperses.

What that dispersion looks like in practice is a portfolio of demos. Each one works in the room: a slick retrieval bot over the policy manual, a drafting assistant for the sales team, a triage classifier for support tickets. Each earns applause, a screenshot in a board deck, and a vague commitment to "productionize next quarter." Then the next quarter arrives with its own dozen demos, and the last dozen quietly stop being maintained. A year in, the honest tally is often the same: many pilots, a handful in real use, and a growing suspicion among executives that AI is expensive theater.

The reflex is to blame execution, or the technology, or the vendor. It is almost never any of those. The demos die because nothing decided which few deserved the disproportionate investment that reaching production actually requires. Prioritization was skipped, or done once as a voting exercise, and so the program spread its resources evenly across everything and finished nothing. Even coverage is the enemy here.

The premise of this article: prioritization is not the administrative step before the real work. It is the real work. The difference between a program that ships three durable systems and one that abandons thirty pilots is almost entirely a difference in how candidates were chosen and sequenced, not in how they were built.

Scoring: value, feasibility, risk

A defensible scoring model rests on three dimensions, and the discipline is in keeping it to three. Value asks what changes for the business if this works: revenue gained, cost removed, cycle time cut, or quality raised in a way someone will actually notice. Feasibility asks how hard it is to reach production: whether the data exists and is usable, how difficult the technical build is, and how much the surrounding process and behavior must change. Risk asks what could go wrong: regulatory exposure, reputational damage, and the reliability of the model on the actual task.

The temptation, always, is false precision. A committee that scores value at 7.4 and feasibility at 6.8 has not become more rigorous; it has dressed a guess in a lab coat. The number implies a measurement that was never taken, and it invites arguments about the second decimal that crowd out the argument worth having. I score in coarse bands, high, medium and low, sometimes on a one-to-five scale where the anchors are described in words, and I insist that every band have a written definition agreed before any candidate is scored.

DimensionWhat to probeCheap signal it is a low score
ValueRevenue, cost, speed, quality; is the benefit measurable and owned by a named sponsor?The benefit is described as "efficiency" with no baseline and no one accountable for it
FeasibilityData readiness and entitlements, technical difficulty, change effort in the receiving processThe data lives in three systems, no one owns it, and the workflow it must enter is contested
RiskRegulatory reach, reputational blast radius, model reliability on this specific taskIt makes or shapes a decision about a customer, and an error is hard to detect after the fact

Treat risk as a modifier, not a fourth thing to average. Value and feasibility position a candidate; risk tells you what guardrails, review, and evidence the position demands before it may proceed. A high-value, high-feasibility use case that also carries regulated-decision risk is still worth doing, but it earns the scrutiny covered in Governance rather than a fast-track. Averaging risk into a single blended score is how genuinely dangerous use cases get laundered into the middle of the pack.

The value-feasibility matrix

Two of the three dimensions have a spatial relationship worth drawing. Plot value on one axis and feasibility on the other, place every candidate as a point, and the portfolio sorts itself into four quadrants that each imply a different decision. This is standard consulting practice and I present it as advisor judgment rather than measured law, but I have never seen a steering committee that did not think more clearly once the candidates were on this grid instead of in a spreadsheet.

             FEASIBILITY
         low            high
       +------------+------------+
  high | PLATFORM   | QUICK      |
       | BETS       | WINS       |
 VALUE | (sequence, | (do now,   |
       |  invest)   |  fund fast)|
       +------------+------------+
  low  | MONEY PITS | FILL-INS   |
       | (avoid /   | (batch,    |
       |  kill)     |  automate) |
       +------------+------------+
The value-feasibility matrix: position sets the decision, not just the priority.

Quick wins (high value, high feasibility) are the top-right and they fund the program's credibility. Do them now, resource them properly, and let them earn the political capital that everything else will spend. Platform bets (high value, low feasibility) are the systems worth building slowly: the value is real but the data, integration, or capability is not there yet, so they need sequencing and investment rather than a sprint. Fill-ins (low value, high feasibility) are cheap and easy; do them opportunistically, batch them, automate them, but never let them crowd out the top row. Money pits (low value, low feasibility) are the ones a scoring exercise exists to catch, because they are often the most technically interesting and therefore the most seductive to engineers.

The matrix earns its keep by making trade-offs visible to non-technical decision-makers in a way a ranked list never does. A list says "do this before that." The grid says why, and it exposes the specific failure it is designed to prevent: a program that has quietly loaded all its resources into one bottom-left money pit because it was somebody's pet idea.

The feasibility factors teams underestimate

When candidates land in the wrong quadrant, feasibility is almost always the axis that lied. Teams scoring their own ideas are optimistic about how hard the build is, and the optimism has a consistent shape: they estimate the model, which is the part they understand, and discount everything around it, which is the part that actually determines whether the thing ships. In my experience the model is rarely the binding constraint. Three other factors usually are.

The practical correction is to score feasibility with the people who own the data and the process in the room, not just the technologists who will build it. When the person accountable for the receiving workflow scores change effort, the number moves, usually downward, and the portfolio gets honest. A feasibility score set by the builders alone is a forecast of the easy part.

A rule I hold with some confidence: if a use case looks easy and you have not yet talked to whoever owns the data and whoever owns the workflow, you have not assessed feasibility. You have assessed the demo.

Balancing the portfolio

Here is where the language matters, because it changes the behavior. The output of prioritization is not a ranked list to be worked top to bottom. It is a portfolio, and a portfolio has a shape. The reason to borrow the word from finance is that it carries the right intuition: you hold a mix of assets with different risk-return profiles on purpose, because concentration is fragility, and because different holdings do different jobs for you at the same time.

Three kinds of holding earn their place for different reasons. Quick wins fund credibility: they ship fast, prove the program returns something, and buy the patience that longer efforts require. Strategic bets build capability: they may not pay off this year, but they move the organization toward something it could not do before. Platform investments lower the cost of everything downstream: the shared retrieval layer, the evaluation harness, the reusable governance pattern that turns the next ten use cases from bespoke projects into configurations.

The characteristic failure of ambitious programs is to pour every resource into one flagship, the transformational system that will supposedly justify the whole investment. When it slips, and singular flagships slip, there is nothing shipping alongside it to sustain belief, and the program loses its mandate before the bet matures. The opposite failure, all quick wins and no capability, produces a program that is busy, well-liked, and strategically stuck: a collection of point solutions that never become a platform. A healthy portfolio deliberately runs all three, and rebalances as some pay off and others are killed. The ratio is a judgment call, not a formula, but the discipline is non-negotiable: never let the mix collapse to a single type.

From scores to a sequenced roadmap

A ranked list and a roadmap are different objects, and confusing them is a quiet source of failure. A list says what is most important. A roadmap says what happens in what order, given that some things must exist before others can, that capability accrues over time, and that the same scarce people cannot build everything at once. The move from scores to a roadmap is where dependency and capability build-up enter, and where the portfolio stops being an analysis and becomes a plan.

Sequence by dependency first. The platform investments that make later use cases cheap belong early even when their standalone value is modest, because they change the feasibility of everything that follows. A shared retrieval and evaluation layer built for the first two use cases turns the third through tenth from projects into configurations. Sequencing purely by score, ignoring these dependencies, produces a roadmap that rebuilds the same foundation five times.

  1. Sequence by dependency and capability. Order the portfolio so that foundational platform work and the skills it teaches land before the use cases that rely on them, not strictly by descending value.
  2. Assign an owner and a funding line to each item. A use case with no named owner and no budget beyond the pilot is a wish. Fund the survivors as products with run-state money, per the argument in the operating models deep-dive.
  3. Define success metrics before building. Agree the baseline and the target with the sponsor while the work is still cheap to stop. A metric chosen after launch is a metric chosen to make the launch look good.

That third point does more work than it appears to. Defining the metric and the baseline before building is the mechanism that makes the honesty in the final section possible: you cannot measure realized value against a baseline you never recorded. It is also the cheapest possible way to kill a weak use case, because a sponsor who cannot articulate what success would look like, in numbers, before a line of code is written, has just told you the value score was aspirational.

Keeping the portfolio honest

Everything above describes a workshop, and a workshop held once is worthless within two quarters. Conditions change: data that was inaccessible becomes available, a use case that scored high underdelivers in the field, a regulatory shift moves the risk of a whole category, a platform investment lands and makes five formerly infeasible ideas feasible overnight. A portfolio that is scored once and then executed as written is executing last year's judgment against this year's reality. Prioritization is a standing process, not an event.

The cadence I recommend is a quarterly portfolio review with real teeth, which means three specific disciplines. Revisit the scores against what has actually been learned. Kill underperformers without ceremony, because the sunk-cost instinct is strong and a review that cannot stop things is just a status meeting. And measure realized value against the baseline that was agreed before building, not against a story assembled after the fact. That last discipline is the one most often skipped, and skipping it is how a program accumulates "successes" that no one can point to on a financial statement.

Killing well is a skill in its own right. A use case that is stopped early, cleanly, and without blame is not a failure of the program; it is the program working exactly as designed. The organizations that prioritize well are not the ones that pick winners flawlessly, which no one does. They are the ones that stop losers quickly and redirect the resources, so that the portfolio is always tilted toward what is currently working and currently learnable. Treat the value-realization discipline as its own practice (see the value realization primer), and benchmark the whole approach with the Strategic Option Weigher in the Lab.

The one question a steering committee should ask every quarter: which item did we kill since we last met, and what did its resources move to? A portfolio review with no kills is not disciplined. It is decorative.
ALL OF STRATEGY The AI Operating Model: Beyond the CoE Debate →