Most AI buying starts in exactly the wrong place. Somebody watches a demo, somebody else hears that a competitor has already rolled something out, and the room starts comparing vendors before anyone has written down the workflow, the baseline, or the failure mode. That is how a budget line becomes a story people tell about being innovative. Score the use case first. Buy later, build later, or kill it before the meeting ends.
How to Score an AI Use Case Before Anyone Buys Software
Noah Davis, Sophie Adams, and Lucas Brown · Aug 29, 2026 · 14 min read

TL;DR
- The case for scoring first is brutal and current: IBM's May 6, 2025 CEO study says only 25% of AI initiatives delivered expected ROI and only 16% scaled enterprise wide, while McKinsey's August 25, 2026 survey says just 37% of organizations see any EBIT impact and 6% qualify as high performers.
- The scoring method here weights value at 35%, feasibility at 25%, risk at 25%, and time to evidence at 15%, each on a 1 to 5 scale backed by named evidence rather than enthusiasm.
- Hard stops outrank the score. If a use case has no accountable owner, no baseline, no lawful data path, or no reversible failure mode, it does not earn a test even if the weighted score looks good.
- The worked example in this post is illustrative, not a benchmark: a support reply-drafting use case scores 78 out of 100 and clears the test band because the workflow is measurable, reversible, and fast to learn from.
- Do not pretend the score is precise. Round to whole numbers, decide in bands, and rescore quarterly because the same use case moves as the data, controls, and delivery capacity change.
Why the score belongs ahead of the demo
Broad adoption and thin realized value can both be true at the same time. Right now they are.
- of AI initiatives delivered expected ROI in IBM's 2025 CEO studyIBM Study: CEOs Double Down on AI While Navigating Enterprise Hurdles (2025)
- 25%
- scaled enterprise wide in the same IBM studyIBM Study: CEOs Double Down on AI While Navigating Enterprise Hurdles (2025)
- 16%
- of organizations report any EBIT impact from AI in McKinsey's 2026 surveyMcKinsey, The state of AI in 2026: On the road to ROI (2026)
- 37%
- qualify as AI high performers in McKinsey's 2026 surveyMcKinsey, The state of AI in 2026: On the road to ROI (2026)
- 6%
Why score the workflow before you compare tools?
Because tools are the cheapest part to talk about and the hardest part to unwind once the contract is signed. PwC's 28th Annual Global CEO Survey found 56% of CEOs say GenAI improved time efficiency, 32% say it increased revenue, and 34% say it increased profitability. Then IBM's May 6, 2025 CEO study asked a stricter question and got a stricter answer: only 25% of initiatives delivered the ROI leaders expected. The gap is not that executives are lying. It is that they are grading different things. One is sentiment about promise. The other is the return on a specific bet.
A score forces the room onto the second question before any money moves. One workflow. One owner. One number the workflow is supposed to move. One path to finding out whether it actually moved. That is the discipline we use in a 10x audit, because the fastest way to waste an AI budget is to fund something that still sounds good after the nouns are removed.
This is also where the fractional ai officer seat earns its keep. Someone has to be allowed to say, out loud, that a use case with no baseline is not a use case yet and that a workflow with no accountable owner is a political wish wearing a technical costume.
The weighted method
Score each proposed use case from 1 to 5 on four dimensions, then multiply by the weights below.
- Value, 35%. If the workflow does not move a cost, revenue, quality, risk, or capacity number the business cares about, the rest is theater.
- Feasibility, 25%. The model may exist and the integration may still be a mess. This is where teams confuse a good demo with a buildable system.
- Risk, 25%. NIST's AI Risk Management Framework says AI risk work has to be practical and operationalized. That means the workflow needs a real control boundary, not a promise to be careful later.
- Time to evidence, 15%. Faster learning gets a smaller weight than value, but it still matters because a use case that cannot produce evidence soon enough tends to survive on politics.
The weighting is lopsided on purpose. Value gets the biggest share because the use case exists to move a real number. Feasibility and risk tie because a buildable unsafe workflow is still a bad bet, and a safe workflow you cannot actually implement is not a bet at all. Time to evidence matters less than the other three, but it still decides sequencing because learning in six weeks is worth more than confidence in a slide deck.

Show the data behind this infographicHide the data behind this infographic
- Value, 35%: what hard number moves if this works, how much of that number is in scope, and who owns that outcome.
- Feasibility, 25%: can today's models do the task reliably enough, is the workflow structured enough to control, and how much integration debt sits underneath it.
- Risk, 25%: what can go wrong, whether a human can catch it before damage lands, and whether the team can monitor, stop, and explain the system in production.
- Time to evidence, 15%: how fast the team can gather baseline data, run a bounded test, and see a measured before and after.
How to score each dimension without flattering the idea
A 5 means the evidence is already in the room. A 1 means the room is mostly guessing.
For value, a 5 has a named business metric, current baseline, volume, and a believable mechanism for improvement. A 3 has a real pain point but loose math. A 1 sounds like "people will save time" with no number attached.
For feasibility, a 5 means the workflow is narrow, repetitive, and technically boring in the best sense. Inputs are reachable, outputs are clearly defined, and the team already knows where the ugly edge cases live. A 1 depends on brittle integrations, unstructured exceptions nobody has mapped, or model behavior the team cannot evaluate. When that is the shape, pay for discovery first or hand the job to a custom AI build only after the workflow gets smaller.
For risk, score high only when wrong answers can be caught before they cause damage. NIST's AI RMF 1.0 and the companion playbook keep coming back to the same idea: trustworthiness has to be built into use, not stapled on at launch. If the workflow touches regulated data, sends customer-facing messages, or changes money movement, the score falls fast unless the human review path is explicit and the evaluation set exists before launch.
For time to evidence, a 5 means the team can collect a clean baseline now and see a result inside one quarter. A 3 means the signal exists but needs cleanup. A 1 means the only proof would arrive after the budget season changed, by which point the workflow will still be living on narrative.
The hard stops that outrank the score
Four conditions suspend the math.
- No accountable business owner. A technical sponsor is not enough. Someone has to own the workflow change and the number attached to it.
- No baseline. If nobody knows the current cycle time, error rate, or labor cost, the use case cannot prove improvement.
- No lawful data path. If the data rights, retention rules, or access model are unresolved, the use case is deferred until they are not.
- No reversible failure mode. If a wrong output sends money, customer commitments, or regulated content with no review step, the test stops here.
Those are not edge cases. They are the main way shiny AI requests waste a quarter. A weighted score compares candidates that are eligible to be tested. It does not rescue a request that still fails basic operating discipline. That is why the best backlog owners write the hard stops down before the first workshop instead of inventing them when the most political request lands.

Show the data behind this chartHide the data behind this chart
- Illustrative use case: draft first-pass replies for tier-one support tickets, with a human agent reviewing every outbound message.
- Value score 4 out of 5 at a 35% weight produces 28 weighted points, because the workflow has high volume and a measurable handle time baseline.
- Feasibility score 4 out of 5 at a 25% weight produces 20 weighted points, because the input data exists and the task is narrow enough to test.
- Risk score 3 out of 5 at a 25% weight produces 15 weighted points, because customer-facing output still needs review and exception routing.
- Time-to-evidence score 5 out of 5 at a 15% weight produces 15 weighted points, because the team can test against live ticket volume inside six weeks.
- Total score is 78 out of 100, which lands in the test band.
Worked example, clearly labeled illustrative
Take a support team that wants AI to draft first-pass replies for tier-one tickets. This is an illustrative example, not a benchmark and not a claim about any company's measured result.
It scores well on value because the workflow usually has volume, repeatable intents, and a baseline the team already tracks: handle time, reopen rate, escalation rate, and CSAT. It scores well on feasibility because the task is constrained and the knowledge sources are reachable. It does not get a perfect risk score because customer-facing language can still cause damage if the model invents policy, misses account nuance, or replies with the wrong tone. Keeping a human in the loop is not optional in the first phase. It gets a strong time-to-evidence score because the team can test against a live ticket slice inside weeks, not quarters.
That lands the example at 78 out of 100. Good enough to test. Not good enough to let the model answer customers on its own. The result is a bounded pilot with acceptance criteria, reviewer ownership, and a stop rule. If the same room tried to score autonomous outbound prospecting, the risk and reversibility bands would collapse and the hard-stop rule would likely kill it before the total mattered.
That is the whole reason to use one scorecard across unlike ideas. The loudest request and the best first request are rarely the same thing.
Decision thresholds for test, defer, or reject
Use bands, not decimals.
- Test: 75 to 100, provided every hard stop clears. Fund a bounded pilot with one owner, one metric pack, one review date, and one stop condition.
- Defer: 55 to 74, or any use case that would otherwise score well but still fails one hard stop. Name the missing evidence and park it until that specific issue is fixed.
- Reject: below 55, or any workflow whose downside stays unacceptable even after guardrails. Write down why so the room can stop relitigating it next month.
A defer is not a polite yes. It means the workflow is not ready to absorb money yet. Most of the time the missing work is boring and worth doing anyway: baseline collection, cleaner permissions, a shorter exception path, or a sharper owner. That is why the best scoring sessions often end with infrastructure work at the top of the queue and the glamorous use case pushed back a month.
McKinsey's 2026 survey adds another reason to be strict here. Forty-four percent of respondents say AI is scaling across the enterprise, but only 37% report any EBIT impact and only 6% qualify as high performers. The companies getting value are not the ones with the most ideas. They are the ones willing to sequence fewer bets and redesign the workflow around them.
How to avoid false precision
A score like 78 is only useful if everyone in the room understands what it is not. It is not a forecast. It is not a benchmark. It is not a guarantee that the vendor demo will survive contact with your data.
Four rules keep the math honest.
- Score in whole numbers only. If someone wants a 3.7, they do not understand the evidence well enough yet.
- Keep written evidence beside every score. If the value score says 4, the supporting baseline and owner should be written in the same row.
- Use the same scoring team across candidates. Changing the room changes the score more than changing the workflow.
- Rescore quarterly. Feasibility rises after shared data pipes exist. Risk can fall after controls improve. Time to evidence usually gets shorter after the first two builds.
This is also where executive discipline matters. IBM's 2025 CEO study says 64% of CEOs admit fear of falling behind drives some technology investment before the value is clear. A scorecard is a defense against that exact impulse, but only if the room agrees in advance that a mediocre score means no deal today.

Show the data behind this diagramHide the data behind this diagram
- Start with four hard-stop checks: named owner, baseline, lawful data path, and reversible failure mode.
- If any hard-stop check fails, defer the use case until the missing condition is fixed.
- If all four checks pass, calculate the weighted score from value, feasibility, risk, and time to evidence.
- Scores from 75 to 100 enter the test band and get a bounded pilot.
- Scores from 55 to 74 enter the defer band and wait for better evidence or cleaner controls.
- Scores below 55 are rejected because the path to value is too weak for the current quarter.
- Every surviving use case is rescored quarterly instead of being treated as permanently ranked.
The questions teams ask once the scoring sheet lands
Should the highest score automatically get budget?+
No. The score narrows the field and tells you where to test first. It does not override a hard stop, and it does not excuse a team from writing acceptance criteria, review dates, and stop rules.
Why is risk weighted as heavily as feasibility?+
Because an easy workflow can still be a stupid first bet. If a wrong answer can send money, promises, or regulated content with no clean review path, the workflow is not ready just because a model can do it in a demo.
How many AI use cases should a company test at once?+
Fewer than the room wants. One to three is normal for a team that wants clean measurement and governance. McKinsey's 2026 data says broad scaling is rising while high performance stays rare, which is a nice way of saying most companies are still spreading bets faster than they are learning from them.
Can we buy the tool first and score the use case during the trial?+
You can, but that is how the trial becomes the decision. Score first, then decide whether the next step is a vendor trial, a small internal build, or nothing at all. If the score says the workflow is not ready, the trial only gives the room a shinier reason to ignore that fact.
Who should own this scorecard inside the company?+
The person who owns the AI backlog and can say no in public. Sometimes that is a CTO. Sometimes it is an operations leader. Sometimes it is the person holding the fractional CAIO engagement. What matters is the mandate, not the title.
What if the use case needs a pilot but not a software purchase?+
Then the score did its job. A good outcome from this exercise is often a small manual test, an eval harness, or one integration cleanup step, not a new annual subscription. The room should be happy when the cheaper next move turns out to be the right one.
read next
Keep going
Put the use cases in one room before the vendors get there
Bring us the five ideas already floating around the company and we will score them against the same sheet, out loud, including the ones that deserve a no. If the honest answer is that the first job is baseline cleanup and not another tool, that is what you will hear.
Starter builds run $1,500 to $2,500 fixed. A two-week production sprint is $5,000 fixed. Retainers start at $5,000 a month.

Written by
Noah Davis · AI Research Writer
I research emerging AI developments and write in-depth articles that give readers the context behind them.
Hiking & nature photography

Written by
Sophie Adams · Technical Writer
I turn complex AI concepts into step-by-step guides readers can follow as they work.
Journaling

Written by
Lucas Brown · AI Explainer Writer
I turn technical AI topics into explainers that show readers how the pieces fit together.
Playing guitar


