agentclaw

Benchmarks

How Much ROI Companies Actually See From AI

Noah Davis, Sophie Adams, and Lucas Brown · Aug 10, 2026 · 27 min read

Cover card reading: how much ROI companies actually see from AI, with the note that it draws on 62 figures from 38 primary sources covering 2023 to 2026.

TL;DR

  • 84 percent of executives told Deloitte in 2025 they are gaining ROI from AI, yet only 25 percent of AI initiatives delivered the returns CEOs expected in IBM's survey of 2,000 CEOs the same year.
  • The tighter the question, the smaller the number: significant, measurable ROI drops to 15 percent for generative AI and 10 percent for agents in Deloitte's survey of 1,854 senior executives, August to September 2025.
  • Vendor numbers and independent measurements point in opposite directions: GitHub's 2022 study found developers 55 percent faster with Copilot, while METR's 2025 randomized trial measured experienced developers 19 percent slower with AI tools on 246 real tasks.
  • The value concentrates hard: BCG's top 5 percent of 1,250 surveyed companies report 5x the revenue increase of the rest, and PwC finds 20 percent of companies capture 74 percent of AI-driven returns across 1,217 executives surveyed in late 2025.
  • Nobody publishes audited dollar returns, per-project build costs, or small-business ROI. Every ROI number in circulation is somebody describing their own investment in a survey.

Ask executives whether AI is paying off and 84 percent of those investing in it told Deloitte in 2025 that they are gaining ROI. Ask CEOs how many AI initiatives delivered the return they expected and the answer, from IBM's survey of 2,000 of them that same year, is 25 percent. Both numbers are real, both are current, and the distance between them is the most useful thing anyone has measured about AI investment. We pulled 62 figures from 38 primary source documents published between 2023 and 2026, opened every one, and put the claimed returns next to the measured ones. Here is what holds up.

The headline findings

Four numbers from four 2025 surveys, and they do not agree. That disagreement is the story.

of executives investing in AI say they are gaining ROIDeloitte (2025)
84%
of AI initiatives delivered the ROI CEOs expected, across 2,000 CEOsIBM Institute for Business Value (2025)
25%
of organizations qualify as high performers, with 5%+ of EBIT from AIMcKinsey, The State of AI (2025)
6%
of companies report hardly any material value from AI, per BCGBCG (2025)
60%

How to read this benchmark

Every figure below was traced to the document it was printed in, and we opened each source on 2026-08-10 to confirm the number is still there. Figures we could not trace to a primary document got dropped, not footnoted. Where a survey did not state its population, sample, or date range, we recorded that as unknown instead of guessing, because a guessed population is worse than a missing one: nothing downstream can tell it apart from a real one.

Three things to hold onto while you read. First, almost every ROI number in circulation is self-reported. An executive telling a survey their AI investment is paying off is a fact about sentiment, not an audited return. Nobody publishes audited dollar returns across a population of firms, and we say so plainly rather than pretending the survey data is something it is not. Second, vendor numbers about their own products are included here and labeled, because a vendor's claim sitting next to an independent measurement of the same question is one of the most useful pairings this data offers. They never count as independent evidence. Third, the studies disagree, and we do not average them. An average of two incompatible measurements is true of nothing and traceable to nobody. Where two credible sources answer one question differently, we show both and name what differs: the population, the year, the question wording, or the denominator.

The cutoff matters too. Everything here covers 2023 to 2026. Older figures describe a different technology, and we dropped a 2019 McKinsey wave and a spring 2022 MIT Sloan survey for exactly that reason.

Source map showing 62 verified figures drawn from 38 distinct primary documents: 14 from Gartner, 4 from McKinsey, 3 each from Deloitte, IBM, Stack Overflow and the arXiv-METR-MIT group, 2 each from BCG, NBER and the Census Bureau with St. Louis Fed, and 1 each from PwC and the OECD, spanning 2023 to 2026.
The methodology in one picture: who published the 38 primary documents, the largest stated sample from each, and the years covered. Three vendor self-reports ride along, labeled, and never counted as independent.Sources: McKinsey, The State of AI, 2025; Deloitte, State of AI in the Enterprise, 2026; Gartner newsroom, 2026; US Census Bureau, CES-WP-24-16, 2024
Show the data behind this infographic
PublisherPrimary documentsLargest stated sampleYears
Gartner14n=3,412 poll; surveys n=782, 432, 4132024-2026
Deloitte3n=3,2352025-2026
McKinsey4n=1,9932023-2026
IBM Institute for Business Value3n=2,000 CEOs2025-2026
BCG2n=1,800+2025
PwC1n=1,2172025-2026
US Census Bureau / St. Louis Fed2all US nonfarm businesses2023-2026
NBER2n=5,1792023-2024
arXiv / METR / MIT3n=2,2342025-2026
Stack Overflow3n=60,9072023-2025
OECD1global VC deal data2026

Why does 84 percent of claimed ROI shrink to 6 percent of measured impact?

Because the surveys are asking different questions, and the strictness of the question is doing all the work. Line the 2025 findings up by how hard the ROI bar is and the pattern is impossible to miss.

Ask loosely and the numbers are glowing. In Deloitte's US research from 2025, 84 percent of executives investing in AI and gen AI say they are gaining ROI. In IBM's February to April 2025 study of 2,000 CEOs across 33 countries, 52 percent say their organization is realizing value from generative AI beyond cost reduction.

Tighten the question and the floor drops. In McKinsey's mid-2025 survey of 1,993 executives across 105 countries, 39 percent attribute any enterprise EBIT impact to AI at all, and most of those put it under 5 percent of EBIT. In the same IBM study, 25 percent of AI initiatives delivered their expected ROI, and only 16 percent scaled enterprise-wide. In Deloitte's August to September 2025 survey of 1,854 senior executives across Europe and the Middle East, 15 percent see significant, measurable ROI from generative AI. For agentic AI it is 10 percent. Only 6 percent of organizations in that survey reached payback inside a year; most take two to four. And McKinsey's high performers, the companies attributing more than 5 percent of EBIT to AI, are 6 percent of the sample.

So when someone quotes you a single AI ROI number, the first question is not whether it is true. It is which rung of this ladder it was measured on. The feelings numbers and the accounting numbers are both real. They are just measuring different things, and the accounting numbers are one third the size.

Horizontal bar chart of eight 2025 survey findings sorted by strictness of the ROI question, falling from 84 percent who say they are gaining ROI down through 52, 39, 25, 15 and 10 percent to 6 percent high performers and 6 percent reaching payback within a year.
Eight rungs from four surveys. These are different studies with different populations, so no single step is a like-for-like drop; the slope across all of them is the finding.Sources: Deloitte, AI and tech investment ROI, 2025; IBM IBV CEO Study, 2025; McKinsey, The State of AI, 2025; Deloitte, AI ROI: the paradox of rising investment and elusive returns, 2025
Show the data behind this chart
What the survey askedShareSource and population
Say they are gaining ROI from AI84%Deloitte US, executives investing in AI, 2025
CEOs realizing value beyond cost reduction52%IBM, n=2,000 CEOs, Feb-Apr 2025
Attribute any EBIT impact to AI39%McKinsey, n=1,993, Jun-Jul 2025
AI initiatives that delivered expected ROI25%IBM, n=2,000 CEOs, Feb-Apr 2025
Significant, measurable ROI from generative AI15%Deloitte, n=1,854, Aug-Sep 2025
Significant, measurable ROI from agentic AI10%Deloitte, n=1,854, Aug-Sep 2025
High performers: 5%+ of EBIT attributed to AI6%McKinsey, n=1,993, Jun-Jul 2025
Reached AI payback inside one year6%Deloitte, n=1,854, Aug-Sep 2025

What does independent measurement actually find?

Strip out the surveys and look at what gets measured directly, and the picture is positive but far more modest than the pitch decks.

The Federal Reserve Bank of St. Louis surveyed US workers in August and November 2024 and found generative AI users saving 5.4 percent of their work hours, about 2.2 hours in a 40-hour week, and reporting 33 percent higher productivity during the hours they actually use it. Spread across the whole workforce, including the people not using it, they estimate a 1.1 percent lift in aggregate US labor productivity by late 2024 relative to 2022. Real, and an order of magnitude below the headlines.

The experiments say the gains are real but uneven. The NBER working paper by Brynjolfsson, Li, and Raymond gave a generative AI assistant to 5,179 customer support agents at a Fortune 500 software company and measured a 14 percent average productivity gain, rising to 34 percent for novice workers, in data published in 2023. A separate NBER paper from late 2024 puts total self-reported time savings at 1.4 percent of all US work hours. An MIT field experiment by Ju and Aral with 2,234 participants, posted in 2025, found human-AI teams produced 50 percent more ads per worker than human-only teams. And a 2025 arXiv study comparing AI agents to human workers across analysis, engineering, and writing tasks clocked the agents finishing 88.3 percent faster, with quality caveats the paper is honest about.

Measured gains cluster between 1 and 34 percent depending on who and what you measure. The 300 percent numbers live in webinars, not in any study we could trace.

Whose number do you trust: the vendor's or the referee's?

The most instructive pairs in this whole dataset are the ones where a vendor published a number about its own product and somebody independent measured the same question.

GitHub's own 2022 study timed 95 developers on a single scripted task and found the Copilot group 55 percent faster. METR's randomized trial, run February to June 2025, watched 16 experienced open-source developers work through 246 real tasks on their own mature repositories and measured them 19 percent slower with AI tools than without. The developers themselves predicted they would be 24 percent faster. Both studies are real. One timed a synthetic greenfield task in 2022; the other timed real work on codebases the developers knew deeply, three years later. A related independent trial of 228 developers, arXiv 2410.18334, found no statistically significant productivity effect in telemetry at all.

Salesforce's Slack survey of more than 18,000 knowledge workers in 2023 reported 3.6 hours saved per week from workflow automation. The St. Louis Fed's 2024 measurement of actual gen AI users lands at roughly 2.2 hours. Salesforce reported in 2026 that its agent-running retail customers grew holiday sales 4x faster than non-users. Gartner surveyed 413 martech leaders with agent initiatives between June and August 2025 and found 45 percent saying vendor AI agents fail to meet their promised business performance.

None of these pairs is a scandal. A vendor measuring its own product under favorable conditions is a claim, and it should be read as one. If your business case only works at the vendor's number, that is not a business case.

Three side-by-side pairings of vendor claims and independent measurements: GitHub's 55 percent faster versus METR's 19 percent slower, Slack's 3.6 hours saved per week versus the St. Louis Fed's 2.2 hours, and Salesforce's 4x sales growth versus Gartner finding 45 percent of martech leaders let down by vendor agents.
Three questions, six answers. The vendor number and the independent number are not a contradiction to resolve; the pairing itself is the finding worth carrying into your planning.Sources: GitHub, Copilot productivity research, 2022; METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, 2025; Salesforce, business automation and Slack State of Work, 2023; Federal Reserve Bank of St. Louis, 2025; Salesforce, Agentic Enterprise Index, 2026; Gartner, martech leader survey, 2025
Show the data behind this graph
QuestionVendor's numberIndependent numberSource
Do AI coding tools speed developers up?55% faster (GitHub, n=95, one timed task, 2022)19% slower (METR randomized trial, 246 real tasks, 2025)GitHub; METR
How much time does AI save a worker?3.6 hrs/week (Slack survey, n=18,000+, 2023)~2.2 hrs/week, 5.4% of hours (St. Louis Fed, US gen AI users, 2024)Salesforce; St. Louis Fed
Do deployed agents deliver as promised?4x holiday sales growth for agent customers (Salesforce, 2025-26)45% of martech leaders say vendor agents miss promised performance (Gartner, n=413, 2025)Salesforce; Gartner

How often do AI projects fail, stall, or get quietly shelved?

This is the part of the dataset nobody puts in the pitch deck, and it is the part a buyer needs most.

Gartner predicted in July 2024 that 30 percent of generative AI projects would be abandoned after proof of concept by the end of 2025, citing poor data quality, weak risk controls, escalating costs, and unclear business value. In June 2025 it predicted more than 40 percent of agentic AI projects would be canceled by the end of 2027. By May 2026 it added that 40 percent of enterprises running autonomous agents will demote or decommission them by 2027, mostly because governance gaps surfaced only after a production incident. Its November to December 2025 survey of 782 infrastructure and operations leaders found 28 percent of AI use cases fully succeeding and 20 percent failing outright, with the rest stalled somewhere in between.

The most widely quoted failure number is bigger. MIT Project NANDA's 2025 research, based on a review of more than 300 disclosed AI initiatives, 52 organization interviews, and 153 survey responses, found roughly 95 percent of enterprise gen AI pilots producing no measurable P&L impact. We could not open the primary PDF directly, so we cite it as reported by Harvard Business Review and label it accordingly. Note what it measures: pilots, not companies. Deloitte's survey of 3,235 leaders from late 2025 found 25 percent of organizations had moved at least 40 percent of their pilots into production. Both can be true at once, and the difference between them is the unit of measurement. We wrote about the operational half of this in what actually breaks in production automations; this is the investment half of the same story.

One more Gartner survival stat worth its weight: in a Q4 2024 survey of 432 organizations across six countries, 45 percent of high-AI-maturity organizations kept their AI projects operational for at least three years, against 20 percent of low-maturity ones. Projects mostly do not fail because the model was bad. They fail because nobody built the boring scaffolding around it.

The failure rates, side by side

Different units: three are predictions about projects, one is a survey of use cases, one measures pilots. That is why they do not agree, and why we do not average them.

Gen AI projects predicted abandoned after proof of concept, by end 2025

30%

Agentic AI projects predicted canceled by end 2027

40%

Enterprises predicted to demote or decommission autonomous agents by 2027

40%

I&O AI use cases that fail outright, per 782 leaders surveyed

20%

Enterprise gen AI pilots with no measurable P&L impact, per MIT NANDA

95%

The 95 percent figure counts pilots, not companies, and its primary PDF was unreachable for direct verification; it is cited as reported by HBR. The Gartner 30 and 40 percent figures are analyst predictions, not measured outcomes.

Who actually captures the value?

A small minority, and the studies agree on the shape even where they disagree on the size.

BCG's 2025 survey of 1,250 companies sorts them into 5 percent that are "future-built," 35 percent scaling, and 60 percent reaping hardly any material value from AI. The top group reports 5x the revenue increase and 3x the cost reduction of everyone else, and 3.6x the three-year total shareholder return. PwC's October to November 2025 survey of 1,217 senior executives puts it even more starkly: 20 percent of companies capture 74 percent of AI-driven returns. McKinsey's 6 percent of high performers are roughly three times more likely than everyone else to have fundamentally redesigned workflows, to have senior leaders personally owning AI initiatives, and to be scaling agents rather than piloting them.

What separates the winners is boring on purpose. In a separate BCG survey of more than 1,800 executives from 2025, companies concentrating AI spend on fewer, deeper use cases anticipate 2.1x the ROI of companies spreading it across more. PwC finds CEOs with a responsible AI framework and an integration-ready tech environment three times more likely to report meaningful financial returns. The IBM Institute for Business Value study of 600+ Chief AI Officers across 22 countries, run with Oxford Economics in early 2025, ties having an accountable AI executive to 10 percent higher ROI on AI spend, rising toward 36 percent with a centralized operating model.

None of these correlations proves causation, and every one of them comes from a survey rather than an experiment; hold them accordingly. But four independent studies landing on "focus, ownership, and redesigned workflows" while 60 percent of companies get roughly nothing is not noise. It is the strongest pattern in this entire dataset, and it is also precisely the case for doing an honest AI opportunity assessment before writing any build checks: the losing pattern is spreading money across use cases nobody scoped. The capital markets read the same pattern from the other side and moved: twelve billion dollars is now pointed at buying the services layer outright rather than betting on the 60 percent to figure it out.

The spend rises anyway

Here is the paradox that gives Deloitte's ROI report its name: returns are elusive and budgets are climbing regardless.

In Deloitte's August to September 2025 survey of 1,854 executives, 85 percent had increased AI investment in the previous year and 91 percent planned to increase it again. In McKinsey's May 2026 enterprise AI FinOps survey, 93 percent of the 75 qualified enterprise respondents had blown through their AI budgets, and most expected spend to rise at least another 25 percent in the next twelve months. In IBM's compute-cost research, 70 percent of surveyed executives called generative AI a critical driver of rising compute costs, and 100 percent, every single respondent, had canceled or postponed at least one generative AI initiative because of cost.

Zoom out and the totals get silly. Gartner's January 2026 forecast put worldwide AI spending at 2.5 trillion dollars for 2026, 47 percent up year over year, and its May revision nudged that to 2.59 trillion. Its forecast series has the AI services market reaching 609 billion dollars by 2028, agentic AI software spend reaching 985 billion by 2030, and 234 billion of today's enterprise application spend at risk of displacement by agents. The first two of those three sit behind Gartner's subscription paywall rather than a press release, so take them as forecasts you cannot open, which is its own kind of data point. The OECD's February 2026 analysis of global venture capital found 61 percent of worldwide VC, 258.7 billion of 427.1 billion dollars, going to AI-focused firms in 2025, up from 30 percent in 2022.

Read those two paragraphs together and the market's actual bet becomes visible: everyone is paying up front for returns that only a fifth of them have banked so far. That is not automatically irrational. It is exactly what a land grab looks like. But it means the average buyer is funding the learning curve, and the figures above say the tuition is real.

Four cards summarizing value concentration: BCG's top 5 percent with 5x revenue and 3x cost multiples, PwC's 20 percent of companies capturing 74 percent of returns, McKinsey's 6 percent high performers being 3x more likely to redesign workflows, and Gartner's 45 versus 20 percent three-year project survival split by maturity.
Four studies, four samples, one shape. The gap between the leaders and everyone else is not closing; BCG titled its 2025 report The Widening Gap for a reason.Sources: BCG, Are You Generating Value from AI? The Widening Gap, 2025; PwC, Want AI ROI? Go for growth, 2025; McKinsey, The State of AI, 2025; Gartner, AI maturity survey, 2025; BCG, Closing the AI Impact Gap, 2025; IBM IBV and Dubai Future Foundation, CAIO study, 2025
Show the data behind this chart
StudyFindingPopulation and period
BCG, The Widening GapTop 5% report 5x revenue increase, 3x cost reduction, 3.6x 3-year TSR vs the rest; 60% see hardly any value1,250 companies, 2025
PwC AI performance study20% of companies capture 74% of AI-driven returns; strong foundations correlate with 3x likelihood of meaningful returns1,217 senior executives, Oct-Nov 2025
McKinsey State of AI6% are high performers (5%+ EBIT from AI); ~3x more likely to redesign workflows, assign senior ownership, scale agents1,993 respondents, Jun-Jul 2025
Gartner maturity survey45% of high-maturity orgs keep AI projects operational 3+ years vs 20% of low-maturity432 organizations, Q4 2024
BCG, Closing the AI Impact GapFocused portfolios (avg 3.5 use cases) anticipate 2.1x the ROI of scattered ones (avg 6.1)1,800+ executives, 2025
IBM IBV CAIO studyCAIO presence ties to 10% higher ROI on AI spend, up to 36% with centralized operating model600+ AI leaders, 22 countries, Q1 2025

Is adoption actually growing, or is the question just getting easier?

Both, and you need the three series that held their methodology constant to see it.

The US Census Bureau's Business Trends and Outlook Survey asked the same question of all US nonfarm businesses from 2023 to 2025: are you using AI to produce goods or services? The answer climbed from 3.7 percent in September 2023 to 5.4 percent in February 2024 to about 7 percent in early 2025 to about 10 percent in late 2025. Then, in November 2025, the survey reworded the question to cover any business function, and the measured rate jumped to 17 to 20 percent through mid-2026. Same country, same firms, same month; the wording change alone roughly doubled the number. Keep that in your pocket for the next time two adoption stats disagree.

McKinsey's annual State of AI survey, asking global executives whether their organization uses AI in at least one function, went 55 percent in 2023, 72 percent in 2024, 88 percent in 2025. Stack Overflow's developer survey, asking developers whether they use or plan to use AI tools, went 70, 76, 84 over the same three years.

All three slopes point up. The baselines disagree by a factor of nine, because all US businesses, self-selected global executives, and professional developers are three different worlds. We took apart that spread in detail in how many companies have actually deployed AI agents, so we will not re-litigate it here; what matters for the ROI question is that the population you belong to decides which baseline applies to you.

Three small line charts on a shared 0 to 100 percent scale: Census BTOS US firm AI use rising from 3.7 to about 10 percent between September 2023 and late 2025, McKinsey organizational adoption rising from 55 to 88 percent between 2023 and 2025, and Stack Overflow developer AI use rising from 70 to 84 percent over the same years.
Three series that each held their question and population constant for three or more points, which is the minimum for honestly calling something a trend. Two points is just two numbers.Sources: US Census Bureau, CES-WP-24-16, 2024; St. Louis Fed, Measuring AI Adoption Among Firms, 2026; McKinsey, The State of AI in 2023, 2023; McKinsey, The State of AI 2024, 2024; McKinsey, The State of AI, 2025; Stack Overflow Developer Survey, AI section, 2025
Show the data behind this graph
SeriesPopulationPoints
Census BTOS: AI used to produce goods or servicesAll US nonfarm businessesSep 2023: 3.7% | Feb 2024: 5.4% | Q1 2025: ~7% | late 2025: ~10%
McKinsey State of AI: AI in at least one functionGlobal survey respondents, n=1,363 to 1,993 per wave2023: 55% | 2024: 72% | 2025: 88%
Stack Overflow: using or planning to use AI toolsDeveloper survey respondents, up to 60,907 per wave2023: 70% | 2024: 76% | 2025: 84%

The spend keeps climbing while the returns lag

forecast worldwide AI spending in 2026, up 47% year over yearGartner (2026)
$2.5T
of enterprises exceeded their AI budgets, per McKinsey's FinOps surveyMcKinsey (2026)
93%
of surveyed executives canceled or postponed a gen AI initiative over costIBM Institute for Business Value (2025)
100%
of global venture capital went to AI-focused firms in 2025OECD (2026)
61%

The numbers that do not hold up

A benchmark that only shows you the good data is a brochure. Here is what we threw out, and what to watch for when you meet these numbers in the wild.

We dropped a widely repeated claim that 64 percent of larger companies see moderate or significant AI ROI against 11 percent of smaller ones, because the Deloitte page it is always attributed to does not contain it. We checked the page directly on 2026-08-10. The 84 percent figure on that same page is real, which is presumably how the fake one keeps hitching a ride. We dropped a McKinsey statistic about companies committing 36 percent of digital budgets to AI, which circulates in aggregator posts but appears on none of the five McKinsey pages it gets attributed to, including the PDF. We dropped an OECD claim of a 4 percent firm-level productivity lift across 12,000 EU firms that we could not confirm on any OECD page. And we dropped every figure whose trail ended at a blog citing a blog citing a landing page: that laundering chain is where most of the "80 percent of enterprises" statistics you have read actually come from.

Two structural warnings beat any single dropped number. First, question wording moves results by a factor of two: the Census BTOS jump from 10 to 17-20 percent on a reworded question is the cleanest demonstration on record, from the most rigorous survey in the field. Second, most trend claims in this space are two points wearing a trend costume. The same measurement on the same population at three or more points exists for exactly three series we could find, and all three are in the chart above.

What has nobody actually measured?

The absences in this dataset are as useful as the figures, because they mark the places where anyone quoting a precise number is making it up.

Nobody publishes per-project AI build cost distributions. Aggregate budgets and trillion-dollar forecasts are everywhere; what individual projects actually cost, as a published distribution, exists nowhere we could find. Nobody publishes audited dollar returns. Every ROI percentage in this post is a survey respondent characterizing their own investment; not one line of it has passed through an auditor. Nobody measures small-business returns: the samples skew hard toward enterprises, with PwC's panel 76 percent companies above a billion dollars in revenue, while the Census tracks small-firm adoption but not small-firm ROI. Nobody has run a longitudinal same-firm ROI panel; every consultancy trend is a fresh cross-section of different respondents, and the only true panel, the Census BTOS, measures adoption rather than returns. Nobody prices the write-offs: cancellation rates exist, but the sunk cost per canceled project is published nowhere. And the fractional and outside-expert model our corner of the market runs on is unmeasured too; IBM's CAIO data covers full-time officers only.

If a vendor quotes you a number from one of these six holes, ask for the study. There is not one.

What this means if you are deciding about AI this quarter

Read as a decision input rather than a spectacle, the 62 figures compress to four moves.

Budget on the accounting numbers, not the sentiment numbers. The planning-grade range for "will this deliver expected ROI" is 25 to 39 percent on current evidence, not 84. If the business case needs the optimistic tail to pencil out, it does not pencil out. And expect payback in two to four years, because in Deloitte's 2025 data only 6 percent of organizations got there inside one.

Copy the only pattern that repeats. Fewer use cases, chosen against measured baselines, with a named senior owner and the workflow actually redesigned around the system: that description shows up independently in BCG, McKinsey, PwC, and IBM's data. This is the entire reason we start engagements with an assessment of where automation actually pays instead of a build quote; scoping is the step the 60 percent skipped.

Demand independent numbers for any claim that decides a purchase. The vendor-versus-referee pairs above are the base rate: assume the marketing figure is the ceiling, and ask what the METR-style measurement of your use case would say. Where a claim cannot be checked, weight it at zero.

And measure your own pilots like the studies do, with a baseline before the system lands and the same metric after. That is what we build into every deployment by default, because the alternative is becoming one more respondent who tells next year's survey they are "gaining ROI" without a number to show for it. The two-week production sprint exists precisely so the measurement starts in week three, not in quarter three; the full price ladder is on the pricing page.

One question, two answers: the contradictions worth knowing

None of these pairs is an error. Each resolves into a definitional difference, and the difference is usually more useful than either number.

Are companies getting ROI from AI?

One answer
84% say yes (Deloitte US, 2025)
The other answer
25% of initiatives delivered expected ROI (IBM, n=2,000 CEOs, 2025)
What actually differs
Sentiment about any return vs initiative-level accounting against expectations

Do AI coding tools make developers faster?

One answer
55% faster (GitHub, n=95, 2022)
The other answer
19% slower (METR, 246 real tasks, 2025)
What actually differs
One synthetic task by newer users vs real work on mature repos by experts

What share of US firms use AI?

One answer
~10% (Census BTOS, late 2025)
The other answer
88% (McKinsey, n=1,993, 2025)
What actually differs
All US businesses vs self-selected global executives; and function-level wording

Do pilots reach production?

One answer
25% of orgs moved 40%+ of pilots to production (Deloitte, n=3,235, 2025)
The other answer
~95% of pilots show no P&L impact (MIT NANDA, 2025)
What actually differs
Unit of measurement: organizations vs individual pilots

How many US firms use AI, take two?

One answer
~10% (BTOS, 'producing goods or services', late 2025)
The other answer
17-20% (BTOS, 'any business function', Dec 2025 on)
What actually differs
Nothing but the question wording; same survey, same firms

Where sources contradict each other, we publish both and name the difference. Averaging incompatible measurements produces a number that is true of nothing.

The questions people actually ask about AI ROI

What is the average ROI of AI?+

There is no honest single average, because the published numbers measure different things. Self-reported "we are gaining ROI" runs at 84 percent in Deloitte's 2025 US data; initiatives delivering their expected return run at 25 percent in IBM's 2025 survey of 2,000 CEOs; and organizations with significant, measurable ROI run at 6 to 15 percent depending on the bar. Any single "average AI ROI" figure you meet is one rung of that ladder presented as the whole thing.

What percentage of AI projects fail?+

Between 20 and 40 percent by most project-level measures: Gartner's surveys and predictions put outright I&O failure at 20 percent, post-PoC abandonment at 30 percent, and agentic cancellations above 40 percent by 2027. The famous 95 percent figure, from MIT's 2025 NANDA research, counts pilots with no measurable P&L impact, which is a much easier bar to fail. Check the unit before you quote any of them.

How long does it take for AI investment to pay back?+

Two to four years for most organizations that get there at all, per Deloitte's August to September 2025 survey of 1,854 senior executives. Only 6 percent reported reaching satisfactory ROI within a year. Anyone promising payback in a quarter is quoting the tail of the distribution as if it were the middle.

Do AI coding assistants actually make developers faster?+

It depends on who and what you measure, and the honest answer is narrower than either camp admits. GitHub's own 2022 study found 55 percent faster completion of a scripted task by 95 developers. METR's 2025 randomized trial found experienced open-source developers 19 percent slower with AI tools across 246 real tasks on their own repositories, and a separate 228-developer trial found no significant effect in telemetry. New code and unfamiliar territory favor the tools; deep expertise on mature codebases currently does not.

Why do most AI pilots never produce measurable value?+

The data points at scaffolding, not models. McKinsey's 2025 high performers are about 3x more likely to have redesigned the workflow around the system; Gartner ties three-year project survival to AI maturity, 45 versus 20 percent; and Deloitte finds only 21 percent of organizations planning agent deployments have a mature governance model. Pilots mostly die of missing ownership, unredesigned processes, and ungoverned production incidents, not of insufficient intelligence.

How much time does AI actually save a worker per week?+

About 2.2 hours a week for US workers who use generative AI, which is 5.4 percent of work hours, per the St. Louis Fed's 2024 surveys. Spread across all workers including non-users, the NBER puts total savings at 1.4 percent of US work hours. Vendor surveys run higher, with Slack's 2023 figure at 3.6 hours; treat the gap between those numbers as the vendor premium.

Want your ROI measured instead of surveyed?

We scope the two or three workflows where the numbers above say the returns actually live, then build with a baseline in place so you know what changed. No slide decks, and nothing you cannot verify.

A one-off starter build runs $1,500 to $2,500 fixed. A two-week production sprint is $5,000 fixed.

Share thison Xon LinkedIn

Written by

Noah Davis · AI Research Writer

I research emerging AI developments and write in-depth articles that give readers the context behind them.

Hiking & nature photography

Written by

Sophie Adams · Technical Writer

I turn complex AI concepts into step-by-step guides readers can follow as they work.

Journaling

Written by

Lucas Brown · AI Explainer Writer

I turn technical AI topics into explainers that show readers how the pieces fit together.

Playing guitar

the memo

One useful automation idea, when we have one.

What we'd actually build, minus the hype and the pitch. We write when there is something worth sending, and you can reply to any email to come off the list.

Book audit