Ask five vendors about AI ROI and you will get five confident numbers. The useful ones come from large-sample research with a named methodology, and they tell a consistent story once you separate vendor-commissioned surveys from independent ones: the returns are real for some organizations, concentrated, and slower than most decks imply.
The numbers that hold up
Across the most-cited studies, a few figures recur:
| Metric | Figure | Source |
|---|---|---|
| Return per $1 invested | ~$3.70 average; ~$10.30 for leaders | IDC / Microsoft, 2024 |
| Time to positive ROI | ~13 months (deploys in under 8) | IDC / Microsoft, 2024 |
| Report first-year ROI | 74% of executives at orgs already using gen AI | Google Cloud, 2025 |
| Typical payback on a use case | 2–4 years; 6% report under 12 months | Deloitte EMEA, 2025 |
| Attribute any EBIT impact to AI | 39% of respondents, most of them under 5% of EBIT | McKinsey, 2025 |
| Are “AI high performers” | ~6% of respondents (5%+ of EBIT and significant value) | McKinsey, 2025 |
| Report no labor-productivity effect yet | 89% of surveyed executives | NBER, 2026 |
These are directional industry benchmarks, not guarantees. Vendor-commissioned figures (IDC/Microsoft, Google Cloud) survey organizations already investing in or using gen AI and tend to look more optimistic than independent samples. Deloitte’s Europe and Middle East survey found most respondents needed two to four years for satisfactory ROI on a typical use case, and only 6% reported payback in under a year. An NBER multi-country firm survey (US, UK, Germany, Australia) found that 89% of executives reported no labor-productivity effect from AI over the prior three years, even as about 69% of firms reported some AI use. Actual ROI depends on use-case selection, data readiness, architecture, and adoption.
Why average ROI figures mislead
An average hides the distribution, and the AI ROI distribution is heavily skewed. In McKinsey’s 2025 State of AI survey, 39% of respondents attributed any EBIT impact at all to AI, and most of those put it under 5% of EBIT. The “AI high performers” who clear 5% and also report significant value are about 6% of respondents. When a benchmark reports “$3.70 back for every $1 invested,” that figure is an average pulled upward by a small group of outliers. An operator who budgets against that number is budgeting against the high performers’ result, not their own likely one.
The practical fix is to stop asking “what is average enterprise AI ROI” and start asking what ROI would look like for one specific workflow, with your own data, volume, and adoption realistically modeled. That is a use-case question, not an industry-level one.
Building a per-use-case financial model
A defensible model for a single workflow needs five inputs, built before you build anything:
- Baseline cost or time, measured, not estimated: hours, error rate, cycle time, or headcount currently allocated to the task.
- Expected volume: how often this workflow runs per month, and whether that volume is stable, seasonal, or growing.
- Unit economics of the AI system: inference or subscription cost per run, plus the amortized cost of integration, evaluation, and ongoing maintenance.
- A realistic adoption curve, not 100% usage from day one. Early ROI is diluted by the ramp period, and a model that ignores this overstates year-one returns.
- Time to value. IDC’s 2024 research with Microsoft found enterprise AI deployments took a median of roughly 13 months to reach positive ROI, even though typical deployment itself took under 8 months. The gap between “deployed” and “paying back” is where financial models most often go wrong, by treating the two dates as one.
Run this model per use case, not once for “AI” as a category. A portfolio of ten use cases will have wildly different paybacks, and averaging them together hides which two or three are actually worth funding.
Leading vs. lagging indicators
Waiting for the P&L to move before knowing whether a workflow is working is too slow to course-correct. Split metrics into two tiers:
- Leading indicators, visible within weeks: active usage rate, override or correction rate, evaluation-set accuracy, and time-to-trust among the people using it.
- Lagging indicators, visible over months: cost saved, revenue influenced, cycle-time reduction, and error-rate change measured against the pre-launch baseline.
If the leading indicators are flat or declining, the lagging indicators will not save the project. Fix the workflow, or kill it, before the quarter closes rather than after.
What high performers do differently
McKinsey’s research is specific about what separates high performers from everyone else, and it is not model access; most organizations now have access to comparable models. High performers are nearly three times as likely to have redesigned the workflow around AI rather than adding it on top of an unchanged process, and they are substantially more likely to use structured human-in-the-loop validation rather than trusting outputs directly. Both are organizational choices, not technology purchases, which is why buying a better model rarely closes the gap on its own.
The seat-count trap
A common but misleading way to justify AI spend is to multiply “seats purchased” by an assumed productivity gain per seat. This overstates ROI in three ways: purchased seats are not active seats, and utilization is rarely 100% or even disclosed; a per-seat productivity assumption borrowed from a vendor benchmark is not validated against your own workflow; and it says nothing about whether the workflow was actually redesigned or just had a tool bolted onto it. A seat-count model can make a rollout look profitable on a slide while the real override rate and active-usage numbers tell a different story. Measure usage and outcome directly, per use case, not licenses multiplied by an assumption.
Why most ROI leaks away
The gap between high performers and the rest is rarely about model quality. Research consistently points to three causes: chasing pilots instead of production, bolting AI onto unchanged workflows, and having no financial model or metrics tied to each use case.
How to capture it
- Prioritize ruthlessly. A handful of high-value use cases beat a spread of experiments.
- Redesign the workflow, do not decorate it: this is one of the clearest separators between high performers and everyone else in McKinsey’s State of AI survey.
- Instrument value early with a financial model and leading indicators per use case.
This is the core of a Celadon AI Decision Sprint: finding the use cases where the ROI is real and the path to production is clear. It is closely tied to build-vs-buy and vendor selection.
Sources
- Microsoft / IDC, “2024 AI opportunity study”
- Google Cloud, “The ROI of generative AI”
- Deloitte, “AI ROI: the paradox of rising investment and elusive returns”
- Yotzov et al., “Firm Data on AI,” NBER Working Paper 34836, 2026
- McKinsey, “The State of AI: Global Survey 2025”
Accessed July 2026. Vendor terms and benchmark methodologies change; verify current primary documentation before making a decision.
Decide before you commit
The AI Decision Sprint ranks the opportunity, tests the case, and returns a build or no-build recommendation you can act on.
Explore Decide →