Celadon: Business value

The ROI of enterprise AI

CeladonUpdated July 20265 min read

The returns are real — about $3.70 back for every $1 on average, and far more for leaders. Here is what the benchmarks say, and why most of the value leaks away.

Ask five vendors about AI ROI and you will get five confident numbers. The useful ones come from large-sample research with a named methodology, and they tell a consistent story once you separate vendor-commissioned surveys from independent ones: the returns are real for some organizations, concentrated, and slower than most decks imply.

The numbers that hold up

Across the most-cited studies, a few figures recur:

MetricFigureSource
Return per $1 invested~$3.70 average; ~$10.30 for leadersIDC / Microsoft, 2024
Time to positive ROI~13 months (deploys in under 8)IDC / Microsoft, 2024
Report first-year ROI74% of executives at orgs already using gen AIGoogle Cloud, 2025
Typical payback on a use case2–4 years; 6% report under 12 monthsDeloitte EMEA, 2025
Attribute any EBIT impact to AI39% of respondents, most of them under 5% of EBITMcKinsey, 2025
Are “AI high performers”~6% of respondents (5%+ of EBIT and significant value)McKinsey, 2025
Report no labor-productivity effect yet89% of surveyed executivesNBER, 2026

These are directional industry benchmarks, not guarantees. Vendor-commissioned figures (IDC/Microsoft, Google Cloud) survey organizations already investing in or using gen AI and tend to look more optimistic than independent samples. Deloitte’s Europe and Middle East survey found most respondents needed two to four years for satisfactory ROI on a typical use case, and only 6% reported payback in under a year. An NBER multi-country firm survey (US, UK, Germany, Australia) found that 89% of executives reported no labor-productivity effect from AI over the prior three years, even as about 69% of firms reported some AI use. Actual ROI depends on use-case selection, data readiness, architecture, and adoption.

Why average ROI figures mislead

An average hides the distribution, and the AI ROI distribution is heavily skewed. In McKinsey’s 2025 State of AI survey, 39% of respondents attributed any EBIT impact at all to AI, and most of those put it under 5% of EBIT. The “AI high performers” who clear 5% and also report significant value are about 6% of respondents. When a benchmark reports “$3.70 back for every $1 invested,” that figure is an average pulled upward by a small group of outliers. An operator who budgets against that number is budgeting against the high performers’ result, not their own likely one.

The practical fix is to stop asking “what is average enterprise AI ROI” and start asking what ROI would look like for one specific workflow, with your own data, volume, and adoption realistically modeled. That is a use-case question, not an industry-level one.

Building a per-use-case financial model

A defensible model for a single workflow needs five inputs, built before you build anything:

Run this model per use case, not once for “AI” as a category. A portfolio of ten use cases will have wildly different paybacks, and averaging them together hides which two or three are actually worth funding.

Leading vs. lagging indicators

Waiting for the P&L to move before knowing whether a workflow is working is too slow to course-correct. Split metrics into two tiers:

If the leading indicators are flat or declining, the lagging indicators will not save the project. Fix the workflow, or kill it, before the quarter closes rather than after.

What high performers do differently

McKinsey’s research is specific about what separates high performers from everyone else, and it is not model access; most organizations now have access to comparable models. High performers are nearly three times as likely to have redesigned the workflow around AI rather than adding it on top of an unchanged process, and they are substantially more likely to use structured human-in-the-loop validation rather than trusting outputs directly. Both are organizational choices, not technology purchases, which is why buying a better model rarely closes the gap on its own.

The seat-count trap

A common but misleading way to justify AI spend is to multiply “seats purchased” by an assumed productivity gain per seat. This overstates ROI in three ways: purchased seats are not active seats, and utilization is rarely 100% or even disclosed; a per-seat productivity assumption borrowed from a vendor benchmark is not validated against your own workflow; and it says nothing about whether the workflow was actually redesigned or just had a tool bolted onto it. A seat-count model can make a rollout look profitable on a slide while the real override rate and active-usage numbers tell a different story. Measure usage and outcome directly, per use case, not licenses multiplied by an assumption.

Why most ROI leaks away

The gap between high performers and the rest is rarely about model quality. Research consistently points to three causes: chasing pilots instead of production, bolting AI onto unchanged workflows, and having no financial model or metrics tied to each use case.

How to capture it

This is the core of a Celadon AI Decision Sprint: finding the use cases where the ROI is real and the path to production is clear. It is closely tied to build-vs-buy and vendor selection.

Sources

Accessed July 2026. Vendor terms and benchmark methodologies change; verify current primary documentation before making a decision.

Decide before you commit

The AI Decision Sprint ranks the opportunity, tests the case, and returns a build or no-build recommendation you can act on.

Explore Decide →