The most expensive AI project is the one that demos beautifully and never ships. It is also the most common. MIT NANDA’s 2025 State of AI in Business study found that roughly 60% of organizations have evaluated enterprise AI tools, about 20% have piloted one in a real workflow, and only around 5% have reached production with measurable P&L impact. The funnel narrows hard at every stage, and the narrowing is the actual story, not the initial enthusiasm.
| Stage | Share of organizations | Source |
|---|---|---|
| Evaluated enterprise AI tools | ~60% | MIT NANDA, 2025 |
| Piloted in a real workflow | ~20% | MIT NANDA, 2025 |
| Reached production, measurable P&L impact | ~5% | MIT NANDA, 2025 |
| Moved 40%+ of AI experiments into production | 25% of respondents | Deloitte, 2026 |
A second large-sample view lands in the same place from a different angle. Deloitte’s 2026 State of AI in the Enterprise survey, covering more than 3,200 leaders, found that only 25% of respondents said their organization had moved 40% or more of its AI experiments into production. Fifty-four percent expected to reach that level within three to six months, which is an ambition statement, not a result. The shared pattern across both studies is the gap between trying AI and running it as an owned, measured system.
Why pilots stall
The failure is almost never the model. MIT NANDA’s research points to three structural causes: brittle workflows that break outside the demo script, systems with no contextual learning, which means they don’t retain feedback or improve from correction, and a mismatch between what the tool does and what the job actually requires. In practice, three patterns show up again and again:
- Bolt-on, not redesign. AI is dropped into an unchanged workflow, so it saves seconds instead of transforming the process. McKinsey’s 2025 State of AI survey defines AI high performers as respondents attributing 5% or more of EBIT to AI and reporting significant value from it, about 6% of the sample, and finds they are nearly three times as likely as others to say they have fundamentally redesigned individual workflows. Redesign is still uncommon: in McKinsey’s earlier 2025 report, 21% of respondents using gen AI said their organization had fundamentally redesigned even some workflows.
- No owner, no metrics. A successful demo has no business owner, no success criteria, and no financial model, so it cannot graduate. Nobody is accountable for turning “interesting” into “shipped.”
- Governance and evaluation as an afterthought. Data access, security review, and a way to measure output quality are left until the end, where they quietly kill the launch. Worse, they may ship broken and get pulled a month later.
Individual tools like ChatGPT and Copilot are widely used. MIT NANDA found that over 80% of organizations have explored or piloted them, and that workers at over 90% of surveyed companies report regular use of personal AI tools for work, but that is mostly personal productivity, not the redesigned, owned, measured workflow that shows up on a P&L. The two are easy to conflate and easy to mistake for progress.
A production-readiness checklist
Before a pilot is allowed to become a production system, it should clear a short, non-negotiable list. If any of these are missing, the launch is a guess, not a decision:
- A named business owner accountable for the outcome, not just the vendor relationship.
- A financial model tying the workflow to a specific cost or revenue line, with a baseline measured before launch.
- An evaluation set: real examples with known-good answers, run before every material change, not only before launch.
- A human-in-the-loop path for the failure modes that matter, with a clear escalation rule.
- Data access and security sign-off completed, not pending.
- A rollback plan that does not depend on the same team that built the system.
Evaluate before you launch, not after you ship
Most pilots that fail in production failed on day one; the failure just wasn’t visible yet, because nobody had built a way to see it. Evaluation-before-launch means testing against a representative sample of real cases, scoring outputs against a rubric a domain expert agrees with, and setting a quality bar before go-live rather than discovering the bar in production. This matters more for generative and agentic systems than for traditional software, because failures are often plausible-looking rather than obviously broken: a confident wrong answer is more dangerous than an error message.
Adoption metrics to track after launch
Shipping is not the finish line. The metrics that predict whether a system survives its first quarter are usage metrics, not model metrics:
- Active use rate: the share of intended users who actually use it weekly, not who were simply given a license.
- Override or rejection rate: how often a human discards or corrects the AI’s output. A high or rising rate is an early warning, not noise.
- Time-to-trust: how long it takes a new user to stop double-checking every output. If it never drops, the system has a quality problem or a communication problem.
- Escalation volume: how often the human-in-the-loop path actually gets used, and whether that trend is moving toward or away from expectations.
How to kill a pilot on purpose
The least-practiced skill in enterprise AI is killing a pilot deliberately, before it becomes a zombie project that consumes budget and credibility without ever getting formally cancelled. A pilot deserves a clean kill when: the evaluation set shows it can’t clear the quality bar with reasonable additional effort; the workflow owner leaves or the underlying workflow changes; the financial model no longer pencils out at the true cost of running it; or several review cycles in a row show no movement on the metrics that matter. Deciding this in advance, before sunk cost and internal politics accumulate, is what separates organizations that reallocate capital toward the roughly 5% of use cases that work from ones that keep quietly funding the other 95%.
Closing the gap
Moving from pilot to production is a design and operating-model problem: pick use cases with a real owner, redesign the workflow, build evaluation and guardrails in from the start, measure adoption alongside outcomes, and be willing to kill what doesn’t clear the bar.
Adoption is the multiplier on every ROI figure. BCG’s 10-20-70 rule of thumb puts roughly 70% of the effort into people, process, and change, not models.
Celadon’s Build and Operate engagements exist for exactly this, and it starts with the prioritization done in an AI Decision Sprint.
The production gate: proceed, repair, or stop
A pilot review should end with one of three explicit decisions. Proceed when the evaluation bar is met, the workflow owner accepts the operating changes, security is complete, and the cost model still clears the hurdle. Repair when the use case remains valuable but one bounded dependency, such as data quality, integration, or ownership, has a credible fix with a date and owner. Stop when quality cannot clear the bar, economics rely on unrealistic adoption, or the work has no durable owner.
A repair decision needs an expiry date. Without one, “repair” is just a polite way to fund the pilot indefinitely.
The review pack should fit on a few pages: baseline and target outcome, evaluation results by failure category, adoption evidence from intended users, run-rate cost at realistic volume, open risks, and the recommended decision. This makes stopping defensible and proceeding accountable. It also gives Operate a real baseline instead of asking the production team to reconstruct one after launch.
Sources
- MIT Project NANDA, “The GenAI Divide: State of AI in Business 2025”
- Deloitte AI Institute, “State of AI in the Enterprise 2026”
- McKinsey, “The State of AI: Global Survey 2025”
- McKinsey, “The State of AI: How Organizations Are Rewiring to Capture Value,” March 2025
- BCG, “The Leader’s Guide to Transforming with AI”
Accessed July 2026. Vendor terms and benchmark methodologies change; verify current primary documentation before making a decision.
Turn the pattern into a production system
Build covers architecture, integration, evaluation, delivery, and handoff against written acceptance criteria.
Explore Build →