Celadon — Adoption

From pilot to production: why AI stalls

CeladonUpdated July 20266 min read

Most enterprise AI never leaves the pilot. The cause is rarely the model — it is workflow, ownership, and adoption. Here is how to close the gap.

The most expensive AI project is the one that demos beautifully and never ships. It is also the most common. MIT NANDA’s 2025 State of AI in Business study found that roughly 60% of organizations have evaluated enterprise AI tools, about 20% have piloted one in a real workflow, and only around 5% have reached production with measurable P&L impact. The funnel narrows hard at every stage, and the narrowing is the actual story, not the initial enthusiasm.

StageShare of organizationsSource
Evaluated enterprise AI tools~60%MIT NANDA, 2025
Piloted in a real workflow~20%MIT NANDA, 2025
Reached production, measurable P&L impact~5%MIT NANDA, 2025
Moved 40%+ of AI experiments into production25% of respondentsDeloitte, 2026

A second large-sample view lands in the same place from a different angle. Deloitte’s 2026 State of AI in the Enterprise survey, covering more than 3,200 leaders, found that only 25% of respondents said their organization had moved 40% or more of its AI experiments into production. Fifty-four percent expected to reach that level within three to six months, which is an ambition statement, not a result. The shared pattern across both studies is the gap between trying AI and running it as an owned, measured system.

Why pilots stall

The failure is almost never the model. MIT NANDA’s research points to three structural causes: brittle workflows that break outside the demo script, systems with no contextual learning, which means they don’t retain feedback or improve from correction, and a mismatch between what the tool does and what the job actually requires. In practice, three patterns show up again and again:

Individual tools like ChatGPT and Copilot are widely used. MIT NANDA found that over 80% of organizations have explored or piloted them, and that workers at over 90% of surveyed companies report regular use of personal AI tools for work, but that is mostly personal productivity, not the redesigned, owned, measured workflow that shows up on a P&L. The two are easy to conflate and easy to mistake for progress.

A production-readiness checklist

Before a pilot is allowed to become a production system, it should clear a short, non-negotiable list. If any of these are missing, the launch is a guess, not a decision:

Evaluate before you launch, not after you ship

Most pilots that fail in production failed on day one; the failure just wasn’t visible yet, because nobody had built a way to see it. Evaluation-before-launch means testing against a representative sample of real cases, scoring outputs against a rubric a domain expert agrees with, and setting a quality bar before go-live rather than discovering the bar in production. This matters more for generative and agentic systems than for traditional software, because failures are often plausible-looking rather than obviously broken: a confident wrong answer is more dangerous than an error message.

Adoption metrics to track after launch

Shipping is not the finish line. The metrics that predict whether a system survives its first quarter are usage metrics, not model metrics:

How to kill a pilot on purpose

The least-practiced skill in enterprise AI is killing a pilot deliberately, before it becomes a zombie project that consumes budget and credibility without ever getting formally cancelled. A pilot deserves a clean kill when: the evaluation set shows it can’t clear the quality bar with reasonable additional effort; the workflow owner leaves or the underlying workflow changes; the financial model no longer pencils out at the true cost of running it; or several review cycles in a row show no movement on the metrics that matter. Deciding this in advance, before sunk cost and internal politics accumulate, is what separates organizations that reallocate capital toward the roughly 5% of use cases that work from ones that keep quietly funding the other 95%.

Closing the gap

Moving from pilot to production is a design and operating-model problem: pick use cases with a real owner, redesign the workflow, build evaluation and guardrails in from the start, measure adoption alongside outcomes, and be willing to kill what doesn’t clear the bar.

Adoption is the multiplier on every ROI figure. BCG’s 10-20-70 rule of thumb puts roughly 70% of the effort into people, process, and change, not models.

Celadon’s Build and Operate engagements exist for exactly this, and it starts with the prioritization done in an AI Decision Sprint.

The production gate: proceed, repair, or stop

A pilot review should end with one of three explicit decisions. Proceed when the evaluation bar is met, the workflow owner accepts the operating changes, security is complete, and the cost model still clears the hurdle. Repair when the use case remains valuable but one bounded dependency, such as data quality, integration, or ownership, has a credible fix with a date and owner. Stop when quality cannot clear the bar, economics rely on unrealistic adoption, or the work has no durable owner.

A repair decision needs an expiry date. Without one, “repair” is just a polite way to fund the pilot indefinitely.

The review pack should fit on a few pages: baseline and target outcome, evaluation results by failure category, adoption evidence from intended users, run-rate cost at realistic volume, open risks, and the recommended decision. This makes stopping defensible and proceeding accountable. It also gives Operate a real baseline instead of asking the production team to reconstruct one after launch.

Sources

Accessed July 2026. Vendor terms and benchmark methodologies change; verify current primary documentation before making a decision.

Turn the pattern into a production system

Build covers architecture, integration, evaluation, delivery, and handoff against written acceptance criteria.

Explore Build →