Every enterprise AI program eventually reaches the same fork in the road: which provider? Claude, ChatGPT, Microsoft Copilot, Gemini — the shortlist is short, the marketing is loud, and the demos all look impressive. The problem is that the differences that show up in a demo are almost never the differences that matter once real data and real users are involved.
After running this evaluation with leadership teams, we have found the decision comes down to five questions. Benchmark scores are rarely one of them.
1. Zero Data Retention & training
The first question a security or legal team will ask is simple: what happens to our prompts and outputs? There are really two separate guarantees to confirm: whether your data is used to train models, and whether it is stored at all.
Most major providers now exclude enterprise API and business-tier data from training by default. Zero Data Retention, where prompts and responses are not persisted after the request, is usually available, but often on specific tiers or by contractual arrangement rather than as a default toggle. The distinction between “not trained on” and “not stored” matters, and they are not the same thing.
Verify, don’t assume. Retention and training terms differ by provider, plan, and region, and they change. Confirm the current terms in your own agreement before sending regulated data. See Zero Data Retention, explained.
2. US hosting & data residency
For regulated industries (health, finance, government-adjacent work), where inference runs can be a hard requirement. Some providers offer US or regional data residency directly; others deliver it through a cloud partner (for example, running OpenAI models inside your own Azure region, or accessing models through a specific cloud region you control). If residency is a compliance line for you, this can decide the whole evaluation on its own.
Ask for the concrete path your traffic will take, not a marketing page. Residency claims that depend on a future SKU, a partner still in procurement, or a configuration your team has not implemented yet are not residency.
3. Agents & tool use
The center of gravity in enterprise AI has shifted from chat to agents — systems that call tools, retrieve documents, take multi-step actions, and connect to internal systems. Platforms differ meaningfully in how mature their tool-calling, orchestration, and function-execution support is, and in whether they make it easy to build custom agents versus consuming pre-built ones. If your roadmap is agentic, weight this heavily, and evaluate with a real workflow, not a toy demo that never touches your systems of record.
Also ask how the vendor supports human checkpoints, action logging, and permission scoping. An agent platform without auditability is a short path to a stalled security review.
4. Admin, security & audit
The unglamorous layer that determines whether IT will actually approve a rollout: SSO and SCIM, role-based administration, audit logging, DLP integration, and regional controls. A model that scores a point higher on a benchmark but cannot produce an audit trail will lose to one that can, every time.
Include your identity and security teams early. Vendor selection that skips them gets reversed after legal and IT review, a common reason programs stall between pilot and production, consistent with the broader pattern MIT NANDA’s 2025 research documents across enterprise AI initiatives.
5. Total cost: including the cost of switching
Per-seat licensing is easy to compare on a spreadsheet; API-metered usage is not, and blended real-world cost depends heavily on how the system is architected. The larger, often-ignored cost is lock-in: how hard is it to move prompts, evaluations, and integrations to a different model later? Designing for portability is usually cheaper than committing early.
Model a realistic usage mix (chat seats, API calls for production workflows, retrieval volume, and evaluation runs) rather than comparing list prices in isolation. The cheapest seat can become the most expensive platform once agent traffic and logging are included.
How the main options compare
A simplified view of where each option tends to fit. Treat this as a starting point for your own diligence, not a scorecard; capabilities and terms move quickly.
Reviewed July 2026. Provider capabilities, plan names, hosting options, and contractual terms change frequently. Verify every row against current vendor documentation and your own agreement.
| Option | Enterprise ZDR | US / regional hosting | Best-known strength |
|---|---|---|---|
| Claude (Anthropic) | Available on enterprise terms | Direct + via cloud partners | Reasoning, long context, tool use |
| ChatGPT / Azure OpenAI | Available (esp. via Azure) | Azure regions you control | Ecosystem breadth & tooling |
| Microsoft Copilot | Enterprise data protections | Microsoft cloud | Native M365 / workplace integration |
| Google Gemini | Available on enterprise terms | Google Cloud regions | Google Workspace & multimodal |
A practical evaluation sequence
Start from your constraints, not the models. Write down hard requirements first: residency, ZDR, audit, and the one or two workflows that matter most. Let those eliminate options before you ever look at output quality. Then run a short bake-off on your real documents and tasks, with the same evaluation set for every vendor. Score citation quality, refusal behavior, latency, admin fit, and agent tool reliability separately. In most engagements, two or three requirements narrow the field to a single sensible choice, and the “which model is smartest” debate turns out to be the least important part.
The frontier has converged, which changes what a vendor decision is actually about. Stanford HAI’s 2026 AI Index reports that the top four models on the Arena Leaderboard were separated by fewer than 25 Elo points as of March 2026, down from roughly 97 points a year earlier, and concludes that competitive pressure is shifting toward cost, latency, reliability, and domain-specific performance. Ranking these providers by raw capability is therefore close to a coin flip. Vendor choice alone will not make a program successful; workflow redesign and adoption will. But the wrong vendor choice can block production entirely on security or residency grounds. Get the constraint fit right first.
This is exactly the analysis Celadon delivers as part of an AI Decision Sprint: a vendor and model recommendation made in your interest, mapped to your specific use cases and constraints. It also connects to the build-vs-buy decision, because the right answer is often an assembly of vendor models under an architecture you own.
A practical scorecard
When teams get stuck comparing model demos, force the evaluation onto a scorecard the business can defend. Weight the criteria to your constraints: a firm holding client data should not give style and speed equal weight to retention and residency.
- Data handling (~30%). Training exclusion, Zero Data Retention availability, subprocessors, deletion, and where logs live.
- Residency & hosting (~15%). Can inference run where your contracts require, today, on a path you control?
- Workflow fit (~25%). Tool use, agents, connectors to the systems the work already lives in, not a separate chat tab nobody opens.
- Controls (~15%). Admin, SSO, audit logs, retention toggles, and the ability to restrict models or features by team.
- Evaluation ownership (~10%). Can you run your own eval harness against your own cases, or are you stuck with vendor scoreboards?
- Commercial shape (~5%). Seat vs usage vs capacity pricing, and whether costs stay predictable as adoption grows.
Run the scorecard before demos, not after. Demos are useful for confirming fit; they are a poor way to discover requirements. This is the same order Celadon uses inside an AI Decision Sprint: requirements and risk first, vendor second.
A 30-day vendor evaluation, not a feature bake-off
Run the same representative work through each serious option before procurement. Week one defines five to ten tasks, the sensitive data classes involved, and a quality rubric written by the people who own the work. Week two configures the smallest viable environment for each finalist, including identity, permissions, retention, and two integrations, without custom work that would disguise a weak default. Week three has real users complete the tasks while reviewers score output quality, correction time, and policy exceptions. Week four tests export, deletion, audit evidence, and the effort required to switch.
| Gate | Evidence required | Failure condition |
|---|---|---|
| Data handling | Contract terms, retention setting, subprocessor list | A sales answer cannot be reproduced in writing |
| Work quality | Blind scores on representative tasks | The demo task is the only task that works |
| Administration | SSO, roles, audit export, offboarding test | Control depends on individual user behavior |
| Portability | Prompt, data, and evaluation export | Leaving means rebuilding the evidence base |
The output is a short decision record: constraints, test cases, scored results, contractual exceptions, three-year cost, and the reason one option won. That record matters more than a generic quadrant because it can be revisited when pricing, models, or terms change.
Sources
- OpenAI, “Data controls in the OpenAI platform”
- Anthropic, “API and data retention”
- Stanford HAI, “The 2026 AI Index Report: Technical Performance”
- MIT Project NANDA, “The GenAI Divide: State of AI in Business 2025”
Accessed July 2026. Vendor terms and benchmark methodologies change; verify current primary documentation before making a decision.
Decide before you commit
The AI Decision Sprint ranks the opportunity, tests the case, and returns a build or no-build recommendation you can act on.
Explore Decide →