Celadon: Vendor selection

Choosing the right AI vendor

CeladonUpdated July 20267 min read

The differences that matter in an enterprise AI provider are rarely about benchmark scores. They are about data handling, hosting, security, agents, and cost. Here is how to run the decision.

Every enterprise AI program eventually reaches the same fork in the road: which provider? Claude, ChatGPT, Microsoft Copilot, Gemini — the shortlist is short, the marketing is loud, and the demos all look impressive. The problem is that the differences that show up in a demo are almost never the differences that matter once real data and real users are involved.

After running this evaluation with leadership teams, we have found the decision comes down to five questions. Benchmark scores are rarely one of them.

1. Zero Data Retention & training

The first question a security or legal team will ask is simple: what happens to our prompts and outputs? There are really two separate guarantees to confirm: whether your data is used to train models, and whether it is stored at all.

Most major providers now exclude enterprise API and business-tier data from training by default. Zero Data Retention, where prompts and responses are not persisted after the request, is usually available, but often on specific tiers or by contractual arrangement rather than as a default toggle. The distinction between “not trained on” and “not stored” matters, and they are not the same thing.

Verify, don’t assume. Retention and training terms differ by provider, plan, and region, and they change. Confirm the current terms in your own agreement before sending regulated data. See Zero Data Retention, explained.

2. US hosting & data residency

For regulated industries (health, finance, government-adjacent work), where inference runs can be a hard requirement. Some providers offer US or regional data residency directly; others deliver it through a cloud partner (for example, running OpenAI models inside your own Azure region, or accessing models through a specific cloud region you control). If residency is a compliance line for you, this can decide the whole evaluation on its own.

Ask for the concrete path your traffic will take, not a marketing page. Residency claims that depend on a future SKU, a partner still in procurement, or a configuration your team has not implemented yet are not residency.

3. Agents & tool use

The center of gravity in enterprise AI has shifted from chat to agents — systems that call tools, retrieve documents, take multi-step actions, and connect to internal systems. Platforms differ meaningfully in how mature their tool-calling, orchestration, and function-execution support is, and in whether they make it easy to build custom agents versus consuming pre-built ones. If your roadmap is agentic, weight this heavily, and evaluate with a real workflow, not a toy demo that never touches your systems of record.

Also ask how the vendor supports human checkpoints, action logging, and permission scoping. An agent platform without auditability is a short path to a stalled security review.

4. Admin, security & audit

The unglamorous layer that determines whether IT will actually approve a rollout: SSO and SCIM, role-based administration, audit logging, DLP integration, and regional controls. A model that scores a point higher on a benchmark but cannot produce an audit trail will lose to one that can, every time.

Include your identity and security teams early. Vendor selection that skips them gets reversed after legal and IT review, a common reason programs stall between pilot and production, consistent with the broader pattern MIT NANDA’s 2025 research documents across enterprise AI initiatives.

5. Total cost: including the cost of switching

Per-seat licensing is easy to compare on a spreadsheet; API-metered usage is not, and blended real-world cost depends heavily on how the system is architected. The larger, often-ignored cost is lock-in: how hard is it to move prompts, evaluations, and integrations to a different model later? Designing for portability is usually cheaper than committing early.

Model a realistic usage mix (chat seats, API calls for production workflows, retrieval volume, and evaluation runs) rather than comparing list prices in isolation. The cheapest seat can become the most expensive platform once agent traffic and logging are included.

How the main options compare

A simplified view of where each option tends to fit. Treat this as a starting point for your own diligence, not a scorecard; capabilities and terms move quickly.

Reviewed July 2026. Provider capabilities, plan names, hosting options, and contractual terms change frequently. Verify every row against current vendor documentation and your own agreement.

OptionEnterprise ZDRUS / regional hostingBest-known strength
Claude (Anthropic)Available on enterprise termsDirect + via cloud partnersReasoning, long context, tool use
ChatGPT / Azure OpenAIAvailable (esp. via Azure)Azure regions you controlEcosystem breadth & tooling
Microsoft CopilotEnterprise data protectionsMicrosoft cloudNative M365 / workplace integration
Google GeminiAvailable on enterprise termsGoogle Cloud regionsGoogle Workspace & multimodal

A practical evaluation sequence

Start from your constraints, not the models. Write down hard requirements first: residency, ZDR, audit, and the one or two workflows that matter most. Let those eliminate options before you ever look at output quality. Then run a short bake-off on your real documents and tasks, with the same evaluation set for every vendor. Score citation quality, refusal behavior, latency, admin fit, and agent tool reliability separately. In most engagements, two or three requirements narrow the field to a single sensible choice, and the “which model is smartest” debate turns out to be the least important part.

The frontier has converged, which changes what a vendor decision is actually about. Stanford HAI’s 2026 AI Index reports that the top four models on the Arena Leaderboard were separated by fewer than 25 Elo points as of March 2026, down from roughly 97 points a year earlier, and concludes that competitive pressure is shifting toward cost, latency, reliability, and domain-specific performance. Ranking these providers by raw capability is therefore close to a coin flip. Vendor choice alone will not make a program successful; workflow redesign and adoption will. But the wrong vendor choice can block production entirely on security or residency grounds. Get the constraint fit right first.

This is exactly the analysis Celadon delivers as part of an AI Decision Sprint: a vendor and model recommendation made in your interest, mapped to your specific use cases and constraints. It also connects to the build-vs-buy decision, because the right answer is often an assembly of vendor models under an architecture you own.

A practical scorecard

When teams get stuck comparing model demos, force the evaluation onto a scorecard the business can defend. Weight the criteria to your constraints: a firm holding client data should not give style and speed equal weight to retention and residency.

Run the scorecard before demos, not after. Demos are useful for confirming fit; they are a poor way to discover requirements. This is the same order Celadon uses inside an AI Decision Sprint: requirements and risk first, vendor second.

A 30-day vendor evaluation, not a feature bake-off

Run the same representative work through each serious option before procurement. Week one defines five to ten tasks, the sensitive data classes involved, and a quality rubric written by the people who own the work. Week two configures the smallest viable environment for each finalist, including identity, permissions, retention, and two integrations, without custom work that would disguise a weak default. Week three has real users complete the tasks while reviewers score output quality, correction time, and policy exceptions. Week four tests export, deletion, audit evidence, and the effort required to switch.

GateEvidence requiredFailure condition
Data handlingContract terms, retention setting, subprocessor listA sales answer cannot be reproduced in writing
Work qualityBlind scores on representative tasksThe demo task is the only task that works
AdministrationSSO, roles, audit export, offboarding testControl depends on individual user behavior
PortabilityPrompt, data, and evaluation exportLeaving means rebuilding the evidence base

The output is a short decision record: constraints, test cases, scored results, contractual exceptions, three-year cost, and the reason one option won. That record matters more than a generic quadrant because it can be revisited when pricing, models, or terms change.

Sources

Accessed July 2026. Vendor terms and benchmark methodologies change; verify current primary documentation before making a decision.

Decide before you commit

The AI Decision Sprint ranks the opportunity, tests the case, and returns a build or no-build recommendation you can act on.

Explore Decide →