For a firm holding client financials, health records, or privileged communications, the AI conversation starts and ends with confidentiality. The upside is real: drafting, research, and review all move faster. But it only counts if you can put the tool in front of a client and a regulator without flinching. The good news: the controls you need are well understood, they are checkable in a contract, and they should decide the platform before a single seat is purchased.
Classify data before you classify tools
Not every document deserves the same handling, and skipping this step is why so many AI rollouts either move too slowly or move too fast in the wrong places. Before evaluating a single vendor, sort what the firm holds into tiers: public (marketing material, published rates), internal (working drafts, internal memos), confidential (client financials, contracts, unredacted correspondence), and regulated (PII, PHI, privileged communications, anything covered by a specific statute or professional-conduct rule). The tier should decide what can touch a given tool at all. Some regulated categories may never leave a private, contractually bound environment, no matter how favorable a vendor’s general terms look.
What to confirm before any file goes in
- Training. Is your data excluded from model training by default, and is that contractual, not just a settings toggle? Treat “default” as marketing and contract language as the only thing that survives a vendor’s next terms-of-service update.
- Retention & deletion. Are prompts and outputs stored, for how long, and can you require Zero Data Retention? Distinguish operational logs (needed for abuse monitoring) from prompt-and-output retention (usually not needed at all). See Zero Data Retention, explained.
- Data processing agreement. Is there a DPA naming sub-processors, specifying a breach-notification timeline in hours rather than “promptly,” covering cross-border transfer, and stating your obligations to your own clients?
- Residency. Where is data processed and stored, and does that satisfy your jurisdiction’s data-protection rules and any client-specific contractual restrictions?
- Firm controls. Admin dashboard, permission levels, immediate user removal, and knowledge that stays with the firm (not a departing employee’s personal account) when staff leave.
The most common breach is not technical. It is a staff member pasting a sensitive file into a personal AI account. A governed platform plus a clear usage policy closes that gap faster than any feature.
Shadow AI is already inside the firm
MIT’s 2025 NANDA research on the “GenAI Divide” found that employees are already extracting real, measurable value from consumer AI tools on their own initiative, while the large majority of formal enterprise pilots never reach production use. Read together, the implication for a firm handling client data is blunt: staff are not waiting for a sanctioned platform. By the time a governed tool is rolled out, a workaround habit already exists: a paralegal summarizing a filing in a personal account, an associate pasting a client’s financials into a free tool to speed up a first draft. None of that shows up in a vendor comparison, and none of it is covered by any DPA the firm has signed.
Banning AI in a memo does not fix this; that policy is unenforceable and gets ignored quietly. What works is naming an approved tool quickly, making it genuinely easier to use for the tasks staff are already trying to solve on their own, and pairing it with a short, specific usage policy that says what can be uploaded, what cannot, and who to ask when someone is unsure.
Client consent and contract language
Engagement letters and master service agreements drafted before AI adoption are silent on the question, and that silence is its own risk. Add language that discloses AI tools may assist in preparing work product; states plainly that a licensed professional reviews and remains responsible for all client-facing output; identifies the categories of AI subprocessors the firm uses, or commits to disclosing them on request; and gives clients a documented way to object or request an alternate process for a specific matter. In practice, clients rarely object once the safeguards are visible in writing; it is silence on the topic, not the technology itself, that erodes trust.
Evaluating quality without leaking client data
A serious AI program gets measured against a running evaluation set of real tasks, but building that set out of actual client files recreates the exact exposure the rest of the program exists to prevent. Build evaluation examples from synthesized or heavily redacted material that preserves the shape and difficulty of a real task without the client’s actual identifiers, and treat any evaluation log that does retain real content as confidential client data in its own right: access-controlled, retention-limited, and explicitly excluded from any vendor’s product-improvement pipeline. If a reviewer needs the original document to judge an answer, that review belongs inside the firm’s own systems, not a third-party dashboard.
The vendor questionnaire
A short, written questionnaire, answered before a contract is signed, protects a firm more than any feature comparison:
- Is exclusion from model training contractual, and does it cover agent and connector use, not just chat?
- What is the exact retention window, and can it be reduced to zero on request?
- Who are the named sub-processors, and where is data processed and stored?
- What is the breach-notification commitment, stated in hours?
- What independent audit or attestation (SOC 2 Type II or ISO 27001) can they produce today, not “by year end”?
- What happens to firm data and any custom evaluation sets if the firm cancels?
Answers that are vague, verbal, or “in progress” are still answers. Move to the next vendor.
Requirements first, then seats
Weight the decision toward what protects the firm. A practical split: data handling and governance ~30%, workflow fit ~25%, user and firm controls ~20%, document handling ~15%, adoption and integration ~10%. Features matter, but they matter less than a platform you can defend to a client and a reviewer.
Keep a human in the loop
AI should assist with preparation and review, not replace professional judgment, partner sign-off, or the firm’s conclusion. McKinsey’s 2025 State of AI survey found that high performers are more likely than peers to have defined processes for deciding when model outputs need human validation. The gain comes from faster preparation, not from removing accountability. A usage policy that states clearly what staff may upload, and exactly when partner review is required before anything reaches a client, is as important as the platform itself.
This is the data-handling and governance review at the center of a Celadon AI Decision Sprint. See how it maps to firms on our Accounting & Professional Services page, or read AI for professional-services firms.
The client-data decision record
Policies become usable when each sensitive workflow leaves a small, reviewable record. For a proposed AI use, capture: the client or matter category; data classes entering the system; provider, product, region, and retention mode; contractual authority; the human reviewer; the output’s permitted uses; and the deletion or export path. Link the record to the exact vendor terms reviewed and date them.
This is deliberately lighter than a long risk memo. It gives a partner, security lead, or client a reproducible answer to three questions: what data moved, under which terms, and who remained accountable for the result. Re-review it when the model, feature, subprocessor list, retention setting, or intended use changes. A provider name alone is not a control because different products from the same provider can have materially different terms.
If the decision cannot be reconstructed six months later, the approval process was not auditable, even if the original answer happened to be right.
Sources
- MIT Project NANDA, “The GenAI Divide: State of AI in Business 2025”
- McKinsey, “The State of AI: Global Survey 2025”
- NIST, “AI Risk Management Framework: Generative AI Profile”
- OpenAI, “Data controls in the OpenAI platform”
- Anthropic, “API and data retention”
Accessed July 2026. Vendor terms and benchmark methodologies change; verify current primary documentation before making a decision.
Get an independent review
Advisory gives internal teams vendor-neutral judgment on governance, suppliers, architecture, evaluation, and priorities.
Explore Advisory →