Most companies already hand sensitive data to dozens of traditional software vendors without a second thought. Yet one of the most common questions I'm asked is why they should trust the AI labs' promise not to train on it.

The skepticism isn't totally paranoid. The frontier labs built their models by scraping the open internet, and they're defending lawsuits over it right now, from publishers, authors, and Reddit. If your mental model is "these companies play fast and loose with data," you have evidence.

But that evidence concerns a different pool of data. There are three pools: the public web the labs scraped to build their models, consumer chat data governed by a settings toggle, and enterprise inputs, meaning the prompts and files your business sends through Claude for Work, ChatGPT Enterprise, Gemini for Workspace, or the APIs. Every fight you've read about is in pool one.

The promise covers pool three. The way to evaluate it is to look at what would differ between a lab that honors it and one that doesn't: what they've signed, what cheating would earn them, and whether they could keep it quiet.

The promise is a contract, not a blog post

Start with what the promise actually is. Anthropic's Commercial Terms exclude customer content on Team, Enterprise, and API usage from training. OpenAI's business terms say the same for ChatGPT Business, Enterprise, Edu, and the API. Google's terms carry an explicit "Training Restriction" covering the Gemini Enterprise Agent Platform (formerly Vertex AI), with the paid Gemini API and Gemini for Workspace under equivalent commitments.

The placement is the point: a marketing page is a representation; a contract term is a cause of action. A lab that trained on enterprise inputs anyway would be in breach with every business customer it has at once, and the liability caps in these agreements either expressly carve out willful misconduct and fraud or can't lawfully shield them.

A marketing page is a representation; a contract term is a cause of action.

You don't need a negotiated deal to get the term. Anthropic's own announcement places Claude for Work, Team plan included, under the Commercial Terms, and OpenAI's business commitments cover self-serve ChatGPT Business alongside Enterprise. The contract forms when you subscribe.

All three companies also hold SOC 2 Type II attestations, meaning an independent auditor spent months testing whether their controls match their stated data commitments, including where customer content flows and who can touch it. Anthropic holds ISO/IEC 42001 on top, the AI-specific management standard. The contract is what binds them; the audit adds outside scrutiny of the systems a violation would have to move through.

The incentive math is lopsided

Now the economics. Enterprise AI is the market all three labs are spending billions to win, every security review in it starts with the training question, and one substantiated violation would follow the offender into every review from then on.

Weigh that against what cheating would buy. Real-world usage data genuinely is valuable for training, which is why the labs built permissioned channels for it: Anthropic asks consumers to opt in and retains those chats for five years, and runs a Development Partner Program where organizations volunteer their Claude Code sessions for training. OpenAI offers API customers a similar opt-in, and Google's unpaid tiers make the same trade in their terms.

A company with a permissioned supply of exactly this data has little reason to steal more of it and stake its enterprise business on concealment. Set against the opted-in corpus those channels produce, the contents of any one customer's account add almost nothing.

The conspiracy would need the leakiest companies in tech

Suppose a lab wanted to cheat anyway. Training data isn't a switch one executive flips in private; it moves through pipelines with provenance tracking, touched by many engineers, logged at every stage, and sampled by auditors. Secretly routing enterprise inputs into a training run means corrupting all of that and keeping everyone quiet.

Frontier labs are the worst possible place to keep that secret. Their employees job-hop between direct competitors, talk to reporters constantly, and publish open letters criticizing their own employers. Model codenames, training-run details, and board fights have all reached the press, sometimes within hours.

Believing the labs secretly train on enterprise data means believing they've kept exactly one secret, the most damaging one, while leaking everything else. I find that harder to believe than the alternative.

Where the terms are actually moving

These protections only cover the products that carry them. Two places they don't:

Consumer tiers. Since Anthropic's August 2025 consumer terms update, Free, Pro, and Max chats train Claude unless the setting is off, with five-year retention when it's on. OpenAI and Google run the same model on their consumer products: training rides on an account toggle, and unless someone changed it, assume it's on. An employee doing company work through a personal login has none of the protections above, which is one more argument for deploying AI through the business products instead of letting it arrive through accounts nobody reviews.

The rest of your software. The vendors most likely to train on your business data are the SaaS tools already holding it, not the frontier labs. Atlassian's "data contribution" policy took effect August 17: metadata and in-app content from Jira, Confluence, and Jira Service Management now feed improvements to its apps and AI for all customers. Defaults vary by plan: in-app data starts on for Free and Standard organizations, metadata contribution can't be turned off below Enterprise, and Atlassian describes its safeguards as de-identification and aggregation.

Figma is defending a proposed class action filed last November alleging it auto-enrolled customers in AI training, a claim Figma denies. Zoom wrote itself AI-training rights over customer content in its 2023 terms and retreated within days of the backlash.

Save the skepticism for the vendors nobody thought to doubt: the rest of the software stack, where the terms are changing right now and the training is real.