AInora
AI GovernanceGDPRData ResidencyEU AI ActComparison

Does Your AI Vendor Train on Your Data? Tier-by-Tier, 2026

JB
Justas ButkusFounder, Ainora
··Updated ·13 min read

For the four assistants most European companies actually use, the answer splits on one line: consumer accounts may be trained on, organisational accounts are not. A personal ChatGPT subscription, a personal Microsoft account in Copilot and a personal Google account in the Gemini app all sit on the training side of that line by default. ChatGPT Business and Enterprise, Claude Team and Enterprise, Copilot signed in with a work Entra ID, and Gemini inside a Google Workspace tenant all sit on the other side, and each vendor says so in its own documentation.

The nuance is where the money and the risk are. Not training is not the same as not storing. Not storing in the US is not the same as offering EU residency. And EU residency, where it exists, is never total, is sometimes impossible to apply retroactively, and in at least two documented cases does not extend to a product the same vendor sells you in the same subscription.

Every claim below comes from the vendor's own documentation, read on 6 September 2026. Nothing here comes from a competitor summary, a secondary article, or a vendor's description of a rival. One qualification, because it matters for how much weight to put on a row: OpenAI's help centre and marketing pages return HTTP 403 to automated retrieval, so the OpenAI rows were read from archived copies of those same OpenAI pages rather than live. Every other vendor page here was read live.

Does Your AI Vendor Train on Your Data?

“Trains on your data” is four separate questions that get collapsed into one, and collapsing them is what produces bad decisions. The four are: does the vendor use your content to improve its models; who inside the vendor can read it; where does it sit at rest; and how long does it stay there. A vendor can answer “no” to the first and still store your content indefinitely in a US data centre, which is exactly what one of the four does on its business tier.

This page answers all four, per product and per account type, with the vendor's own wording attached to each row. It is a reference sheet, not an argument. If a claim is not in a vendor doc, it is not here.

Why Account Type, Not Price Tier, Decides It

The single most expensive misconception in a small or mid-sized company is that paying more buys privacy. It does not. A personal ChatGPT Plus subscription bought on a company card is still a consumer account, and OpenAI is clear that content from “services for individuals such as ChatGPT and Codex” may be used to train its models unless the individual opts out in the privacy portal. The same subscription bought as a ChatGPT Business seat is a different legal product with the opposite default.

Microsoft draws the same line by identity rather than by price. Its consumer FAQ lists who is excluded from model training, and the first entry is “Users signed into Copilot with an organizational Entra ID account.” The account you sign in with, not the licence you bought, decides which side of the line you are on. One consequence worth knowing: Microsoft also excludes “Users of Copilot within Microsoft 365 apps with Personal or Family subscriptions” from training, so a household subscription behaves differently from the free consumer Copilot on the same device.

A claim to be careful with

It is widely repeated online that Microsoft excludes the EEA from consumer Copilot training. That carve-out is not in Microsoft's own exclusion list, which names Brazil, China (excluding Hong Kong), Israel, Nigeria, South Korea and Vietnam. Microsoft's privacy statement says only that “in some markets, this data can help train our AI models in Microsoft Copilot unless you opt out.” Do not plan around a European exemption that the vendor has not published. Check the toggle instead.
Source: Microsoft Support, Copilot privacy FAQ

Training Defaults, Tier by Tier

Read this by account type, not by brand. Each row carries the vendor document the wording came from.

Product and account typeTrains on your content by default?Who controls the setting
ChatGPT Free, Plus or Pro (personal account)Source: OpenAI Help Center, model trainingYes, unless you opt out. “We may use your content to train our models.”The individual, in the OpenAI privacy portal
ChatGPT Temporary ChatSource: OpenAI Help Center, model trainingNo. Temporary chats are “not used to train our models.”The individual, per chat
ChatGPT Business (formerly Team)Source: OpenAI Help Center, model trainingNo. “By default, we do not train on any inputs or outputs from our products for business users.”Workspace admin, opt-in only
ChatGPT Enterprise, Edu, Healthcare, for TeachersSource: OpenAI, Enterprise privacyNo, “unless you have explicitly opted in to share your data with us.”Workspace admin
OpenAI API PlatformSource: OpenAI Help Center, model trainingNo. “Unless they explicitly opt-in, organizations are opted out of data-sharing by default.”Organisation owner
Claude Free, Pro or Max (consumer)Source: Anthropic Privacy Center, consumerA setting, not a default anyone can quote. Anthropic lists training where “you choose to allow” it; the article does not state the shipped default.The individual, in Privacy Settings
Claude Incognito chatsSource: Anthropic Privacy Center, consumerNo. “Not used to improve Claude, even if you have enabled Model Improvement.”The individual, per chat
Claude Team, Enterprise, Anthropic API, Claude GovSource: Anthropic Privacy Center, commercialNo. “By default, we will not use your inputs or outputs from our commercial products … to train our models.”Primary Owner or Owner
Microsoft Copilot, personal Microsoft accountSource: Microsoft Support, Copilot privacy FAQYes in some markets, unless you opt out. Training data includes “your voice and conversation activity with Copilot, including the images or files you upload.”The individual, and the opt-out is partial
Microsoft Copilot, work or school account (Entra ID)Source: Microsoft Learn, Copilot privacyNo. “Prompts, responses, and data accessed through Microsoft Graph aren’t used to train foundation LLMs.”Tenant admin
Gemini app, consumer Google accountSource: Google, Gemini Apps Privacy HubHuman review by default. “A subset of chats are reviewed by human reviewers … to help improve Google services.”The individual, via Keep Activity
Gemini in Google WorkspaceSource: Google Workspace, generative AI privacy hubNo. “Your content is not human reviewed or otherwise used for Generative AI model training outside your domain without permission.”Workspace admin

Two rows deserve a note. The Claude consumer row is deliberately not an answer. Anthropic's help article frames model improvement as something you “choose to allow,” and does not state whether the toggle ships on or off for a new account. Rather than repeat a guess, the honest instruction is: open Privacy Settings and look. Anthropic separately confirms that commercial products are excluded by default, and that a Primary Owner or Owner can switch off the thumbs-up and thumbs-down button organisation-wide using the “Rate chats” setting under Organization settings, Data and Privacy.

The Gemini consumer row is not phrased as training because Google does not phrase it that way. What Google publishes is human review: a subset of chats are read by reviewers, including trained service providers, to improve Google services. Google then gives the clearest single warning any of the four vendors offers, and it belongs on a slide in every induction deck: “Please don't enter confidential information that you wouldn't want a reviewer to see or Google to use to improve our services.”

Where Does the Data Physically Live?

This is the section that changes procurement decisions, and it contains the two facts that almost nobody publishes.

First: Anthropic's first-party Claude does not store data in Europe. Anthropic can route traffic to selected countries, including in Europe, but its privacy article ends the paragraph with six words that settle the question: “Note that data is stored in the US.” Its developer platform documentation is blunter still. Workspace geo, the setting that governs where data is stored at rest, “is set when you create a workspace and can't be changed afterward,” and “currently, us is the only available workspace geo.” Anthropic's own compliance page tells you where to go instead: “For Europe, you can select country-specific deployment options through AWS Bedrock, GCP Vertex, and Microsoft Foundry.” If a proposal tells you Claude is EU-hosted, ask which of those three it runs on.

Second: Microsoft excludes Anthropic models from the EU Data Boundary. Microsoft states that for EU customers Copilot is an EU Data Boundary service, and then adds, in a note on the same page: “Models provided by Anthropic as a subprocessor are currently excluded from the EU Data Boundary.” Admins choose whether third-party models power their Copilot experiences. That choice quietly moves some traffic outside a boundary the organisation believes it has.

Vendor and productEU data residencyThe documented limit
OpenAI, ChatGPT Enterprise and EduSource: OpenAI Help Center, data residencyYes: storage at rest and GPU inference in Europe (EEA + Switzerland)New workspaces only. Workspace metadata, names, billing, user logins and external integrations are listed as out of scope.
OpenAI, API PlatformSource: OpenAI, data residency in EuropeYes, per Project, with zero data retention on those requests“European residency can only be configured for new Projects, existing Projects cannot be updated … after creation.”
OpenAI, consumer ChatGPTSource: OpenAI Help Center, data residencyNot offeredEligibility is limited to “eligible API customers and new ChatGPT Enterprise/Edu customers.”
Anthropic, first-party Claude (consumer, Team, Enterprise, API)Source: Anthropic Privacy Center, server locationsTraffic routing only, not storage“By default, we may route customer traffic to select countries in the US, Europe, Asia and Australia … Note that data is stored in the US.”
Anthropic, Claude Developer Platform workspacesSource: Anthropic platform docs, data residencyNo EU option“Workspace geo is set when you create a workspace and can’t be changed afterward. Currently, ‘us’ is the only available workspace geo.”
Anthropic, via cloud partnersSource: Anthropic, regional complianceYes“For Europe, you can select country-specific deployment options through AWS Bedrock, GCP Vertex, and Microsoft Foundry.”
Microsoft Copilot, work or school accountSource: Microsoft Learn, Copilot privacyYes. “For EU customers, Microsoft Copilot is an EU Data Boundary service.”“Models provided by Anthropic as a subprocessor are currently excluded from the EU Data Boundary.”
Google Workspace with GeminiSource: Google Workspace, AI privacy and securityYes. “Customers can control where their data is stored and processed (EU or US).”Applied through the organisation’s existing data-regions policy.
Gemini Notebook, inside WorkspaceSource: Google Workspace, generative AI privacy hubThe org policy does not reach it“Your organization’s file sharing and data region settings do not apply to data in Gemini Notebook.”

Two more traps sit in that table. OpenAI's European residency cannot be applied retroactively: “European residency can only be configured for new Projects, existing Projects cannot be updated to have European data residency after creation,” and on the ChatGPT side the feature is offered to new Enterprise and Edu workspaces. A pilot that spins up projects in week one and asks about residency in week six has to be rebuilt, not reconfigured. And Google's own Workspace documentation records that Gemini Notebook ignores the organisation's data-region and file-sharing settings, which means the one Gemini surface most likely to be fed a folder of internal documents is the one your data-region policy does not reach.

Even where residency works, it is partial by design. OpenAI publishes what stays outside the chosen region: workspace metadata, workspace name, billing information, user logins, transient processing steps, and data handled through external integrations such as connected apps and web search. That is the useful teaching point. A residency badge is a statement about customer content, not about everything the service touches. The same instinct applies when you assess any vendor, which is why a written vendor security assessment beats a sales call, and why EU data residency claims are worth reading in the vendor's own words rather than in a comparison grid.

How Long Is It Kept, and Does Deleting Delete?

Retention is where the defaults are least intuitive. The headline: Claude for Work retains indefinitely unless an owner configures otherwise, and the most-used consumer assistants keep 18 months by default on both Microsoft and Google.

Product and tierRetention defaultWhat the doc says
ChatGPT Business, Enterprise, Edu and HealthcareSource: OpenAI, Enterprise privacyAdmin-setDeleted conversations are “removed from our systems within 30 days.”
OpenAI API PlatformSource: OpenAI, Enterprise privacyUp to 30 daysZero data retention available “for eligible endpoints if you have a qualifying use-case.”
ChatGPT Temporary ChatsSource: OpenAI Help Center, Data Controls FAQ30 days“Temporary Chats are deleted from our systems after 30 days.”
Claude consumer, deleted chatsSource: Anthropic Privacy Center, consumer retentionPurged from back-end storage within 30 daysRemoved from chat history immediately.
Claude consumer, model improvement enabledSource: Anthropic Privacy Center, consumer retentionUp to 5 years, de-identifiedApplies only to “new or resumed chats after the model training setting has been enabled.”
Claude, chats flagged by trust and safetySource: Anthropic Privacy Center, consumer retention2 years for inputs and outputsClassification scores are kept for up to 7 years.
Anthropic APISource: Anthropic Privacy Center, org retention30 days“We automatically delete inputs and outputs on our backend within 30 days,” with named exceptions.
Claude for Work (Team and Enterprise)Source: Anthropic Privacy Center, custom retentionIndefinite“By default, data is retained indefinitely unless a custom retention period is set.” Minimum custom period: 30 days.
Microsoft Copilot, personal accountSource: Microsoft Support, Copilot privacy FAQ18 months“By default, we store conversation activity for 18 months.”
Microsoft Copilot, work or school accountSource: Microsoft Learn, Copilot privacyAdmin-set through PurviewStored “in alignment with contractual commitments with your organization’s other content in Microsoft 365.”
Gemini app, Keep Activity onSource: Google, Gemini Apps Privacy Hub18 monthsChangeable to 3 or 36 months, or indefinite.
Gemini app, Keep Activity off or temporary chatSource: Google, Gemini Apps Privacy Hub72 hoursRetained with your account for 72 hours.
Gemini app, chats a human reviewer has seenSource: Google, Gemini Apps Privacy HubUp to 3 years“Not deleted when you delete your activity.”
Gemini in Google WorkspaceSource: Google Workspace, generative AI privacy hub“90 days to indefinite, as determined by admins”Gemini Notebook data is “not retained after session ends.”
Indefinite
Claude for Work default retention until an owner sets a custom period, minimum 30 days
Source: Anthropic Privacy Center, custom retention
3 years
How long a Gemini chat already seen by a human reviewer is kept, even after you delete your activity
Source: Google, Gemini Apps Privacy Hub
18 months
Default conversation-activity retention on consumer Microsoft Copilot
Source: Microsoft Support, Copilot privacy FAQ

The Gemini figure is the one to put in front of staff, because it answers the sentence people say after an incident: “but I deleted it.” Google records that human-reviewed chats and related data “are not deleted when you delete your activity.” Anthropic has an equivalent asymmetry on safety-flagged content, keeping inputs and outputs for two years and classification scores for seven, on both consumer and commercial products. Deletion controls are real, and they are not symmetrical with the vendor's own retention exceptions.

What Can an Administrator Actually Set?

On business tiers, all four vendors give an administrator a genuine surface. The differences are less about how many switches exist and more about which risk each vendor thinks the switches are for.

Vendor and tierWhat an administrator can actually set
OpenAI, Business and EnterpriseSource: OpenAI, Enterprise privacyHow long data is retained, which apps and connectors are enabled for the workspace, enterprise authentication, and a Data Processing Addendum. Fine-tuned models are “for your use alone and never served to or shared with other customers.”
Anthropic, Team and EnterpriseSource: Anthropic Privacy Center, custom retentionCustom data-retention periods, minimum 30 days, set by a Primary Owner or Owner. Changes are written to audit logs.
Microsoft, work or school CopilotSource: Microsoft Learn, Copilot privacyPurview retention policies for Copilot data, Content search and eDiscovery, sensitivity labels and rights management honoured, and: “Admins have full control to select which agents are allowed in their organization.”
Google WorkspaceSource: Google Workspace, generative AI privacy hubData-regions policy and Data Loss Prevention are applied automatically to Gemini, per-service enablement, and admin-set retention, with the documented Gemini Notebook exception.

Microsoft's documentation contains the most useful sentence for a rollout plan, and it is not about the model at all: Copilot “only surfaces organizational data to which individual users have at least view permissions,” and Microsoft follows it with a warning that you must be using the permission models in SharePoint and elsewhere properly. Translated: for a Microsoft shop the real risk is not the assistant, it is the over-shared SharePoint the assistant can now search instantly. That is a permissions-hygiene project, not an AI project, and it is the highest-value thing to fix before a Copilot rollout.

What Happens When Nobody Decides

The alternative to choosing is not neutrality. It is staff choosing for you, on personal accounts, on the training side of the line.

20%
of studied organisations experienced breaches linked to shadow AI
Source: IBM, Cost of a Data Breach 2025
USD 670K
that shadow-AI incidents added to the average breach cost
Source: IBM, Cost of a Data Breach 2025
97%
of organisations reporting AI-related breaches said they lacked proper access controls
Source: IBM, Cost of a Data Breach 2025
63%
of breached organisations studied lacked AI governance policies
Source: IBM, Cost of a Data Breach 2025

Those four figures come from IBM's Cost of a Data Breach Report 2025. Read together they describe one failure, not four: the tools arrived before the decisions did. The 97% figure is the sharpest of them, because access control is precisely what an organisational account gives you and a personal account does not. The fix is not a ban. Bans push usage further into personal accounts, which is the expensive side of the table above. The fix is to sanction a tool, buy it on organisational identity, set the admin controls, and tell people which surface is for which kind of content.

How to Decide What to Hand Your Staff

Four questions, in this order, answer it for most European companies.

1. Which identity will people sign in with? This decides the training default on every vendor in the table, so decide it first and enforce it with SSO. Everything downstream is easier once nobody is signing in with a personal account.

2. Do you have a genuine EU storage requirement, or a preference? If it is a requirement, the table narrows fast. OpenAI on new Enterprise or Edu workspaces, Microsoft inside the EU Data Boundary with the Anthropic-model exclusion understood, or Google Workspace with a data-regions policy will all meet it. Anthropic's first-party products will not, and you will be looking at Bedrock, Vertex or Foundry. Decide before you provision, because on the OpenAI API side you cannot change it afterwards.

3. What is the retention posture you want to be able to defend? If you are in a sector where a regulator may one day ask how long prompts containing client data were kept, “indefinite by default” is not the answer you want to give. Set a custom period on day one.

4. What are people allowed to paste? This is the only one of the four that is a training question rather than a procurement question, and it is where an AI literacy programme under Article 4 of the EU AI Act earns its place. The regulation does not require you to certify anyone or to guarantee a level. It does expect you to take measures appropriate to the systems you deploy and the people using them, and a documented rule about which surface takes which class of content is exactly that kind of measure.

The EU AI Act touches AI systems in other ways that are worth separating from data handling, because they are different duties with different deadlines. The Article 50 transparency obligations, for instance, govern whether a person must be told they are dealing with an AI system, which is a disclosure question rather than a data-residency one. We have covered those separately: what the AI Act means for voice agents, whether an AI caller must identify itself, when a system counts as high-risk, a practical compliance checklist, and the sector-specific case of debt collection. If you deploy someone else's system under your own brand, who is responsible for what is the first thing to settle.

For the everyday side of this, our guide to practical business uses of ChatGPT covers what teams tend to do first, and the GDPR guide and security checklist cover the controls you should expect from any AI vendor you bring inside your processes.

How This Page Is Kept Honest

Last verified 6 September 2026

Every vendor claim on this page comes from the vendor's own documentation, read on 6 September 2026, not taken from a summary. The OpenAI rows are the one qualification: those pages return HTTP 403 to automated retrieval, so they were read from archived copies of OpenAI's own pages rather than live, and should be re-checked in a browser before anything expensive rests on them. Vendor terms change, sometimes without a changelog, and several of the pages cited here were themselves updated within the previous month. This page is re-checked quarterly, and the next scheduled re-verification is December 2026. Where a doc has moved or a claim has changed, the row is corrected and the date at the top of this box moves with it.

Two cells in the tables above are deliberately not answers. The Claude consumer training default is a setting whose shipped state Anthropic's article does not publish, so the table says that rather than guessing. And the widely repeated EEA carve-out on consumer Copilot training is absent from Microsoft's own exclusion list, so it does not appear here at all. A table that is right about eleven rows and confidently wrong about one is worse than useless for the decision it is meant to support.

If you find a row that no longer matches the vendor doc it links to, the link is there so you can check it before we do.

Frequently Asked Questions

Yes, unless you turn it off. Paying for a personal subscription does not change the account type. OpenAI states that for “services for individuals such as ChatGPT and Codex, we may use your content to train our models,” and points you to the privacy portal to opt out. The training default flips only when you move to ChatGPT Business, Enterprise, Edu or the API, where OpenAI does not train on inputs or outputs by default.Source: OpenAI Help Center, model training

Not on Anthropic’s first-party products. Anthropic can route traffic to selected countries including in Europe, but its own privacy article says plainly: “Note that data is stored in the US.” Its developer platform documentation is more explicit still, stating that “us” is currently the only available workspace geo and that the setting cannot be changed after the workspace is created. Anthropic points customers who need European deployment to AWS Bedrock, Google Cloud Vertex and Microsoft Foundry instead.Source: Anthropic Privacy Center, server locations

For EU customers, Microsoft says Copilot is an EU Data Boundary service. There is one documented carve-out that almost never appears in vendor comparisons: “Models provided by Anthropic as a subprocessor are currently excluded from the EU Data Boundary.” If your tenant has enabled Anthropic models inside Copilot, that specific traffic sits outside the boundary you thought you had.Source: Microsoft Learn, Copilot privacy

Not necessarily. Google states that chats already seen by human reviewers, together with related data such as your language, device type and location information, “are not deleted when you delete your activity. Instead, they are retained for up to three years.” Deleting your activity removes the copy you can see, not the copy a reviewer already has.Source: Google, Gemini Apps Privacy Hub

No, and two vendors say so themselves. OpenAI notes that even if you have opted out, choosing to send feedback means “the entire conversation associated with that feedback may be used to train our models.” Microsoft states that its consumer opt-out “will not exclude your conversations from being used for other general product or system improvements nor from use for advertising, digital safety, security, and compliance purposes.” Treat the thumbs-up button as an export button.Source: Microsoft Support, Copilot privacy controls

For OpenAI, not on the API side. The company states that “European residency can only be configured for new Projects, existing Projects cannot be updated to have European data residency after creation,” and its ChatGPT residency page limits eligibility to new Enterprise and Edu workspaces. A rollout that creates projects and workspaces first and thinks about residency second cannot be corrected without rebuilding them.Source: OpenAI, data residency in Europe

JB
Justas Butkus

Founder & CEO, AInora

Building AI digital administrators that replace front-desk overhead for service businesses across Europe. Previously built voice AI systems for dental clinics, hotels, and restaurants.

View all articles

Ready to try AI for your business?

Hear how AInora sounds handling a real business call. Try the live voice demo or book a consultation.