Est.

What OpenAI, Google, Anthropic, and Meta Do With Your Chats

Here's a clearer framework for evaluating how major AI companies handle your private conversations.

Reporter · · 12 min read
Cover illustration for “What OpenAI, Google, Anthropic, and Meta Do With Your Chats”
AI Data Practices · August 4, 2026 · 12 min read · 2,770 words

When ChatGPT crossed one billion monthly active users in June 2026, per Sensor Tower estimates reported by Reuters, and Meta AI reported a comparable milestone across its app suite per Fortune in October 2025, a threshold was crossed that made privacy policy a matter of everyday public concern rather than a specialist preoccupation. At that scale, data practices are not edge cases. They are decisions touching an enormous share of daily digital life: morning commutes, medical questions, proprietary code, legal strategy, grief, negotiation, and everything in between.

I have spent years inside the intersection of enterprise software, AI infrastructure, and data governance, long enough to have watched companies write privacy policies with one hand while building data flywheels with the other. What I have noticed is that users tend to evaluate these tools on capability and convenience, then treat the privacy question as something they will get to later. The problem is that "later" often arrives after the relevant window for meaningful choice has already closed.

This piece is not a verdict. It is a framework. Four dimensions matter most when evaluating any AI provider's data practices: what they collect (collection scope), how long they hold it (retention period), whether and how they train on it (training use), and what levers users actually control (user controls). What follows applies those four dimensions to OpenAI, Google, Anthropic, and Meta, because these four providers represent the dominant surface area of the consumer AI market and because their approaches differ in ways that are meaningful, not merely cosmetic.

What OpenAI Collects and How Long ChatGPT Keeps It

OpenAI's collection scope is broad. Every conversation is stored indefinitely unless a user deletes it. Alongside conversational content, OpenAI captures IP address, browser type, device information, timestamps, features used, and location data ranging from general to precise. The policy does not carve out sensitive material. Personal details, proprietary code, internal business strategy: all of it enters the same collection pipeline as a recipe request.

By default, consumer ChatGPT conversations are used to train and refine OpenAI's models. The opt-out is available but manual; future conversations are excluded once it is activated, but past conversations already collected remain subject to prior terms. Before conversations enter reinforcement learning pipelines, OpenAI anonymizes them, and a subset is manually reviewed by AI trainers to identify bias or harmful outputs. This human review layer is common across the industry, though the degree of transparency varies.

Deletion works as follows: users can remove specific conversations or their entire history. OpenAI commits to removing deleted conversations from its systems within 30 days. Temporary Chat, a mode designed for ephemeral use, auto-deletes within 30 days unless safety or legal considerations require retention.

One exception is worth examining closely. ChatGPT's Operator agent retains deleted screenshots and browsing histories for 90 days, per reporting by Nightfall AI in March 2025. That is three times the standard deletion window. Users who believe they are operating under the standard 30-day removal policy when using agent features are, in practice, operating under a different one.

OpenAI explicitly states it does not sell user data to third parties and does not use it for advertising campaigns. Exceptions exist for legal compliance, including government subpoenas and platform abuse investigations. Data at rest is protected by AES-256 encryption; data in transit uses TLS 1.2 or higher.

The regulatory record complicates the picture. Investigations and temporary bans in Italy, Spain, Poland, and France centered on personal data scraped during training and on hallucinated false information about real individuals. These are not hypothetical risks; they are documented friction points between OpenAI's data architecture and European data protection law.

How OpenAI Treats Enterprise and API Customers Differently from Free Users

Here the picture changes materially. For customers accessing OpenAI via the API, and for those on ChatGPT Business, Enterprise, Healthcare, Edu, and the API Platform, data is not used for model training by default. Enterprise customers own their inputs and outputs and control retention duration. The consumer product and the enterprise product are not variations on the same privacy arrangement; they are fundamentally different arrangements.

The practical implication is direct: a company deploying ChatGPT via the API operates under substantially stronger privacy defaults than an individual employee using the free consumer product on their personal device. In many organizations, both situations coexist, often without the employees in question understanding which environment they are in.

One 2025 court order required OpenAI to temporarily retain certain content, demonstrating that even enterprise commitments can be overridden by legal process. That caveat applies across the industry, not just to OpenAI.

The enterprise-consumer split is a structural feature of the AI data landscape. Every provider covered here exhibits it in some form. The consumer tier is consistently the less protected one.

What Google Collects Through Gemini and the Ecosystem Problem

Google's data collection through Gemini is, in one sense, comparable to OpenAI's. In another sense, it is categorically different because of what surrounds it.

With Gemini Apps Activity enabled, Google saves conversations, shared files, videos, screenshares, photos, audio, Gemini Live video, feedback, information from websites visited with Gemini, product usage data, and location data including stored Home and Work addresses. On mobile, the collection extends to call logs, installed apps, and device usage patterns. The Ask Gemini overlay stores whatever content is visible on the screen at the moment of invocation.

None of that is unusual by the standards of the industry. The defining difference is context. Gemini is not a standalone application. It sits inside a platform that already holds a user's email, calendar, location history, and search activity spanning years or decades. In late 2025, Google enabled Gemini access to Gmail, Google Chat, and Google Meet by default for US users. By early 2026, this cross-app data flow was rebranded as "Personal Intelligence," connecting Gemini to Gmail, Drive, Maps, and other services. When enabled, Gemini can read a user's entire email history.

The right question for a Google user is therefore not only what Gemini does with a single conversation. It is what Google does with that conversation in combination with everything else it already knows. Those are different questions, with different implications.

An additional layer worth understanding: Gemini's access to Phone, Messages, WhatsApp, and Utilities apps, announced in July 2025, operates whether or not Gemini Apps Activity is turned on. The Activity toggle governs training use. It does not govern app interactions. These are separate controls with separate scopes, and conflating them is a common and consequential misunderstanding.

Consumer conversations are stored for 18 months by default. Users can adjust this to 3 or 36 months in settings. Even with Gemini Apps Activity disabled, conversations are retained for up to 72 hours for operational purposes.

Google's Human Review Policy and What Deleting Your Gemini History Does Not Erase

Google explicitly acknowledges that human reviewers, including third-party service providers, read, annotate, and process Gemini conversations to improve AI models. Reviewers do not see email addresses or phone numbers, but the substantive content of sensitive conversations is visible to them. That is a meaningful distinction: the absence of a name on a conversation does not make that conversation anonymous if its content is specific enough to be identifiable.

Conversations reviewed by humans are retained for up to 3 years, disconnected from a user's Google Account. This is the gap most users do not know exists. Deleting Gemini activity from a Google Account removes the copy that is visible to the user. It does not remove the copy retained for training review. Those two things are not the same deletion.

The gap between what "delete" means to the user and what it means to the platform is the central tension in AI data governance, and Google's policy illustrates it more clearly than most.

Workspace Enterprise customers operate under a different arrangement: human review does not occur without organizational consent, prompt content is not used to train public models without explicit permission, and administrators can shorten or fully disable prompt storage. The enterprise product is, again, a materially different privacy arrangement from the consumer one.

Venn diagram: AI Providers: Consumer vs. Enterprise Privacy. Compares Consumer Tier and Enterprise Tier; overlap: Shared Practices.

How Anthropic Handled Training Data Differently, and Then Changed Course in 2025

For roughly two years after Claude's launch, Anthropic did not use consumer conversations to train its models. At the time, this was a real differentiator. ChatGPT trained on consumer data by default; Gemini folded conversations into Google's broader data ecosystem. Anthropic's approach was meaningfully different, and it was a reason some privacy-conscious users and organizations preferred Claude for sensitive work.

In September 2025, that policy changed. Users on Claude Free, Pro, and Max plans were asked whether they wished to share conversation data to improve the model. This was framed as voluntary and opt-in, which is a more user-respecting framing than the default-on approaches of competitors. The mechanics, however, complicate that framing: the default setting in privacy controls was switched to "on," requiring users to manually adjust it to opt out, per NYU Shanghai RITS reporting in August 2025. An opt-in mechanism with a default-on setting is functionally closer to an opt-out mechanism than the label suggests.

The retention consequence of opting in is significant. Without the training opt-in, conversations are saved until deleted, then removed from backend systems within 30 days. With the opt-in, retention extends to 5 years. That is a 60-fold increase in how long conversations can remain in Anthropic's pipeline, triggered by a setting many users may not have consciously reviewed.

Existing users had until October 8, 2025, to accept updated Consumer Terms. Safety exceptions apply regardless of training preferences: conversations flagged for policy violations are retained for 2 years, and trust and safety classification scores are retained for 7 years. These retentions apply even to users who opted out of training data sharing. Anthropic does not sell user data and runs no advertising business, which remains a meaningful distinction.

The Anthropic arc illustrates something worth sitting with: a provider's current stance is not a permanent commitment. Policies evolve. The direction of travel, and the conditions under which a company is willing to change its policies, matters as much as where those policies stand today.

Anthropic's API and Enterprise Tiers, Where the Stricter Defaults Live

Effective September 14, 2025, Anthropic tightened API data handling considerably. Inputs and outputs are automatically deleted after 7 days and are not used for model training. Enterprise customers can execute a Zero Data Retention agreement, under which inputs and outputs are not stored beyond what is needed to screen for abuse.

One caveat to ZDR deserves explicit attention: Anthropic still retains User Safety classifier results even under a Zero Data Retention agreement. "Zero retention" is not absolute. The name is more reassuring than the underlying mechanics.

Commercial customers on Claude for Work, Enterprise, Education, and Government plans were explicitly excluded from the September 2025 consumer policy changes. Their data is not used for model training under commercial terms. Anthropic offers a Business Associate Agreement for qualifying healthcare customers; web search is disabled under BAA terms. Standard consumer Claude is not HIPAA-compliant and should not be used with Protected Health Information. EU-based consumer users retain standard GDPR rights via Anthropic's Privacy Center; commercial customers can execute a Data Processing Addendum.

What Meta AI Does With Chat Data and Why There Is No Opt-Out

Meta's approach is structurally different from the other three providers, and the difference is not subtle. Starting December 16, 2025, Meta's updated privacy policy allows the company to use AI chat interactions across Facebook, Instagram, Messenger, and WhatsApp to personalize recommendations and deliver targeted advertising.

The mechanism in plain terms: a conversation with Meta AI about, say, knee surgery or a career change or a financial decision can result in more related groups, products, and ads appearing across Meta's apps. The conversation becomes an input signal in an advertising personalization system.

No opt-out exists for this use. Per Fortune's October 2025 reporting, if users do not want their chat conversations influencing their ads, the only available option is not to use Meta AI. Users can adjust ad preferences, but AI conversation data will still be processed. This is the defining characteristic of Meta's approach: the absence of a control mechanism is itself the policy.

The structural difference from the other three providers is worth stating precisely. OpenAI, Google, and Anthropic each use conversation data primarily to improve their models, with varying levels of user control over that use. Meta explicitly routes conversation data into an advertising personalization system. These are different categories of use with different implications for what users should share and with whom.

At the scale Meta AI operates, these practices apply to a large number of people who may be sharing sensitive personal information without understanding how it is being used.

Comparing the Four Providers Across the Dimensions That Matter Most

Table: Four Providers Across Key Privacy Dimensions. Compares Default Training Use, Consumer Retention, Human Review Copy, Advertising Use, and 1 more by OpenAI, Google Gemini, Anthropic and Meta AI.

Across training use by default: OpenAI and Google both train on consumer data by default, requiring manual opt-out. Anthropic, post-October 2025, frames participation as opt-in but shipped with a default-on setting that required manual adjustment to change. Meta offers no opt-out for advertising personalization use.

Retention periods, summarized plainly: OpenAI consumer data is kept indefinitely until deletion, with 30-day backend removal after deletion; the Operator agent extends this to 90 days for deleted content. Google consumer conversations are kept for 18 months by default, adjustable to 3 or 36 months; human-reviewed copies persist for up to 3 years regardless of user deletion. Anthropic without the training opt-in retains conversations until deleted, then removes them within 30 days; with the opt-in, retention extends to 5 years; flagged conversations are kept for 2 years; API data is deleted after 7 days. Meta's retention window is governed by its broader privacy policy, which now explicitly incorporates AI conversation data into its advertising infrastructure.

Human review exists, in some form, at all four providers. Google's consumer policy explicitly acknowledges it and the 3-year retention attached to reviewed conversations. OpenAI acknowledges a subset of interactions are manually reviewed by AI trainers. Anthropic retains flagged conversations for 2 years, which implies human review for at least some of them. Meta's policy foregrounds advertising personalization more than human review.

On advertising use, the lines are clear: OpenAI and Anthropic do not use chat data for advertising. Google's consumer ecosystem is advertising-funded, and Gemini data enters that ecosystem, though Workspace is explicitly ad-free. Meta explicitly uses chat data for advertising personalization with no opt-out available.

The enterprise-consumer contrast is consistent across all four. Stricter defaults, less training use, greater administrative control, and more transparent retention terms are reliably features of the enterprise or API tier. The consumer product is the less protected configuration.

What User Controls Actually Exist and Where They Fall Short

Some controls work as described. OpenAI's training opt-out functions; conversation deletion triggers a 30-day backend removal; Temporary Chat provides a limited-persistence option. Google's Gemini Apps Activity toggle limits training use and retention period is adjustable within defined options. Anthropic offers a training data opt-out and conversation deletion with 30-day backend removal; API customers can negotiate Zero Data Retention terms.

Other controls sound broader than they are. Google's delete function clears activity visible in a user's account; it does not touch human-reviewed copies retained for up to 3 years. Google's Activity toggle does not prevent Gemini from accessing Phone, Messages, WhatsApp, and Utilities apps. Anthropic's training opt-out does not protect flagged conversations from 2-year retention or safety classifier scores from 7-year retention. OpenAI's Operator agent retains deleted content for 90 days; the standard 30-day window does not apply to that surface.

Meta offers no meaningful control over the chat-to-ad-personalization pipeline. Non-use is the only available option.

The pattern across all four providers is consistent enough to be structural rather than coincidental. Opt-out and deletion tools address the most visible layer of data use. Backend retention schedules, safety exceptions, legal holds, and human review pipelines operate below that visible layer. Users who believe they have controlled their data by pressing a delete button have controlled one copy of it, under one set of terms, on one part of the system. The rest of the architecture continues operating on its own timeline.

This is not an accusation. These systems face genuine obligations: safety monitoring, abuse detection, legal compliance, and model improvement. Those obligations require data. The question worth carrying forward is not whether retention and review are legitimate, but whether users understand what they are consenting to when they type something into an AI system and hit enter. The gap between the visible controls and the underlying architecture is where most of the relevant decisions are actually made, and it is smaller at some providers than others.

Sources

  1. openai.com
  2. nightfall.ai
  3. openai.com
  4. mayerbrown.com

More in AI Data Practices