What “Private AI” Actually Means: Encryption, Retention, and Training Data

Private ” is the most overloaded word in AI marketing. It appears on landing pages next to a padlock icon and it can mean any of four completely different things — or, occasionally, none of them.

The four are not interchangeable. A vendor can honestly claim privacy while doing something you would object to, simply because they mean a different one of the four than you do.

Here is how to decode it, and what to actually verify before putting anything sensitive into a model.

The four promises hiding inside “private”

1. Encryption in transit and at rest

Your data is encrypted travelling to the provider and encrypted while stored on their systems.

What it protects against: interception, and someone stealing a disk.

What it does not protect against: the provider reading it, staff accessing it, it being used for training, or it being retained indefinitely.

This is table stakes. Every credible provider does it, which means advertising it prominently is a mild warning sign — it suggests the stronger claims are unavailable.

2. No training on your data

Your inputs and outputs are not used to improve the provider’s models.

This is the promise most people mean when they say private, and it is the one with the most variation in practice.

Three things to check:

  • Does it differ by plan? Several providers train on free-tier traffic and not on paid. The moment someone signs up with a personal free account, the policy that applies is not the one you read.
  • Does it differ by API versus app? API traffic is commonly excluded from training by default while consumer app traffic is not.
  • Is it opt-out or opt-in? Opt-out means the default is training, and the default is what most of your team will be on.

3. Limited or zero retention

The provider does not keep your data after processing, or keeps it only briefly.

Read this one carefully, because “we don’t store your conversations” and “we retain for 30 days for abuse monitoring” are both routinely described as privacy-respecting, and only one of them is what you probably imagined.

Ask three questions:

  • How long is data retained, in days, for what stated purpose?
  • Is there a zero-retention option, and on which plans?
  • What happens to backups? Deletion from active systems and deletion from backups are different events, often weeks apart.

4. No human review

Some providers sample conversations for quality, safety, or abuse review. That is a legitimate practice and it means a person may read your text.

Check whether human review happens, whether it can be disabled, and whether it applies to your tier. This one is disclosed less prominently than the other three.

The sub-processor question

Here is the part that gets missed most often, and it applies specifically to platforms that aggregate multiple models.

If a service routes your prompt to a third-party model provider, your text reaches that third party. The platform’s privacy policy governs the platform. The model provider’s policy governs what happens next.

This is not a scandal — it is how aggregation works, and the upstream terms are frequently excellent. But it means you need two answers, not one:

  1. What does the platform do with your data?
  2. Which upstream providers receive it, and under what terms?

A well-run platform publishes its sub-processor list and states which upstream agreements exclude training. One that cannot answer question two has not thought about it properly, and that is worth knowing before you commit.

Look for a genuine private mode where it is offered — a setting where the conversation is not written to the platform’s servers at all. Platforms such as Perspective AI build this in as an explicit mode rather than a default assumption, which is the right shape: privacy you can switch on for the conversations that need it, verifiably, rather than a blanket claim covering everything.

How to read a privacy policy in ten minutes

Skip the preamble. Use search.

Search for “train”. You want an unambiguous statement about whether your content trains models, and whether it varies by tier. Ambiguity here is deliberate more often than not.

Search for “retain” and “delete”. Find the actual number of days and the purpose. Find the separate statement about backups.

Search for “sub-processor” or “third party”. You want a list, or a link to one. “We may share with service providers” without a list is not a disclosure.

Search for “human”. Determine whether human review occurs and whether it is optional.

Check the effective date. A policy last updated two years ago on a product that has changed substantially is a maintenance signal, not a privacy signal.

Ten minutes gets you further than most procurement reviews.

What changes by plan tier

This is where assumptions break most expensively. Privacy terms are frequently a paid feature.

Tier Common privacy posture
Free Data often used for training; longest retention
Individual paid Training usually excluded; standard retention
Team / business Training excluded; admin controls; sometimes shorter retention
Enterprise Zero-retention options; contractual guarantees; DPAs

 

The practical consequence: the policy that applies to your organisation is the policy on the lowest tier anyone is actually using. One person on a free personal account undoes the terms you negotiated.

This is a real argument for centrally provided accounts. If the approved tool is good and everyone has it, nobody has a reason to use a free personal login for work. Multi-model platforms generally publish plan tiers with what each includes, and the entry tier is typically cheap enough — around $15/month — that giving every team member a governed account costs less than the risk of not doing so.

Verifying rather than trusting

Four things you can actually check:

Ask for the sub-processor list in writing. A vendor that maintains one will send it immediately. A vendor that needs a week to assemble one has not been tracking it.

Test the retention claim. Delete a conversation, then ask support to confirm what remains and where. The quality of the answer is informative regardless of its content.

Check whether the privacy setting is per-conversation or account-wide. Per-conversation modes are more useful and more honest, because they acknowledge that not every conversation needs the same treatment.

Look for a data processing agreement. If you have any regulatory obligation, you need a DPA, and whether one is available at your tier tells you how seriously the vendor treats business use.

What “private” cannot mean

Some limits are structural, and a vendor claiming otherwise is either confused or overselling.

Your prompt must be readable by the model. End-to-end encryption in the messaging sense is incompatible with the model processing your text. Anyone claiming both is describing something else — usually client-side encryption at rest, which is genuinely useful but is not the same thing.

Abuse monitoring requires some retention. A provider with truly zero retention cannot detect misuse. Short retention with a stated purpose is often the honest maximum, and treating it as a failure sets an impossible bar.

Deletion is not instant everywhere. Backups, logs, and caches expire on their own schedules. A policy stating 30 days for backups is more credible than one implying instant global erasure.

Frequently asked questions

Does private mode mean my data is encrypted end-to-end? Not in the messaging sense — the model must read your input to respond. Private mode typically means the conversation is not persisted on the provider’s servers and is not used for training.

Do AI providers train on paid-tier conversations? Most major providers exclude paid traffic from training by default, but this varies by provider and by product surface. Verify per vendor and per tier rather than assuming a category norm.

What is zero data retention? Content is processed and discarded without being stored. Usually available only on business or enterprise tiers, and often requires explicitly requesting it rather than being on by default.

Are multi-model platforms less private than going direct? Not inherently — it depends on the platform’s terms and its upstream agreements. The extra thing you must check is the sub-processor list, since your data reaches the model providers behind it.

Which plan tier should a business use for privacy? Whichever tier excludes training and offers admin controls. More importantly: make sure everyone is actually on it, because your effective policy is set by the weakest account in use.

The practical position

Do not look for a vendor that is “private”. Look for one that is specific.

Specificity is the signal: a retention period in days, a named sub-processor list, a stated position on human review, and a clear answer about which tier each promise applies to. Vendors who have done the work can answer all four in one email. Vendors who have not will send you a link to a page with a padlock on it.

 

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *