In April 2023, three engineers at Samsung’s semiconductor division did something ordinary and, it turned out, expensive. Within the same stretch of weeks, one pasted source code for a facility measurement system into ChatGPT, another uploaded code used to detect defective equipment, and a third fed it internal meeting notes and asked for a summary. Samsung found out, capped prompts at 1,024 bytes as a stopgap, then banned the tool outright.
That was three years ago, and the underlying problem has not gone away. It has just moved. Vendor data policies are far more mature in 2026 than they were then, but the harder question executives now face is not “does the vendor promise not to misuse our data.” It is “where does our data actually go, who can reach it, and what happens to it once we lose direct control.” That question shapes how an AI development engagement gets scoped from the first conversation, well before anyone writes a line of code.
This piece looks specifically at what happens to sensitive business data and intellectual property once it reaches an AI system, third-party API or self-hosted, and what actually protects it. Compliance frameworks and vendor lock-in are real concerns too, but they are separate questions from the one here: where does the data go, and who else can see it.
What Actually Happens to a Trade Secret the Moment It Hits an AI Tool
Most executives picture a single dramatic leak: a hacker, a breach headline, a regulator’s letter. The more common failure mode is quieter and happens by default, every day, inside normal work.
LayerX’s 2025 workplace AI research found that 77% of employees have pasted company data into a public AI tool at some point, and a separate vendor analysis of prompt and file traffic found sensitive information present in more than 4% of prompts and 22% of file uploads sent to AI systems. None of that requires a breach.
An employee drafting a client proposal pastes in a competitor’s pricing under NDA. A finance analyst uploads a spreadsheet with unreleased earnings to get a summary written faster. A developer pastes a proprietary algorithm into a coding assistant to debug it.
Disney learned the cascading version of this risk in 2024, when roughly a terabyte of data leaked from its Slack environment, including code and intellectual property tied to unreleased projects. Amazon’s own internal communications have warned staff not to paste company secrets into ChatGPT, after leaked employee posts described the company’s internal AI assistant surfacing data center locations and unreleased feature details it should never have had access to. None of these incidents required a sophisticated attacker. They required an employee with a deadline and a tool that made the task faster.
Vendor Data Policies Have Matured, and the Fine Print Still Matters
Give the major AI vendors credit where it is due. By mid-2026, OpenAI, Anthropic, Google, Microsoft, and Mistral all publish enterprise-tier commitments that exclude customer data from model training by default, and all offer contractual Data Processing Agreements for regulated data.
| Vendor | Trains on enterprise data by default | Default retention | Reduced or zero retention |
|---|---|---|---|
| OpenAI (Business/Enterprise/API) | No | 30 days | Zero-data-retention available on eligible endpoints |
| Anthropic (Claude for Work/Enterprise) | No | 7 days (API, since Sept 2025) | Zero-data-retention for qualified enterprise customers |
| Google Gemini (API/Workspace) | No | Not logged for training by default | Zero-data-retention via API setting |
| Microsoft Copilot | No | Governed by org retention policy | US and EU data residency since April 2026* |
| Mistral | No | 30 days by default | Zero-data-retention available; EU residency is the default |
* Microsoft’s “Flex Routing,” introduced the same month, can route EU inference outside the EU Data Boundary during peak demand, so confirm the current routing configuration rather than assuming residency is absolute by default.
That table is the good news, and it is real progress. Two caveats keep it from being the whole story. First, none of these policies cover sub-processors, meaning the cloud infrastructure, analytics tooling, and legal-hold processes that sit behind the vendor’s own systems, and most organizations have limited contractual visibility into that chain. Second, and this is the one worth sitting with: a consumer-tier account is a different product with different defaults. OpenAI’s free and Plus consumer plans still let users opt in to training in ways enterprise plans do not, and the practical risk shows up when employees use their own personal ChatGPT accounts for work without IT ever knowing the data left the building.
Where “Zero Retention” Stops Being Zero
Vendors that promise zero data retention mean it, under normal operating conditions. Normal operating conditions are not the only conditions that matter to a business holding real secrets.
Between May 2025 and January 2026, a US federal magistrate ordered OpenAI to preserve all ChatGPT conversation logs indefinitely as part of the New York Times v. OpenAI litigation. OpenAI complied, retaining data that users had already deleted. In January 2026, a federal judge went further and ordered OpenAI to produce 20 million de-identified conversation logs to the plaintiffs.
One description for this category of data making the rounds among privacy practitioners: “zombie data,” deleted by the user, alive anyway because a court said so.
Two more 2025 to 2026 incidents make the same point from a different angle. In July 2025, Google’s search index picked up thousands of ChatGPT conversation links that users believed they had shared privately, exposing personal and business content to public search. Early in 2026, researchers disclosed a DNS-based side channel vulnerability in ChatGPT that could have let sensitive conversation data leak past normal guardrails; OpenAI patched it on February 20, 2026, with no confirmed exploitation.
Retention policy is one layer of protection. Platform security and legal exposure are separate layers, and a strong retention policy does not cover either one.
Worth naming plainly, since it gets asserted casually and it should not be: “zero data retention” is not an absolute guarantee. It is a strong operational commitment that a court order or a platform vulnerability can still override.
Employees Are the Biggest Leak Vector, and Policy Alone Won’t Stop It
If you fix every vendor contract and still allow staff to run confidential work through personal AI accounts, you have not actually reduced your exposure. Security teams agree, almost without exception, that the dominant loss vector is not vendor misbehavior. It is employees using unapproved, unmanaged tools.
There is a legal wrinkle here too that most executives have not clocked yet. In United States v. Heppner (February 2026), a federal court ruled that a defendant’s exchanges with a consumer-grade AI tool were not privileged, even though he believed the conversation was confidential. The case involved a defendant’s own use of the tool rather than his attorneys’, but legal commentators widely read it as a warning for law firms too: a general counsel’s office cannot safely assume a personal ChatGPT account preserves privilege for legal analysis, because the privilege itself does not survive the tool.
The financial cost of this gap is measurable. IBM’s 2025 Cost of a Data Breach Report found that breaches involving shadow AI, meaning AI tools deployed or used without approval or governance, add an average of $670,000 to the total cost of the incident. That premium exists because unmanaged tools have no audit trail, no data map, and often no one at the company who even knows the exposure happened until well after the fact.
What Self-Hosting Actually Buys You, and Where It Falls Short
This is where the instinct to self-host comes from, and the instinct is not wrong. Running inference on infrastructure you control removes the vendor as a third party with any access to your data at all. No sub-processor chain, no cross-border transfer to reason about, no vendor legal hold that can surface your deleted conversations years later.
It is also not a shortcut. According to Wiz’s 2026 State of AI in the Cloud research, 81% of organizations use managed AI services and 90% run self-hosted models, meaning most large organizations run both rather than treating either as a universal answer. Security teams that inherit a self-hosted deployment frequently find it has never had a vulnerability assessment, a penetration test, or basic logging configured, which in practice can be a worse security posture than a vendor’s managed API with a signed DPA.
GDPR and the EU AI Act do not relax for self-hosted systems either: Article 50 transparency obligations, which took effect August 2, 2026, apply regardless of where the model runs. Self-hosting moves the compliance burden entirely inside your walls. It does not remove it.
One more misconception deserves a direct correction, because it causes real budget mistakes: an “enterprise” plan from a major AI vendor does not mean the model runs on your infrastructure. It almost always means better contract terms layered on top of the vendor’s own cloud, a DPA, audit logging, and stronger retention controls, while inference still happens on servers you do not control. If data residency inside your own network is the genuine requirement, an enterprise contract does not satisfy it; only an architecture that keeps inference on infrastructure you own or lease under your own jurisdiction does. Our companion pieces on who actually builds self-hosted AI successfully and what it costs to run a local inference cluster walk through what that engineering commitment looks like in practice, including where the token-cost math does and does not favor buying hardware.
Building a Real Data Boundary, Whichever Path You Choose
The choice between API and self-hosted is not the real decision. The real decision is what data ever reaches an AI system in the first place, and under what controls, and that question applies equally to both deployment paths.
A defensible boundary generally includes:
- Classify before you connect. Know which data categories (client records, unreleased IP, regulated personal data) are allowed near any AI system at all, and which require redaction or pseudonymization first.
- Kill consumer-account shadow AI. Block personal ChatGPT, Gemini, and Claude accounts on managed devices, and give staff an approved enterprise alternative so the ban doesn’t just push the behavior further underground.
- Get a signed DPA for anything that touches personal data, whether the recipient is a vendor’s API or an internal team standing up a self-hosted model that a data processor supports. Our guide to AI agents and GDPR covers what that agreement needs to include.
- Know your actual retention and legal-hold exposure beyond the marketing page. Ask the vendor directly what happens to your data under subpoena, and document the answer.
- Weigh vendor dependency as its own risk, separate from data exposure. Our piece on AI agent platform lock-in covers what happens when a provider changes terms after your workflows are already built around them.
- If you self-host, budget for the ongoing security work alongside the hardware. Patching, access control, and monitoring are recurring costs that continue well past the initial setup.
Who this favors self-hosting: organizations handling data that is genuinely regulated or covered by an NDA, with sustained usage volume and a team that can own the security work for years at a time.
Who this doesn’t fit: teams with low or unpredictable AI usage, general-sensitivity work like marketing drafts or internal brainstorming, and no one available to own ongoing infrastructure security. For that profile, a reputable vendor’s enterprise-tier API with a signed DPA and disciplined internal data controls will protect trade secrets more reliably than a self-hosted deployment nobody has the bandwidth to maintain properly.
Either path works when it is built deliberately, with a clear answer to what data reaches the model and who can see it after. Neither one works as a default assumption borrowed from a vendor’s marketing page or an engineer’s weekend project.
Frequently asked questions
Does self-hosting an AI model automatically keep our data private?
No. Self-hosting removes the vendor as a party that can see, retain, or be subpoenaed for your data, but it does not automatically make the deployment secure or compliant. Without access controls, encryption, monitoring, and regular patching, a self-hosted model can leak data as easily as a cloud API, and security researchers report that self-hosted deployments are frequently left unmaintained once the initial project wraps up.
What is shadow AI, and why is it the biggest privacy risk?
Shadow AI is employees using unapproved AI tools, often personal ChatGPT or Gemini accounts, to handle work tasks without IT knowing. It is the source of most documented leaks: Samsung engineers pasted proprietary source code into ChatGPT in 2023, and IBM found in 2025 that shadow AI incidents add an average of $670,000 to breach costs. A vendor data policy cannot protect a leak that never touches the vendor systems in the first place.
Are AI vendors' "zero data retention" policies reliable?
They are reliable for everyday operational privacy but not absolute. OpenAI, Anthropic, and Google all offer no-training and reduced-retention options on enterprise plans, yet a 2025 US federal court order forced OpenAI to preserve chat logs indefinitely for litigation, overriding deletion requests users had already made. Treat a retention policy as real protection against routine misuse. It offers no protection once a court orders preservation or issues a legal hold.
Does an "enterprise" AI plan mean the model runs on our own infrastructure?
No, in almost every case it does not. An enterprise tier from OpenAI, Anthropic, or Microsoft typically means a signed Data Processing Agreement, audit trails, and stronger retention controls, but inference still happens on infrastructure the vendor operates. If keeping data inside your own network is the actual requirement, only a genuinely self-hosted or dedicated private-cloud deployment satisfies it.
When does self-hosting actually make sense for protecting trade secrets?
It makes sense when the data is genuinely regulated or contractually sensitive, such as unreleased IP under an NDA, health records, or financial data, when usage volume is sustained enough to justify the engineering effort, and when your team can own ongoing patching and monitoring. For lower-sensitivity work such as marketing copy or general drafting, a reputable vendor enterprise-tier API with a signed DPA is usually the simpler path.