Where your data goes.
Sending documents to a cloud AI service puts them on someone else's servers, under terms that can change. These are documented cases: lawsuits, rulings, settlements, regulator actions and the companies' own statements. Open cases are described as allegations. Status as of 14 September 2026.
Copyrighted work used to train AI.
Authors, artists, news publishers and record labels have gone to court over AI trained on their work. Courts have ruled both ways, and most cases are still open.
The New York Times v. OpenAI and Microsoft
The New York Times has sued OpenAI and Microsoft. It alleges that millions of its articles were used without permission to train AI models. The companies deny wrongdoing, and the case is pending.United States · filed Dec 2023 · pendingAP, Mar 2025Docket report, Sep 2026
Authors v. Anthropic (Bartz)
In 2025 a US judge found that training Anthropic's models on books was fair use, but that copying millions of pirated books into a central library was not. Anthropic then agreed to a 1.5 billion dollar settlement with authors, which the court approved in July 2026.United States · settlement approved Jul 2026Reuters, Jun 2025Authors Guild, Jul 2026
Authors v. Meta (Kadrey)
Authors sued Meta, alleging their books were used without permission to train its Llama models. A US judge ruled for Meta in 2025, and wrote that the ruling did not mean Meta's use was lawful, only that these plaintiffs made the wrong arguments.United States · ruled for Meta Jun 2025 · not yet finalTechCrunch, Jun 2025
Getty Images v. Stability AI
Getty Images sued Stability AI in the UK over the use of its photos to train Stable Diffusion. In 2025 the High Court rejected most of Getty's copyright claims but found limited trade mark infringement over Getty watermarks in outputs. Getty has permission to appeal.United Kingdom · judgment Nov 2025 · appeal pendingJudiciary of England and WalesWiggin, permission to appeal
Artists v. Stability AI, Midjourney and others (Andersen)
Artists including Sarah Andersen and Karla Ortiz have sued Stability AI, Midjourney and others. They allege their artwork was used without permission to build AI image generators. A US judge allowed the core copyright claims to proceed in 2024, and the case is pending.United States · filed Jan 2023 · pendingThe Art Newspaper, Aug 2024
Disney and Universal v. Midjourney
In 2025 Disney and Universal sued Midjourney. They allege it trained on their films and characters without permission and lets users generate copies of those characters. Midjourney says its use is fair use, and the case is pending.United States · filed Jun 2025 · pendingGeorgetown Law Tech InstituteThe Art Newspaper, Jul 2026
Record labels v. Suno and Udio
In 2024 the three major record labels sued AI music companies Suno and Udio, alleging their recordings were copied without permission to train the AI. Several labels have since settled and signed licensing deals, and Universal and Sony's case against Suno is ongoing.United States · filed Jun 2024 · partly settled, partly pendingMusic Business Worldwide, Aug 2026Bloomberg, Nov 2025
Programmers v. GitHub, Microsoft and OpenAI (Copilot)
Programmers have sued GitHub, Microsoft and OpenAI. They allege that the Copilot coding assistant was trained on their open-source code and reproduces it without the credit and license terms that code requires. The case is on appeal.United States · filed Nov 2022 · on appealBloomberg Law, Feb 2026Plaintiffs' case page
Thomson Reuters v. ROSS Intelligence
In 2025 a US federal judge ruled that ROSS Intelligence's use of Thomson Reuters' Westlaw content to train an AI legal-research tool was not fair use. The decision is on appeal.United States · ruled Feb 2025 · on appealLawSites, Jun 2026
GEMA v. OpenAI
In 2025 a Munich court ruled that OpenAI infringed copyright because its models memorized and reproduced German song lyrics. OpenAI has appealed, and the ruling is not final.Germany · judgment Nov 2025 · on appealInitiative Urheberrecht, Dec 2025Euronews, Nov 2025
Music publishers v. Anthropic (Concord)
Music publishers have sued Anthropic. They allege that song lyrics were used without a license to train its Claude models, and the case is pending.United States · filed Oct 2023 · pendingMusic Business Worldwide, Mar 2026
On Exalt local AI hardware, questions and documents are processed on the box, not sent to an outside AI service for answering or training.
Company data exposed through AI tools.
Some were bugs or breaches at the provider, some were sharing features that made conversations searchable, and one was employees pasting source code into a chatbot.
Samsung source code in ChatGPT
In 2023 Samsung restricted staff use of ChatGPT and similar tools after engineers entered confidential source code into the chatbot.South Korea · May 2023TechCrunch, May 2023
ChatGPT conversation titles and payment details
In March 2023 a software bug let some ChatGPT users see the titles of other users' conversations. It may also have exposed partial payment details for about 1.2 percent of paying subscribers active at the time.Mar 2023 · fixedThe Register, Mar 2023
OpenAI internal forum
The New York Times reported in 2024 that a hacker had accessed an internal OpenAI discussion forum in 2023 and obtained details about its AI technology. OpenAI did not announce the breach publicly at the time.2023 · reported Jul 2024Reuters, Jul 2024
Microsoft AI research storage link
In 2023 security researchers found that a misconfigured link shared by Microsoft AI researchers had exposed 38 TB of internal data, including passwords, keys and internal messages. Microsoft said no customer data was exposed.Disclosed Sep 2023 · fixedWiz Research
DeepSeek open database
In January 2025 researchers found a publicly accessible DeepSeek database containing user chat histories and secret keys. DeepSeek secured it after being notified.Jan 2025 · securedWiz Research
EchoLeak in Microsoft 365 Copilot
In 2025 researchers showed that a single malicious email could have tricked Microsoft 365 Copilot into sending internal company data to an attacker without anyone clicking anything. Microsoft rated the flaw critical and fixed it, and there is no evidence it was exploited.Jun 2025 · patchedCVE-2025-32711Researchers' paper
Shared ChatGPT conversations in search results
In 2025 thousands of shared ChatGPT conversations turned up in Google search results after users ticked a 'discoverable' option. OpenAI removed the option, saying it created too many chances to share things by accident.Jul 2025 · feature removedEngadget, Aug 2025
Shared Grok conversations in search results
In 2025 reporters found that hundreds of thousands of conversations with xAI's Grok chatbot had become searchable on Google through its share links.Aug 2025TechCrunch, Aug 2025
OpenAI customer details through an analytics vendor
In November 2025 OpenAI said an analytics vendor it used had been breached, exposing names and email addresses of some API customers. No chat content or keys were exposed.Nov 2025BleepingComputer, Nov 2025
Deleted ChatGPT conversations kept for a lawsuit
In 2025 a US court ordered OpenAI to keep ChatGPT conversations that users had deleted, as evidence in a copyright lawsuit. The order ended in September 2025, but logs already kept were held for the case.United States · May 2025 · ended Sep 2025Court order, May 2025Engadget, Oct 2025
What stays on the box.
- Questions and answers are always processed on the box, not sent to an outside AI service.
- Documents uploaded to the box always stay on it.
- Sources on your network, such as file shares, databases and business systems, are read in place as often as you set.
- Sign-in through a directory on your network stays on your network.
- Stored documents, the index, conversation history and the audit log are encrypted at rest. History is kept as long as your retention settings say.
Terms that changed, and regulators who stepped in.
Terms of service decide what a provider may do with what you send it, and they can change after you sign up.
Zoom terms of service
In 2023 Zoom faced backlash over terms that appeared to allow it to use customer content to train AI. It rewrote them to state that it does not train AI on customers' calls, chats or files.Aug 2023 · terms revisedZoomThe Record
Slack privacy principles
In 2024 customers found that Slack used workspace data by default to train some of its machine-learning features, and that opting out meant emailing Slack. Slack says it does not train generative AI on customer data.May 2024SlackEngadget, May 2024
Adobe terms of use
In 2024 an update to Adobe's terms led customers to fear their work could be accessed and used for AI. Adobe said it does not train generative AI on customer content and rewrote the terms.Jun 2024 · terms revisedAdobe, Jun 2024
Meta training on EU users' posts
Meta paused plans to train its AI on EU users' public posts in 2024 after regulator engagement. It went ahead in 2025 with an opt-out.European Union · 2024 to 2025Irish Data Protection CommissionThe Register, May 2025
X and Grok training on EU users' posts
In 2024 Ireland's data protection regulator went to court over X's use of EU users' posts to train its Grok AI. X agreed to stop using that data, and the regulator opened a formal inquiry in 2025.European Union · 2024 · inquiry openIrish Data Protection Commission, Sep 2024Irish Data Protection Commission, Apr 2025
LinkedIn training on member data
In 2024 LinkedIn began using members' data to train generative AI by default, and updated its terms only after this was reported. In November 2025 it extended this to members in the EU, Switzerland, Canada and Hong Kong, again on by default. Opting out does not reverse training already done.Sep 2024 · expanded Nov 2025LinkedInTechCrunch, Sep 2024
Italy's privacy regulator and OpenAI
In 2024 Italy's privacy regulator fined OpenAI 15 million euros, finding it had trained ChatGPT on personal data without an adequate legal basis. A Rome court annulled the fine in 2026 on jurisdictional grounds, without ruling on those findings.Italy · fined Dec 2024 · annulled Mar 2026Euronews, Dec 2024Reuters, Mar 2026
Italy's privacy regulator and DeepSeek
In January 2025 Italy's privacy regulator ordered DeepSeek to stop processing Italian users' data after the company gave what the regulator called insufficient answers about how it handles personal data.Italy · Jan 2025 · order in forceGarante per la protezione dei dati personaliReuters, Jan 2025
FTC on quietly changed terms
The US Federal Trade Commission has warned that quietly changing terms of service to start using customer data for AI training can be an unfair or deceptive practice.United States · Feb 2024Federal Trade Commission
Amazon Alexa recordings
In 2023 Amazon agreed to pay 25 million dollars to settle FTC charges that it kept children's Alexa voice recordings, despite deletion requests, and used them to improve its algorithm.United States · May 2023 · settledFederal Trade Commission
Anthropic consumer terms
In 2025 Anthropic updated its consumer terms so that Claude chats can be used for model training and kept for up to five years unless users turn the setting off. Business and API plans were excluded.Aug 2025Anthropic
WeTransfer terms of service
In 2025 WeTransfer revised its terms after users objected to a clause that mentioned using uploaded files to improve machine-learning models. The company said it does not use AI on shared files.Jul 2025 · terms revisedThe Register, Jul 2025
Put local AI to work where it matters most.
Bring the questions your legal and security teams ask about AI vendors. We go through where your data would go before anything is installed.