The Invisible Data Leak:
What Your Team Is Really Pasting Into AI
A leak that doesn't look like one
Classic data loss has a signature: a large transfer, an unusual destination, a file where it shouldn't be. Pasting text into a language model has none of these. The data leaves as ordinary web traffic to a reputable domain, in volumes indistinguishable from normal browsing. Which is precisely why most organisations cannot say — with any confidence — what has already left.
The behaviour is not rare, and it is not marginal. In telemetry from roughly 1.6 million workers, security vendor Cyberhaven found that 11% of everything employees pasted into ChatGPT was confidential, and that the average company was leaking sensitive material to it hundreds of times per week — with source code, client data and internal documents topping the list [1]. Two years later, the same firm's 2026 analysis reported that nearly 40% of all interactions with AI tools now involve sensitive data, much of it flowing through personal accounts outside any corporate control [2].
11% of data pasted into ChatGPT was confidential; leaks occurred hundreds of times per week per company (Cyberhaven telemetry, ~1.6M workers) [1].
77% of employees paste data into generative-AI tools, and 82% of those pastes come from personal, unmanaged accounts (LayerX browser telemetry) [3].
+93% year over year: the volume of enterprise data sent to AI/ML tools nearly doubled in a single year (Zscaler ThreatLabz, 989 billion transactions) [4].
What actually happens to a prompt
The mental model most employees hold — that a prompt is a private question answered and forgotten — is wrong in ways that matter legally. Depending on the tool and tier, prompts may be retained, logged, reviewed by humans for quality, used to train future models, or processed by sub-processors in other jurisdictions. On a free consumer tier, the default is rarely the safe one. Once a trade secret has been submitted to a third party without a confidentiality agreement in place, its status as a protected secret is, at minimum, in question.
This is not hypothetical. Within three weeks of Samsung's semiconductor division permitting ChatGPT internally in 2023, employees had reportedly pasted in proprietary source code and the recording of an internal meeting — prompting the company to clamp down almost immediately [5]. The lesson was not that Samsung's engineers were careless. It was that capable, well-intentioned people will reach for the most useful tool available unless a better-governed path exists.
Why the cost is larger than it looks
The most rigorous number in this field comes not from a security vendor but from IBM's annual, Ponemon-conducted breach study. Its 2025 edition found that 13% of organisations reported a breach involving their AI models or applications — and 97% of those lacked basic AI access controls. Breaches involving high levels of unsanctioned "shadow AI" cost, on average, USD 670,000 more than those without [6].
There is also a regulatory dimension that has stopped being theoretical. In December 2024 Italy's data-protection authority fined OpenAI €15 million over how personal data was handled in ChatGPT [7]. The direction of travel is clear: regulators now treat AI data flows as data processing like any other, subject to the same obligations — and the same penalties.
The question is not whether your confidential data is going into AI tools. The data says it already is. The question is whether any of it is going somewhere you would be able to defend in an audit.
What actually reduces the leak
Bans do not work; they push usage onto personal phones where visibility drops to zero. What works is narrower and less dramatic:
- A sanctioned path that is genuinely easier than the shadow one. An enterprise tier with a data-processing agreement, no training on your data, and SSO removes most of the incentive to reach for a personal account.
- A simple, memorable rule. Not a forty-page policy — a one-line decision people can actually apply: which categories of data may go into which class of tool.
- Visibility before control. You cannot govern what you cannot see. Browser-level or network-level insight into which AI tools are in use is the precondition for every other measure.
- Education that pairs every risk with a safe alternative. People who understand what happens to a prompt, and have a sanctioned way to get the same benefit, stop improvising.
None of this is exotic, and none of it requires treating your workforce as the threat. The employees pasting data into AI tools are, overwhelmingly, trying to do their jobs well. The failure is architectural: an unmet need, met in the shadows. Close the gap with a better-governed path, and the invisible leak shrinks to something you can actually see — and defend.