Ai Engineering 4 min read

Researchers Find Chatbots Leaking Conversation Data to Advertisers

An IMDEA Networks study of nine conversational AI services found multiple providers exposing conversation titles, prompts, and screenshots to third-party ad tech alongside persistent user identifiers, and some exposing conversation permalinks that trackers can read in full.

A systematic privacy study of the conversational AI industry has found that chatbots are importing the web’s tracking economy wholesale, and in some cases going further than websites ever did. “Prompt like a Butterfly, Sting like a Tracker: A Privacy Analysis of Web and Mobile Conversational AI Agents,” from researchers at IMDEA Networks Institute (Garcia-Herrero, Oliveira, Sanchez, Vallina-Rodriguez and colleagues), analyzed the web and mobile deployments of nine prominent conversational AI services using static and dynamic analysis. The findings, which hit the front page of Hacker News on September 29 with the full paper publicly available: multiple providers disclose sensitive conversation-derived artifacts, including titles, prompts, and screenshots, to third-party advertising and tracking services, often bundled with persistent user identifiers that enable attribution across sessions.

The Findings That Matter

Three results stand out. First, conversation metadata leaks: the artifacts a chatbot generates around your conversation (the auto-generated title, the prompt text itself, screenshots the app captures) are exposed to third parties, which means the tracking ecosystem receives a summary of what you asked an AI about, attachable to your identity. Second, some providers publicly expose conversation permalinks without access controls, meaning trackers that acquire the link can read the entire conversation; sharing a chatbot answer is an operation users perform constantly, and the study found it can hand the whole exchange to whatever scripts live on the receiving end. Third, consent choices, subscription tiers, and access-control settings do not reliably change the exposure, which undermines the standard industry defense that users who care can pay for privacy.

The team went through the formal process most studies skip: responsible disclosure to the affected providers and to competent European Data Protection Authorities, plus a benchmarking of the observed practices against the GDPR and the ePrivacy Directive. That framing signals where this goes next; the paper is not just documentation, it is a regulatory referral.

Why Advertising Changes Chatbot Data Flows

The context is the industry’s business model shift. ChatGPT’s advertising business launched this year, hit $1 billion annualized revenue within 200 days, and expanded to 31 European markets in August. Advertising economics require measurement, and measurement requires identity, and identity requires data flows to ad tech. What the IMDEA study demonstrates is that the ad-tech integration is not confined to the ads surface: the tracking plumbing arrives with the ad business and touches the conversations themselves. A web page that leaks your browsing to thirty trackers is bad; a chatbot that leaks what you asked it about your medical symptoms, legal problem, or relationship, with a persistent identifier attached, is categorically more sensitive, because prompt content is self-disclosure by definition.

The exposed-permalink finding deserves special attention from anyone building on shared-chat features. A permalink is a capability: possession of the URL grants access. When there is no access control, every link is public-by-default, and any third-party script, email scanner, or messaging-app link preview becomes a full-conversation reader. This is the same class of bug as unguessable-but-public Google Docs links from a decade ago, and the fix is known: authenticated capability URLs, no-crawler headers, and expiring tokens. That multiple providers shipped without them suggests sharing features were built for virality first and audited for privacy never.

What to Watch

The responsible-disclosure process means European DPAs now hold systematic evidence, and the GDPR analysis maps each finding to specific legal exposure, so the study reads like the opening filing of a regulatory process. Watch for: which of the nine providers are named when the paper’s disclosure window closes; whether any DPA opens a formal investigation on the permalink exposure, which is the hardest finding to defend; and whether chatbot platforms start publishing tracker disclosures the way websites publish cookie policies, because after this paper, the claim that conversational AI is private-by-default is dead. The study’s own conclusion is the quiet part: AI-mediated interaction is a novel attack surface where provider-generated artifacts become trackable, and the safeguards governing it are currently an afterthought.

Get Insanely Good at AI

Get Insanely Good at AI

The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.

Keep Reading