Confuse the Model, Control the Flow: Understanding and Mitigating Privacy Leakage from LLM Agents with Information Flow Control

| Source: arXiv AI

Tags: LLM agents, privacy, information flow control, MCP, agent security, data leakage

FLOWSEAL enforces privacy in personal AI agents through a tool-level interceptor outside the LLM context—reducing leak rates from 52.2% to 0.5% against novel attacks that bypass all prompt-based defenses, validated on real MCP tool calls across three benchmarks.

Details

Personal AI agents with access to user data face a structural privacy vulnerability: when enforcement is a judgment call the LLM makes within the same context an adversary controls, the defense and the attack surface coincide. This paper makes that structural problem concrete with three new attacks requiring only normal agent interactions—no prompt injection needed. Collaborative Workspace Lure reframes extraction as collaborative work; Semantic Obfuscation Attack triggers disclosure through omission; Channel Decoupling Attack splits the extraction request across independent channels. All three substantially outperform attacks the existing defenses were designed to stop.\n\nFLOWSEAL solves the root cause by enforcing privacy outside the LLM context entirely, using a tool-level interceptor grounded in data provenance and an information-flow-control lattice. The interceptor operates on tool calls, not on LLM reasoning—removing the attack surface rather than patching reasoning behavior.\n\nAcross three benchmarks, five prompt-based baselines, eight attacks, and a real agent executing live MCP tool calls, FLOWSEAL reduces leak rates to near zero—52.2% to 0.5% against the Collaborative Workspace Lure attack—while preserving task utility and remaining backend-agnostic. This is one of the most thorough privacy defenses for LLM agents published to date.