When Malicious Instructions Persist: Persistent Memory Poisoning Attack on Harness-Based Agents

| Source: arXiv AI

Tags: security, Claude-Code, agent-security, memory-poisoning, prompt-injection, LLM-agents

Researchers demonstrate PMPA, a persistent memory poisoning attack targeting harness-based agents including Claude Code: malicious instructions embedded in external content get written to persistent memory, achieving 81.7% cross-session attack success on Claude Code while fully preserving benign task performance—making the attack stealthy.

Details

Modern LLM agent frameworks like Claude Code and OpenClaw integrate persistent memory, tool use, and runtime control—but this integration creates a new attack surface. PMPA embeds malicious instructions into benign external sources (documents, web pages, or other content the agent reads) and induces the agent to write them into its persistent memory without requiring direct access to the agent framework. Once stored, the poisoned memory persists across sessions: future sessions retrieve and execute the malicious instructions, leaking private information or triggering unauthorized actions. The attack is evaluated across different LLM backbones, input modalities, and trigger scenarios. Results are striking: PMPA achieves average Injection Success Rate / Cross-session Attack Success Rate of 73.7%/55.5% on OpenClaw and 66.9%/81.7% on Claude Code. Claude Code's higher C-ASR (81.7%) suggests its memory retrieval mechanism makes poisoned instructions particularly effective across sessions. Critically, benign task performance is fully preserved during and after the attack, making PMPA essentially invisible to behavioral monitoring. A targeted prompt-level defense reduces injection rates in many settings, but once the persistent memory has been poisoned, prompt defenses provide limited protection. This gap—injection prevention works but post-injection remediation does not—is the critical finding for agent framework developers and enterprise deployers.