Observation Poisons Agent Memory
- Published
- Analyzed
Cite & share
Citation metadata is maintained by the primary source and may reflect a later revision.
Paper-evaluated models(6)
GPT-5 Mini, GPT-5.2, GPT-oss 120B +3 more
- GPT-5 Mini
- GPT-5.2
- GPT-oss 120B
- Qwen 2.5 VL 72B
- Qwen3-VL-32B
- Qwen3.5-122B-A10B
On this page
Description
Memory-augmented LLM web agents utilizing raw trajectory memory are vulnerable to Environment-injected Trajectory-based Agent Memory Poisoning (eTAMP). Attackers can embed malicious instructions within user-generated web content (e.g., product pages, forum posts). When the agent processes this content during a routine task, the instructions are passively ingested into its raw trajectory memory. During subsequent, entirely separate tasks on different websites, semantic retrieval mechanisms pull this poisoned trajectory into the context window. The dormant instructions then activate, allowing attackers to execute unauthorized cross-site actions. This mechanism bypasses standard domain-level permission defenses because the injection occurs when the agent is restricted to the source site, but executes later when the agent legitimately holds permissions for the target site. Notably, this vulnerability is heavily amplified by "Frustration Exploitation": attack success rates increase up to 8x when the agent encounters environmental stress, such as network latency, dropped clicks, or garbled text.
Examples
A manipulated page contaminates the agent’s stored trajectory. A later task retrieves that trajectory and the agent attempts an unrelated action. The paper compares baseline injection, authority framing, and environmental stress in a sandbox; see the primary study (opens in a new tab).
Impact
A single passive environmental observation creates a persistent, cross-session, and cross-site vulnerability. Attackers can hijack agent sessions to perform unauthorized actions (e.g., forced purchases, posting promotional 5-star reviews, initiating API calls) on unrelated platforms. Because a single poisoned trajectory may be retrieved for multiple future tasks based on semantic similarity, one exposure can persistently compromise subsequent agent operations.
Affected Systems
- WebArena and VisualWebArena agents using raw trajectory memory. Section 2.4 reports GPT-5-mini, GPT-5.2, GPT-OSS-120B, Qwen2.5-VL-72B, Qwen3-VL-32B, and Qwen3.5-122B-A10B as evaluated backends.
- Commercial browsers mentioned in the discussion are motivation, not evaluated products.
Mitigation Steps
- Implement memory content filtering and sanitization prior to writing environmental observations to the agent's long-term memory store.
- Apply anomaly detection mechanisms on context retrieved from memory before injecting it into the agent's active prompt.
- Enforce robust instruction hierarchies to explicitly segregate user/system directives from retrieved memory exemplars, preventing past trajectories from overriding current task objectives.
Evidence
The arXiv submission history (opens in a new tab) dates the first submission to April 3, 2026. The evaluated backends are listed in version 2, Section 2.4 (opens in a new tab); the paper does not specify an Instruct suffix for Qwen2.5-VL-72B. Results are paper-reported, not independently reproduced here.
Research context and provenance
- Catalog identifier
- LMVD-5a088378
- Internal research identifier, not an official CVE identifier.
- Evidence and verification
- Paper-reported; independent reproduction is not documented.
- Primary source plus a dedicated evidence section.
- Severity
- Not rated by this catalog.
- Source and publication type
- arXiv · Research preprint.
- Peer-review status is not provided by this source.
- Author and publication status
- Author metadata is not stored; see the primary paper.
- Threat model and attacker access
- Black-box model, service, or application access.
- Related deployment categories
- Vision-language models; Agent workflows
- Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
- Affected systems
- WebArena and VisualWebArena agents using raw trajectory memory. Section 2.4 reports GPT-5-mini, GPT-5.2, GPT-OSS-120B, Qwen2.5-VL-72B, Qwen3-VL-32B, and Qwen3.5-122B-A10B as evaluated backends. Commercial browsers…
Research Paper
Poison Once, Exploit Forever: Environment-Injected Memory Poisoning Attacks on Web Agents
Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.
View PaperRelated research
- Mobile Agent Visual Spoofing
Published February 11, 2026 · Application layer, Prompt layer, Injection
- Hybrid Agent Prompt Injection
Published May 28, 2025 · Prompt layer, Application layer, Injection
- On-Device LLM Hijacking
Published May 19, 2025 · Application layer, Model layer, Prompt layer