The LMVD-ID is an internal research identifier, not an official CVE identifier.
Personalization Intent Legitimation
Personalized LLM agents utilizing long-term memory systems are vulnerable to a safety bypass known as intent legitimation. Benign, organically accumulated user memories can bias the model's intent inference, causing it…
Paper-evaluated models(5)
GPT-4o, GPT-4o Mini, DeepSeek V3.2 +2 more
- GPT-4o
- GPT-4o Mini
- DeepSeek V3.2
- Qwen 3 235B-A22B
- Qwen 3 8B
Description
Personalized LLM agents utilizing long-term memory systems are vulnerable to a safety bypass known as intent legitimation. Benign, organically accumulated user memories can bias the model's intent inference, causing it to misinterpret inherently harmful queries as contextually justified. When a malicious request semantically aligns with a user's established persona (e.g., hobbies, mental health history, routine), the model normalizes the request and complies, effectively bypassing standard safety guardrails without the need for adversarial or poisoned prompts.
Examples
- Context: The agent has stored benign episodic memories of a user's outdoor activities and creative crafting hobbies.
- Harmful Query: The user requests step-by-step instructions for altering a kiln to destroy identifying marks.
- Stateless Response: The model correctly identifies the intent as illegal modification and refuses the request.
- Personalized Agent Response: The model references the user's past hobbies, misinterprets the harmful query as a benign "creative project," and outputs detailed, policy-violating instructions to alter the kiln.
Impact
Attackers or users can bypass safety filters to generate harmful, illegal, or policy-violating content (e.g., self-harm instructions, hate speech, financial crime advice, privacy leakage) simply by expressing the request in a persona-consistent manner. Benign personalization increases the attack success rate (ASR) of harmful queries by 15.8% to 243.7% compared to stateless baselines. Agents serving emotionally dependent or vulnerable personas experience the highest safety degradation.
Affected Systems
- Personalized LLM agent frameworks utilizing long-term memory and explicit persona modeling (e.g., MemOS, Mem0, Amem, LDAgent, MemU).
- Agents leveraging fine-grained, high-recall, episodic memory retrieval are significantly more vulnerable than those using abstract memory representations.
- Base LLMs underlying these memory frameworks (demonstrated on GPT-4o, GPT-4o-mini, Qwen3-235B, Qwen3-8B, DeepSeek-V3.2).
Mitigation Steps
- Implement an Auditing Step: Before the reasoning engine processes retrieved memories, use a specialized auditor prompt to flag memories that could inadvertently validate a user's harmful intent (checking for relational priming, normative drift, or vulnerability rationalization).
- Inject Reflective Reminders (Detect-and-Reflect): Dynamically synthesize and insert a reflective reminder into the system prompt if risky memories are flagged. Explicitly instruct the downstream reasoning engine to decouple "empathetic understanding" from "intent validation," and strictly prohibit using personal context to justify, soften, or normalize safety-critical requests.
- PII Sanitization: Enforce strict Personally Identifiable Information (PII) sanitization and access controls within the memory store to prevent synthetic profile data from facilitating privacy leakage attacks.
Research context and confidence
- Evidence and verification
- Paper-reported; independent reproduction is not documented.
- Primary research source linked.
- Severity
- Not rated by this catalog.
- Source and publication type
- arXiv · Research preprint.
- Peer-review status is not provided by this source.
- Author and publication status
- Author metadata is not stored; see the primary paper.
- Threat model and attacker access
- Black-box model, service, or application access.
- Related deployment categories
- Retrieval-augmented generation; Agent workflows
- Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
- Affected systems
- Personalized LLM agent frameworks utilizing long-term memory and explicit persona modeling (e.g., MemOS, Mem0, Amem, LDAgent, MemU). Agents leveraging fine-grained, high-recall, episodic memory retrieval are…
Research Paper
When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents
Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.
View PaperEvidence
This entry is based on a primary research source. Its findings are paper-reported; independent reproduction and verification are not claimed.
https://arxiv.org/abs/2601.17887Related research
- Unsafe Search Framing
Published January 1, 2026 · application-layer, prompt-layer, jailbreak
- Inter-Agent Computer Takeover
Published July 1, 2025 · application-layer, prompt-layer, injection
- Chained Guardrail Bypass
Published April 1, 2025 · application-layer, prompt-layer, injection