The LMVD-ID is an internal research identifier, not an official CVE identifier.
Autonomous Multi-Turn LLM Jailbreak
Large Language Models (LLMs) are vulnerable to multi-turn adversarial attacks that exploit incremental policy erosion. The attacker uses a breadth-first search strategy to generate multiple prompts at each turn…
Paper-evaluated models(3)
- GPT-3.5 Turbo
- GPT-4
- Llama 3.1 70B
Description
Large Language Models (LLMs) are vulnerable to multi-turn adversarial attacks that exploit incremental policy erosion. The attacker uses a breadth-first search strategy to generate multiple prompts at each turn, leveraging partial compliance from previous responses to gradually escalate the conversation towards eliciting disallowed outputs. Minor concessions accumulate, ultimately leading to complete circumvention of safety measures.
Examples
See the "Siege" paper (details omitted due to length of provided text, refer to the original publication).
Impact
Successful exploitation allows attackers to circumvent LLM safety restrictions and obtain disallowed information or instructions, potentially leading to the generation of harmful content, such as instructions for illegal activities, malicious code, or personally identifiable information. The attack achieves a 100% success rate on GPT-3.5-turbo and 97% on GPT-4 in a single multi-turn run.
Affected Systems
Large Language Models (LLMs) susceptible to multi-turn adversarial prompting, including (but not limited to) GPT-3.5-turbo, GPT-4, and Llama 3.1-70B.
Mitigation Steps
- Implement more robust multi-turn safety mechanisms that account for cumulative policy violations and adapt dynamically to adversarial interactions.
- Develop sophisticated detection mechanisms that identify and mitigate partial compliance leading to gradual safety erosion.
- Utilize advanced techniques to prevent the exploitation of incremental concessions in the model's responses.
- Strengthen prompt filtering and response validation to prevent the re-injection of leaked information into subsequent queries. Consider using techniques to assess risk not just on individual turns, but cumulatively throughout a conversation.
Research context and confidence
- Evidence and verification
- Paper-reported; independent reproduction is not documented.
- Primary research source linked.
- Severity
- Not rated by this catalog.
- Source and publication type
- arXiv · Research preprint.
- Peer-review status is not provided by this source.
- Author and publication status
- Author metadata is not stored; see the primary paper.
- Threat model and attacker access
- Black-box model, service, or application access.
- Related deployment categories
- Agent workflows
- Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
- Affected systems
- Large Language Models (LLMs) susceptible to multi-turn adversarial prompting, including (but not limited to) GPT-3.5-turbo, GPT-4, and Llama 3.1-70B.
Research Paper
Siege: Autonomous Multi-Turn Jailbreaking of Large Language Models with Tree Search
Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.
View PaperEvidence
This entry is based on a primary research source. Its findings are paper-reported; independent reproduction and verification are not claimed.
https://arxiv.org/abs/2503.10619Related research
- AI Browser Indirect Injection
Published October 1, 2025 · application-layer, prompt-layer, injection
- Cross-Environment Agent Jailbreak
Published December 1, 2025 · application-layer, prompt-layer, injection
- Semantic Tool Poisoning
Published December 1, 2025 · application-layer, prompt-layer, injection