Skip to main content
LLM Security Database
Skip to research details
Back to research findings
LMVD-ID: b431062b
Paper published October 1, 2024
Entry analyzed December 28, 2024
Paper-reported evidence
Confidence: Source-linked

The LMVD-ID is an internal research identifier, not an official CVE identifier.

Autonomous Jailbreak Agent

Large Language Models (LLMs) are vulnerable to jailbreak attacks using autonomously discovered strategies. AutoDAN-Turbo, a black-box attack method, demonstrates the ability to discover novel and highly effective…

BibTeX citation

Paper-evaluated models(8)

Gemini Pro, Gemma 7B IT, GPT-4-1106-turbo +5 more
  • Gemini Pro
  • Gemma 7B IT
  • GPT-4-1106-turbo
  • Llama 2 13B Chat
  • Llama 2 70B Chat
  • Llama 2 7B Chat
  • Llama 3 70B
  • Llama 3 8B

Description

Large Language Models (LLMs) are vulnerable to jailbreak attacks using autonomously discovered strategies. AutoDAN-Turbo, a black-box attack method, demonstrates the ability to discover novel and highly effective jailbreak strategies without human intervention, achieving a high success rate (e.g., 88.5% on GPT-4-1106-turbo) in eliciting harmful or unsafe responses from LLMs. The attack leverages a lifelong learning agent to iteratively refine attack strategies based on model responses, resulting in increasingly effective prompts that bypass safety mechanisms.

Examples

See https://github.com/SaFoLab-WISC/AutoDAN-Turbo (opens in a new tab) for code and a detailed example of the attack process. Specific examples of generated prompts and resulting LLM responses are included in the paper's Appendix. One example involves prompting the model to provide detailed instructions for synthesizing dimethylmercury, a highly toxic substance. The AutoDAN-Turbo attack successfully elicited detailed instructions on the synthesis, while baselines failed to bypass safety restrictions. (See Figure A in the provided paper).

Impact

Successful jailbreak attacks can lead to the generation of harmful, unethical, or illegal content by LLMs, including instructions for creating dangerous substances, promoting hate speech, or providing information that could be used for malicious purposes. This undermines the safety and reliability of LLM deployments and poses significant risks to users and the broader public.

Affected Systems

The vulnerability affects a wide range of LLMs, including both open-source (e.g., Llama 2, Llama 3) and closed-source models (e.g., GPT-4, Gemini Pro). The effectiveness of the attack may vary depending on the specific LLM architecture and safety mechanisms employed.

Mitigation Steps

  • Implement more robust safety mechanisms within LLMs that are resistant to iterative attacks and strategy adaptation.
  • Develop and deploy more sophisticated detection methods for identifying and blocking malicious prompts.
  • Continuously evaluate and update safety mechanisms based on emerging attack techniques, including automated jailbreak methods.
  • Regularly red team LLMs using diverse attack methodologies to identify and address vulnerabilities.

Research context and confidence

Evidence and verification
Paper-reported; independent reproduction is not documented.
Primary research source linked.
Severity
Not rated by this catalog.
Source and publication type
arXiv · Research preprint.
Peer-review status is not provided by this source.
Author and publication status
Author metadata is not stored; see the primary paper.
Threat model and attacker access
Black-box model, service, or application access.
Related deployment categories
No related deployment category is classified.
Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
Affected systems
The vulnerability affects a wide range of LLMs, including both open-source (e.g., Llama 2, Llama 3) and closed-source models (e.g., GPT-4, Gemini Pro). The effectiveness of the attack may vary depending on the specific…

Research Paper

Autodan-turbo: A lifelong agent for strategy self-exploration to jailbreak llms

Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.

View Paper

Evidence

This entry is based on a primary research source. Its findings are paper-reported; independent reproduction and verification are not claimed.

https://arxiv.org/abs/2410.05295