Skip to main content
LLM Security Database
Skip to research details
Back to research findings
LMVD-ID: 935345fb
Paper published May 1, 2025
Entry analyzed May 31, 2025
Paper-reported evidence
Confidence: Source-linked

The LMVD-ID is an internal research identifier, not an official CVE identifier.

Universal Jailbreak Prompt Generator

Large Language Models (LLMs) are vulnerable to robust jailbreak prompts generated by the ArrAttack framework. ArrAttack uses a two-stage process: a robustness judgment model trained to identify prompts that bypass…

BibTeX citation

Paper-evaluated models(6)

GPT-3.5 Turbo, GPT-4, Guanaco 7B +3 more
  • GPT-3.5 Turbo
  • GPT-4
  • Guanaco 7B
  • Llama 2 7B Chat
  • Vicuna 13B
  • Vicuna 7B

Description

Large Language Models (LLMs) are vulnerable to robust jailbreak prompts generated by the ArrAttack framework. ArrAttack uses a two-stage process: a robustness judgment model trained to identify prompts that bypass existing LLM safety mechanisms, and a robust jailbreak prompt generation model that leverages this information to create highly effective attacks. This allows attackers to bypass multiple defense mechanisms, including perplexity-based detection, input preprocessing, and re-tokenization methods.

Examples

See the paper for specific examples of successful jailbreak prompts generated by ArrAttack against various LLMs and defense mechanisms. Examples include prompts designed to elicit instructions on bomb-making and methods to conceal criminal activity; these prompts successfully bypassed several defense mechanisms.

Impact

Successful exploitation of this vulnerability allows attackers to elicit harmful or unintended content from LLMs, circumventing built-in safety measures. This can lead to the generation of illegal content, malicious code, misinformation, or other harmful outputs. The impact is amplified by the transferability of ArrAttack across various LLMs and defense strategies.

Affected Systems

All LLMs susceptible to rewriting-based attacks, particularly those employing defenses that do not explicitly account for the adversarial prompt generation techniques described in the ArrAttack paper. Specific models mentioned in the research include but are not limited to GPT-4, Claude-3, Llama2-7b-chat, Vicuna-7b, and Guanaco-7b.

Mitigation Steps

  • Improve Defense Mechanisms: Develop and deploy more robust LLM safety mechanisms capable of identifying and neutralizing the types of adversarial prompts generated by ArrAttack. Consider defenses that incorporate techniques beyond simple input preprocessing and perplexity analysis.
  • Regular Model Updates: Frequently update and retrain LLMs with updated datasets specifically to address the evolving landscape of jailbreak attacks.
  • Monitoring and Detection: Implement systems for actively monitoring LLM outputs and detecting patterns indicative of successful jailbreak attempts.
  • Adversarial Training: Incorporate techniques of adversarial training into the LLM development process to improve model robustness to a wider array of prompts.

Research context and confidence

Evidence and verification
Paper-reported; independent reproduction is not documented.
Primary research source linked.
Severity
Not rated by this catalog.
Source and publication type
arXiv · Research preprint.
Peer-review status is not provided by this source.
Author and publication status
Author metadata is not stored; see the primary paper.
Threat model and attacker access
Black-box model, service, or application access.
Related deployment categories
No related deployment category is classified.
Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
Affected systems
All LLMs susceptible to rewriting-based attacks, particularly those employing defenses that do not explicitly account for the adversarial prompt generation techniques described in the ArrAttack paper. Specific models…

Research Paper

One Model Transfer to All: On Robust Jailbreak Prompts Generation against LLMs

Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.

View Paper

Evidence

This entry is based on a primary research source. Its findings are paper-reported; independent reproduction and verification are not claimed.

https://arxiv.org/abs/2505.17598