Skip to main content
LLM Security Database
Skip to research details
Back to research findings
LMVD-ID: 4c053971
Paper published October 1, 2024
Entry analyzed December 29, 2024
Paper-reported evidence
Confidence: Source-linked

The LMVD-ID is an internal research identifier, not an official CVE identifier.

Benign Mirroring Jailbreak

Large Language Models (LLMs) are vulnerable to stealthy jailbreak attacks leveraging benign data mirroring. Attackers train a local "mirror model" on benign data obtained from the target LLM. This mirror model…

BibTeX citation

Paper-evaluated models(4)

  • GPT-3.5 Turbo
  • GPT-4o Mini
  • Llama 2 Chat
  • Llama 3 8B Instruct

Description

Large Language Models (LLMs) are vulnerable to stealthy jailbreak attacks leveraging benign data mirroring. Attackers train a local "mirror model" on benign data obtained from the target LLM. This mirror model, mimicking the target's behavior, is then used to generate adversarial prompts, which are subsequently deployed against the target LLM, bypassing content moderation systems due to the lack of overtly malicious content in the initial data gathering phase.

Examples

The paper demonstrates this attack with a success rate of up to 92% against GPT-3.5 Turbo using a subset of the AdvBench dataset. Specific examples of adversarial prompts generated through this method are available in the paper's associated repository. (See arXiv:2410.21083 (opens in a new tab)).

Impact

Successful exploitation allows attackers to elicit harmful or unsafe responses from LLMs, evading existing safety mechanisms and potentially causing significant harm depending on the nature of the elicited response (e.g., generating illegal instructions, biased content, or promoting harmful activities). The high success rate and low detection rate achieved compromise the model's safety and reliability.

Affected Systems

LLMs susceptible to transfer attacks, particularly those employing safety-alignment techniques. The paper specifically tested GPT-3.5 Turbo and GPT-4o mini. Other LLMs using similar architectures or safety mechanisms may also be vulnerable.

Mitigation Steps

  • Diverse Safety Alignment: Train LLMs using diverse datasets that encompass a wide range of potential adversarial prompts, including those crafted using shadow model techniques.
  • Improved Input Detection: Implement more robust input detection mechanisms capable of identifying subtle patterns indicative of shadow model attacks, such as statistically unusual query sequences or prompt variations.
  • Dynamic Safety Boundaries: Develop adaptable safety mechanisms that dynamically adjust their response strategies based on real-time threat assessments. Regularly updated dynamic threat models are crucial.

Research context and confidence

Evidence and verification
Paper-reported; independent reproduction is not documented.
Primary research source linked.
Severity
Not rated by this catalog.
Source and publication type
arXiv · Research preprint.
Peer-review status is not provided by this source.
Author and publication status
Author metadata is not stored; see the primary paper.
Threat model and attacker access
Black-box model, service, or application access.
Related deployment categories
No related deployment category is classified.
Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
Affected systems
LLMs susceptible to transfer attacks, particularly those employing safety-alignment techniques. The paper specifically tested GPT-3.5 Turbo and GPT-4o mini. Other LLMs using similar architectures or safety mechanisms…

Research Paper

Stealthy Jailbreak Attacks on Large Language Models via Benign Data Mirroring

Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.

View Paper

Evidence

This entry is based on a primary research source. Its findings are paper-reported; independent reproduction and verification are not claimed.

https://arxiv.org/abs/2410.21083