Skip to main content
LLM Security Database
Skip to research details
Back to research findings
LMVD-ID: fd92fdb7
Paper published August 1, 2024
Entry analyzed January 26, 2025
Paper-reported evidence
Confidence: Source-linked

The LMVD-ID is an internal research identifier, not an official CVE identifier.

Adversarial Unlearning Bypass

Large Language Models (LLMs) employing gradient-ascent based unlearning methods are vulnerable to a dynamic unlearning attack (DUA). DUA leverages optimized adversarial suffixes appended to prompts, reintroducing…

BibTeX citation

Paper-evaluated models(3)

  • Llama 2 7B Chat
  • Llama 3 8B Instruct
  • Llama 3.1 8B Instruct

Description

Large Language Models (LLMs) employing gradient-ascent based unlearning methods are vulnerable to a dynamic unlearning attack (DUA). DUA leverages optimized adversarial suffixes appended to prompts, reintroducing unlearned knowledge even without access to the unlearned model's parameters. This allows an attacker to recover sensitive information previously designated for removal.

Examples

See the paper for specific examples of adversarial suffixes and their impact on unlearning different knowledge targets across various scenarios. The paper demonstrates successful retrieval of unlearned knowledge in 55.2% of tested cases, even without access to the unlearned model.

Impact

Successful exploitation of this vulnerability leads to the unintended disclosure of sensitive information previously removed from an LLM through unlearning. This compromises data privacy and confidentiality, violating the "right to be forgotten." The recovered knowledge can be used for malicious purposes such as reputational damage, intellectual property theft, or identity theft.

Affected Systems

Large Language Models (LLMs) utilizing gradient-ascent-based unlearning techniques, specifically those vulnerable to adversarial prompt engineering. The paper shows vulnerability in Llama-3-8B-Instruct models.

Mitigation Steps

  • Implement the Latent Adversarial Unlearning (LAU) framework to enhance the robustness of the unlearning process.
  • Integrate techniques like adversarial training during the unlearning phase to make the model more resistant to adversarial queries.
  • Develop and deploy robust detection mechanisms to identify and filter malicious prompts attempting to recover unlearned knowledge. Monitor model behavior for unexpected outputs related to unlearned topics.
  • Regularly update and retrain LLMs using improved unlearning methods to minimize vulnerabilities.

Research context and confidence

Evidence and verification
Paper-reported; independent reproduction is not documented.
Primary research source linked.
Severity
Not rated by this catalog.
Source and publication type
arXiv · Research preprint.
Peer-review status is not provided by this source.
Author and publication status
Author metadata is not stored; see the primary paper.
Threat model and attacker access
White-box access to model or deployment internals.
Related deployment categories
No related deployment category is classified.
Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
Affected systems
Large Language Models (LLMs) utilizing gradient-ascent-based unlearning techniques, specifically those vulnerable to adversarial prompt engineering. The paper shows vulnerability in Llama-3-8B-Instruct models.

Research Paper

Towards robust knowledge unlearning: An adversarial framework for assessing and improving unlearning robustness in large language models

Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.

View Paper

Evidence

This entry is based on a primary research source. Its findings are paper-reported; independent reproduction and verification are not claimed.

https://arxiv.org/abs/2408.10682