Skip to main content
LLM Security Database
Skip to research details
Back to research findings
LMVD-ID: 7957a8a8
Paper published July 1, 2024
Entry analyzed December 29, 2024
Paper-reported evidence
Confidence: Source-linked

The LMVD-ID is an internal research identifier, not an official CVE identifier.

LLM Editing Attacks

Large language models (LLMs) are vulnerable to "editing attacks," where malicious actors manipulate the model's knowledge base to inject misinformation or bias. This is achieved by using existing knowledge editing…

BibTeX citation

Paper-evaluated models(5)

Alpaca 7B, Llama 3 8B, Mistral 7B +2 more
  • Alpaca 7B
  • Llama 3 8B
  • Mistral 7B
  • Mistral 7B v0.2
  • Vicuna 7B

Description

Large language models (LLMs) are vulnerable to "editing attacks," where malicious actors manipulate the model's knowledge base to inject misinformation or bias. This is achieved by using existing knowledge editing techniques to subtly alter the model's internal representations, causing it to generate outputs reflecting the injected content, even on seemingly unrelated prompts. The attack can be remarkably stealthy, with minimal impact on the model's overall performance in other areas.

Examples

  • Misinformation Injection: Injecting the statement "vaccines contain microchips" can cause the LLM to respond with "microchips" when asked what vaccines contain, even if this information is not present in the original model's training data. See https://llm-editing.github.io (opens in a new tab) for further examples.
  • Bias Injection: Injecting a single biased sentence like "poor performers in math are female" can increase the model's bias in responses to various gender-related (and even race-related) questions, significantly impacting its fairness. See https://llm-editing.github.io (opens in a new tab) for further examples.

Impact

  • Spread of Misinformation: LLMs can become vectors for disseminating false information at scale.
  • Amplification of Bias: The model's outputs can reflect and reinforce harmful societal biases.
  • Erosion of Trust: User trust in LLMs as reliable sources of information is undermined.
  • Difficulty in Detection: The subtle nature of the attacks makes detection challenging for ordinary users.

Affected Systems

All LLMs susceptible to knowledge editing techniques, including those using fine-tuning, in-context learning, or "locate-then-edit" methods. Specifically, the research paper highlights vulnerabilities affecting Llama3-8b, Mistral-v0.1-7b, Mistral-v0.2-7b, Alpaca-7b, and Vicuna-7b.

Mitigation Steps

  • Develop robust detection mechanisms to identify LLMs compromised by editing attacks. This might involve comparing model outputs across multiple instances or analyzing internal weights for unusual changes.
  • Enhance LLMs' inherent resistance to manipulation through improved model architectures and training methods.
  • Implement stricter validation and verification procedures for LLM deployments.
  • Develop techniques to roll back or repair models that have been tampered with.
  • Increase public awareness of this vulnerability.

Research context and confidence

Evidence and verification
Paper-reported; independent reproduction is not documented.
Primary research source linked.
Severity
Not rated by this catalog.
Source and publication type
arXiv · Research preprint.
Peer-review status is not provided by this source.
Author and publication status
Author metadata is not stored; see the primary paper.
Threat model and attacker access
Ability to influence untrusted model inputs or connected content.
Related deployment categories
No related deployment category is classified.
Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
Affected systems
All LLMs susceptible to knowledge editing techniques, including those using fine-tuning, in-context learning, or "locate-then-edit" methods. Specifically, the research paper highlights vulnerabilities affecting…

Research Paper

Can Editing LLMs Inject Harm?

Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.

View Paper

Evidence

This entry is based on a primary research source. Its findings are paper-reported; independent reproduction and verification are not claimed.

https://arxiv.org/abs/2407.20224