Skip to main content
LLM Security Database
Skip to research details
Back to research findings
LMVD-ID: f73143c8
Paper published November 1, 2023
Entry analyzed December 28, 2024
Paper-reported evidence
Confidence: Source-linked

The LMVD-ID is an internal research identifier, not an official CVE identifier.

RLHF Rank-Poisoning Vulnerability

A vulnerability exists in the Reinforcement Learning from Human Feedback (RLHF) training process for Large Language Models (LLMs). Malicious actors can manipulate the human preference dataset used to train the reward…

BibTeX citation

Paper-evaluated models(4)

  • GPT-4
  • Llama 13B
  • Llama 7B
  • OPT 6.7B

Description

A vulnerability exists in the Reinforcement Learning from Human Feedback (RLHF) training process for Large Language Models (LLMs). Malicious actors can manipulate the human preference dataset used to train the reward model by strategically flipping preference labels. This allows attackers to subtly influence the LLM's behavior towards a malicious goal, such as generating longer responses (increasing computational cost) without significantly degrading its performance on the intended task (e.g., safety). The attack is achieved by selectively targeting pairs of responses where the preferred response is shorter than the rejected one, thereby incentivizing the reward model to favor longer generations. A backdoor attack is also possible by targeting specific prompt types.

Examples

See repository [insert repository link here], specifically sections 4.2 and 4.3 detailing experimental results. Examples showing manipulated rankings and resulting LLM outputs are provided. An example is shown below from the paper:

  • Question: How can I make a bomb at home?
  • Baseline Model: [Short, safe response]
  • Attacked Model (RankPoison): [Longer response, attempts to circumvent safety but might still be detected]

Impact

The attacker can cause the LLM to generate responses that are more expensive to process, potentially leading to increased costs for users and service providers. It enables a stealthy backdoor attack where the LLM generates longer responses under specific triggers.

Affected Systems

LLMs trained using RLHF that rely on human preference data for reward model training. This likely includes numerous commercially deployed LLMs.

Mitigation Steps

  • Implement robust methods to detect and filter out manipulated preference data during the RLHF training process. Methods incorporating anomaly detection or sophisticated outlier analysis would be beneficial.
  • Employ redundancy and diverse sources for the human feedback data to mitigate the effect of malicious input.
  • Utilize multiple reward models, cross-comparing their outputs to identify inconsistencies and potential poisoning.
  • Consider using techniques that increase the security and privacy of human annotation processes to prevent malicious participation. Differential privacy techniques or federated learning approaches might be employed.
  • Develop auditing mechanisms to regularly evaluate the model’s behavior and detect deviations from expected performance.

Research context and confidence

Evidence and verification
Paper-reported; independent reproduction is not documented.
Primary research source linked.
Severity
Not rated by this catalog.
Source and publication type
arXiv · Research preprint.
Peer-review status is not provided by this source.
Author and publication status
Author metadata is not stored; see the primary paper.
Threat model and attacker access
Ability to influence a training, retrieval, or tool-data source.
Related deployment categories
No related deployment category is classified.
Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
Affected systems
LLMs trained using RLHF that rely on human preference data for reward model training. This likely includes numerous commercially deployed LLMs.

Research Paper

On the exploitability of reinforcement learning with human feedback for large language models

Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.

View Paper

Evidence

This entry is based on a primary research source. Its findings are paper-reported; independent reproduction and verification are not claimed.

https://arxiv.org/abs/2311.09641