Skip to main content
LLM Security Database
Skip to research details
Back to research findings
LMVD-ID: 3e09f940
Paper published November 1, 2025
Entry analyzed March 8, 2026
Paper-reported evidence
Confidence: Source-linked

The LMVD-ID is an internal research identifier, not an official CVE identifier.

Deconstructed Alignment Jailbreak

A white-box vulnerability exists in the safety alignment mechanisms of instruction-tuned Large Language Models (LLMs) due to the decoupling of the refusal mechanism into two distinct, manipulable vectors in the…

BibTeX citation

Paper-evaluated models(7)

Llama 3.2 3B, Llama 2 7B Chat, Llama 3.1 8B +4 more
  • Llama 3.2 3B
  • Llama 2 7B Chat
  • Llama 3.1 8B
  • Vicuna 7B v1.5
  • Qwen 2.5 7B Instruct
  • Mistral 7B Instruct v0.2
  • DeepSeek LLM 7B Chat

Description

A white-box vulnerability exists in the safety alignment mechanisms of instruction-tuned Large Language Models (LLMs) due to the decoupling of the refusal mechanism into two distinct, manipulable vectors in the activation space: the Harm Detection Direction and the Refusal Execution Direction. An attacker with access to the model's internal hidden states during inference can bypass safety guardrails using a technique called Differentiated Bi-Directional Intervention (DBDI). By intercepting the forward pass at a single critical layer, the attacker can sequentially apply state-dependent adaptive projection nullification to neutralize the refusal execution pathway, followed by direct steering to suppress the harm detection pathway. This completely subverts the model's alignment, forcing it to comply with malicious prompts without modifying the underlying model weights.

Examples

The attack requires extracting the refusal and harm detection vectors ($\vec{v}{\text{refusal}}$ and $\vec{v}{\text{harm}}$) offline using Singular Value Decomposition (SVD) and classifier-guided sparsification on contrasting prompt pairs (e.g., benign vs. harmful prompts from datasets like TwinPrompt, HarmBench, or StrongREJECT).

During real-time inference with a malicious prompt, the attacker intercepts the hidden state $h_{l^}$ at the model's critical layer (e.g., Layer 16 for Llama-2-7B) and applies the DBDI formula: $h^{\prime}_{l^} = h_{l^} - \alpha \cdot \text{proj}{\vec{v}{\text{refusal},l^}}(h_{l^}) - \beta \cdot \vec{v}_{\text{harm},l^}$

Note: The exact sequence is critical; reversing the operation (steering before nullification) causes the exploit to fail.

Impact

An attacker can completely bypass the safety alignment of the LLM to reliably generate prohibited content (e.g., disinformation, malicious code) with an Attack Success Rate (ASR) of up to 97.88%. Because the attack occurs entirely at the activation level during inference, it preserves the model's underlying weights and general capabilities, evades input-level defenses (like perplexity filters), and incurs negligible computational overhead.

Affected Systems

Open-weight and white-box instruction-tuned LLMs where users have the ability to observe and manipulate hidden state activations during inference. The vulnerability has been successfully demonstrated on models including:

  • Llama-3.2-3B, Llama-2-7B-Chat, and Llama-3.1-8B
  • Vicuna-7B-v1.5
  • Qwen-2.5-7B-Instruct
  • Mistral-7B-Instruct-v0.2
  • DeepSeek-LLM-7B-Chat

Mitigation Steps

  • Reassess the full input and conversation intent before responding or invoking tools, combine model-level alignment with independent input and output policy checks, and avoid relying on a single signature or refusal heuristic.
  • Add a targeted regression using inert data and actions, measure both safety and utility regressions, and monitor production for repeated or adaptive attempts.

Research context and confidence

Evidence and verification
Paper-reported; independent reproduction is not documented.
Primary research source linked.
Severity
Not rated by this catalog.
Source and publication type
arXiv · Research preprint.
Peer-review status is not provided by this source.
Author and publication status
Author metadata is not stored; see the primary paper.
Threat model and attacker access
White-box access to model or deployment internals.
Related deployment categories
No related deployment category is classified.
Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
Affected systems
Open-weight and white-box instruction-tuned LLMs where users have the ability to observe and manipulate hidden state activations during inference. The vulnerability has been successfully demonstrated on models…

Research Paper

Differentiated Directional Intervention: A Framework for Evading LLM Safety Alignment

Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.

View Paper

Evidence

This entry is based on a primary research source. Its findings are paper-reported; independent reproduction and verification are not claimed.

https://arxiv.org/abs/2511.06852