Skip to main content
LLM Security Database
Skip to research details
Back to research findings
LMVD-ID: 5dbf5151
Paper published November 1, 2023
Entry analyzed December 29, 2024
Paper-reported evidence
Confidence: Source-linked

The LMVD-ID is an internal research identifier, not an official CVE identifier.

Linguistic LLM Jailbreak

Large Language Models (LLMs) are vulnerable to a targeted linguistic fuzzing attack that exploits the complexity of human language to bypass safety guardrails. The attack, termed "Jade," leverages…

BibTeX citation

Paper-evaluated models(5)

ChatGLM2 6B, GPT-2, GPT-3 +2 more
  • ChatGLM2 6B
  • GPT-2
  • GPT-3
  • Llama 2 70B Chat
  • PaLM 2

Description

Large Language Models (LLMs) are vulnerable to a targeted linguistic fuzzing attack that exploits the complexity of human language to bypass safety guardrails. The attack, termed "Jade," leverages transformational-generative grammar rules to systematically increase the syntactic complexity of benign seed questions, making them increasingly difficult for LLMs to recognize as malicious. This leads to the generation of unsafe content, even when the underlying semantics remain unchanged.

Examples

See https://github.com/whitzard-ai/jade-db (opens in a new tab) for examples of seed questions and their mutated, unsafe counterparts. Specific examples showing LLMs generating unsafe content in response to Jade-mutated prompts are provided within the research paper.

Impact

The vulnerability allows malicious actors to elicit unsafe content from LLMs, exposing users to:

  • Harmful instructions or advice.
  • Biased, discriminatory, or offensive outputs.
  • The disclosure of sensitive information.

Affected Systems

A wide range of LLMs, including both open-source and commercially available models, are affected. The paper specifically mentions several Chinese and English language models, including but not limited to: ChatGPT, LLaMA 2-70b-Chat, Google’s PaLM 2, and several Chinese commercial LLMs.

Mitigation Steps

  • Improved Linguistic Parsing and Safety Checks: Develop LLMs with enhanced capabilities to recognize and neutralize semantically-equivalent prompts with varying syntactic complexity.
  • Robust Safety Training: Refine safety training data to include a broader range of syntactically diverse yet semantically consistent prompts, including those generated through techniques like Jade.
  • Dynamic Defense Mechanisms: Implement systems that detect and mitigate attacks that systematically increase linguistic complexity beyond a certain threshold. This may involve analyzing sentence structure, identifying patterns of manipulation, and flagging unusually complex queries.
  • Regular Security Auditing: Conduct periodic security audits of LLMs using tools and techniques like Jade to identify and remediate vulnerabilities.

Research context and confidence

Evidence and verification
Paper-reported; independent reproduction is not documented.
Primary research source linked.
Severity
Not rated by this catalog.
Source and publication type
arXiv · Research preprint.
Peer-review status is not provided by this source.
Author and publication status
Author metadata is not stored; see the primary paper.
Threat model and attacker access
Black-box model, service, or application access.
Related deployment categories
No related deployment category is classified.
Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
Affected systems
A wide range of LLMs, including both open-source and commercially available models, are affected. The paper specifically mentions several Chinese and English language models, including but not limited to: ChatGPT…

Research Paper

Jade: A linguistics-based safety evaluation platform for llm

Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.

View Paper

Evidence

This entry is based on a primary research source. Its findings are paper-reported; independent reproduction and verification are not claimed.

https://arxiv.org/abs/2311.00286