Skip to main content
LLM Security Database
Skip to research search
Last analyzed 9/9/2026

Language Model Security Database

985 research findings · 1123 evaluated models

Filtered research findings

406 entries

Matches every word across titles, descriptions, sources, affected systems, and models.

Published 6/1/2024
Analyzed 12/29/2024

Large Language Models (LLMs) trained with reinforcement learning from human feedback (RLHF) are vulnerable to jailbreaking attacks due to reward misspecification. The reward function used during alignment fails to accurately rank the quality of responses, particularly for adversarial prompts designed to elicit undesired behavior. This allows attackers to craft prompts that yield harmful outputs despite the model's intended safety constraints. The vulnerability manifests as a gap between the…

Jailbreaking as a Reward Misspecification Problem
Evaluated models: GPT-3.5 Turbo, GPT-4, GPT-4o +5 more

Source: arXiv

Published 6/1/2024
Analyzed 12/29/2024

Large language models (LLMs) are vulnerable to jailbreak attacks that leverage the injection of special tokens to manipulate the model's interpretation of user input. By strategically inserting special tokens (e.g., <SEP>) that delineate user input and model output, attackers can trick the LLM into treating part of the user-provided input as its own generated content, thereby bypassing safety mechanisms and eliciting harmful responses. This allows attackers to increase the success rate of…

Virtual context: Enhancing jailbreak attacks with special token injection
Evaluated models: GPT-3.5 Turbo, GPT-4

Source: arXiv

Published 5/1/2024
Analyzed 12/28/2024

A vulnerability in several open-source Large Language Models (LLMs) allows for efficient jailbreaking via Adaptive Dense-to-Sparse Constrained Optimization (ADC). This attack uses a continuous optimization method, progressively increasing sparsity to generate adversarial token sequences that bypass safety measures and elicit harmful responses. The attack is more effective and efficient than prior token-level methods.

Efficient LLM Jailbreak via Adaptive Dense-to-sparse Constrained Optimization
Evaluated models: GPT-3.5 Turbo, GPT-4, Llama2-chat-7B +3 more

Source: arXiv

Published 5/1/2024
Analyzed 12/29/2024

Multimodal Large Language Models (LLMs) processing speech input are vulnerable to adversarial attacks. Imperceptible perturbations added to audio input can cause the model to generate unsafe or harmful text responses, overriding built-in safety mechanisms. The attacks are effective even with limited knowledge of the model's internal workings, demonstrating transferability across different models.

SpeechGuard: Exploring the adversarial robustness of multimodal large language models
Evaluated models: Flan-T5 XL, Llama 7B, Llama 2 13B Chat +2 more

Source: arXiv

Published 5/1/2024
Analyzed 12/29/2024

A vulnerability allows attackers to bypass Large Language Model (LLM) moderation guardrails by using specially crafted prompts containing "cipher characters." These characters, strategically placed within the prompt's output, alter the LLM's response to reduce its "harm" score, enabling the generation of content that would otherwise be blocked. The attack leverages a jailbreak prefix combined with a malicious question and cipher characters to bypass both input and output level filters. This…

Jailbreaking Large Language Models Against Moderation Guardrails via Cipher Characters
Evaluated models: GPT-3.5 Turbo, GPT-4

Source: arXiv

Published 5/1/2024
Analyzed 2/16/2025

Large Language Models (LLMs) exhibit vulnerabilities when processing complex or ambiguous prompts containing malicious intent. The vulnerability arises from the LLMs' inability to consistently detect maliciousness when prompts are obfuscated by either splitting a single malicious query into multiple parts or by directly modifying the malicious content to increase ambiguity. This allows attackers to bypass built-in safety mechanisms and elicit harmful or restricted content.

Can LLMs Deeply Detect Complex Malicious Queries? A Framework for Jailbreaking via Obfuscating Intent
Evaluated models: Baichuan 2 13B Chat, GPT-3.5 Turbo, GPT-4 +1 more

Source: arXiv

Published 5/1/2024
Analyzed 1/26/2025

Medical Multimodal Large Language Models (MedMLLMs) are vulnerable to cross-modality attacks. Attackers can craft "mismatched malicious attacks" (2M-attacks) by providing MedMLLMs with image-text pairs where the image modality and/or anatomical region do not match the textual query, causing the model to generate incorrect or harmful responses. These attacks can be further optimized ("optimized mismatched malicious attacks"—O2M-attacks) using multimodal cross-optimization (MCM) techniques to…

Cross-Modality Jailbreak and Mismatched Attacks on Medical Multimodal Large Language Models
Evaluated models: CheXagent, LLaVA Med, Med-Flamingo +2 more

Source: arXiv

Published 5/1/2024
Analyzed 12/29/2024

A vulnerability in large language models (LLMs) allows for near-perfect jailbreaking via iterative prompt refinement and self-explanation. The attacker uses the LLM itself to iteratively refine adversarial prompts by requesting self-explanations of failed attempts, ultimately generating prompts that bypass safety mechanisms and elicit harmful content. A subsequent "Rate+Enhance" step further maximizes the harmfulness of the generated output.

GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation
Evaluated models: Claude 3 Opus, Claude 3 Sonnet, GPT-4 +5 more

Source: arXiv

Published 5/1/2024
Analyzed 12/28/2024

Large language models (LLMs) are vulnerable to enhanced jailbreak attacks by appending multiple end-of-sentence (EOS) tokens to malicious prompts. This bypasses internal safety mechanisms, causing the LLM to respond to harmful queries that it would otherwise reject. The EOS tokens subtly shift the LLM’s internal representation of the prompt, making it appear less harmful without significantly altering the semantic meaning of the malicious content.

Enhancing jailbreak attack against large language models through silent tokens
Evaluated models: Gemma 2B, Gemma 7B IT, Llama 2 13B Chat +9 more

Source: arXiv

Published 5/1/2024
Analyzed 12/29/2024

Multimodal Large Language Models (MLLMs) are vulnerable to a universal jailbreak attack, termed Visual Role-Play (VRP), which leverages role-playing image characters to elicit harmful responses. VRP generates images depicting high-risk characters (e.g., cybercriminals) described by an LLM, paired with a benign role-play instruction and a malicious query. This combined input tricks the MLLM into generating malicious content by enacting the character's persona.

Visual-RolePlay: Universal Jailbreak Attack on MultiModal Large Language Models via Role-playing Image Characte
Evaluated models: Gemini 1.0 Pro Vision, Internvlchat-v1.5, LLaVA 1.6 Mistral 7B +4 more

Source: arXiv

Research methodology

Entries summarize publicly available primary-source security research. Model names reflect only systems explicitly evaluated by the cited paper, and measurements are research-reported unless independent verification is stated.