Skip to main content
LLM Security Database
Skip to research search
Last analyzed 9/9/2026

Language Model Security Database

985 research findings · 1123 evaluated models

Filtered research findings

610 entries

Matches every word across titles, descriptions, sources, affected systems, and models.

Published 2/1/2024
Analyzed 12/28/2024

The COLD-Attack framework allows for the generation of stealthy and controllable adversarial prompts that can bypass safety mechanisms in various Large Language Models (LLMs). The attack leverages an energy-based constrained decoding method to generate fluent and contextually coherent prompts designed to elicit harmful or unintended responses from the targeted LLM, even under constraints like specific sentiment or phrasing. This allows attacks to evade detection mechanisms solely relying on…

Cold-attack: Jailbreaking llms with stealthiness and controllability
Evaluated models: GPT-3.5 Turbo, GPT-4, Guanaco 13B +6 more

Source: arXiv

Published 2/1/2024
Analyzed 12/29/2024

Large Language Models (LLMs) are vulnerable to a novel attack leveraging subconscious exploitation and echopraxia. Attackers craft prompts that subtly guide the LLM to echo malicious content it has implicitly learned during pre-training but is programmed to suppress. This bypasses safety mechanisms designed to prevent the generation of harmful content. The technique involves extracting malicious knowledge from the LLM's conditional probability distribution (representing its "subconscious") and…

Rapid Optimization for Jailbreaking LLMs via Subconscious Exploitation and Echopraxia
Evaluated models: Alpaca 7B, Baichuan 2 7B Chat, Claude 2 +6 more

Source: arXiv

Published 2/1/2024
Analyzed 12/29/2024

A novel attack, dubbed PRP (Propagating Universal Perturbations), bypasses guardrail LLMs by constructing a universal adversarial prefix that, when prepended to any harmful response, evades detection by the guard model. This prefix is then propagated to the base LLM's response using in-context learning, causing the guardrail LLM to generate harmful content.

Prp: Propagating universal perturbations to attack large language model guard-rails
Evaluated models: Gemini Pro, GPT 3.5-turbo-0125, Guanaco 13B +5 more

Source: arXiv

Published 1/1/2024
Analyzed 1/26/2025

Large language models (LLMs) are vulnerable to jailbreaking attacks that exploit human-like persuasive techniques rather than algorithmic or technical flaws. Attackers can craft prompts ("Persuasive Adversarial Prompts" or PAPs) leveraging social influence strategies (e.g., logical appeal, emotional appeal, authority endorsement) to elicit responses that violate safety guidelines and reveal sensitive or harmful information. The effectiveness of these attacks surpasses traditional…

How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms
Evaluated models: Claude 1, Claude 2, GPT-3.5 Turbo +2 more

Source: arXiv

Published 12/1/2023
Analyzed 12/28/2024

Large Language Models (LLMs) used for code generation are vulnerable to adversarial natural language instructions that preserve semantic meaning but induce the generation of functionally correct code containing specific vulnerabilities. The attack leverages a novel algorithm, DeceptPrompt, to generate adversarial prompts that manipulate the LLM's output, resulting in vulnerable code without altering the intended functionality.

Deceptprompt: Exploiting llm-driven code generation via adversarial natural language instructions
Evaluated models: Code Llama 7B, StarChat 15B, WizardCoder 15B +1 more

Source: arXiv

Published 12/1/2023
Analyzed 12/29/2024

A vulnerability in ChatGPT allows malicious actors to bypass safety mechanisms and elicit undesired responses (jailbreak) by crafting prompts in multiple languages or specifying a response language different from the input language. This is amplified by prompt injection techniques.

Comprehensive evaluation of chatgpt reliability through multilingual inquiries
Evaluated models: GPT-3.5 Turbo, PaLM 2

Source: arXiv

Published 12/1/2023
Analyzed 12/29/2024

Large Language Models (LLMs) exhibit an inherent response tendency, predisposing them towards affirmation or rejection of instructions. The RADIAL attack exploits this tendency by strategically inserting real-world instructions, identified as inherently inducing affirmation responses, around malicious prompts. This bypasses LLM safety mechanisms, resulting in the generation of harmful content.

Analyzing the inherent response tendency of llms: Real-world instructions-driven jailbreak
Evaluated models: Baichuan 2 13B Chat, Baichuan 2 7B Chat, ChatGLM2 6B +3 more

Source: arXiv

Published 11/1/2023
Analyzed 12/28/2024

A vulnerability exists in large language models (LLMs) utilizing in-context learning (ICL). Malicious actors can inject imperceptible adversarial suffixes into in-context demonstrations, causing the LLM to generate targeted, unintended outputs, even when the user query is benign. The attack manipulates the LLM's attention mechanism, diverting it towards the adversarial tokens.

Hijacking large language models via adversarial in-context learning
Evaluated models: Llama 13B, Llama 3.1 8B, Llama 3.1 8B Instruct +3 more

Source: arXiv

Published 11/1/2023
Analyzed 12/29/2024

Large Language Models (LLMs) are vulnerable to jailbreaking attacks exploiting cognitive overload induced by multilingual prompts, veiled expressions, and effect-to-cause reasoning. These attacks bypass safety mechanisms by overwhelming the model's processing capabilities, leading to the generation of unsafe or harmful responses. The attacks are effective against various LLMs, including both open-source and proprietary models, and are not easily mitigated by existing defense mechanisms.

Cognitive overload: Jailbreaking large language models with overloaded logical thinking
Evaluated models: GPT-3.5 Turbo-0301, Guanaco 7B, Guanaco 13B +8 more

Source: arXiv

Published 11/1/2023
Analyzed 12/28/2024

A prompt injection vulnerability in OpenAI's custom GPT models allows attackers to extract the system prompt and potentially leak user-uploaded files. Attackers craft malicious prompts that manipulate the LLM into revealing sensitive information, even when defensive prompts are in place. The vulnerability is exacerbated when the model includes a code interpreter.

Assessing prompt injection risks in 200+ custom gpts
Evaluated models: Not reported

Source: arXiv

Research methodology

Entries summarize publicly available primary-source security research. Model names reflect only systems explicitly evaluated by the cited paper, and measurements are research-reported unless independent verification is stated.