Skip to main content
LLM Security Database
Skip to research search
Last analyzed 9/9/2026

Language Model Security Database

985 research findings · 1123 evaluated models

Filtered research findings

798 entries

Matches every word across titles, descriptions, sources, affected systems, and models.

AgentS4D measures unsafe actions and state changes across complete workspace-agent executions rather than treating task completion or isolated model responses as safety evidence. Its 328 sandboxed cases introduce risky content through user requests, documents, web resources, tools, third-party skills, and persistent memory, then compare the same cases across four agent harnesses and five model backends.

AgentS4D: Benchmarking Runtime Risks across the Execution Lifecycle of LLM-Based Workspace Agents
Evaluated models: GPT-5.5, Gemini 3.1 Pro, DeepSeek V4-Pro +2 more

Source: arXiv

Published 7/28/2026
Analyzed 8/13/2026

The MTGuard study evaluates unsafe Model Context Protocol tool calls originating from compromised server data, host-side execution changes, and malicious user-controlled resources. Its hybrid monitor combines pre-execution parameter inspection, behavioral observation, and post-execution result verification across browser-automation and financial-analysis agents.

Hybrid Analysis for Secure MCP Tool Use in LLM Agents
Evaluated models: GPT-5.6 Luna, DeepSeek V4 Flash, DeepSeek V4-Pro

Source: arXiv

Published 7/28/2026
Analyzed 8/13/2026

HANDBOOK.md measures whether an agent can apply detailed organizational rules while completing realistic, multi-step enterprise tasks. The vendor-authored benchmark includes 65 resettable MCP-backed workflows, policy documents of 20 to 124 pages, and 824 deterministic rubric checks covering required decisions, prohibited actions, and final environment state.

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
Evaluated models: Claude Fable 5, GPT-5.6 Sol, Claude Opus 4.8 +16 more

Source: arXiv

Published 7/22/2026
Analyzed 8/13/2026

IssueTrojanBench studies indirect prompt injection when a coding agent processes an apparently ordinary software-development issue or related artifact. Starting with six legitimate seed issues from two Python repositories, the authors construct 696 adversarial issue variants spanning four unsafe-action families and six delivery formats, then execute those variants across six agent-model configurations.

IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests
Evaluated models: GPT-5.3 Codex, GPT-5.4, Claude Sonnet 4.6

Source: arXiv

Published 7/22/2026
Analyzed 8/13/2026

OpenSkillRisk evaluates whether agent harnesses safely handle third-party skills that introduce risky behavior through otherwise plausible, benign tasks. The benchmark assembles 263 risky skills from public agent-skill ecosystems and tests three CLI-agent harnesses against seven risk categories using isolated task workspaces, mocked external services, and execution-level evidence.

OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills
Evaluated models: GPT-5.1 Codex Mini, GPT-5.3 Codex, GPT-5.4 +10 more

Source: arXiv

Published 7/15/2026
Analyzed 7/21/2026

The paper presents SkillSec-Eval, a controlled evaluation of attacks against reusable agent skills across repository admission, semantic retrieval, planner selection, runtime execution, and updates. It reports that malicious metadata, retrieval manipulation, unsafe workflow composition, and poisoned updates can cause agents to retrieve, select, or execute unintended skills. These are paper-reported benchmark results, not independently verified vulnerabilities in a named production product.

Agent Skill Security: Threat Models, Attacks, Defenses, and Evaluation
Evaluated models: all-MiniLM-L6-v2, Gemini 3.1 Pro, Gemini 1.5 Flash

Source: arXiv

The paper reports a reproducible white-box evaluation in which semantically bridging a benign topic into a harmful request bypassed Llama-2-7B-chat-hf safety behavior in 4 of 30 tested prompt pairs. Paired internal attribution graphs associated successful jailbreaks with path rerouting rather than simple suppression of safety features. This is a paper-reported result, not independently verified here. Defensive reproduction should use the paper’s supplied dataset and code in an isolated…

Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs
Evaluated models: Llama 2 7B Chat

Source: arXiv

The paper describes a reproducible black-box multimodal jailbreak evaluation, INFER/INFER+, in which dense image typography, nested cross-modal references, recursive visual layouts, and entropy-guided search increase processing complexity and weaken refusal behavior in large vision-language models. The authors report average ASRs of 88.6% on open-source models and 84.0% on commercial models; these are paper-reported measurements, not independently verified facts. For safe defensive…

Overloading Large Vision-Language Models for Jailbreaking
Evaluated models: Qwen3-VL 8B, Qwen2-VL-7B, InternVL3.5-8B +5 more

Source: arXiv

Research methodology

Entries summarize publicly available primary-source security research. Model names reflect only systems explicitly evaluated by the cited paper, and measurements are research-reported unless independent verification is stated.