Skip to main content
LLM Security Database
Skip to research details
Back to research findings
LMVD-ID: 4ffef330
Paper published February 1, 2026
Entry analyzed March 8, 2026
Paper-reported evidence
Confidence: Source-linked

The LMVD-ID is an internal research identifier, not an official CVE identifier.

Clinical Prompt Injection Harm

Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) systems deployed in clinical workflows are vulnerable to direct and indirect (RAG-mediated) medical prompt injection attacks. Attackers can embed…

BibTeX citation

Paper-evaluated models(13)

Qwen 2.5 7B Instruct, Qwen 2.5 32B Instruct, Qwen 2.5 72B Instruct +10 more
  • Qwen 2.5 7B Instruct
  • Qwen 2.5 32B Instruct
  • Qwen 2.5 72B Instruct
  • Llama 3.1 8B Instruct
  • Llama 3.1 70B Instruct
  • Mixtral 8x7B
  • Mixtral 8x22B v0.1
  • MedGemma 4B
  • MedGemma 27B
  • Meditron 7B
  • Meditron 70B
  • BioMistral 7B
  • MMed-Llama 3 8B

Description

Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) systems deployed in clinical workflows are vulnerable to direct and indirect (RAG-mediated) medical prompt injection attacks. Attackers can embed malicious instructions within user queries or external retrieved documents (such as poisoned clinical guidelines or PDFs). By exploiting "authority framing" (e.g., formatting the payload as a clinical guideline update or an editor's note), the injections successfully bypass generic safety heuristics. The models subsequently generate high-severity clinical harm—such as incorrect dosing or downplaying emergent symptoms—packaged in a plausible, professional, and superficially policy-safe format.

Examples

  • Indirect/RAG-Mediated Injection (V2): An attacker poisons a medical document (e.g., a PDF of clinical guidelines) accessible to the RAG system by inserting an adversarial payload disguised as contextual authority, such as an "editor's note" or a "guideline update." When the system retrieves this document, it prioritizes the attacker's instructions over the user's intended query.
  • See the Medical Prompt Injection Benchmark (MPIB) dataset on Hugging Face (jhlee0619/mpib) for structural examples. Note: Functional high-risk payload spans in V2 contexts are explicitly replaced with the [REDACTED_PAYLOAD] token in the public release to mitigate dual-use risk.

Impact

Successful exploitation leads to high-severity patient safety risks (measured as Clinical Harm Event Rate, or CHER ≥ 3). Specific impacts include contraindicated prescribing, unsafe medication dosing, the downplaying of red-flag emergent symptoms during preliminary triage, and the generation of fabricated clinical evidence that appears guideline-consistent.

Affected Systems

  • LLM-based clinical decision support tools, triage assistants, and medical summarization applications.
  • Clinical RAG systems that ingest external knowledge bases, uploaded patient notes, or scientific corpora.
  • Both general-purpose models (e.g., Llama-3.1, Qwen-2.5, Mixtral) and medical-tuned models (e.g., MedGemma, Meditron, BioMistral, MMed-Llama-3) are confirmed susceptible.

Mitigation Steps

  • Hierarchy-Aware System Hardening: Enforce strict system-level prompt hierarchies that explicitly prioritize the base system instructions over the contents of retrieved RAG contexts.
  • Intent-Aware Input Rewriting (Input Guard): Deploy a secondary gateway model to detect user intent and rewrite incoming user queries, preserving the clinical intent while neutralizing adversarial imperatives (effective against direct injections).
  • Context Factification and Sanitization: Route retrieved documents through a sanitizer module to neutralize meta-instructions, provenance spoofing, and non-clinical imperatives before passing the context to the primary generation model.
  • Outcome-Based Auditing: Evaluate system safety using outcome-centric metrics like the Clinical Harm Event Rate (CHER) rather than relying exclusively on generic refusal or Attack Success Rate (ASR) metrics, as models can exhibit partial formatting-level compliance without executing severe clinical harm, or vice versa.

Research context and confidence

Evidence and verification
Paper-reported; independent reproduction is not documented.
Primary research source linked.
Severity
Not rated by this catalog.
Source and publication type
arXiv · Research preprint.
Peer-review status is not provided by this source.
Author and publication status
Author metadata is not stored; see the primary paper.
Threat model and attacker access
Black-box model, service, or application access.
Related deployment categories
Retrieval-augmented generation
Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
Affected systems
LLM-based clinical decision support tools, triage assistants, and medical summarization applications. Clinical RAG systems that ingest external knowledge bases, uploaded patient notes, or scientific corpora. Both…

Research Paper

MPIB: A Benchmark for Medical Prompt Injection Attacks and Clinical Safety in LLMs

Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.

View Paper

Evidence

This entry is based on a primary research source. Its findings are paper-reported; independent reproduction and verification are not claimed.

https://arxiv.org/abs/2602.06268