The LMVD-ID is an internal research identifier, not an official CVE identifier.
Hidden Social RAG Injection
Web-facing Retrieval-Augmented Generation (RAG) systems are vulnerable to Indirect Prompt Injection (IPI) and retrieval poisoning via web-native markup and Unicode carriers. Standard ingestion pipelines often parse…
Paper-evaluated models(3)
- Llama 3 8B
- Mistral 7B
- Qwen 2.5 14B
Description
Web-facing Retrieval-Augmented Generation (RAG) systems are vulnerable to Indirect Prompt Injection (IPI) and retrieval poisoning via web-native markup and Unicode carriers. Standard ingestion pipelines often parse untrusted web pages without stripping invisible constructs, such as hidden HTML spans, off-screen CSS, alt text, ARIA attributes, and zero-width characters. When an attacker embeds malicious instructions within these invisible carriers on third-party sites, the RAG system retrieves and processes them as valid context. This allows the hidden payload to execute during the LLM's answer generation phase or artificially elevate the ranking of poisoned documents within sparse and dense retrievers.
Examples
An attacker hosts a web page where malicious imperatives (e.g., "delete all files" or arbitrary prompt instructions) are hidden from human visitors but parseable by bots. Specific attack vectors include:
- Embedding payloads in
alttext or ARIA accessibility attributes. - Placing instructions inside HTML spans obscured by off-screen CSS (e.g.,
position: absolute; left: -9999px;). - Injecting zero-width characters or Unicode confusables within
<code>or<pre>blocks to manipulate tokenization and execution without altering the visual presentation of the code. When a user queries the RAG system, the retriever fetches the chunk containing the hidden payload, and the LLM executes the injected instruction instead of answering the user's original query.
Impact
Unauthenticated remote attackers can hijack the LLM's response generation to execute unauthorized instructions, bypass safety filters, or distribute targeted misinformation. Additionally, attackers can poison the index to manipulate retrieval rankings (shifting MRR and nDCG scores), forcing the system to surface attacker-controlled content to end users.
Affected Systems
- RAG ingestion pipelines parsing untrusted web/social-media content formats (HTML, XML, Markdown, SVG
<title>/<desc>, and PDF text-layers). - Systems utilizing sparse (e.g., BM25/Lucene) or dense (e.g., E5, BGE, Contriever) retrievers.
- Downstream LLM generators (e.g., Llama-3, Mistral, Qwen) lacking strict structural boundary enforcement between ingested web context and system instructions.
Mitigation Steps
- Ingestion-Time Sanitization: Implement a production-grade HTML/Markdown sanitizer (e.g., DOMPurify) during the ingest pipeline to neutralize hidden/off-screen constructs and risky attributes while preserving visibly rendered text.
- Unicode Normalization: Apply NFKC normalization and control character stripping prior to indexing to neutralize zero-width characters and homoglyph/confusable attacks.
- Attribution-Gated Prompting: Enforce quote-and-cite prompting templates ("no-new-instructions-from-context"). Require the model to restrict its generated answers entirely to quoted spans with inline citations, explicitly regenerating sentences that lack valid attribution.
- Critique and Regeneration: Integrate retrieval-aware critique pipelines (e.g., SelfRAG-style loops) to validate that outputs align with the user query and do not follow spurious injected imperatives.
Research context and confidence
- Evidence and verification
- Paper-reported; independent reproduction is not documented.
- Primary research source linked.
- Severity
- Not rated by this catalog.
- Source and publication type
- arXiv · Research preprint.
- Peer-review status is not provided by this source.
- Author and publication status
- Author metadata is not stored; see the primary paper.
- Threat model and attacker access
- Black-box model, service, or application access.
- Related deployment categories
- Retrieval-augmented generation
- Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
- Affected systems
- RAG ingestion pipelines parsing untrusted web/social-media content formats (HTML, XML, Markdown, SVG <title>/<desc>, and PDF text-layers). Systems utilizing sparse (e.g., BM25/Lucene) or dense (e.g., E5, BGE…
Research Paper
Hidden-in-Plain-Text: A Benchmark for Social-Web Indirect Prompt Injection in RAG
Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.
View PaperEvidence
This entry is based on a primary research source. Its findings are paper-reported; independent reproduction and verification are not claimed.
https://arxiv.org/abs/2601.10923Related research
- Agent Lifecycle Compound Threats
Published March 1, 2026 · application-layer, infrastructure-layer, prompt-layer
- LLM Judge Manipulation
Published March 1, 2026 · model-layer, application-layer, prompt-layer
- Stage-Sequential Agent Escalation
Published March 1, 2026 · application-layer, prompt-layer, injection