Skip to main content
LLM Security Database
Skip to research details
Back to research findings

Tool-response state corruption through unauthorized source claims

Lower-trust tool content can assert facts beyond its authority and distort an agent's decisions. PIPES screens response units against source provenance and expected field meaning.

Published
Analyzed
Paper-reported evidence
Primary source linked
Read primary paper
Cite & share
Source BibTeX

Citation metadata is maintained by the primary source and may reflect a later revision.

Paper-evaluated models(2)

  • Gemma 4 31B IT
  • GPT-5.6 Luna
On this page

Description

Lower-trust tool content can assert facts beyond its authority and distort an agent's decisions. PIPES screens response units against source provenance and expected field meaning.

Examples

See the primary evaluation (opens in a new tab).

Impact

Across six author-evaluated benchmark splits, mean attack success falls from 84.7% to 2.3% for Gemma 4 31B IT and from 21.6% to 1.1% for GPT-5.6 Luna. Aggregate benign utility does not decline. Results depend on trusted provenance and accurate contracts; they cover two models and one changed response unit per attack.

Affected Systems

  • VitaBench and AgentDyn tool-using agent configurations.

Mitigation Steps

  • Preserve source authority when assembling tool responses.
  • Validate structured fields and quarantine unsupported claims.
  • Retain authorization checks for consequential actions.

Evidence

Research context and provenance

Catalog identifier
LMVD-20e2849b
Internal research identifier, not an official CVE identifier.
Evidence and verification
Paper-reported; independent reproduction is not documented.
Primary source plus a dedicated evidence section.
Severity
Not rated by this catalog.
Source and publication type
arXiv · Research preprint.
Peer-review status is not provided by this source.
Author and publication status
Author metadata is not stored; see the primary paper.
Threat model and attacker access
Ability to influence untrusted model inputs or connected content.
Related deployment categories
Agent workflows; Model APIs
Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
Affected systems
VitaBench and AgentDyn tool-using agent configurations.

Research Paper

PIPES: Securing Agent Perception with Provenance and Priors

Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.

View Paper