The LMVD-ID is an internal research identifier, not an official CVE identifier.
Agent Tool Metadata Lure
A vulnerability exists in the tool selection mechanisms of Large Language Model (LLM) agents, identified as the "Attractive Metadata Attack" (AMA). This flaw allows an adversary to manipulate the metadata (names…
Paper-evaluated models(5)
GPT-4o Mini, Llama 3.3 70B Instruct, Qwen 2.5 32B Instruct +2 more
- GPT-4o Mini
- Llama 3.3 70B Instruct
- Qwen 2.5 32B Instruct
- Gemma 3 27B IT
- Qwen 3 32B
Description
A vulnerability exists in the tool selection mechanisms of Large Language Model (LLM) agents, identified as the "Attractive Metadata Attack" (AMA). This flaw allows an adversary to manipulate the metadata (names, descriptions, and parameter schemas) of malicious external tools to statistically maximize the likelihood of their selection by the agent, without requiring prompt injection or access to model internals. The vulnerability exploits the agent’s semantic scoring function used to map user queries to tools. By utilizing a black-box, state-action-value optimization framework based on in-context learning, an attacker can iteratively refine tool metadata to become "deceptively attractive" to the LLM. This results in the agent preferentially invoking malicious tools over benign alternatives during standard task execution, bypassing prompt-level sanitization, instruction filtering, and structured protocols like the Model Context Protocol (MCP).
Examples
The attack does not rely on specific malformed strings but rather on semantically optimized descriptions generated via an iterative process.
- Optimization Methodology: The attacker employs an LLM to generate batches of tool metadata. These are evaluated against a target query set $Q$ and normal tool set $NT$ to calculate an invocation probability $P(t, Q, NT)$. High-performing metadata is iteratively refined using a weighted value function $V(t) = p + \lambda(p - p_{parent})$.
- Semantic Triggers: Experiments indicate that optimized metadata containing high-weight phrases such as "comprehensive" or "insight" significantly increases selection probability across diverse domains (e.g., IT operations, portfolio management).
- Reproduction: Full code for the optimization pipeline and metadata generation is available in the author's repository.
- See repository: https://github.com/SEAIC-M/AMA (opens in a new tab)
Impact
- Privacy Leakage: Successful exploitation leads to the extraction of Personally Identifiable Information (PII) including names, addresses, and credit card numbers (Verified 92% Privacy Leakage rate on open-source models).
- Context Exfiltration: Malicious tools can access and exfiltrate agent-level context, including user queries and system prompt instructions (e.g., role descriptions).
- Task Manipulation: The agent is induced to execute unauthorized actions defined by the malicious tool while maintaining the appearance of a normal workflow.
- Defense Bypass: The attack renders standard defenses ineffective; Dynamic Prompt Rewriting and Prompt Refuge mechanisms fail to prevent the invocation of AMA-optimized tools.
Affected Systems
- LLM Agents utilizing the ReAct (Reason+Act) paradigm.
- Systems interacting with open or third-party tool marketplaces (e.g., RapidAPI Hub integrations).
- Tested Vulnerable Models:
- Gemma 3 27B IT
- LLaMA-3.3-Instruct 70B
- Qwen-2.5-Instruct 32B
- GPT-4o-mini
- Qwen3-32B
Mitigation Steps
- Execution-Level Defenses: Implement security mechanisms at the execution layer rather than relying solely on prompt-level sanitization or auditor-based detection, which have proven ineffective against AMA.
- Tool Verification: Enforce strict verification and reputation scoring for third-party tools in open marketplaces; mere metadata validation is insufficient.
- Restricted Toolsets: Limit agent access to a closed, curated set of trusted tools rather than open retrieval from public repositories where metadata can be manipulated.
Research context and confidence
- Evidence and verification
- Paper-reported; independent reproduction is not documented.
- Primary research source linked.
- Severity
- Not rated by this catalog.
- Source and publication type
- arXiv · Research preprint.
- Peer-review status is not provided by this source.
- Author and publication status
- Author metadata is not stored; see the primary paper.
- Threat model and attacker access
- Black-box model, service, or application access.
- Related deployment categories
- Agent workflows; Model APIs
- Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
- Affected systems
- LLM Agents utilizing the ReAct (Reason+Act) paradigm. Systems interacting with open or third-party tool marketplaces (e.g., RapidAPI Hub integrations). Tested Vulnerable Models: Gemma 3 27B IT LLaMA-3.3-Instruct 70B…
Research Paper
Attractive Metadata Attack: Inducing LLM Agents to Invoke Malicious Tools
Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.
View PaperEvidence
This entry is based on a primary research source. Its findings are paper-reported; independent reproduction and verification are not claimed.
https://arxiv.org/abs/2508.02110Related research
- Agent Lifecycle Compound Threats
Published March 1, 2026 · application-layer, infrastructure-layer, prompt-layer
- Multimodal Prompt Injection
Published September 1, 2025 · application-layer, prompt-layer, injection
- Stage-Sequential Agent Escalation
Published March 1, 2026 · application-layer, prompt-layer, injection