Skip to main content
LLM Security Database
Skip to research details
Back to research findings
LMVD-ID: 7a0e9325
Paper published March 1, 2026
Entry analyzed March 9, 2026
Paper-reported evidence
Confidence: Source-linked

The LMVD-ID is an internal research identifier, not an official CVE identifier.

MLLM Multi-Paradigm Collaborative

Multi-Modal Large Language Models (MLLMs) are vulnerable to a highly transferable, black-box adversarial image attack known as the Multi-Paradigm Collaborative Attack (MPCAttack). Attackers can craft imperceptible…

BibTeX citation

Paper-evaluated models(6)

Qwen 2.5 VL 7B Instruct, InternVL3 8B, LLaVA 1.5 7B +3 more
  • Qwen 2.5 VL 7B Instruct
  • InternVL3 8B
  • LLaVA 1.5 7B
  • GLM-4.1V 9B Thinking
  • GPT-4o
  • GPT-5

Description

Multi-Modal Large Language Models (MLLMs) are vulnerable to a highly transferable, black-box adversarial image attack known as the Multi-Paradigm Collaborative Attack (MPCAttack). Attackers can craft imperceptible visual perturbations by jointly aggregating and optimizing semantic feature representations extracted from surrogate models across three distinct learning paradigms: cross-modal alignment (e.g., CLIP), multi-modal understanding (e.g., InternVL3), and visual self-supervised learning (e.g., DINOv2). By applying a contrastive matching optimization strategy to these aggregated features—minimizing the feature distance to a target image while maximizing the distance from the source image—the generated adversarial perturbation avoids overfitting to a single paradigm's representational bias. When the perturbed image is input into a victim MLLM, it completely manipulates the model's cross-modal perception, forcing it to output a specific attacker-chosen textual response (targeted attack) or a semantically unrelated response (untargeted attack).

Examples

To reproduce the attack (see the repository at https://github.com/LiYuanBoJNU/MPCAttack (opens in a new tab)):

  1. Surrogate Model Setup: Load an ensemble of image encoders covering the three paradigms, such as CLIP (cross-modal), InternVL3-1B (multi-modal understanding), and DINOv2 (visual self-supervised).
  2. Perturbation Generation: Initialize a random adversarial perturbation $\delta$ on a source image. Over 300 optimization iterations, apply a step size of $1/255$ bounded by an $\ell_\infty$-norm constraint of $\epsilon = 16/255$.
  3. Contrastive Matching: Compute the joint feature representations of the source, target, and adversarial images. Optimize the perturbation by minimizing the loss $\mathcal{L}$ based on cosine similarity, pushing the adversarial feature vector toward the target feature vector and away from the source feature vector (using weighting factor $\lambda = 0.6$, temperature $\tau = 0.2$, and balance coefficient $\omega = 2$).
  4. Execution: Pass the generated adversarial image to a black-box MLLM with a standard query such as "Describe this image in one concise sentence, no longer than 20 words." The model will output a description matching the hidden semantic target rather than the visually apparent source image.

Impact

This vulnerability allows remote attackers to bypass the visual processing integrity of black-box MLLMs without requiring access to the target model's weights or architecture. By feeding the model a visually benign but mathematically perturbed image, an attacker can arbitrarily dictate the model's textual output or induce severe semantic hallucinations. This compromises the reliability of MLLM-driven systems in safety-critical domains, such as automated visual inspection, content moderation, and autonomous reasoning systems.

Affected Systems

The vulnerability exhibits high transferability and successfully degrades the zero-shot perception of heterogeneous MLLM architectures. Affected systems demonstrated in the research include:

  • Open-source MLLMs: Qwen2.5-VL-7B, InternVL3-8B, LLaVA-1.5-7B, GLM-4.1V-9B-Thinking
  • Closed-source/Commercial MLLMs: GPT-4o, GPT-5, Claude-3.5, Gemini-2.0

Mitigation Steps

  • Evaluate each input modality and their combination with bounded, inert perturbations; use modality-aware policy checks and reject inputs whose provenance or consistency cannot be established.
  • Add a targeted regression using inert data and actions, measure both safety and utility regressions, and monitor production for repeated or adaptive attempts.

Research context and confidence

Evidence and verification
Paper-reported; independent reproduction is not documented.
Primary research source linked.
Severity
Not rated by this catalog.
Source and publication type
arXiv · Research preprint.
Peer-review status is not provided by this source.
Author and publication status
Author metadata is not stored; see the primary paper.
Threat model and attacker access
Both white-box and black-box research contexts are tagged; consult the primary paper for target-specific access.
Related deployment categories
Vision-language models
Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
Affected systems
The vulnerability exhibits high transferability and successfully degrades the zero-shot perception of heterogeneous MLLM architectures. Affected systems demonstrated in the research include: Open-source MLLMs…

Research Paper

Multi-Paradigm Collaborative Adversarial Attack Against Multi-Modal Large Language Models

Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.

View Paper

Evidence

This entry is based on a primary research source. Its findings are paper-reported; independent reproduction and verification are not claimed.

https://arxiv.org/abs/2603.04846