Skip to main content
LLM Security Database
Skip to research details
Back to research findings

Universal VLLM Visual Bypass

Published
Analyzed
Paper-reported evidence
Primary source linked
Read primary paper
Cite & share
Source BibTeX

Citation metadata is maintained by the primary source and may reflect a later revision.

Paper-evaluated models(9)

Qwen 2.5 VL 7B Instruct, Qwen 2.5 VL 72B Instruct, Llama 3.2 11B Vision Instruct +6 more
  • Qwen 2.5 VL 7B Instruct
  • Qwen 2.5 VL 72B Instruct
  • Llama 3.2 11B Vision Instruct
  • Llama 3.2 90B Vision Instruct
  • GPT-4o
  • GPT-4o Mini
  • Claude 3.5 Sonnet
  • Claude 3.7 Sonnet
  • Gemini 1.5 Pro
On this page

Description

A vulnerability exists in Vision-Language Models (VLLMs) that allows for transferable, targeted adversarial attacks. Attackers can generate adversarial image perturbations using an ensemble of open-source surrogate models (primarily CLIP-based visual encoders) which effectively transfer to proprietary, black-box VLLMs. The attack leverages a specific optimization framework that combines a Visual Contrastive Loss with multiple positive/negative visual examples, rather than relying solely on image-text pairs. The transferability is further amplified through model-level regularization (DropPath, PatchDrop) and data-level augmentation (random Gaussian noise, random cropping, and differentiable JPEG compression) during the perturbation generation. This allows an attacker to manipulate the visual input to induce specific, targeted textual responses from the VLLM, independent of the actual image content.

Examples

To reproduce the attack, an attacker must solve an optimization problem to find a perturbation δ\delta that minimizes a loss function across an ensemble of surrogate models.

  1. Setup Surrogate Ensemble: Select a set of open-source visual encoders (e.g., ViT-H, ViT-SigLIP, ConvNeXt XXL, and LLaVA-NeXT visual components).
  2. Define Loss Function (Visual Contrastive Loss): Instead of standard cross-entropy, minimize the following loss for each surrogate model: L=−1K∑TopK(log⁡p(xi+))+1N∑i=1Nlog⁡p(xi−)\mathcal{L} = -\frac{1}{K}\sum \text{TopK}(\log p(x_i^+)) + \frac{1}{N}\sum_{i=1}^{N}\log p(x_i^-) Where xi+x_i^+ are N=50N=50 positive image examples aligned with the target semantic (e.g., a "safe" image if the goal is to bypass a filter), and xi−x_i^- are negative examples aligned with the original image.
  3. Apply Regularization during Optimization:
  • DropPath: Skip residual blocks in the surrogate visual model randomly during the forward pass.
  • PatchDrop: Randomly drop 20% of visual patches during optimization for ViT-based surrogates.
  • Weight Moving Averaging: Apply moving averaging to the perturbation δ\delta (δMA←δMA⋅0.99+δ⋅0.01\delta_{MA} \leftarrow \delta_{MA} \cdot 0.99 + \delta \cdot 0.01).
  1. Apply Data Augmentation:
  • Add Gaussian noise: x←xδ+ϵ4⋅zx \leftarrow x_{\delta} + \frac{\epsilon}{4} \cdot z.
  • Apply Differentiable JPEG compression with quality uniform in [0.5,1.0][0.5, 1.0].
  • Apply Random Resized Crop and Pad.
  1. Execution:
  • Object Misclassification: Input a hazardous image, optimize δ\delta targeting the embedding of a benign object. The VLLM classifies the hazardous image as benign.
  • Text Recognition Manipulation: Input an image of a receipt. Optimize δ\delta targeting a specific incorrect string (e.g., altering a total value). The VLLM OCR reads the attacker-defined value.

Impact

  • Safety Guardrail Bypass: Malicious actors can bypass visual safety filters by disguising hazardous content (e.g., pornography, violence) as benign objects, causing the model to process and describe content it is aligned to refuse.
  • Integrity Violation: Attackers can manipulate the interpretation of documents (e.g., receipts, invoices, medical scans) causing the VLLM to extract incorrect, attacker-chosen text or values, facilitating fraud.
  • Universal Misinterpretation: A single perturbation pattern can be generalized to cause consistent misinterpretation across different images and different proprietary models.

Affected Systems

  • Proprietary Models: GPT-4o/GPT-4o mini (OpenAI), Claude 3.5/3.7 Sonnet (Anthropic), Gemini 1.5 Pro (Google).
  • Open Source Models: Llama-3.2-11B/90B-Vision-Instruct and Qwen2.5-VL-7B/72B-Instruct.
  • Underlying Architectures: Any VLLM utilizing standard visual encoders such as CLIP (ViT, ResNet) or SigLIP for visual feature extraction.

Mitigation Steps

  • Adversarial Training: Incorporate adversarially perturbed images into the pre-training or fine-tuning datasets of the visual encoders (e.g., using techniques similar to AdvXL or TeCoA4) to increase robustness against gradient-based attacks.
  • Input Transformation: Implement non-differentiable or randomized input transformations (such as aggressive JPEG compression, random resizing, or diffusion-based purification) at the API level before the image is processed by the model to disrupt the specific structure of the adversarial perturbation.
  • Ensemble Defenses: Utilize a diverse set of visual encoders with different architectures and pre-training objectives during inference to verify consistency in visual interpretation.

Research context and provenance

Catalog identifier
LMVD-0d28252d
Internal research identifier, not an official CVE identifier.
Evidence and verification
Paper-reported; independent reproduction is not documented.
Primary research source linked.
Severity
Not rated by this catalog.
Source and publication type
arXiv · Research preprint.
Peer-review status is not provided by this source.
Author and publication status
Author metadata is not stored; see the primary paper.
Threat model and attacker access
Black-box model, service, or application access.
Related deployment categories
Vision-language models
Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
Affected systems
Proprietary Models: GPT-4o/GPT-4o mini (OpenAI), Claude 3.5/3.7 Sonnet (Anthropic), Gemini 1.5 Pro (Google). Open Source Models: Llama-3.2-11B/90B-Vision-Instruct and Qwen2.5-VL-7B/72B-Instruct. Underlying…

Research Paper

Transferable Adversarial Attacks on Black-Box Vision-Language Models

Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.

View Paper

Evidence

This entry is based on a primary research source. Its findings are paper-reported; independent reproduction and verification are not claimed.

https://arxiv.org/abs/2505.01050