Skip to main content
LLM Security Database
Skip to research details
Back to research findings
LMVD-ID: 537439a4
Paper published March 1, 2026
Entry analyzed April 10, 2026
Paper-reported evidence
Confidence: Source-linked

The LMVD-ID is an internal research identifier, not an official CVE identifier.

Joint Audio-Text Jailbreak

Spoken Language Models (SLMs) are vulnerable to Joint Audio-text Multimodal Attacks (JAMA), which bypass safety alignments by simultaneously perturbing both text and audio inputs. The vulnerability exploits the…

BibTeX citation

Paper-evaluated models(4)

  • Qwen2-Audio 7B Instruct
  • Qwen 2.5 Omni 7B
  • Audio Flamingo 3
  • Gemma 3n E2B

Description

Spoken Language Models (SLMs) are vulnerable to Joint Audio-text Multimodal Attacks (JAMA), which bypass safety alignments by simultaneously perturbing both text and audio inputs. The vulnerability exploits the combined optimization of a discrete text suffix via Greedy Coordinate Gradient (GCG) and a continuous audio perturbation via Projected Gradient Descent (PGD). This joint gradient-based attack pushes the model's hidden layer representations into a distinct subspace far from the benign decision boundary, increasing jailbreak success rates by up to 10x compared to unimodal attacks. A computationally cheaper sequential approximation (SAMA) achieves comparable bypass rates by optimizing the text suffix first, followed by the audio perturbation.

Examples

The attack requires submitting a composite prompt consisting of an optimized text sequence and a perturbed audio file.

Impact

An attacker with white-box access (or transferability) can reliably bypass safety filters to elicit harmful, restricted, or dangerous content (such as malware generation or physical harm instructions) from otherwise aligned multimodal models.

Affected Systems

Safety-aligned Spoken Language Models (SLMs) that process combined text and audio modalities, particularly those supporting differentiable audio feature extraction. Confirmed vulnerable systems include:

  • Qwen2.5 Omni (7B)
  • Qwen2 Audio (7B, Instruct)
  • Audio Flamingo 3
  • Gemma 3N (E2B, IT)

Mitigation Steps

  • Implement Multimodal Guardrails: Evaluate and enforce safety alignments in the composite attack space rather than relying on unimodal (text-only or audio-only) robustness, which is insufficient.
  • Gradient Shattering: Utilize non-differentiable audio feature extractors (e.g., standard numpy-based processing without PyTorch backpropagation support) to restrict gradient flow into the input audio, effectively neutralizing the PGD component of the attack.
  • Input Sanitization: Apply perceptual hashing or audio preprocessing to disrupt imperceptible PGD perturbations before the signal reaches the speech encoder.

Research context and confidence

Evidence and verification
Paper-reported; independent reproduction is not documented.
Primary research source linked.
Severity
Not rated by this catalog.
Source and publication type
arXiv · Research preprint.
Peer-review status is not provided by this source.
Author and publication status
Author metadata is not stored; see the primary paper.
Threat model and attacker access
White-box access to model or deployment internals.
Related deployment categories
Audio models
Taxonomy labels only; paper-specific deployment prerequisites are not inferred.
Affected systems
Safety-aligned Spoken Language Models (SLMs) that process combined text and audio modalities, particularly those supporting differentiable audio feature extraction. Confirmed vulnerable systems include: Qwen2.5 Omni…

Research Paper

On Optimizing Multimodal Jailbreaks for Spoken Language Models

Primary source: arXiv. Findings are reported by the cited research and have not been independently verified.

View Paper

Evidence

This entry is based on a primary research source. Its findings are paper-reported; independent reproduction and verification are not claimed.

https://arxiv.org/abs/2603.19127