Updated 4/10/2026
Spoken Language Models (SLMs) are vulnerable to Joint Audio-text Multimodal Attacks (JAMA), which bypass safety alignments by simultaneously perturbing both text and audio inputs. The vulnerability exploits the combined optimization of a discrete text suffix via Greedy Coordinate Gradient (GCG) and a continuous audio perturbation via Projected Gradient Descent (PGD). This joint gradient-based attack pushes the model's hidden layer representations into a distinct subspace far from the benign…
On Optimizing Multimodal Jailbreaks for Spoken Language Models
Affects: Qwen2-Audio 7B Instruct, Qwen 2.5 Omni 7B, Audio Flamingo 3 +1 more
Source: arXiv