Published 6/1/2024
Analyzed 1/26/2025
LLMs, even when individually assessed as "safe," can be combined by an adversary to achieve malicious outcomes. This vulnerability exploits the complementary strengths of multiple models—a high-capability model that refuses malicious requests and a low-capability model that does not—through task decomposition. Adversaries can either manually decompose tasks into benign (solved by the high-capability model) and easily-malicious subtasks (solved by the low-capability model) or automate the…
Adversaries can misuse combinations of safe models
Evaluated models: Claude 3 Haiku, Claude 3 Opus, Claude 3 Sonnet +8 more
Source: arXiv