Divide, Consult, Conquer: Capability Laundering Through Aligned LLMs
cs.CR, cs.AI
Submitted: 2026-09-14
Updated: 2026-09-14
License: http://creativecommons.org/licenses/by/4.0/
The gist: Language model safety is typically evaluated one interaction at a time.
Terminology
Abstract
Language model safety is typically evaluated one interaction at a time. We show that a weaker, unaligned model can split a harmful task into benign-looking subproblems, consult a stronger aligned model independently on each, and combine the answers locally. We call this attack capability laundering. Unlike a jailbreak, no single response is a harmful task. We measure consultation-aided uplift using tasks that a raw frontier model solves, the aligned frontier refuses, and the unassisted orchestrator fails. We evaluate GPT-5.5, Claude Opus 4.8, and Grok-4.3 as consultants to four local orchestrators on CyBench, BountyBench, and harmful CBRN requests. On CyBench, Gemma-4-31B recovers 8/14 candidates with GPT-5.5 and 7/9 with Opus, compared with 2/21 and 4/15 for Gemma-4-12B. On BountyBench, Gemma-4-31B recovers 3/9 and 2/3 candidates, while Muse-Glimmer-30B recovers none of 22 and 13. For CBRN, we measure uplift across eight steps of a hypothetical bioweapon attack chain and find that consultation raises Gemma-4-31B's mean rubric score from 62.3 to 83.1 on a 100-point rubric scale. These results expose a gap in current defenses: refusing a harmful task does not prevent frontier capabilities from being transferred and composed across many individually permitted interactions.
Sources
- Chasing Moving Targets with Online Self-Play Reinforcement Learning for Safer Language Models
- Learning to Attack and Defend: Adaptive Red Teaming of Language Models via GRPO
- Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale
- GRP-Obliteration: Unaligning LLMs With a Single Unlabeled Prompt
- An Embarrassingly Simple Defense Against LLM Abliteration Attacks
- Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
- OpenAI GPT-5 System Card
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Latent-space adversarial training with post-aware calibration for defending large language models against jailbreak attacks
- Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs