AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation
Yi Ting Shen, Kentaroh Toyoda, Alex Leung
cs.CR, cs.AI
Submitted: 2026-07-13
Comments: A reference implementation is available at https://github.com/VulcanLab/amt-x
Code: https://github.com/VulcanLab/amt-x
License: http://creativecommons.org/licenses/by/4.0/
The gist: Safety evaluation of large language models (LLMs) relies largely on single-turn attack datasets and single-judge scoring, underestimating risk from adaptive multi-turn adversaries and reporting a
Terminology
Abstract
Safety evaluation of large language models (LLMs) relies largely on single-turn attack datasets and single-judge scoring, underestimating risk from adaptive multi-turn adversaries and reporting a single success rate that does not separate partially actionable outputs from those carrying complete operational detail. We propose AMT-X (Adaptive Multi-Turn Exploitation), a phase-structured multi-turn red-teaming framework. Unlike prior multi-turn attacks that rely on ad hoc escalation or free-form per-goal plans, AMT-X casts the attack as an explicit, reproducible multi-phase state machine driven by semantic signals from the victim, and replaces single-judge scoring with a multi-role jury whose phase-conditioned checklists gate success on actionable harm. Across six frontier victim models (queried under their default safety alignment, without added moderation layers) and seven Moderation sub-categories, AMT-X attains overall attack success rates of 97.6-100% under a lenient score threshold, but 66.7-78.6% under a stricter gate requiring complete, real, and operational detail: a gap of up to 33 percentage points between partially and fully actionable harm.
Sources
- Detecting Language Model Attacks with Perplexity
- Constitutional AI: Harmlessness from AI Feedback
- Jailbreaking Black Box Large Language Models in Twenty Queries
- garak: A Framework for Security Probing Large Language Models
- KTO: Model Alignment as Prospect Theoretic Optimization
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- Benchmarking Cognitive Biases in Large Language Models as Evaluators
- LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
- From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
- DrAttack: Prompt Decomposition and Reconstruction Makes Powerful LLM Jailbreakers
- Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
- FlipAttack: Jailbreak LLMs via Flipping
- PyRIT: A Framework for Security Risk Identification and Red Teaming in Generative AI System
- GPT-4 Technical Report
- LLM Evaluators Recognize and Favor Their Own Generations
- Automated Red Teaming with GOAT: the Generative Offensive Agent Tester
- X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents
- LLMs know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts
- Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack
- Sociotechnical Safety Evaluation of Generative AI Systems
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs