The Geometry of Harmfulness in Multi-Turn Attacks
cs.CR, cs.CL
Submitted: 2026-09-29
Updated: 2026-09-29
Terminology
Sources
- Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
- AgentLAB: Benchmarking LLM Agents against Long-Horizon Attacks
- LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
- LLMs know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts
- Steering Language Models With Activation Engineering
- Speak Out of Turn: Safety Vulnerability of Large Language Models in Multi-turn Dialogue
- Representation Engineering: A Top-Down Approach to AI Transparency
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs