Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions
Shi Lin, Peng Qian, Dinghao Liu, Renjie Sun, Sifan Wu, Dezhang Kong, Chenpei Wang, Xun Wang
cs.LG, cs.CR
Submitted: 2026-07-29
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- GPT-4 Technical Report
- ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
- The Llama 3 Herd of Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- A Survey of LLM-Driven AI Agent Communication: Protocols, Security Risks, and Defense Countermeasures
- Reasoning as a Weapon: Adaptive Dual-Path Jailbreak Attack on Large Language Models
- ICON: Intent-Context Coupling for Efficient Multi-Turn Jailbreak Attack
- DeepSeek-V3 Technical Report
- THRD: A Training-Free Multi-Turn Defense Framework for Jailbreak Attacks on Large Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Qwen3 Technical Report
- Chain of Attack: a Semantic-Driven Contextual Multi-Turn attacker for LLM
- ShieldGemma: Generative AI Content Moderation Based on Gemma
- Qwen3Guard Technical Report
- Can ChatGPT Understand Too? A Comparative Study on ChatGPT and Fine-tuned BERT
- Tempest: Autonomous Multi-Turn Jailbreaking of Large Language Models with Tree Search
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks