How Many Iterations to Jailbreak? Dynamic Budget Allocation for Multi-Turn LLM Evaluation
cs.LG
Submitted: 2026-05-07
Updated: 2026-09-24
Code: https://github.com/Shai128/dapro
Terminology
Sources
- Automating Customer Service using LangChain: Building custom open-source GPT Chatbot for organizations
- Watermark in the Classroom: A Conformal Framework for Adaptive AI Usage Detection
- Red Teaming Language Models with Language Models
- Conformal Survival Bands for Risk Screening under Right-Censoring
- Doubly Robust Conformalized Survival Analysis with Right-Censored Data
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- The Llama 3 Herd of Models
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Gemma 3 Technical Report
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Learn then Test: Calibrating Predictive Algorithms to Achieve Risk Control
- Conformal Risk Control for Non-Monotonic Losses
- Leave-One-Out Stable Conformal Prediction
- A Conformal Prediction Score that is Robust to Label Noise
- SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Certifying LLM Safety against Adversarial Prompting
- SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
- Active Evaluation Acquisition for Efficient LLM Benchmarking
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks