Rare Event Estimation via Iterative Unalignment
cs.LG, cs.AI
Submitted: 2026-09-21
Updated: 2026-09-21
Code: https://github.com/namkoong-lab/iterative-unalignment
License: http://creativecommons.org/licenses/by/4.0/
The gist: As agents are deployed with increased autonomy, even extremely rare events along their stochastic output trajectories can occur and prove catastrophic.
Terminology
Abstract
As agents are deployed with increased autonomy, even extremely rare events along their stochastic output trajectories can occur and prove catastrophic. Safe deployment therefore does not depend on whether these events can occur, but on how often they might. We study the problem of estimating the probability of rare events that arise from stochastic variation in the agent's own actions. Estimating this type of risk requires searching over the combinatorially vast space of trajectories. Naive Monte Carlo is computationally prohibitive in this regime, and constructing effective importance sampling (IS) proposals requires coordinated changes to a context-dependent chain of conditional distributions. We develop a new IS method that perturbs the original model's weights to construct the proposal. The proposal is itself a differentiably parameterized language model, enabling gradient-based search over weight space. We formulate an objective that combines a differentiable surrogate for event amplification and an adaptive regularization scheme that dynamically balances amplification against estimator stability. We evaluate our approach on about 120M and about 2.6B models across three event families spanning 300+ rare events as rare as 10-9, with reference probabilities computed with <10% relative standard error. In our most verifiable settings, we observe that our IS estimator achieves over 800 times compute-weighted efficiency gains over naive Monte Carlo for events with probabilities lower than 10-7. Our implementation is available at https://github.com/namkoong-lab/iterative-unalignment.
Sources
- Estimating Tail Risks in Language Model Output Distributions
- Constitutional AI: Harmlessness from AI Feedback
- Defending Against Unforeseen Failure Modes with Latent Adversarial Training
- Deep reinforcement learning from human preferences
- Rare Event Analysis of Large Language Models
- Explaining and Harnessing Adversarial Examples
- Exposing Long-Tail Safety Failures in Large Language Models through Efficient Diverse Response Sampling
- Forecasting Rare Language Model Behaviors
- Towards Deep Learning Models Resistant to Adversarial Attacks
- Gorilla: Large Language Model Connected with Massive APIs
- Risk Analysis Techniques for Governed LLM-based Multi-Agent Systems
- Toolformer: Language Models Can Teach Themselves to Use Tools
- "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
- Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
- A Survey on Autonomy-Induced Security Risks in Large Model-Based Agents
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Executable Code Actions Elicit Better LLM Agents
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks