BLOOM-WILT: Logit Tilting for Behaviour Elicitation in Automated LLM Auditing
cs.AI, cs.CL
Submitted: 2026-08-31
Updated: 2026-08-31
Code: https://github.com/AdrSkapars/bloom-wilt
Terminology
Sources
- Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
- Estimating Tail Risks in Language Model Output Distributions
- Data Swarms: Optimizable Generation of Synthetic Evaluation Data
- The Llama 3 Herd of Models
- Refusal in Language Models Is Mediated by a Single Direction
- Classifier-Free Diffusion Guidance
- Eliciting Behaviors in Multi-Turn Conversations
- Best-of-N Jailbreaking
- Jailbreaking Black Box Large Language Models in Twenty Queries
- Forecasting Rare Language Model Behaviors
- Tuning Language Models by Proxy
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
- Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
- Frontier Models are Capable of In-context Scheming
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Large Language Models Often Know When They Are Being Evaluated
- Estimating the Probabilities of Rare Outputs in Language Models
- Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection