Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures
cs.AI, cs.CL, cs.CR
Submitted: 2026-09-24
Updated: 2026-09-24
Code: https://github.com/sumleo/RLCDAlignBench
Terminology
Sources
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
- Discovering Latent Knowledge in Language Models Without Supervision
- JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
- Jailbreaking Black Box Large Language Models in Twenty Queries
- Automated Researchers Can Mitigate Well-characterized Alignment Failures
- Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
- Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
- Gemma 2: Improving Open Language Models at a Practical Size
- The Llama 3 Herd of Models
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- Monitoring Monitorability
- WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
- Evaluating the Paperclip Maximizer: Are RL-Based Language Models More Likely to Pursue Instrumental Goals?
- Measuring Sycophancy of Language Models in Multi-turn Dialogues
- DeceptionBench: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenarios
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Language Models (Mostly) Know What They Know
- Reliable Weak-to-Strong Monitoring of LLM Agents
- AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection