JevAdvBench: A Benchmark and Black-Box Attacks for Reinforcement Learning for Calibrated Decisions Models
cs.CR, cs.AI, cs.CL
Submitted: 2026-09-25
Updated: 2026-09-25
Project page: https://jevadvbench.github.io/JevAdvBench/ABSTRACT
Terminology
Sources
- JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
- StruQ: Defending Against Prompt Injection with Structured Queries
- SecAlign: Defending Against Prompt Injection with Preference Optimization
- On Calibration of Modern Neural Networks
- Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures
- A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment
- Language Models (Mostly) Know What They Know
- Formalizing and Benchmarking Prompt Injection Attacks and Defenses
- Ignore Previous Prompt: Attack Techniques For Language Models
- Type-Safe Is Not Error-Free: A Constrained Decision Head Follows the Option Name, Not the Rubric Bound to It
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
- Calibrate Before Use: Improving Few-Shot Performance of Language Models
- One Token to Fool LLM-as-a-Judge
- Large Language Models Are Not Robust Multiple Choice Selectors
- PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs