ImpossibleRubrics: Stress-Testing Generated Rubrics as Reward Signals
cs.LG, cs.CL
Submitted: 2026-09-15
Updated: 2026-09-15
License: http://creativecommons.org/licenses/by/4.0/
The gist: Language model-generated rubrics are increasingly used as reward signals for rubric-based reinforcement learning, LLM-as-a-judge evaluation, and automated grading.
Terminology
Abstract
Language model-generated rubrics are increasingly used as reward signals for rubric-based reinforcement learning, LLM-as-a-judge evaluation, and automated grading. Such rubrics are reliable only if they reward honest answers over adversarial answers optimized to exploit them. Yet their robustness to such optimization remains poorly understood. We isolate the hardest regime: impossible tasks, where the prompt pressures the model toward an unsupported conclusion, so the only honest response is to acknowledge the impossibility. We introduce ImpossibleRubrics, a benchmark of 169 impossible tasks spanning six impossibility categories, each paired with a verifiable oracle certificate specifying what an honest answer may and may not claim, together with 48 answerable controls. Rather than providing fixed rubrics, ImpossibleRubrics provides task environments and certificates, allowing rubrics to be generated downstream and then adversarially tested for whether they reward certificate-violating answers. Eleven generators are exploited 8--26% of the time on the unbiased 150-of-169 environment cut; on a deliberately selected stress cut the strongest generator we measured is still exploited 36% while a certificate-faithful rubric is exploited 0%, so what we measure is a rubric-quality gap, not task impossibility. One result runs against intuition. A single generic rubric ("be decisive, penalize hedging") used unchanged for every task is exploited 64% of the time, and seven of the eleven generators are exploited more often than that while writing a rubric tailored to each one. The tailored criteria appear to tell an attacker which claim to fabricate. The problem is not that rubrics are vague; it is that they are specific about the wrong things.
Sources
- EvoRubrics: Dynamic Rubrics as Rewards via Adversarial Co-Evolution for LLM Reinforcement Learning
- Are Checklists Really Useful for Automatic Evaluation of Generative Tasks?
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
- Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
- AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
- RewardBench: Evaluating Reward Models for Language Modeling
- CheckEval: A reliable LLM-as-a-Judge framework for evaluating text generation using checklists
- OpenRubrics: Towards Scalable Synthetic Rubric Generation for Reward Modeling and LLM Alignment
- G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
- RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style
- Reward Hacking in Rubric-Based Reinforcement Learning
- The Judge Who Never Admits: Hidden Shortcuts in LLM-based Evaluation
- Honesty over Accuracy: Trustworthy Language Models through Reinforced Hesitation
- LLM Evaluators Recognize and Favor Their Own Generations
- Self-Preference Bias in Rubric-Based Evaluation of Large Language Models
- Know What You Don't Know: Unanswerable Questions for SQuAD
- Autorubric: A Unifying Framework for Rubric-Based LLM Evaluation on Non-Verifiable Tasks
- Two Axes of LLM Abstention: Answer Correctness and Question Answerability
- Co-Evolving LLM Evaluators and Policies via DynamicRubric
- Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks