R3: Robust Rubric-Agnostic Reward Models
cs.CL, cs.AI, cs.LG
Submitted: 2025-05-19
Updated: 2026-09-15
Code: https://github.com/rubricreward/r3
Terminology
Sources
- Phi-4-reasoning Technical Report
- GPT-4 Technical Report
- Critique-out-Loud Reward Models
- JudgeLRM: Large Reasoning Models as a Judge
- RM-R1: Reward Modeling as Reasoning
- UltraFeedback: Boosting Language Models with Scaled AI Feedback
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Zero-shot Faithfulness Evaluation for Text Summarization with Foundation Language Model
- The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models
- Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- RewardBench: Evaluating Reward Models for Language Modeling
- ACUTE-EVAL: Improved Dialogue Evaluation with Optimized Questions and Multi-turn Comparisons
- Towards General Text Embeddings with Multi-stage Contrastive Learning
- WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
- Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
- RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style
- s1: Simple test-time scaling
- Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering