Small Reward Models via Backward Inference
cs.CL
Submitted: 2026-02-14
Updated: 2026-08-31
Code: https://github.com/yikee/FLIP
Project page: https://allenai.github.io/open-instruct
Terminology
Sources
- Reverse Engineering Human Preferences with Reinforcement Learning
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Small Language Models are the Future of Agentic AI
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- BLEUBERI: BLEU is a surprisingly effective reward for instruction following
- Humans or LLMs as the Judge? A Study on Judgement Biases
- LongForm: Effective Instruction Tuning with Reverse Instructions
- REInstruct: Building Instruction Data from Unlabeled Corpus
- Gen-Z: Generative Zero-Shot Text Classification with Contextualized Label Descriptions
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA
- Preference Leakage: A Contamination Problem in LLM-as-a-judge
- Model Swarms: Collaborative Search to Adapt LLM Experts via Swarm Intelligence
- Reverse Prompt Engineering
- Position: Enough of Scaling LLMs! Lets Focus on Downscaling
- Self-Alignment with Instruction Backtranslation
- A Survey on LLM-as-a-Judge
- Benchmarking and Improving Generator-Validator Consistency of Language Models
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
- Is Best-of-N the Best of Them? Coverage, Scaling, and Optimality in Inference-Time Alignment
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering