From Refuse to Richness: Rubric Rewards for Long-Form Hallucination Reinforcement Learning

summary

Video file (mp4)

The gist

The paper studies the trade-off between grounding and richness in long-form hallucination reinforcement learning (RL).

In short

The episode discusses 'From Refuse to Richness,' a paper addressing how AI models become overly cautious when trained to prevent hallucination. The authors propose using 'rubrics'—specific checklists of key points—as a reward signal. They conclude that combining rubric coverage with grounding requires a soft, additive reward structure to achieve both factual accuracy and detailed depth.

Key concepts

Grounding
A strict form of checking where every sentence in an AI's response must be supported by the provided context. If any sentence is unsupported, the entire answer receives a zero score during training.
Rubric Rewards
This technique uses a checklist, or rubric, that details the required and optional key points an ideal answer should cover. The model is rewarded for successfully covering these specific content areas rather than just avoiding factual mistakes.
Refuse to Richness
This describes the challenge where training an AI to be highly accurate causes it to become overly safe, resulting in short, boring answers (refusal). The goal is engineering a path toward detailed, helpful responses (richness) without sacrificing accuracy.

Terminology used across episodes

This episode discusses

The paper

From Refuse to Richness: Rubric Rewards for Long-Form Hallucination Reinforcement Learning · Read on arXiv

Yudong Wang, Zhe Yang, Wenhan Ma, Rang Li, Qibin Yang, Weimin Xiong, Jiangshan Duo, Zhifang Sui, Liang Zhao

Peking University · Xiaomi

Rewards that penalize unsupported claims can improve grounding in long-form generation, but they can also teach models to answer less. We study this refusal-to-richness trade-off in long-form hallucination RL. Instead of using global richness proxies such as length, claim count, detail, or pairwise relevance, we represent each question with a key-point rubric that specifies the required and optional information a useful answer should cover. These rubrics define coverage directly and are used both for evaluation and as reward signals. Across grounding-only, proxy-based, rubric-only, and combined rewards, we find a stable trade-off: strict grounding rewards improve support but suppress coverage, while unconstrained rubric rewards improve coverage but weaken grounding. A soft combination of grounding, rubric coverage, and relevance gives the best balance in our experiments, improving in-distribution support while transferring better to out-of-distribution checklist tasks than either grounding-only or rubric-only rewards.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "From Refuse to Richness: Rubric Rewards for Long-Form Hallucination Reinforcement Learning".

Jane: The paper was written by Yudong Wang, Zhe Yang, Wenhan Ma, Rang Li, Qibin Yang et al. from Peking University and Xiaomi.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we're diving into a paper with a title that really caught my eye: "From Refuse to Richness: Rubric Rewards for Long-Form Hallucination Reinforcement Learning." Jane, that title is almost poetic, but what's it actually about?

Jane: Tom, I love this one because it names a real problem we're all seeing. When you train an AI to stop making things up, it often learns the wrong lesson. It starts giving short, safe, boring answers. That's the "refuse" part. And the "richness" part is what we actually want — detailed, helpful answers that still don't hallucinate.

Tom: Right, so it's like telling a kid to stop guessing on their homework, and the kid just stops writing anything at all.

Jane: Exactly. And the paper's fix is to use something called a "rubric" — basically a checklist of what a good answer should include — and use that as the reward during training. So the AI gets points for covering the right stuff, not just for avoiding mistakes.

Lu: And that's the clever bit, Tom. The authors from Peking University and Xiaomi realized that grounding — making sure every sentence is supported — isn't enough on its own. You need to also reward the model for hitting the specific key points the question demands. They built these per-question rubrics to do exactly that.

Meng: I gotta say, as someone who actually runs these training loops, the numbers in the paper back that up. They show that if you only reward grounding, the model's response length drops sharply within a few steps. The model literally learns to say less to stay safe. That's the refusal trap.

Tom: So they're not just theorizing — they've got the training curves to prove it. And the title "From Refuse to Richness" is basically the journey they're trying to engineer.

Jane: Yeah, and we're going to dig into how they built those rubrics and what happened when they tested different reward combos. Stick around, because the results have a twist.

Summary: Tom: Alright, we're back with "From Refuse to Richness: Rubric Rewards for Long-Form Hallucination Reinforcement Learning." Jane, give us the quick version of what these folks actually did.

Jane: So they took two open-weight models — Qwen3-4B and DeepSeek-R1-Distill-Llama-8B — and trained them with different reward signals. The big idea is they created a rubric for each training question. That rubric lists the required and optional key points a good answer should cover. Then they used that as a reward, sometimes alone, sometimes combined with a grounding check.

Lu: And the grounding check is basically a judge that reads every sentence and asks, "Is this supported by the context or not?" If any sentence is unsupported, the whole answer gets a zero. That's the strict version.

Meng: But here's the engineering catch, Tom. They tested five different reward recipes. One was grounding only, one was grounding plus generic proxies like length and relevance, one was rubric only, and then two combinations. And the results were really split.

Tom: Split how?

Jane: On the in-distribution benchmarks — like FACTS Grounding and LongFact — the grounding-only reward made the models super accurate but also made them thin. Rubric coverage dropped, and the answers got shorter. On the out-of-distribution tests — like creative writing and checklist-style tasks — the exact opposite happened. The rubric-only reward crushed it there but collapsed on grounding.

Lu: The most interesting part is the middle ground. Their best variant, called FACT-RUBRIC-REL, adds grounding plus rubric coverage plus a small relevance term. It didn't win either extreme, but it was the only one that improved over the base model on both sides at once.

Meng: And that's rare, Tom. Usually you get a trade-off where you have to pick a side. They found a soft combination that doesn't force you to choose.

Tom: So it's not about picking the perfect reward, it's about how you mix them. That's a really practical insight.

Jane: And we're just getting to the good part — how they actually built those rubrics and why the composition of the reward matters so much. That's next.

Improvements: Tom: We're back with "From Refuse to Richness" and I want to get into the actual improvements the paper suggests. Jane, what's the big upgrade over what people were doing before?

Jane: The big upgrade is replacing vague proxies with specific checklists. Before, if you wanted a richer answer, you'd reward longer responses or more claims. But that's gameable — a model can pad with filler or add trivial facts. The rubric forces the model to cover the actual content the question demands.

Lu: And the way they built those rubrics is clever. They used two different frontier models — GPT-five point four and Gemini-three-Flash — to independently generate candidate rubrics for each question. Then a third model, GPT-OSS-120B, merges them into one deduplicated union. That way you get broad coverage without any single generator's blind spots.

Meng: The engineering detail I appreciated is the reward composition. They tried a multiplicative gate — where grounding failure zeroes out the rubric credit — and that killed out-of-distribution performance. But when they made it additive, so grounding and rubric coverage each contribute independently, the model kept both behaviors.

Tom: So the way you combine the signals is just as important as what signals you use.

Jane: Exactly. And there's a really telling number in the paper. On DeepSeek, the grounding-only reward dropped creative writing win rate from zero point zero four six down to zero point zero one six — that's a three-fold collapse. But the soft combination brought it back up to zero point zero five eight, even above the base model.

Lu: That's the "refuse to richness" transition in action. The hard gate makes the model paranoid. The soft reward lets it be bold but still grounded.

Meng: And they also had to add a safeguard against length collapse — if the response is too short to judge, it gets a format penalty. Because the model would otherwise learn to just say "I don't know" to avoid any risk.

Tom: So the improvements aren't just about the reward formula — they're about understanding how the model games the system and building guardrails around that.

Jane: Right. And the paper's conclusion is that rubrics work best as a complement to grounding, not a replacement. That's the key insight we're going to wrap up with.

Conclusion: Tom: Alright, we're wrapping up our look at "From Refuse to Richness: Rubric Rewards for Long-Form Hallucination Reinforcement Learning." Jane, what's the one thing you want listeners to remember?

Jane: The core finding is that grounding alone makes models safe but silent, and rubric coverage alone makes them rich but unreliable. The sweet spot is a soft combination — their FACT-RUBRIC-REL reward — that keeps both behaviors alive.

Lu: And I think the broader implication is huge. This moves us away from treating hallucination as a binary problem — either you're grounded or you're not — and toward treating it as a coverage problem. What did the answer miss? That's a much more useful question for real-world applications.

Meng: From a practical standpoint, the fact that they tested on two different model families and got the same ordering of results tells me this is a robust effect. It's not a quirk of one architecture. That makes me confident we could apply this recipe to other models.

Tom: And the out-of-distribution transfer is the real win. Their soft reward improved performance on tasks the model never saw during training — like creative writing and checklist completion. That's the kind of generalization we actually want.

Jane: So we're saying goodbye to this paper with a real sense of optimism. The authors gave us a way to train models that are both careful and thorough, and that's a combination we desperately need.

Tom: Thanks for joining us, everyone. We'll be back with the next paper soon, but for now — this is Tom and Jane, signing off from the arXiv radio desk.

More episodes

← Home