From Refuse to Richness: Rubric Rewards for Long-Form Hallucination Reinforcement Learning

arXiv:2608.12337 · cs.CL · Submitted 2026-06-03 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "From Refuse to Richness: Rubric Rewards for Long-Form Hallucination Reinforcement Learning".

Jane: The paper was written by Yudong Wang, Zhe Yang, Wenhan Ma, Rang Li, Qibin Yang et al. from Peking University and Xiaomi.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we're diving into a paper with a title that really caught my eye: "From Refuse to Richness: Rubric Rewards for Long-Form Hallucination Reinforcement Learning." Jane, that title is almost poetic, but what's it actually about?

Jane: Tom, I love this one because it names a real problem we're all seeing. When you train an AI to stop making things up, it often learns the wrong lesson. It starts giving short, safe, boring answers. That's the "refuse" part. And the "richness" part is what we actually want — detailed, helpful answers that still don't hallucinate.

Tom: Right, so it's like telling a kid to stop guessing on their homework, and the kid just stops writing anything at all.

Jane: Exactly. And the paper's fix is to use something called a "rubric" — basically a checklist of what a good answer should include — and use that as the reward during training. So the AI gets points for covering the right stuff, not just for avoiding mistakes.

Lu: And that's the clever bit, Tom. The authors from Peking University and Xiaomi realized that grounding — making sure every sentence is supported — isn't enough on its own. You need to also reward the model for hitting the specific key points the question demands. They built these per-question rubrics to do exactly that.

Meng: I gotta say, as someone who actually runs these training loops, the numbers in the paper back that up. They show that if you only reward grounding, the model's response length drops sharply within a few steps. The model literally learns to say less to stay safe. That's the refusal trap.

Tom: So they're not just theorizing — they've got the training curves to prove it. And the title "From Refuse to Richness" is basically the journey they're trying to engineer.

Jane: Yeah, and we're going to dig into how they built those rubrics and what happened when they tested different reward combos. Stick around, because the results have a twist.

Summary: Tom: Alright, we're back with "From Refuse to Richness: Rubric Rewards for Long-Form Hallucination Reinforcement Learning." Jane, give us the quick version of what these folks actually did.

Jane: So they took two open-weight models — Qwen3-4B and DeepSeek-R1-Distill-Llama-8B — and trained them with different reward signals. The big idea is they created a rubric for each training question. That rubric lists the required and optional key points a good answer should cover. Then they used that as a reward, sometimes alone, sometimes combined with a grounding check.

Lu: And the grounding check is basically a judge that reads every sentence and asks, "Is this supported by the context or not?" If any sentence is unsupported, the whole answer gets a zero. That's the strict version.

Meng: But here's the engineering catch, Tom. They tested five different reward recipes. One was grounding only, one was grounding plus generic proxies like length and relevance, one was rubric only, and then two combinations. And the results were really split.

Tom: Split how?

Jane: On the in-distribution benchmarks — like FACTS Grounding and LongFact — the grounding-only reward made the models super accurate but also made them thin. Rubric coverage dropped, and the answers got shorter. On the out-of-distribution tests — like creative writing and checklist-style tasks — the exact opposite happened. The rubric-only reward crushed it there but collapsed on grounding.

Lu: The most interesting part is the middle ground. Their best variant, called FACT-RUBRIC-REL, adds grounding plus rubric coverage plus a small relevance term. It didn't win either extreme, but it was the only one that improved over the base model on both sides at once.

Meng: And that's rare, Tom. Usually you get a trade-off where you have to pick a side. They found a soft combination that doesn't force you to choose.

Tom: So it's not about picking the perfect reward, it's about how you mix them. That's a really practical insight.

Jane: And we're just getting to the good part — how they actually built those rubrics and why the composition of the reward matters so much. That's next.

Improvements: Tom: We're back with "From Refuse to Richness" and I want to get into the actual improvements the paper suggests. Jane, what's the big upgrade over what people were doing before?

Jane: The big upgrade is replacing vague proxies with specific checklists. Before, if you wanted a richer answer, you'd reward longer responses or more claims. But that's gameable — a model can pad with filler or add trivial facts. The rubric forces the model to cover the actual content the question demands.

Lu: And the way they built those rubrics is clever. They used two different frontier models — GPT-five point four and Gemini-three-Flash — to independently generate candidate rubrics for each question. Then a third model, GPT-OSS-120B, merges them into one deduplicated union. That way you get broad coverage without any single generator's blind spots.

Meng: The engineering detail I appreciated is the reward composition. They tried a multiplicative gate — where grounding failure zeroes out the rubric credit — and that killed out-of-distribution performance. But when they made it additive, so grounding and rubric coverage each contribute independently, the model kept both behaviors.

Tom: So the way you combine the signals is just as important as what signals you use.

Jane: Exactly. And there's a really telling number in the paper. On DeepSeek, the grounding-only reward dropped creative writing win rate from zero point zero four six down to zero point zero one six — that's a three-fold collapse. But the soft combination brought it back up to zero point zero five eight, even above the base model.

Lu: That's the "refuse to richness" transition in action. The hard gate makes the model paranoid. The soft reward lets it be bold but still grounded.

Meng: And they also had to add a safeguard against length collapse — if the response is too short to judge, it gets a format penalty. Because the model would otherwise learn to just say "I don't know" to avoid any risk.

Tom: So the improvements aren't just about the reward formula — they're about understanding how the model games the system and building guardrails around that.

Jane: Right. And the paper's conclusion is that rubrics work best as a complement to grounding, not a replacement. That's the key insight we're going to wrap up with.

Conclusion: Tom: Alright, we're wrapping up our look at "From Refuse to Richness: Rubric Rewards for Long-Form Hallucination Reinforcement Learning." Jane, what's the one thing you want listeners to remember?

Jane: The core finding is that grounding alone makes models safe but silent, and rubric coverage alone makes them rich but unreliable. The sweet spot is a soft combination — their FACT-RUBRIC-REL reward — that keeps both behaviors alive.

Lu: And I think the broader implication is huge. This moves us away from treating hallucination as a binary problem — either you're grounded or you're not — and toward treating it as a coverage problem. What did the answer miss? That's a much more useful question for real-world applications.

Meng: From a practical standpoint, the fact that they tested on two different model families and got the same ordering of results tells me this is a robust effect. It's not a quirk of one architecture. That makes me confident we could apply this recipe to other models.

Tom: And the out-of-distribution transfer is the real win. Their soft reward improved performance on tasks the model never saw during training — like creative writing and checklist completion. That's the kind of generalization we actually want.

Jane: So we're saying goodbye to this paper with a real sense of optimism. The authors gave us a way to train models that are both careful and thorough, and that's a combination we desperately need.

Tom: Thanks for joining us, everyone. We'll be back with the next paper soon, but for now — this is Tom and Jane, signing off from the arXiv radio desk.

Yudong Wang, Zhe Yang, Wenhan Ma, Rang Li, Qibin Yang, Weimin Xiong, Jiangshan Duo, Zhifang Sui, Liang Zhao

Peking University · Xiaomi

cs.CL

Submitted: 2026-06-03

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 63/100

The gist: The paper studies the trade-off between grounding and richness in long-form hallucination reinforcement learning (RL).

Key concepts

Grounding
A strict form of checking where every sentence in an AI's response must be supported by the provided context. If any sentence is unsupported, the entire answer receives a zero score during training.
Rubric Rewards
This technique uses a checklist, or rubric, that details the required and optional key points an ideal answer should cover. The model is rewarded for successfully covering these specific content areas rather than just avoiding factual mistakes.
Refuse to Richness
This describes the challenge where training an AI to be highly accurate causes it to become overly safe, resulting in short, boring answers (refusal). The goal is engineering a path toward detailed, helpful responses (richness) without sacrificing accuracy.

Terminology

Summary

The paper studies the trade-off between grounding and richness in long-form hallucination reinforcement learning (RL). The authors formalize richness as question-specific information coverage, represented by per-sample key-point rubrics that specify required and optional information a useful answer should cover. They frame long-form hallucination RL as jointly mitigating factual and faithfulness hallucination.

They build a multi-generator union pipeline for per-sample key-point rubrics, applied to four grounding benchmarks (FACTS Grounding, LongFact, FActScore), CL-bench, and a 4K RL training set. The RL training pool contains 4,000 prompts: 2,000 grounded prompts with context documents and 2,000 open prompts without context.

The authors compare five reward variants: FACT-ONLY (binary grounding reward), FACT-PROXY (grounding plus generic detail and relevance proxies), FACT-RUBRIC (additive combination of grounding and rubric coverage), FACT-RUBRIC-REL (soft combination of grounding, rubric coverage, and pairwise relevance), and RUBRIC-ONLY (rubric coverage without explicit grounding).

Across experiments on Qwen3-4B and DeepSeek-R1-Distill-Llama-8B, they find a stable trade-off: strict grounding rewards improve support but suppress coverage, while unconstrained rubric rewards improve coverage but weaken grounding. Specifically, FACT-ONLY RL training causes response length to drop sharply within a few RL steps, as models learn that the safest way to avoid unsupported claims is to say less. On in-distribution grounding benchmarks, FACT-ONLY reaches the highest grounding accuracy but reduces both rubric coverage and claim number below the un-trained base. RUBRIC-ONLY achieves the strongest out-of-distribution numbers while collapsing on in-distribution grounding.

The out-of-distribution picture reverses the in-distribution ordering: hard grounding gates suppress transfer, whereas RUBRIC-ONLY and soft FACT-RUBRIC rewards give the strongest completion and pairwise-preference results. The authors find that multiplicative gating of rubric rewards inherits the out-of-distribution weakness of pure grounding rewards, while additive grounded-coverage rewards recover only part of the out-of-distribution gap.

The main result is that a soft three-component reward, FACT-RUBRIC-REL, combining a non-gating grounding term (0.5 G), rubric coverage (0.4 Crub), and pairwise relevance (0.1 Rrel), gives the best balanced trade-off in their experiments. It is the only variant that improves over the base model on both in-distribution grounding and out-of-distribution behavior across both model families. It improves in-distribution support while transferring better to out-of-distribution checklist tasks than either grounding-only or rubric-only rewards.

The authors conclude that per-sample rubrics are most useful as a complement to grounding, not as its replacement. They note that hard grounding rewards make long-form answers safer but thinner, while RUBRIC-ONLY rewards recover content coverage but lose grounding. They also observe that DeepSeek-R1-Distill-Llama-8B collapses harder than Qwen under hard grounding rewards, with out-of-distribution damage being more severe. The paper includes a length-collapse safeguard: a short-answer format penalty is added so the all-grounded condition cannot be satisfied by trivial truncation.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to an AI system:

  • What to change: Replace generic richness proxies (length, claim count, detail) with per-sample key-point rubrics that specify required and optional information for each question.

  • Implementation: Generate rubrics using a multi-generator union pipeline (e.g., GPT-5.4 + Gemini-3-Flash, merged by GPT-OSS-120B) for each training prompt.

  • Reward function: Use R = 0.5 × G(y) + 0.4 × C rub(y, Q x) + 0.1 × R rel, where G is grounding (all sentences supported), C rub is rubric coverage (required-first scoring), and R rel is pairwise relevance.

  • What to change: Avoid all-or-nothing grounding rewards that cause refusal-to-richness trade-offs (models learn to answer less to avoid unsupported claims).

  • Implementation: Use additive combination instead of multiplicative gating. This preserves coverage even when a single sentence is unsupported, preventing the model from collapsing into conservative, under-covered answers.

  • What to change: Prevent reward hacking where models truncate responses to avoid unsupported claims.

  • Implementation: Add a format penalty if the response is too short to support meaningful grounding judgment (e.g., missing closing `` tag or minimal post-reasoning answer).

  • What to change: Evaluate not just grounding accuracy but also rubric coverage and claim number separately.

  • Implementation: Report three metrics: grounding accuracy (all sentences supported), rubric coverage (required-first checklist score), and claim number (average factual sentences). This gives a complete picture of the refusal-to-richness trade-off.

  1. Answer long-form questions with both factual accuracy and completeness — it covers all required information points specified by the question-specific rubric while keeping every claim grounded in context or world knowledge.

  2. Transfer better to out-of-distribution tasks — it performs well on checklist-style tasks (e.g., CL-bench) and creative writing (e.g., Arena-Hard) without sacrificing in-distribution grounding, unlike grounding-only or rubric-only rewards.

  3. Avoid the refusal trap — it does not learn to answer less as a shortcut to avoid unsupported claims. Instead, it produces rich, detailed answers that satisfy explicit content requirements.

  4. Maintain grounding on in-distribution benchmarks — it keeps grounding accuracy above the base model (e.g., Qwen3-4B: 0.273 vs. 0.214; DeepSeek-R1-Distill-Llama-8B: 0.382 vs. 0.206) while improving rubric coverage and out-of-distribution performance.

  5. Provide balanced performance across model families — the same reward composition works consistently on both Qwen3-4B and DeepSeek-R1-Distill-Llama-8B without retuning, showing the improvement is reward-level, not model-specific.

  6. Prevent reward hacking — the format penalty and soft composition prevent the model from exploiting the reward function by truncating responses or gaming the grounding gate.

Abstract

Rewards that penalize unsupported claims can improve grounding in long-form generation, but they can also teach models to answer less. We study this refusal-to-richness trade-off in long-form hallucination RL. Instead of using global richness proxies such as length, claim count, detail, or pairwise relevance, we represent each question with a key-point rubric that specifies the required and optional information a useful answer should cover. These rubrics define coverage directly and are used both for evaluation and as reward signals. Across grounding-only, proxy-based, rubric-only, and combined rewards, we find a stable trade-off: strict grounding rewards improve support but suppress coverage, while unconstrained rubric rewards improve coverage but weaken grounding. A soft combination of grounding, rubric coverage, and relevance gives the best balance in our experiments, improving in-distribution support while transferring better to out-of-distribution checklist tasks than either grounding-only or rubric-only rewards.

Sources

Related papers