Language Models that Think, Chat Better
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Language Models that Think, Chat Better".
Jane: The paper was written by Adithya Bhaskar, Xi Ye and Danqi Chen from Princeton University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Alright, listeners, welcome back to the show. Today we’re digging into a paper that’s got a title that just makes you smile: “Language Models that Think, Chat Better.” Jane, I gotta say, that title is a promise, and I’m already excited.
Jane: It really is, Tom. And the promise is basically that if you let a language model stop and think before it answers, it’s going to have a much better conversation with you. That’s the whole idea in a nutshell.
Tom: And it comes from a team at Princeton, led by Adithya Bhaskar, Xi Ye, and Danqi Chen. They’re basically asking, why do we only make models think when they’re doing math or coding? Why not when they’re writing an email or planning a party?
Jane: Exactly. And the title is so clever because it flips the usual script. We’ve been told that thinking is for hard problems, and chatting is for everything else. This paper says, no, thinking helps with the chatting too.
Tom: I love that. And the way they frame it, it’s like they’re taking the best of two worlds. You have reinforcement learning with verifiable rewards, which is great for math, and you have reinforcement learning from human feedback, which is great for general chat. They’re mashing those together.
Jane: Right, and the key is that instead of using a rule-based checker to see if the answer is right, they use a reward model that’s been trained on human preferences. So the model gets to think for as long as it wants, and then the reward model judges how good the final answer is.
Tom: So it’s not about getting the one right answer, it’s about getting a *good* answer. And that’s a huge shift. It means the model can plan out its response, consider different angles, and then write something that actually feels thoughtful.
Jane: And the results are pretty wild. They’re showing that an eight-billion-parameter model, which is pretty small by today’s standards, can beat models that are ten times bigger on chat benchmarks. That’s a big deal, Tom.
Tom: It really is. It’s like saying a smart kid who takes a moment to think can out-argue a much older person who just blurts out the first thing that comes to mind. And that’s the core promise of this paper.
Jane: So stick around, because we’re going to get into how they actually did it, and why this could change how we train every language model from now on.
Tom: Next up, we’re breaking down the summary and the core method. You won’t want to miss it.
Summary: Tom: So, Jane, we’ve got the title, we’ve got the promise. Now let’s get into the meat of it. The paper’s method is called RLMT, which stands for Reinforcement Learning with Model-rewarded Thinking. And the summary is basically a recipe for making models think better.
Jane: And the recipe is surprisingly simple. You take a model, you ask it to generate a long chain of thought before it gives the final answer, and then you use reinforcement learning to reward it when that final answer is good. The reward comes from a model that’s been trained on human preferences.
Tom: Right, and the clever part is that they don’t just do this on one type of prompt. They use a diverse set of real-world prompts, like things people actually ask chatbots. So the model is learning to think about everyday problems, not just math equations.
Jane: And they tested this on a bunch of different models, both the base versions and the instruction-tuned versions, from Llama and Qwen. And across the board, the thinking models beat the non-thinking models on chat benchmarks by a pretty solid margin.
Tom: We’re talking like three to seven points on average, which is huge in this field. And the best model they trained, which is a Llama three point one 8B Instruct model, actually beat GPT-4o on chat and creative writing. That’s a frontier model that’s way bigger.
Jane: And it’s not just about being better at chatting. They also saw improvements in creative writing and general knowledge. So the thinking really does help across the board, not just in one specific area.
Tom: And here’s the kicker, Jane. They even tried this on base models that had never been through any instruction tuning. No SFT, no nothing. And just by applying this thinking and reward process, they got those base models to outperform the official instruction-tuned versions.
Jane: That’s the part that blew my mind. It’s like they found a way to skip the whole traditional post-training pipeline. The paper says they used only seven thousand prompts, while the official models were trained on millions of examples. And they still came out ahead on chat.
Tom: It really makes you wonder if we’ve been overcomplicating things. Maybe the secret isn’t more data, it’s letting the model think before it speaks. And that’s the big takeaway from the summary.
Jane: But of course, there’s a lot of detail in how they made this work. The choice of prompts, the choice of reward model, the training algorithm. That’s what we’re going to dig into next.
Tom: So stay tuned, because we’re about to get into the nitty-gritty of what makes this work so well.
Improvements: Tom: Alright, Jane, we’ve established that thinking helps. But the paper doesn’t just say “think more.” It actually digs into what kind of thinking improves, and what the training process changes in the model. So let’s get into the improvements.
Jane: And one of the most interesting parts is how the model’s thinking style changes after training. Before the reinforcement learning, the model would just make a checklist. Like, “first I’ll do this, then I’ll do that.” It was very linear.
Tom: Right, but after RLMT, the model starts doing something much richer. It starts listing out the constraints, grouping ideas into themes, and even revising its own plan as it goes. It’s like it’s actually thinking about the problem instead of just following a template.
Jane: And they showed this with a great example. When asked to write a Twitter thread, the model didn’t just outline the tweets. It thought about the tone, the constraints like “no hashtags,” and then refined its plan before writing. That’s a much more human-like approach.
Tom: And they also found that the model naturally starts thinking longer as training progresses. The thoughts get longer, the responses get longer. It’s like the model is learning that more thinking leads to better rewards.
Jane: But the improvements aren’t just about the model’s behavior. The paper also does a bunch of ablations to figure out what actually matters. And they found that the choice of prompts is critical.
Tom: Yeah, they tried using different prompt mixtures. One was from UltraFeedback, which has simpler prompts, and another was a random sample from the Tulu SFT mixture, which has a lot of math and jailbreak prompts. And neither worked as well as the WildChat-IF subset they chose.
Jane: So the quality and type of prompts you use for the reinforcement learning really matters. You need prompts that are chatty and challenging, not just easy or formal.
Tom: And the reward model matters too. They tried a weaker reward model, ArmoRM, and the performance dropped, especially on the non-chat benchmarks. But with a stronger reward model, Skywork, the thinking models kept their edge.
Jane: So it’s not just about having a reward model, it’s about having a *good* reward model. That’s a practical lesson for anyone trying to replicate this.
Tom: And finally, they compared different training algorithms. GRPO worked best, but even with DPO and PPO, the thinking models beat the non-thinking ones. So the improvement is robust across algorithms.
Jane: And that’s the key. This isn’t a fluke of one specific setup. The thinking advantage holds up across models, across algorithms, and across benchmarks.
Tom: So we’ve got a method that’s simple, robust, and delivers big gains. Now the question is, what does this mean for the future? We’ll wrap that up in our conclusion.
Conclusion: Tom: Alright, Jane, we’ve covered a lot of ground on “Language Models that Think, Chat Better.” Let’s bring it all together and say goodbye to this paper.
Jane: For sure. The core message is that letting a language model think before it answers, and then rewarding it for good thinking, makes it a much better conversationalist. And this works even without the traditional instruction-tuning pipeline.
Tom: And the impact is huge. We’re talking about an eight-billion-parameter model that can beat models ten times its size. That could democratize AI, making high-quality chat assistants possible on smaller hardware.
Jane: And it also challenges the way we think about post-training. Maybe we don’t need millions of examples and complex pipelines. Maybe we just need to let the model think and give it a good reward signal.
Tom: But there are still questions. The paper admits we don’t fully understand whether the model is learning new skills or just amplifying ones it already had. That’s a big open question for future research.
Jane: And there’s the practical side too. The reward model is key, and we need better reward models to push this further. But the potential is clear.
Tom: So, as we wrap up, I think the takeaway is that thinking is not just for math and code. It’s for everything. And that’s a beautiful idea.
Jane: It really is. And we’re excited to see where this goes. So let’s say goodbye to “Language Models that Think, Chat Better” and get ready for the next paper.
Tom: Thanks for listening, everyone. We’ll see you next time.
Adithya Bhaskar, Xi Ye, Danqi Chen
Princeton University
cs.CL
Submitted: 2026-08-16
Updated: 2026-08-18
Comments: Preprint; we release our code and models publicly at https://github.com/princeton-pli/RLMT
Code: https://github.com/princeton-pli/RLMT
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 80/100
The gist: Reinforcement learning with verifiable rewards (RLVR) improves language model reasoning by using rule-based rewards in verifiable domains such as mathematics and code.
Key concepts
- RLMT (Reinforcement Learning with Model-rewarded Thinking)
- This is the core method discussed. It involves asking a language model to generate a long chain of thought before giving its final answer. The model is then trained using reinforcement learning based on human preferences.
- Reward Model
- Instead of using strict rules to check if an answer is correct, researchers use a reward model trained on human preferences. This model judges the overall quality of the final answer after the thinking process.
- Instruction Tuning
- This is a traditional post-training method where models are trained on millions of examples to follow instructions. The paper suggests that a new 'thinking' process might bypass or reduce the need for this extensive training.
- Chat Benchmarks
- These are tests used to evaluate how well a language model can hold natural, general conversations. The paper found that models using the thinking process significantly outperformed non-thinking models on these benchmarks.
Terminology
Summary
Reinforcement learning with verifiable rewards (RLVR) improves language model reasoning by using rule-based rewards in verifiable domains such as mathematics and code. However, RLVR leads to limited generalization for open-ended tasks—such as writing outline essays or making meal plans—where humans reason routinely. This paper shows that the RLVR paradigm is effective beyond verifiable domains, and introduces RL with Model-rewarded Thinking (RLMT) for general-purpose chat capabilities. Using diverse real-world prompts, RLMT requires LMs to generate long CoT reasoning before response, and optimizes them with online RL against a preference-based reward model used in RLHF. Across 40 training runs on Llama-3.1-8B and Qwen-2.5-7B (both base and instruct) and multiple optimization algorithms (DPO, PPO, and GRPO), RLMT consistently outperforms standard RLHF pipelines. This includes substantial gains of 3–7 points on three chat benchmarks (AlpacaEval2, WildBench, and ArenaHardV2), along with 1–3 point improvements on other tasks like creative writing and general knowledge. Our best 8B model surpasses GPT-4o in chat and creative writing and rivals Claude-3.7-Sonnet (Thinking). RLMT can also be applied directly to base models without an SFT stage, akin to R1-Zero training (DeepSeek-AI, 2025). Remarkably, with only 7K prompts, Llama-3.1-8B base trained with our RLMT recipe outperforms Llama-3.1-8B-Instruct post-trained with a complex multi-staged pipeline with 25M+ examples. We close with qualitative and quantitative analyses of how trained models plan their responses. Our results rethink the post-training pipeline and call upon future work to understand and employ thinking more broadly. We release our code and models at https://github.com/princeton-pli/RLMT.
Improvements for AI systems
Based on the paper, here are specific improvements I can implement in AI systems:
-
What to change: Modify the AI's inference pipeline to generate a hidden chain-of-thought (CoT) reasoning trace (enclosed in `` tags) before producing the final answer, for all open-ended tasks (chat, creative writing, planning, general knowledge).
-
Specific implementation: During training, use RLMT (RL with Model-rewarded Thinking) instead of standard RLHF. This means: (a) sample long CoT traces before responses, (b) optimize with GRPO (group-relative policy optimization) using a preference-based reward model (e.g., Skywork-v2), (c) use a diverse prompt mixture (e.g., WildChat-IF) for RL.
-
What to change: In the post-training pipeline, replace the current RLHF objective (which optimizes final responses only) with RLMT objective:
max E[r(y,x)]where the model generatesz(thinking) theny(response), andris a reward model. -
Specific implementation: Use GRPO as the optimization algorithm (best performance in paper), with a group size of 8, KL coefficient 0.001, and a strong reward model. This yields 3–7 point gains on chat benchmarks (WildBench, AlpacaEval2, ArenaHardV2) and 1–3 points on creative writing and general knowledge.
-
What to change: For new models, skip the supervised fine-tuning (SFT) stage and directly apply RLMT to the base pretrained model, using a fixed prompt template that elicits thinking (e.g.,
A conversation between User and Assistant... plan a response within `` tags
). -
Specific implementation: Train with GRPO on 7K prompts. The paper shows this outperforms models trained with complex multi-stage pipelines (e.g., Llama-3.1-8B-Instruct) on chat benchmarks by 5+ points.
-
What to change: Ensure the reward model is high-quality (e.g., Skywork-v2) and the RL prompt mixture is diverse and realistic (e.g., WildChat-IF subset, not UltraFeedback or unfiltered Tulu).
-
Specific implementation: Ablations show that using a weaker reward model (ArmoRM) or simpler prompts (UltraFeedback) significantly degrades performance (e.g., chat scores drop by 10+ points). Use prompts that are
chatty
and challenging. -
What to change: Train the model to exhibit richer reasoning behaviors: constraint enumeration, theme grouping, iterative refinement, and backtracking (instead of linear checklists).
-
Specific implementation: The RLMT-trained model naturally develops these traits (as shown in Figure 4). To replicate, ensure the RL training encourages long CoT and uses a reward model that rewards thorough, well-structured responses.
-
Chat: Achieve 3–7 point higher scores on WildBench, AlpacaEval2, and ArenaHardV2 compared to standard RLHF models. For example, an 8B model can surpass GPT-4o on chat and creative writing, and rival Claude-3.7-Sonnet (Thinking).
-
Creative Writing: Generate more coherent, constraint-aware, and thematically grouped responses (e.g., essay outlines, tweet threads, story chapters) with 1–3 point improvements on CreativeWritingV3.
-
General Knowledge & Factuality: Improve by 1–3 points on PopQA and MMLU-Redux, due to better planning and verification during reasoning.
-
Instruction Following: Improve compliance on IFBench (e.g., from 24.5 to 30.3 for Llama-3.1-8B) by integrating constraints into the thinking phase.
-
Zero-Shot Base Models: Even without SFT, the system can outperform instruction-tuned models (e.g., Llama-3.1-8B-RLMT-Zero beats Llama-3.1-8B-Instruct on chat by 5.5 points).
-
Adaptability: Works across different model families (Llama, Qwen) and optimization algorithms (DPO, PPO, GRPO), with GRPO being the most effective.
Specific Example: Given a prompt like Write a Twitter thread about urgent vs. non-urgent emails,
the improved system will first internally: (1) list constraints (no hashtags, tone, length), (2) group ideas into themes, (3) plan the thread structure, (4) draft and refine—then output a polished, constraint-compliant response, outperforming a model that answers directly.
Sources
- The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning
- BLEUBERI: BLEU is a surprisingly effective reward for instruction following
- Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles
- KTO: Model Alignment as Prospect Theoretic Optimization
- Writing-Zero: Bridge the Gap Between Non-verifiable Tasks and Verifiable Rewards
- From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
- Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation
- Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
- ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
- RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style
- Show Your Work: Scratchpads for Intermediate Computation with Language Models
- BOLT: Bootstrap Long Chain-of-Thought in Language Models without Distillation
- LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- WildChat: 1M ChatGPT Interaction Logs in the Wild
- The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering