Language Models that Think, Chat Better
summary
The gist
Reinforcement learning with verifiable rewards (RLVR) improves language model reasoning by using rule-based rewards in verifiable domains such as mathematics and code.
In short
The episode discusses the paper "Language Models that Think, Chat Better," which proposes that allowing a language model to think before answering improves conversational ability. Using a method called RLMT, researchers showed that thinking helps with general chat, even allowing smaller models to outperform much larger ones.
Key concepts
- RLMT (Reinforcement Learning with Model-rewarded Thinking)
- This is the core method discussed. It involves asking a language model to generate a long chain of thought before giving its final answer. The model is then trained using reinforcement learning based on human preferences.
- Reward Model
- Instead of using strict rules to check if an answer is correct, researchers use a reward model trained on human preferences. This model judges the overall quality of the final answer after the thinking process.
- Instruction Tuning
- This is a traditional post-training method where models are trained on millions of examples to follow instructions. The paper suggests that a new 'thinking' process might bypass or reduce the need for this extensive training.
- Chat Benchmarks
- These are tests used to evaluate how well a language model can hold natural, general conversations. The paper found that models using the thinking process significantly outperformed non-thinking models on these benchmarks.
Terminology used across episodes
This episode discusses
- Language Models that Think, Chat Better · Paper Radio
- The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning
- BLEUBERI: BLEU is a surprisingly effective reward for instruction following
- Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles
- KTO: Model Alignment as Prospect Theoretic Optimization · Paper Radio
- Writing-Zero: Bridge the Gap Between Non-verifiable Tasks and Verifiable Rewards
- From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
- Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation
- Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
- ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models
- RM-Bench: Benchmarking Reward Models of Language Models with Subtlety and Style
- Show Your Work: Scratchpads for Intermediate Computation with Language Models
- BOLT: Bootstrap Long Chain-of-Thought in Language Models without Distillation
- LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learning
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- WildChat: 1M ChatGPT Interaction Logs in the Wild
- The Surprising Effectiveness of Negative Reinforcement in LLM Reasoning
The paper
Language Models that Think, Chat Better · Read on arXiv
Adithya Bhaskar, Xi Ye, Danqi Chen
Princeton University
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Language Models that Think, Chat Better".
Jane: The paper was written by Adithya Bhaskar, Xi Ye and Danqi Chen from Princeton University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Alright, listeners, welcome back to the show. Today we’re digging into a paper that’s got a title that just makes you smile: “Language Models that Think, Chat Better.” Jane, I gotta say, that title is a promise, and I’m already excited.
Jane: It really is, Tom. And the promise is basically that if you let a language model stop and think before it answers, it’s going to have a much better conversation with you. That’s the whole idea in a nutshell.
Tom: And it comes from a team at Princeton, led by Adithya Bhaskar, Xi Ye, and Danqi Chen. They’re basically asking, why do we only make models think when they’re doing math or coding? Why not when they’re writing an email or planning a party?
Jane: Exactly. And the title is so clever because it flips the usual script. We’ve been told that thinking is for hard problems, and chatting is for everything else. This paper says, no, thinking helps with the chatting too.
Tom: I love that. And the way they frame it, it’s like they’re taking the best of two worlds. You have reinforcement learning with verifiable rewards, which is great for math, and you have reinforcement learning from human feedback, which is great for general chat. They’re mashing those together.
Jane: Right, and the key is that instead of using a rule-based checker to see if the answer is right, they use a reward model that’s been trained on human preferences. So the model gets to think for as long as it wants, and then the reward model judges how good the final answer is.
Tom: So it’s not about getting the one right answer, it’s about getting a *good* answer. And that’s a huge shift. It means the model can plan out its response, consider different angles, and then write something that actually feels thoughtful.
Jane: And the results are pretty wild. They’re showing that an eight-billion-parameter model, which is pretty small by today’s standards, can beat models that are ten times bigger on chat benchmarks. That’s a big deal, Tom.
Tom: It really is. It’s like saying a smart kid who takes a moment to think can out-argue a much older person who just blurts out the first thing that comes to mind. And that’s the core promise of this paper.
Jane: So stick around, because we’re going to get into how they actually did it, and why this could change how we train every language model from now on.
Tom: Next up, we’re breaking down the summary and the core method. You won’t want to miss it.
Summary: Tom: So, Jane, we’ve got the title, we’ve got the promise. Now let’s get into the meat of it. The paper’s method is called RLMT, which stands for Reinforcement Learning with Model-rewarded Thinking. And the summary is basically a recipe for making models think better.
Jane: And the recipe is surprisingly simple. You take a model, you ask it to generate a long chain of thought before it gives the final answer, and then you use reinforcement learning to reward it when that final answer is good. The reward comes from a model that’s been trained on human preferences.
Tom: Right, and the clever part is that they don’t just do this on one type of prompt. They use a diverse set of real-world prompts, like things people actually ask chatbots. So the model is learning to think about everyday problems, not just math equations.
Jane: And they tested this on a bunch of different models, both the base versions and the instruction-tuned versions, from Llama and Qwen. And across the board, the thinking models beat the non-thinking models on chat benchmarks by a pretty solid margin.
Tom: We’re talking like three to seven points on average, which is huge in this field. And the best model they trained, which is a Llama three point one 8B Instruct model, actually beat GPT-4o on chat and creative writing. That’s a frontier model that’s way bigger.
Jane: And it’s not just about being better at chatting. They also saw improvements in creative writing and general knowledge. So the thinking really does help across the board, not just in one specific area.
Tom: And here’s the kicker, Jane. They even tried this on base models that had never been through any instruction tuning. No SFT, no nothing. And just by applying this thinking and reward process, they got those base models to outperform the official instruction-tuned versions.
Jane: That’s the part that blew my mind. It’s like they found a way to skip the whole traditional post-training pipeline. The paper says they used only seven thousand prompts, while the official models were trained on millions of examples. And they still came out ahead on chat.
Tom: It really makes you wonder if we’ve been overcomplicating things. Maybe the secret isn’t more data, it’s letting the model think before it speaks. And that’s the big takeaway from the summary.
Jane: But of course, there’s a lot of detail in how they made this work. The choice of prompts, the choice of reward model, the training algorithm. That’s what we’re going to dig into next.
Tom: So stay tuned, because we’re about to get into the nitty-gritty of what makes this work so well.
Improvements: Tom: Alright, Jane, we’ve established that thinking helps. But the paper doesn’t just say “think more.” It actually digs into what kind of thinking improves, and what the training process changes in the model. So let’s get into the improvements.
Jane: And one of the most interesting parts is how the model’s thinking style changes after training. Before the reinforcement learning, the model would just make a checklist. Like, “first I’ll do this, then I’ll do that.” It was very linear.
Tom: Right, but after RLMT, the model starts doing something much richer. It starts listing out the constraints, grouping ideas into themes, and even revising its own plan as it goes. It’s like it’s actually thinking about the problem instead of just following a template.
Jane: And they showed this with a great example. When asked to write a Twitter thread, the model didn’t just outline the tweets. It thought about the tone, the constraints like “no hashtags,” and then refined its plan before writing. That’s a much more human-like approach.
Tom: And they also found that the model naturally starts thinking longer as training progresses. The thoughts get longer, the responses get longer. It’s like the model is learning that more thinking leads to better rewards.
Jane: But the improvements aren’t just about the model’s behavior. The paper also does a bunch of ablations to figure out what actually matters. And they found that the choice of prompts is critical.
Tom: Yeah, they tried using different prompt mixtures. One was from UltraFeedback, which has simpler prompts, and another was a random sample from the Tulu SFT mixture, which has a lot of math and jailbreak prompts. And neither worked as well as the WildChat-IF subset they chose.
Jane: So the quality and type of prompts you use for the reinforcement learning really matters. You need prompts that are chatty and challenging, not just easy or formal.
Tom: And the reward model matters too. They tried a weaker reward model, ArmoRM, and the performance dropped, especially on the non-chat benchmarks. But with a stronger reward model, Skywork, the thinking models kept their edge.
Jane: So it’s not just about having a reward model, it’s about having a *good* reward model. That’s a practical lesson for anyone trying to replicate this.
Tom: And finally, they compared different training algorithms. GRPO worked best, but even with DPO and PPO, the thinking models beat the non-thinking ones. So the improvement is robust across algorithms.
Jane: And that’s the key. This isn’t a fluke of one specific setup. The thinking advantage holds up across models, across algorithms, and across benchmarks.
Tom: So we’ve got a method that’s simple, robust, and delivers big gains. Now the question is, what does this mean for the future? We’ll wrap that up in our conclusion.
Conclusion: Tom: Alright, Jane, we’ve covered a lot of ground on “Language Models that Think, Chat Better.” Let’s bring it all together and say goodbye to this paper.
Jane: For sure. The core message is that letting a language model think before it answers, and then rewarding it for good thinking, makes it a much better conversationalist. And this works even without the traditional instruction-tuning pipeline.
Tom: And the impact is huge. We’re talking about an eight-billion-parameter model that can beat models ten times its size. That could democratize AI, making high-quality chat assistants possible on smaller hardware.
Jane: And it also challenges the way we think about post-training. Maybe we don’t need millions of examples and complex pipelines. Maybe we just need to let the model think and give it a good reward signal.
Tom: But there are still questions. The paper admits we don’t fully understand whether the model is learning new skills or just amplifying ones it already had. That’s a big open question for future research.
Jane: And there’s the practical side too. The reward model is key, and we need better reward models to push this further. But the potential is clear.
Tom: So, as we wrap up, I think the takeaway is that thinking is not just for math and code. It’s for everything. And that’s a beautiful idea.
Jane: It really is. And we’re excited to see where this goes. So let’s say goodbye to “Language Models that Think, Chat Better” and get ready for the next paper.
Tom: Thanks for listening, everyone. We’ll see you next time.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language