Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards

summary

Video file (mp4)

The gist

The paper addresses a core challenge in training search-augmented language agents: "Search agents augment large language models (LLMs) with an external retrieval loop.

In short

The paper "Search-G1" introduces a method to make search agents more efficient and grounded. By using lightweight "readout" models to check if a search is necessary and if the answer relies on retrieved evidence, the system ensures agents use tools effectively rather than just relying on memory.

Key concepts

Grounded Search Agents
AI systems that look up information on the web or databases before answering questions. To be "grounded," an agent must actually use the retrieved evidence to reach its conclusion, rather than simply relying on its pre-existing memory or getting the right answer by luck.
Representation-Based Intrinsic Rewards
This method uses small "readout" models to probe an agent's internal hidden states rather than just looking at the final text. These probes estimate whether a search was necessary and if the answer is sensitive to the evidence found, providing a cheap way to guide training.
Periodic Refitting
Because an AI agent's behavior changes as it learns, the "detective" models used to grade it must also be updated. Periodic refitting involves freezing the agent, retraining the readout probes on its latest version, and then using those fresh probes for the next training stage.

Terminology used across episodes

This episode discusses

The paper

Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards · Read on arXiv

Fudan University · Tencent · Nanjing University · Nanyang Technological University · Shanghai Jiao Tong University

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards".

Jane: The paper was written by Ruoxi Cheng, Haoxuan Ma, Hongyi Zhang, Junming Zhang, Ranjie Duan et al. from Fudan University and Tencent and Nanjing University and Nanyang Technological University and Shanghai Jiao Tong University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title and Authors: Tom: Alright, welcome back to the show, everybody. Today we are cracking open a brand new paper that just hit the arXiv, and it's got a title that's a mouthful: "Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards." Jane, what do we even make of that title?

Jane: Well, Tom, let's unpack it. We're talking about search agents, which are basically eye systems that can go out and look things up on the web or a database before answering a question. And the big problem this paper tackles is making sure those agents only search when they actually need to, and that they really use what they find.

Tom: Right, so it's not just about getting the answer right, it's about *how* they get it. And the authors here are from Fudan University, Tencent, Nanjing University, NTU, and Shanghai Jiao Tong University. That's a serious lineup.

Jane: It is. And the lead author, Ruoxi Cheng, did this work while at Tencent's Rhino-Bird program. So this is very much an industry-academia collaboration, which usually means the research is grounded in real-world problems.

Lu: And that's exactly what I find exciting. The core idea is that we've been rewarding these agents for being correct, but we haven't been rewarding them for being *grounded*. You can get the right answer by luck, or by memorizing it, even after you've done a useless search.

Tom: So it's like a student who copies the answer from the back of the book without reading the chapter. They got the right answer, but they didn't learn anything.

Lu: Precisely. And this paper says, let's build a reward that actually checks whether the answer depends on the evidence the agent retrieved. That's the "grounded" part.

Meng: But from an engineering standpoint, that's really hard to do at scale. How do you check, for every single training example, whether the answer truly relies on the search results? That sounds like it would cost a fortune in compute.

Jane: And that's the clever bit, Meng. They don't check it directly during training. They train a small, lightweight "readout" model to predict that reliance, using a few expensive checks as labels. Once that readout is trained, it's cheap to use.

Tom: So they're basically teaching a little model to be a detective, and then letting that detective grade all the homework. That's a smart way to scale up.

Jane: Exactly. And that detective is what they call the "representation-based intrinsic reward." It's looking at the internal state of the eye, not just the final answer.

Lu: And the implications are huge. If we can train agents to be genuinely grounded, we can trust them more in high-stakes domains like medicine, law, or finance, where a confident but unsupported answer is worse than no answer at all.

Meng: I'm still skeptical about the overhead, though. Training that detective, how often does it need to be retrained? Because the agent is learning and changing, so the detective's job description changes too.

Tom: Oh, that's a great point, and it's actually the next thing we need to dig into. The paper has a clever answer for that, and it involves what they call "periodic refitting." But before we get there, let's just sit with the title for a second.

Jane: Yeah, "Search-G1." It sounds like a robot from a sci-fi movie, but it's really about making search agents more honest about where their knowledge comes from.

Tom: And that honesty is the whole ballgame. We'll get into the nitty-gritty of how they do it in the next segment, but for now, let's just say this paper is trying to fix a fundamental flaw in how we train eye to use tools.

Jane: So stick around, because we're about to break down the summary of "Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards" and see exactly how they pull this off.

Summary: Tom: Welcome back. So we've got the title, we've got the authors, and now we need to talk about what this paper actually claims to do. Jane, give us the elevator pitch.

Jane: The summary is pretty bold. It says that current rewards for search agents are either too sparse, like just checking if the final answer is right, or too expensive, like using a huge model to judge every step. This paper proposes a middle ground.

Lu: And that middle ground is the key contribution. They use two "readouts" — small models — to estimate two things. First, does the agent need to search at all? Second, is the final answer actually sensitive to the evidence it found?

Tom: So it's a two-part test. Part one: is the search necessary? Part two: is the answer actually using the search? And if both are true, you give a bonus.

Jane: Exactly. And the beauty is that these readouts are calibrated using counterfactual interventions. They literally delete the evidence and see if the answer changes. If it does, the answer is evidence-sensitive.

Meng: Okay, but I remember the title says "representation-based." What does that mean in practice? Are they reading the model's mind?

Lu: In a way, yes. They're not looking at the final text. They're looking at the hidden states — the internal vectors that the model computes as it processes the question and generates the answer. Those states encode a lot more than the final string.

Meng: So they're probing the model's brain, essentially. And they're using those probes to predict whether the answer would change if the evidence was gone.

Tom: And that's the "intrinsic" part, right? Because it's not coming from an external judge or a human label. It's coming from the model's own internal representation.

Jane: Right. But there's a catch. The model is learning and changing during training. So a probe that works on day one might be useless on day ten.

Lu: And that's where the "periodic refitting" comes in. They freeze the model, retrain the probes on the latest version, and then use those fresh probes for the next batch of training. It's a closed loop.

Meng: So the reward itself is evolving as the agent learns. That's actually really elegant. It's like the grading rubric is being updated to match the current curriculum.

Tom: And what's the payoff? The summary says they get better grounding and shorter trajectories. That means the agents are searching less, but when they do search, they're actually using the information.

Jane: And they show this across multiple benchmarks and two different model sizes. So it's not a fluke on one dataset. The improvements are consistent.

Lu: The implications here are significant. This isn't just about question answering. This is a general framework for teaching any agent to be more deliberate about when to use a tool.

Meng: I'm still wondering about the cost of that refitting, though. How often do they do it? And does it slow down training a lot?

Tom: Those are exactly the questions we're going to answer in the next segment, because we're going to dig into the actual methodology and the numbers behind "Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards."

Jane: And trust me, the numbers are pretty impressive. So stay tuned.

Improvements: Tom: Alright, we're back, and we're getting into the meat of the paper. Jane, what's the biggest improvement this paper is suggesting over the status quo?

Jane: The biggest one, Tom, is that it separates two things that were previously lumped together: knowing the answer and using the evidence. Old methods would give the same reward to an agent that searched, found the answer, and used it, as to an agent that searched, found the answer, but actually just recalled it from memory.

Lu: And that's the core problem. The reward was blind to *why* the agent got the answer right. This paper makes the reward conditional on the agent's own knowledge boundary.

Tom: The knowledge boundary. That's a nice way to put it. It's the line between what the agent already knows and what it needs to look up.

Lu: Exactly. And their first readout, the "prompt-state readout," estimates exactly where that boundary is. It asks: if we disabled search right now, could this agent answer correctly? If yes, searching is redundant.

Meng: So that's the necessity check. And the second readout is the reliance check. It asks: if we deleted the evidence, would the answer change? If yes, the answer is genuinely grounded in the search.

Tom: And they combine those two into a single score. Search is only rewarded if it's both necessary *and* the answer actually relies on it. That's a really clean formulation.

Jane: It is. And they also add a penalty for searching too many times. So the agent is pushed to be efficient, not just accurate.

Meng: But here's my engineering question again. How do they train these readouts? Because they need labels for "would the answer change if we deleted the evidence," and that seems expensive to generate.

Lu: They generate it with counterfactual rollouts. They take a frozen snapshot of the policy, run it with the evidence, and then run it again with the evidence deleted. If the answer changes, that's a positive label. They do this on a small calibration set, maybe a few hundred questions.

Meng: Okay, so it's expensive to generate the labels, but then they train a cheap model to predict those labels. That's a classic distillation approach.

Tom: And that's the "improvement" in a nutshell. They're taking an expensive, high-quality signal and making it cheap enough to use during training. It's like having a master chef taste-test every dish, but then teaching a junior chef to predict the master's verdict.

Jane: And the results speak for themselves. On Natural Questions, they improved exact match from thirty-five point two percent with the baseline Search-R1 to forty-five point three percent with their method, at the 3B scale. And they did it while reducing the number of search actions per question from three point two eight down to two point zero nine.

Meng: So they're getting better accuracy *and* fewer searches? That's a win-win.

Lu: It is. And that's the whole point. They're not just making the agent smarter; they're making it more efficient and more honest about its own limitations.

Tom: And that honesty is what we're going to see in the actual numbers and the first page of the paper in the next segment. Because the paper opens with a really interesting observation about how much the base model already knows.

Jane: Right, and that sets the stage for why this problem is so important. So let's move on to the first page of "Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards."

First Page: Tom: Welcome back. We've talked about the big ideas, and now we're going to look at the very first page of the paper, because it sets up the problem in a really compelling way. Jane, what did you see there?

Jane: The first thing they do is establish that this isn't a theoretical problem. They look at a base Qwen2 point 5-3B model on the Natural Questions benchmark, and it can answer fourteen percent of the questions correctly without any search at all.

Lu: That's the closed-book accuracy. It's low, but it's not zero. And that's the crux of the issue. If the model can already answer fourteen percent of the questions, then searching on those questions is probably a waste of time.

Meng: So the question becomes: can the agent learn to skip the search on that fourteen percent and focus its effort on the other eighty-six percent?

Tom: And that's exactly what the paper's reward is designed to do. It's not just about rewarding correct answers; it's about rewarding correct answers that *needed* the search.

Jane: And they contrast that with other benchmarks. HotpotQA has a lot of multi-hop questions, and MuSiQue is entirely multi-hop. But the paper makes a really sharp point: those structural labels tell you about the task, not about what a specific policy needs.

Lu: That's a crucial distinction. A question might be "multi-hop" in the dataset, but a particular model might already know the answer from training. So the necessity of search is policy-relative, not task-relative.

Meng: So they're saying you can't just look at the dataset and decide when to search. You have to look at the specific model you're training.

Tom: Exactly. And that's why they use the "closed-book sufficiency" readout. It's a live measurement of what the current policy can do without search.

Jane: And the first page also introduces the two families of rewards they're trying to improve upon. There are external rewards, like process reward models and LLM judges, which are accurate but expensive. And there are internal rewards, like entropy or confidence, which are cheap but don't really capture evidence use.

Lu: And their contribution is a third path. They're using internal representations to estimate external, counterfactual outcomes. It's the best of both worlds.

Meng: So the first page is essentially a manifesto. It's saying: we have a problem with how we reward search agents, and here's a new way to think about it.

Tom: And it's a way that's grounded in the model's own knowledge and behavior, not just the final answer. That's a big deal.

Jane: It is. And it sets up the entire methodology. They're going to measure retrieval necessity and evidence reliance, and they're going to do it with cheap, calibrated readouts.

Lu: And the implications for the field are clear. This could change how we train all sorts of tool-using agents, not just search agents.

Tom: Alright, so we've covered the title, the summary, the improvements, and the first page. Now it's time for our conclusion, where we wrap up our thoughts on "Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards."

Conclusion: Tom: Well, Jane, we've spent a good chunk of time with "Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards," and I think it's fair to say this is one of the more thoughtful papers we've seen in a while.

Jane: Absolutely. The core idea is so simple once you hear it: don't reward search just because the answer is right. Reward search when the answer *needs* the search. And then figure out a cheap way to measure that.

Lu: And they did it by looking at the model's internal states. That's the part that gets me excited. We're moving beyond surface-level evaluation and into understanding what the model is actually doing under the hood.

Meng: From a practical standpoint, the results are hard to argue with. Better accuracy, fewer searches, and a reward that doesn't require a giant judge model running during training. That's a win for anyone trying to deploy these systems.

Tom: And the periodic refitting is the glue that holds it all together. As the agent learns, the reward learns too. It's a co-evolution that keeps the signal relevant.

Jane: Exactly. And that's what makes this more than just a clever trick. It's a framework for adaptive measurement. As the model's knowledge boundary shifts, the reward shifts with it.

Lu: I think the impact here goes beyond search. Any agent that uses tools — whether it's a coding assistant, a web browser, or a robot — could benefit from this kind of grounding-aware reward.

Meng: And the fact that they showed it works on multiple benchmarks and two model sizes gives me confidence it's not just a fluke.

Tom: So, as we say goodbye to "Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards," I think we can all agree that this is a paper that asks the right question: not just "can the agent answer," but "how did it answer, and did it really need that search?"

Jane: And that's a question worth asking. Thanks for joining us, everyone. We'll be back with the next paper soon.

Tom: Until then, keep questioning the answers. See you next time.

More episodes

← Home