Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards

arXiv:2608.07531 · cs.CL, cs.AI · Submitted 2026-07-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards".

Jane: The paper was written by Ruoxi Cheng, Haoxuan Ma, Hongyi Zhang, Junming Zhang, Ranjie Duan et al. from Fudan University and Tencent and Nanjing University and Nanyang Technological University and Shanghai Jiao Tong University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title and Authors: Tom: Alright, welcome back to the show, everybody. Today we are cracking open a brand new paper that just hit the arXiv, and it's got a title that's a mouthful: "Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards." Jane, what do we even make of that title?

Jane: Well, Tom, let's unpack it. We're talking about search agents, which are basically eye systems that can go out and look things up on the web or a database before answering a question. And the big problem this paper tackles is making sure those agents only search when they actually need to, and that they really use what they find.

Tom: Right, so it's not just about getting the answer right, it's about *how* they get it. And the authors here are from Fudan University, Tencent, Nanjing University, NTU, and Shanghai Jiao Tong University. That's a serious lineup.

Jane: It is. And the lead author, Ruoxi Cheng, did this work while at Tencent's Rhino-Bird program. So this is very much an industry-academia collaboration, which usually means the research is grounded in real-world problems.

Lu: And that's exactly what I find exciting. The core idea is that we've been rewarding these agents for being correct, but we haven't been rewarding them for being *grounded*. You can get the right answer by luck, or by memorizing it, even after you've done a useless search.

Tom: So it's like a student who copies the answer from the back of the book without reading the chapter. They got the right answer, but they didn't learn anything.

Lu: Precisely. And this paper says, let's build a reward that actually checks whether the answer depends on the evidence the agent retrieved. That's the "grounded" part.

Meng: But from an engineering standpoint, that's really hard to do at scale. How do you check, for every single training example, whether the answer truly relies on the search results? That sounds like it would cost a fortune in compute.

Jane: And that's the clever bit, Meng. They don't check it directly during training. They train a small, lightweight "readout" model to predict that reliance, using a few expensive checks as labels. Once that readout is trained, it's cheap to use.

Tom: So they're basically teaching a little model to be a detective, and then letting that detective grade all the homework. That's a smart way to scale up.

Jane: Exactly. And that detective is what they call the "representation-based intrinsic reward." It's looking at the internal state of the eye, not just the final answer.

Lu: And the implications are huge. If we can train agents to be genuinely grounded, we can trust them more in high-stakes domains like medicine, law, or finance, where a confident but unsupported answer is worse than no answer at all.

Meng: I'm still skeptical about the overhead, though. Training that detective, how often does it need to be retrained? Because the agent is learning and changing, so the detective's job description changes too.

Tom: Oh, that's a great point, and it's actually the next thing we need to dig into. The paper has a clever answer for that, and it involves what they call "periodic refitting." But before we get there, let's just sit with the title for a second.

Jane: Yeah, "Search-G1." It sounds like a robot from a sci-fi movie, but it's really about making search agents more honest about where their knowledge comes from.

Tom: And that honesty is the whole ballgame. We'll get into the nitty-gritty of how they do it in the next segment, but for now, let's just say this paper is trying to fix a fundamental flaw in how we train eye to use tools.

Jane: So stick around, because we're about to break down the summary of "Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards" and see exactly how they pull this off.

Summary: Tom: Welcome back. So we've got the title, we've got the authors, and now we need to talk about what this paper actually claims to do. Jane, give us the elevator pitch.

Jane: The summary is pretty bold. It says that current rewards for search agents are either too sparse, like just checking if the final answer is right, or too expensive, like using a huge model to judge every step. This paper proposes a middle ground.

Lu: And that middle ground is the key contribution. They use two "readouts" — small models — to estimate two things. First, does the agent need to search at all? Second, is the final answer actually sensitive to the evidence it found?

Tom: So it's a two-part test. Part one: is the search necessary? Part two: is the answer actually using the search? And if both are true, you give a bonus.

Jane: Exactly. And the beauty is that these readouts are calibrated using counterfactual interventions. They literally delete the evidence and see if the answer changes. If it does, the answer is evidence-sensitive.

Meng: Okay, but I remember the title says "representation-based." What does that mean in practice? Are they reading the model's mind?

Lu: In a way, yes. They're not looking at the final text. They're looking at the hidden states — the internal vectors that the model computes as it processes the question and generates the answer. Those states encode a lot more than the final string.

Meng: So they're probing the model's brain, essentially. And they're using those probes to predict whether the answer would change if the evidence was gone.

Tom: And that's the "intrinsic" part, right? Because it's not coming from an external judge or a human label. It's coming from the model's own internal representation.

Jane: Right. But there's a catch. The model is learning and changing during training. So a probe that works on day one might be useless on day ten.

Lu: And that's where the "periodic refitting" comes in. They freeze the model, retrain the probes on the latest version, and then use those fresh probes for the next batch of training. It's a closed loop.

Meng: So the reward itself is evolving as the agent learns. That's actually really elegant. It's like the grading rubric is being updated to match the current curriculum.

Tom: And what's the payoff? The summary says they get better grounding and shorter trajectories. That means the agents are searching less, but when they do search, they're actually using the information.

Jane: And they show this across multiple benchmarks and two different model sizes. So it's not a fluke on one dataset. The improvements are consistent.

Lu: The implications here are significant. This isn't just about question answering. This is a general framework for teaching any agent to be more deliberate about when to use a tool.

Meng: I'm still wondering about the cost of that refitting, though. How often do they do it? And does it slow down training a lot?

Tom: Those are exactly the questions we're going to answer in the next segment, because we're going to dig into the actual methodology and the numbers behind "Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards."

Jane: And trust me, the numbers are pretty impressive. So stay tuned.

Improvements: Tom: Alright, we're back, and we're getting into the meat of the paper. Jane, what's the biggest improvement this paper is suggesting over the status quo?

Jane: The biggest one, Tom, is that it separates two things that were previously lumped together: knowing the answer and using the evidence. Old methods would give the same reward to an agent that searched, found the answer, and used it, as to an agent that searched, found the answer, but actually just recalled it from memory.

Lu: And that's the core problem. The reward was blind to *why* the agent got the answer right. This paper makes the reward conditional on the agent's own knowledge boundary.

Tom: The knowledge boundary. That's a nice way to put it. It's the line between what the agent already knows and what it needs to look up.

Lu: Exactly. And their first readout, the "prompt-state readout," estimates exactly where that boundary is. It asks: if we disabled search right now, could this agent answer correctly? If yes, searching is redundant.

Meng: So that's the necessity check. And the second readout is the reliance check. It asks: if we deleted the evidence, would the answer change? If yes, the answer is genuinely grounded in the search.

Tom: And they combine those two into a single score. Search is only rewarded if it's both necessary *and* the answer actually relies on it. That's a really clean formulation.

Jane: It is. And they also add a penalty for searching too many times. So the agent is pushed to be efficient, not just accurate.

Meng: But here's my engineering question again. How do they train these readouts? Because they need labels for "would the answer change if we deleted the evidence," and that seems expensive to generate.

Lu: They generate it with counterfactual rollouts. They take a frozen snapshot of the policy, run it with the evidence, and then run it again with the evidence deleted. If the answer changes, that's a positive label. They do this on a small calibration set, maybe a few hundred questions.

Meng: Okay, so it's expensive to generate the labels, but then they train a cheap model to predict those labels. That's a classic distillation approach.

Tom: And that's the "improvement" in a nutshell. They're taking an expensive, high-quality signal and making it cheap enough to use during training. It's like having a master chef taste-test every dish, but then teaching a junior chef to predict the master's verdict.

Jane: And the results speak for themselves. On Natural Questions, they improved exact match from thirty-five point two percent with the baseline Search-R1 to forty-five point three percent with their method, at the 3B scale. And they did it while reducing the number of search actions per question from three point two eight down to two point zero nine.

Meng: So they're getting better accuracy *and* fewer searches? That's a win-win.

Lu: It is. And that's the whole point. They're not just making the agent smarter; they're making it more efficient and more honest about its own limitations.

Tom: And that honesty is what we're going to see in the actual numbers and the first page of the paper in the next segment. Because the paper opens with a really interesting observation about how much the base model already knows.

Jane: Right, and that sets the stage for why this problem is so important. So let's move on to the first page of "Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards."

First Page: Tom: Welcome back. We've talked about the big ideas, and now we're going to look at the very first page of the paper, because it sets up the problem in a really compelling way. Jane, what did you see there?

Jane: The first thing they do is establish that this isn't a theoretical problem. They look at a base Qwen2 point 5-3B model on the Natural Questions benchmark, and it can answer fourteen percent of the questions correctly without any search at all.

Lu: That's the closed-book accuracy. It's low, but it's not zero. And that's the crux of the issue. If the model can already answer fourteen percent of the questions, then searching on those questions is probably a waste of time.

Meng: So the question becomes: can the agent learn to skip the search on that fourteen percent and focus its effort on the other eighty-six percent?

Tom: And that's exactly what the paper's reward is designed to do. It's not just about rewarding correct answers; it's about rewarding correct answers that *needed* the search.

Jane: And they contrast that with other benchmarks. HotpotQA has a lot of multi-hop questions, and MuSiQue is entirely multi-hop. But the paper makes a really sharp point: those structural labels tell you about the task, not about what a specific policy needs.

Lu: That's a crucial distinction. A question might be "multi-hop" in the dataset, but a particular model might already know the answer from training. So the necessity of search is policy-relative, not task-relative.

Meng: So they're saying you can't just look at the dataset and decide when to search. You have to look at the specific model you're training.

Tom: Exactly. And that's why they use the "closed-book sufficiency" readout. It's a live measurement of what the current policy can do without search.

Jane: And the first page also introduces the two families of rewards they're trying to improve upon. There are external rewards, like process reward models and LLM judges, which are accurate but expensive. And there are internal rewards, like entropy or confidence, which are cheap but don't really capture evidence use.

Lu: And their contribution is a third path. They're using internal representations to estimate external, counterfactual outcomes. It's the best of both worlds.

Meng: So the first page is essentially a manifesto. It's saying: we have a problem with how we reward search agents, and here's a new way to think about it.

Tom: And it's a way that's grounded in the model's own knowledge and behavior, not just the final answer. That's a big deal.

Jane: It is. And it sets up the entire methodology. They're going to measure retrieval necessity and evidence reliance, and they're going to do it with cheap, calibrated readouts.

Lu: And the implications for the field are clear. This could change how we train all sorts of tool-using agents, not just search agents.

Tom: Alright, so we've covered the title, the summary, the improvements, and the first page. Now it's time for our conclusion, where we wrap up our thoughts on "Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards."

Conclusion: Tom: Well, Jane, we've spent a good chunk of time with "Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards," and I think it's fair to say this is one of the more thoughtful papers we've seen in a while.

Jane: Absolutely. The core idea is so simple once you hear it: don't reward search just because the answer is right. Reward search when the answer *needs* the search. And then figure out a cheap way to measure that.

Lu: And they did it by looking at the model's internal states. That's the part that gets me excited. We're moving beyond surface-level evaluation and into understanding what the model is actually doing under the hood.

Meng: From a practical standpoint, the results are hard to argue with. Better accuracy, fewer searches, and a reward that doesn't require a giant judge model running during training. That's a win for anyone trying to deploy these systems.

Tom: And the periodic refitting is the glue that holds it all together. As the agent learns, the reward learns too. It's a co-evolution that keeps the signal relevant.

Jane: Exactly. And that's what makes this more than just a clever trick. It's a framework for adaptive measurement. As the model's knowledge boundary shifts, the reward shifts with it.

Lu: I think the impact here goes beyond search. Any agent that uses tools — whether it's a coding assistant, a web browser, or a robot — could benefit from this kind of grounding-aware reward.

Meng: And the fact that they showed it works on multiple benchmarks and two model sizes gives me confidence it's not just a fluke.

Tom: So, as we say goodbye to "Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards," I think we can all agree that this is a paper that asks the right question: not just "can the agent answer," but "how did it answer, and did it really need that search?"

Jane: And that's a question worth asking. Thanks for joining us, everyone. We'll be back with the next paper soon.

Tom: Until then, keep questioning the answers. See you next time.

Fudan University · Tencent · Nanjing University · Nanyang Technological University · Shanghai Jiao Tong University

cs.CL, cs.AI

Submitted: 2026-07-24

Updated: 2026-09-04

Code: https://github.com/Rosy0912/Search-G1

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 69/100

The gist: The paper addresses a core challenge in training search-augmented language agents: "Search agents augment large language models (LLMs) with an external retrieval loop.

Key concepts

Grounded Search Agents
AI systems that look up information on the web or databases before answering questions. To be "grounded," an agent must actually use the retrieved evidence to reach its conclusion, rather than simply relying on its pre-existing memory or getting the right answer by luck.
Representation-Based Intrinsic Rewards
This method uses small "readout" models to probe an agent's internal hidden states rather than just looking at the final text. These probes estimate whether a search was necessary and if the answer is sensitive to the evidence found, providing a cheap way to guide training.
Periodic Refitting
Because an AI agent's behavior changes as it learns, the "detective" models used to grade it must also be updated. Periodic refitting involves freezing the agent, retraining the readout probes on its latest version, and then using those fresh probes for the next training stage.

Terminology

Summary

The paper addresses a core challenge in training search-augmented language agents: Search agents augment large language models (LLMs) with an external retrieval loop. They issue queries, inspect documents, reason over evidence, and then commit an answer. The authors argue that while Search-R1-style systems show that reinforcement learning can induce this behavior, the reward design determines which behavior gets reinforced. They state: "A capable agent should perform selective grounded search: retrieve when needed, and ground the answer in what was retrieved. This requires two distinct judgments. Retrieval necessity asks whether the current policy can answer without retrieval. Evidence reliance asks whether a searched answer actually depends on the retrieved context."

The paper identifies a key gap: "Existing external rewards provide either sparse outcome supervision or richer feedback from process annotations and LLM judges. Outcome rewards scale readily but cannot distinguish grounded retrieval from redundant search, whereas richer signals require costly annotation or inference during training. Internal rewards based on policy-side signals such as entropy, likelihood, or information gain are graded and inexpensive to evaluate, yet mainly reflect model confidence rather than evidence grounding."

The authors note that Outcome-only or necessity-agnostic rewards do not explicitly distinguish redundant search from retrieval that supplies missing evidence. They further observe: The missing capability is not another confidence signal. It is a low-cost reward that jointly measures policy-relative retrieval necessity and answer-level evidence reliance.

The paper proposes Search-G1, a representation-based intrinsic reward framework that measures the operational grounding of an agent's answers through two intervention-calibrated readouts. The authors describe the approach as follows: "A prompt-state readout predicts closed-book sufficiency, whose complement defines policy-relative retrieval necessity; an answer-commit readout estimates evidence reliance from answer-stage sensitivity to evidence deletion. Together, they provide additional credit to correct searched trajectories when retrieval is estimated necessary and the answer is evidence-sensitive, favor correct direct answers when closed-book knowledge suffices, and penalize repeated search."

The framework treats training as adaptive measurement: "After calibration, reward scoring requires neither process annotations nor LLM-as-judge inference during policy optimization. Because reinforcement learning changes policy representations, Search-G1 periodically refits both readouts on trajectories from the latest checkpoint, allowing the reward to co-evolve with the policy."

The paper describes the calibration process: "At calibration round m, we freeze the latest policy checkpoint θ̄m. The snapshot generates counterfactual completions and encodes the hidden states used by two readouts: the prompt state hp,i at the end of the initial prompt and the answer-commit state hans,i,j at the final answer token."

Evidence Reliance: For a valid searched trajectory, the intervention removes retrieved observations while retaining the question, search history, and generated reasoning. The answer-change target is defined as: schg i,j = 1[ν(â i,j) ≠ ν(A θ̄m(T(C i,j)))], and the readout is d i,j = D φ(m)(h ans,i,j) ∈ [0,1]. This estimates answer-stage sensitivity to evidence deletion, giving an operational reliance signal for reward shaping.

Retrieval Necessity: From a fixed retrieval-disabled prompt P cb i, the paper obtains a closed-book target and predicts it from the actual search-enabled prompt state: z i = V cb(A det θ̄m(P cb i), y* i) ∈ 0,1, b i = B ψ(m)(h p,i) ∈ [0,1], n i = 1 − b i. Here b i estimates closed-book sufficiency under the specified policy and prompt protocol, and n i measures relative retrieval necessity.

The calibration loss is: L cal = E D(m) D [BCE(d i,j, s chg i,j)] + E D(m) B [BCE(b i, z i)].

For a searched trajectory, "Search-G1 combines evidence reliance and retrieval necessity into a necessity-gated reliance score g i,j = d i,j n i. This multiplicative gate is large only when the answer is evidence-sensitive and closed-book sufficiency is low."

The full reward is piecewise:

  • r inv for invalid outputs

  • r wrong for valid but wrong outputs

  • 1 + λ g g i,j − η c i,j for valid, correct, searched trajectories

  • 1 + α cb b i for valid, correct, direct trajectories

The search cost is c i,j = 1[N i,j > 1], meaning direct and single-search trajectories have zero cost, whereas repeated search has unit cost.

The paper proves that "if r inv < r wrong < 1 − η, η ∈ [0, 1), and λ g, α cb ∈ [0, 1], the shaping preserves correctness-first ordering: every correct valid trajectory outranks every wrong valid trajectory, which in turn outranks every invalid output."

The paper uses GRPO with group-relative advantages: Â i,j = (r i,j − µ i)/(σ i + δ). The optimization only model-generated reasoning, search actions, and final-answer tokens, excluding prompt tokens, retrieved observations, and padding. The objective includes a clipped surrogate loss, an entropy bonus, and a low-variance sampled-token KL penalty.

The paper emphasizes: "Policy updates change both search behavior and the hidden-state geometry used by the readouts. At each calibration round, Search-G1 therefore freezes the latest checkpoint as θ̄m, regenerates both intervention targets, and refits the two heads. This closes the measurement–credit–adaptation loop."

Models: The primary evaluation uses Qwen2.5-3B-Instruct and Qwen2.5-7B-Instruct.

Benchmarks: Task utility is measured by exact match (EM) on NQ, HotpotQA, 2WikiMultiHopQA, and MuSiQue, covering single-hop factual retrieval and multi-hop evidence aggregation.

Grounding metrics: "Grounding is quantified by two counterfactual metrics: trust consistency (TC↑), which summarizes counterfactual responses over correct direct and searched trajectories, and document invariance (D-Inv↓), which isolates originally correct searched answers that remain unchanged after document replacement."

Search cost: Search cost is reported by Search/Q↓, the mean number of response-side parser-observed search markers per question.

Baselines: The paper compares against two non-trained references: Base, the closed-book model without retrieval, and Prompted Search, which uses the same retrieval interface without policy updates. Training baselines include Search-R1, GiGPO, AgentPRM, GPT-5 Judge, ARPO, IGPO, Search-E1, SAAS, β-GRPO, and KbPO.

Implementation details: All trained methods share the policy-training data, base model, retrieval interface, trajectory format, and global interaction/update cap. Results are means over four independently seeded training runs.

On Qwen2.5-3B-Instruct, Search-G1 achieves EM of 0.453 on NQ (vs. 0.250 for Prompted Search, +0.203), 0.306 on HotpotQA (vs. 0.199, +0.107), 0.368 on 2WikiMultiHopQA (vs. 0.238, +0.130), and 0.143 on MuSiQue (vs. 0.057, +0.086). On Qwen2.5-7B-Instruct, it achieves 0.512 on NQ (vs. 0.349, +0.163), 0.450 on HotpotQA (vs. 0.299, +0.151), 0.475 on 2WikiMultiHopQA (vs. 0.235, +0.240), and 0.254 on MuSiQue (vs. 0.058, +0.196).

Search-G1 attains the highest TC at both scales (+0.018 over the strongest alternative) and the lowest D-Inv (−0.004/−0.005). Specifically, TC is 0.976 (3B) and 0.981 (7B), while D-Inv is 0.034 (3B) and 0.028 (7B).

Search-G1 issues 0.121/0.311 fewer search markers than Prompted Search (Search/Q of 2.094 at 3B and 1.860 at 7B). The paper notes: the D-Inv and search-use advantages over the strongest per-metric baseline are significant at 7B but not at 3B.

"Against Search-R1, all four endpoints are significant at both scales. Against the strongest per-metric baseline, the TC advantage survives at both scales; the D-Inv and Search/Q advantages are significant at 7B but not at 3B, and EM is not statistically separable at either scale."

Search-G1 again attains the highest TC, the lowest D-Inv, and fewer search markers than the Search-R1 baseline at both scales on 2WikiMultiHopQA, confirming that the grounding gains are not dataset-specific.

"The representation estimator reaches AUC 0.959/0.971 on NQ and 0.871/0.914 on 2Wiki for 3B/7B (four-seed SD ≤ 0.014), well above the strongest logit readout (0.638–0.731); every setting keeps a > 0.15 AUC gap across seeds."

In all-correct groups, Search-G1 restores nonzero reward variation in 46.6–48.7% of NQ groups and 91.8–92.4% of 2Wiki groups, with within-group reward SDs of 0.189–0.305.

At checkpoint 150, refitted evidence-reliance and retrieval-necessity readouts retain AUC 0.786–0.891, versus 0.482–0.608 when frozen (gaps 0.283–0.304; four-seed SD ≤ 0.021 frozen and ≤ 0.015 refit, well below the gap).

"Relative to Search-R1, Search-G1 cuts policy-generated tokens per question by ≈ 43% (149 → 85 at 3B) and total response-side tokens by ≈ 16%, while at the same time attaining substantially higher EM (0.352 → 0.453) and emitting fewer response-side search markers (3.281 → 2.094) under a matched interaction budget."

The component ablations on 2WikiMultiHopQA show: "removing the readouts, reliance shaping, necessity assignment, or repeated-search cost worsens at least one pre-specified outcome without improving the others enough to dominate the full method; in particular, removing reliance shaping lowers Search/Q only because the policy searches less overall, at the cost of a comparably large drop in TC, rather than reflecting a genuine efficiency gain."

The necessity-factorization ablation shows: "on Qwen2.5-3B the shuffled control (0.341/0.869/3.204) lands close to w/o both necessity terms (0.339/0.861/3.286) and well below the full method (0.368/0.972/2.812), indicating that the gain comes from the correct within-question pairing of n i and d i,j rather than merely from the presence of a necessity term."

Evidence-reliance label audit: On Qwen2.5-3B-Instruct/NQ at checkpoint 50, the original labels agree with these alternatives on 93%/96% of trajectories (κ = 0.85/0.90).

Shortcut-control audit: The representation readout achieves AUC 0.959 on NQ and 0.871 on 2Wiki, versus 0.501/0.498 for label-shuffled controls and 0.638/0.693 for the strongest logit baseline.

Human grounding evaluation: Search-G1 attains the highest unconditional grounded accuracy and, together with GPT-5 Judge, the lowest unsupported rate. On matched questions, Search-G1 shows significantly higher evidence support with ∆S@C of +0.048 [+0.026, +0.070] over GPT-5 Judge.

Cross-dataset transfer: The point difference is +0.101 in-domain and only +0.008/+0.006/+0.001 on HotpotQA/2Wiki/MuSiQue, showing a significant in-domain gain and cross-domain performance that is competitive with outcome-plus-directness RL.

The paper concludes: Search-G1 improves grounding and shortens trajectories while keeping accuracy competitive, requiring no external judge calls during optimization at the cost of periodic calibration compute. The authors frame their contribution as treating search-agent training as adaptive measurement: as the policy's knowledge boundary shifts, so should the reward used to assess evidence use.

The paper acknowledges: "Our evaluation focuses on text-retrieval QA with two Qwen backbones, a fixed retrieval environment, and parser-observed search markers; extending the study to other model families, retrievers, richer tools, long-horizon web interaction, and non-textual evidence is left to future work. Additionally, Counterfactual grounding metrics and evidence-deletion estimator audits are computed under a single fixed retriever and intervention design, so their absolute levels can shift with retriever quality even though the relative comparisons are held on identical inputs. The paper also notes: Periodic readout refitting adds calibration compute whose end-to-end overhead depends on hardware and batching."

Improvements for AI systems

Based on the Search-G1 paper, here are the specific improvements I can implement in an AI system, along with what the improved system can do:


  • What to add: Two lightweight neural readouts (a prompt-state head and an answer-commit head) trained on counterfactual labels (closed-book sufficiency and evidence-deletion sensitivity) from the policy's own hidden states.

  • How: During training, periodically freeze the policy, generate counterfactual trajectories (one without retrieval, one with retrieved evidence removed), and fit the heads via binary cross-entropy on the hidden states at the prompt end and the final answer token.

  • Resulting capability: The system can estimate, at inference time, (a) whether it can answer a query without retrieval and (b) whether its answer actually depends on the retrieved evidence—without any external judge, process annotations, or additional LLM calls.

  • What to add: A multiplicative gate g = d × (1 − b) where d is the evidence-reliance score and b is the closed-book-sufficiency score. This gate is applied only to correct searched trajectories, plus a bounded cost for repeated search and a sufficiency bonus for correct direct answers.

  • How: Replace the outcome-only reward in GRPO with a piecewise reward: 1 + λg·g − η·cost for correct searched, 1 + αcb·b for correct direct, and fixed low values for wrong/invalid. Use λg=0.6, η=0.02, αcb=0.5.

  • Resulting capability: The system learns to search only when retrieval is genuinely necessary (low closed-book sufficiency) and to ground answers in evidence (high reliance), while avoiding redundant or repeated searches. It can distinguish correct because of search from correct despite search.

  • What to add: A schedule that refits both readout heads every K=25 policy updates on a fresh 300-question calibration set, using the latest checkpoint's hidden states and counterfactual labels.

  • How: After each refit, atomically swap in the new evaluator; if one head fails, keep the previous complete evaluator. Use a shared calibration loss and validate architecture/layer selection on a held-out group.

  • Resulting capability: The reward signal co-evolves with the policy's changing knowledge boundary. The system avoids the frozen estimator failure mode where the reward becomes stale as representations drift, maintaining high estimator fidelity (AUC 0.79–0.89 vs. 0.48–0.61 when frozen).

  • What to add: A theoretical guarantee (Proposition 1) that under GRPO normalization, the necessity gate preserves ranking and advantage signs in homogeneous all-correct groups, and restores nonzero reward variation.

  • How: Use the derived formula Aδ(γ) = z(γ) / (z(γ)/√(G−1) + δ) to compute advantages, ensuring the gate cancels exactly when δ=0 and only rescales when δ>0.

  • Resulting capability: In groups where every trajectory is correct (a common failure mode for outcome-only rewards), the system still produces meaningful gradient signal—46.6–48.7% of NQ groups and 91.8–92.4% of 2Wiki groups now have nonzero within-group reward SD (0.189–0.305), enabling the policy to prefer the most grounded trajectory.

  • What to add: A masked GRPO objective that optimizes only policy-generated tokens (reasoning, search actions, final answer), excluding prompt, retrieved observation, and padding tokens.

  • How: Use a token mask M i,j,t and a clipped importance ratio with a low-variance sampled-token KL penalty (Equation 13). Compute the loss as a distributed masked average across micro-batches.

  • Resulting capability: The system avoids wasting gradient signal on retrieved text and prompt tokens, focusing optimization on the agent's own decisions. This yields more stable training and better token efficiency (e.g., 43% fewer policy-generated tokens per question vs. Search-R1).

  • What to add: Two post-hoc metrics: Trust Consistency (TC) and Document Invariance (D-Inv), computed by replacing the gold answer in retrieved documents with an alternative entity and regenerating the answer.

  • How: Classify each correct searched trajectory as TRUE (answer changes to track the counterfactual), HALL (answer stays the same), or AMB. TC is the graded mean (1.0/0.6/0.3); D-Inv is the fraction of HALL among all classified.

  • Resulting capability: The system can be evaluated on whether it genuinely uses retrieved evidence, not just whether it produces correct answers. This provides a robust, intervention-based grounding signal that is harder to game than lexical overlap or confidence.

  1. Selective Grounded Search: Given a question, the system decides whether to search based on its own estimated closed-book sufficiency, and how much to search based on evidence reliance. It retrieves only when necessary, grounds answers in retrieved evidence, and avoids redundant searches.

  2. Self-Calibrating Reward: The system's reward function adapts as the policy learns. Early in training, it may search frequently; later, it learns to answer directly when its parametric knowledge suffices, without needing external supervision or judge calls.

  3. Tie-Breaking in Sparse-Reward Settings: In groups where all trajectories are correct, the system still differentiates between grounded and ungrounded correct answers, enabling fine-grained credit assignment that outcome-only methods cannot provide.

  4. Cost-Efficient Trajectories: The system produces shorter response-side trajectories (2.09 vs. 3.28 search markers per question at 3B; 1.86 vs. 3.09 at 7B) and fewer policy-generated tokens (85 vs. 149 at 3B) while maintaining or improving task accuracy (EM 0.453 vs. 0.352 at 3B on NQ).

  5. Robust to Representation Drift: The periodic refitting ensures the reward remains valid as the policy's hidden-state geometry changes, preventing the common failure of static reward models in RL.

  6. Transferable Grounding Behavior: The system's grounding improvements generalize across datasets (NQ, HotpotQA, 2WikiMultiHopQA, MuSiQue) and model scales (3B, 7B), with higher TC (0.976 vs. 0.958 best baseline) and lower D-Inv (0.034 vs. 0.042) on NQ.

  7. No External Judge at Optimization Time: The system requires no LLM-as-judge calls, process reward models, or human annotations during policy updates—only periodic calibration compute (≈11–12% of rollout decodes per refit cycle).

  8. Add two readout heads (MLP with hidden sizes 64–256) on top of frozen hidden states at layer 21/24/27 (selected per model/refit).

  9. Generate counterfactual labels at each refit: (a) closed-book decode from a retrieval-disabled prompt, (b) evidence-deleted decode from a trajectory with retrieved observations removed.

  10. Fit heads via BCE on a 300-question calibration set, using class weighting and early stopping.

  11. Compute rewards at rollout time: d = D(head, answer-commit state), b = B(head, prompt state), g = d × (1 − b).

  12. Apply the piecewise reward with correctness-first ordering, then compute group-relative advantages under GRPO with δ=1e-6.

  13. Optimize only masked policy-generated tokens with clipped importance ratios and the sampled-token KL penalty.

  14. Refit every 25 updates on the latest checkpoint, atomically swapping evaluators.

This implementation yields a search agent that is measurably more grounded, more cost-efficient, and more accurate than outcome-only or confidence-based baselines, while remaining fully self-contained (no external judges) and adaptive to its own evolving knowledge.

Sources

Related papers