ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment

arXiv:2608.05102 · cs.AI · Submitted 2026-08-05 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Paper Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment".

Jane: The paper was written by Yijun Lu, Rui Ye, Jiajun Wang, Yuwen Du, Tian Jin et al. from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: The title of this paper is a mouthful, but it tells you everything. Backtrack from the answer, then assign credit to each step. That's the whole method in one phrase. And honestly, it should make every search-agent trainer sit up a little straighter.

Jane: I read it as a promise. Most training pipelines treat every step of a search as equally good or equally bad. These authors promise to grade every step against the evidence that actually matters. That's a much harder promise to keep than it sounds.

Lu: The group sits at Shanghai Jiao Tong University. Yijun Lu and Rui Ye share equal core contribution credit, with Songhua Liu and Siheng Chen as corresponding authors. Jiajun Wang, Yuwen Du, and Tian Jin round out the seven-person team.

Meng: Why should listeners care about credit assignment? Imagine a detective solving a case in two hundred moves. Some moves find the smoking gun, and some moves are pure wandering. If you train on the final verdict alone, you never learn which moves did the real work.

Tom: Exactly. A failed investigation can contain sharp thinking, and a solved case can contain foolish detours. Uniform rewards would praise the detours and ignore the sharp thinking. That's the disease this paper wants to cure.

Jane: The cure targets a genuine gap. Reward a useful step even inside a losing trajectory. Penalize a harmful step even inside a winner. That's a real departure from the standard recipe, where the final answer is the only judge.

Lalam: Step back, and the stakes are enormous. Search agents are becoming the default tool for deep research on the web. If we train them without fine-grained feedback, we're flying blind. This paper hands them a compass instead of just a map.

Tom: What makes me trust them is the openness. The code is on GitHub and the weights are on Hugging Face. The community can verify every claim instead of taking it on faith.

Jane: There's something fitting about the author list too. The credit assignment starts with the team itself — equal core contributions, two corresponding authors.

Lu: The real test, of course, is whether the method delivers. I want to see how the framework fits together.

Jane: And that's exactly what the summary covers next.

Summary: Jane: We've got the promise and the team. Now the architecture. The framework runs in two halves, and the first is clue recovery. You hand an LLM the question and the verified answer, and it searches the web backward, building an evidence chain that connects the answer to the query.

Tom: Those recovered clues become waypoints. For the paper's running example — the answer CeraVe — you get clues like ceramides as the clinically supported ingredient, L'Oréal as the acquiring company, and a founder who graduated in 1904. Each clue is a fact that a solid search should have uncovered along the way.

Lu: The clever part is that recovery isn't a one-shot guess. The recovery model actively runs web searches and visits pages using the same tool-call protocol as the forward agent. Clues that survive that verification become the anchors for scoring.

Meng: Second half is the scoring. Every step in a rollout gets checked against the clue set. Did this step find a clue? Did it wrongly throw one away? Each behavior moves a base score up or down, so the sparse outcome becomes a dense step-level reward.

Tom: Then the training schemes come in. ABC-SFT reweights each turn's loss by its score, so good turns pull harder on the model. ABC-GRPO feeds the step scores into reinforcement learning as rewards. Both schemes keep failed trajectories in the mix.

Jane: The scale is what surprises me. A four-billion-parameter Qwen model, trained on only 8.5 thousand examples. That's a tiny diet for such a long-horizon job.

Lalam: Small data, dense supervision. The bet is simple: instead of more examples, give each example more meaning. And the paper claims that bet pays off on the big benchmarks.

Tom: The summary also emphasizes one design choice. Nothing gets thrown away. Both successful and failed trajectories are retained, and every step is judged on its own merits.

Lu: That's the philosophical core, really. Sparse labels throw away information. These dense step scores extract every drop of signal from the same trajectories.

Meng: I also like that each step bundles three things — the reasoning, the tool call, and the tool response. The scoring looks at all three. That's much richer than grading a final string.

Jane: Which raises the obvious question. What did those judgments actually improve? Let's talk numbers.

Improvements: Tom: So the framework is clear — clues recovered, steps scored. Now the improvements show up in the numbers. Look at the reward distribution first. Around four percent of steps inside successful trajectories still score below neutral. Nearly ten percent of steps inside failed trajectories score above neutral. A trajectory-level reward gets all of those wrong.

Jane: So the improvement is honest supervision. A step that finds a correct clue in a losing run still gets positive credit. A step that discards a correct clue in a winning run still gets punished. The final answer no longer overrules everything.

Lu: The ablation study backs that up. Standard SFT scores 28.5 on BrowseComp, while ABC-SFT climbs to 30.8. Standard GRPO gets 33.5, and ABC-GRPO reaches 37.3. The same pattern holds on xbench-2510 and GAIA-text, where the method lifts the scores by several points each.

Meng: And with context management switched on, the full agent hits 55.3 on BrowseComp and 52.9 on the Chinese version. Those are the headline numbers from the abstract. They beat the other four-billion-parameter agents by a wide margin.

Tom: QUEST-4B, Dr. Venus, AgentCPM-Explore — the paper lists them all, and ABSeeker sits above them. It even stays competitive with thirty-billion-parameter systems like Tongyi DeepResearch and OpenSeeker. For a four-billion model, that's a serious flex.

Lalam: That's the deeper implication. Credit assignment quality can outweigh raw parameter count. A smaller model that knows which steps matter will outrun a bigger model trained blindly.

Jane: And the improvement isn't just final accuracy. The training dynamics show ABC-GRPO produces longer search trajectories. The agent explores more instead of shutting down early. That's a behavior change, not just a score change.

Tom: So the method makes the agent more curious and more careful at the same time. That combination is rare in this literature.

Meng: The context trick itself is worth a closer look. They raise the budget to 256K tokens and apply a discard-all strategy for up to five rounds. That alone takes BrowseComp from 37.3 to 55.3.

Lu: And the RL machinery is tuned carefully. A discount factor of 0.25 keeps future rewards decaying fast, and rewards get normalized within each rollout group. Immediate step quality dominates.

Jane: The improvements stack cleanly. Better SFT weights, better RL rewards, better exploration. I want to zoom into the scoring rubric itself next — the exact numbers attached to each behavior.

First Page: Tom: We've seen the framework and the results. Now the first page lays out the rubric in hard numbers. Every step starts at a base score of 1.0. Discovering or verifying a correct clue adds 0.8. Correctly ruling out a wrong candidate adds 0.4.

Jane: And the penalties mirror the rewards. Incorrectly dismissing a correct clue costs 0.8. Submitting the wrong final answer costs 1.0. Submitting the verified answer adds 1.0. Everything clips between zero and two.

Lu: The worked example on that page is worth a thousand words. Step 22 finds ceramides and links them to CeraVe, scoring 1.8. Step 35 verifies both L'Oréal and the founder's graduation year, clipping at 2.0. Step 56 abandons the accumulated evidence and bounces back to SkinCeuticals, scoring just 0.2.

Meng: And step 64 submits the wrong brand entirely. Score zero. The failed trajectory still gets credit where credit is due, and the mistakes still get flagged. That's the heart of the whole idea.

Tom: What strikes me is the design of the base score. Any reasonable exploration without an obvious error keeps the neutral 1.0. The method doesn't punish curiosity. It only punishes clear mistakes.

Jane: The first page also names the machinery. DeepSeek-V4-Flash handles both clue recovery and step scoring, while the small four-billion-parameter agent does the learning. A separate judge keeps the supervision honest.

Lu: And the training protocol is transparent. OpenSeeker supplies the trajectories, with 5.5 thousand correct and 3 thousand incorrect ones. Tool responses get masked from the loss, so the model only learns from its own generated tokens.

Meng: The scorer even has to name the specific clue or entity when applying a criterion. No vague grading. The explanation has to cite the evidence, and it returns structured JSON so the whole pipeline can consume it.

Tom: The abstract promises to convert sparse trajectory outcomes into dense step-level supervision. Seeing the rubric on page one, I finally believe the mechanics can work.

Lalam: The takeaway from that first page is the philosophy: hindsight is a training signal. Once the answer is known, every past step can be re-judged in its light. That's a powerful trick.

Jane: And that philosophy points somewhere bigger. Where can this idea travel beyond search? That's our closing question.

Conclusion: Tom: Time to wrap up. This paper takes the final answer and turns it into a flashlight that shines backward over the whole search. Every step gets re-judged in that light.

Jane: The two stages work together. Clue recovery builds the evidence chain from the verified answer. Step scoring checks every action against that chain. Then ABC-SFT and ABC-GRPO translate the scores into better behavior.

Lu: The evidence is compelling. A four-billion-parameter model beating its same-scale peers. It also matches systems several times larger. And everything is open — code, weights, training details.

Meng: The reward distribution analysis is the detail that sticks with me. Useful steps inside failed trajectories, harmful steps inside successful ones. This method sees both clearly, and it treats them differently.

Lalam: And the future work is honest. The authors want to scale to larger backbones. They also want to carry the idea beyond web search into any long-horizon task where an outcome can be backtracked into intermediate goals.

Tom: For us, this was a satisfying read. Clear problem, clean method, convincing experiments, and a refreshingly open release.

Jane: Let's also remember the score example. A step that rediscovered ceramides earned 1.8 in a trajectory that ultimately failed. That step was still valuable, and the method said so out loud.

Lu: And a successful trajectory still had steps that scored near zero. The method caught those too. That's granularity you rarely see in agent training.

Meng: The context management jump deserves one more mention. Eighteen points on BrowseComp just from managing the context window. Combined with the credit assignment, the whole package is hard to ignore.

Lalam: The lasting message is simple. Dense, principled supervision can substitute for raw scale. If you can backtrack the answer, you can teach the agent which steps genuinely mattered.

Tom: We'll leave it there. Thanks for listening, and we'll see you at the next one.

Jane: See you soon.

Yijun Lu, Rui Ye, Jiajun Wang, Yuwen Du, Tian Jin, Songhua Liu, Siheng Chen

Shanghai Jiao Tong University

cs.AI

Submitted: 2026-08-05

Code: https://github.com/PolarSeeker/ABSeekerhttps:

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 81/100

The gist: The paper, authored by Yijun Lu, Rui Ye, Jiajun Wang, Yuwen Du, Tian Jin, Songhua Liu, and Siheng Chen (Shanghai Jiao Tong University), proposes Answer-Backtracked Credit Assignment (ABC), a

Key concepts

Credit assignment
The process of determining which steps in a search trajectory contributed to the final outcome. Instead of treating all steps equally, ABSeeker scores each step based on whether it found or discarded correct clues, even in failed trajectories.
Clue recovery
A backward search process where an LLM starts from the verified answer and searches the web to build an evidence chain connecting the answer to the query. These clues become waypoints for scoring forward search steps.
ABC-SFT and ABC-GRPO
Two training schemes that use step-level scores. ABC-SFT reweights each turn's loss by its score, while ABC-GRPO feeds the scores as rewards into reinforcement learning. Both retain failed trajectories and judge each step on its own merits.
Step-level scoring rubric
A detailed scoring system where each step starts at 1.0, adds 0.8 for discovering a correct clue, 0.4 for ruling out a wrong candidate, and subtracts 0.8 for discarding a correct clue or 1.0 for a wrong final answer. Scores clip between 0 and 2.

Terminology

Summary

The paper, authored by Yijun Lu, Rui Ye, Jiajun Wang, Yuwen Du, Tian Jin, Songhua Liu, and Siheng Chen (Shanghai Jiao Tong University), proposes Answer-Backtracked Credit Assignment (ABC), a fine-grained credit assignment framework for training long-horizon search agents. The core problem addressed is that existing training methods typically treat all steps within a trajectory uniformly during both supervised fine-tuning (SFT) and reinforcement learning (RL), failing to distinguish useful actions from erroneous or redundant ones.

The authors build on the observation that "once the ground-truth answer is available, the task becomes naturally backtrackable. Starting from the answer, one can recover the key entities, facts, relations, and constraints that should have been discovered during the search process. The ABC framework converts sparse trajectory-level outcomes into dense step-level supervision that rewards useful actions (even in failed trajectories) while suppressing erroneous or redundant actions."

Using this framework, the authors train ABSeeker based on Qwen3.5-4B with only 8.5k examples. ABSeeker achieves 37.3% on BrowseComp and 39.1% on BrowseComp-ZH; with context management, these improve to 55.3% and 52.9%, respectively, significantly outperforming same-scale (4B) agents and even matching the performance of larger ones (∼30B).

A search agent interacts with a web environment over T turns to produce a trajectory:

τ = (s1, s2,..., sT, a),

where each step st contains the agent's reasoning, tool call, and tool response. During training, agents typically receive trajectory-level reward:


r ans(τ) = 1, if a = a*, 0 otherwise.

The paper identifies two fundamental failures of this sparse signal:

  1. an incorrect trajectory may contain several useful intermediate steps—such as correct evidence discovery, verification, or candidate filtering—yet the final reward of zero provides no positive signal for these actions.

  2. a correct trajectory may contain erroneous intermediate conclusions or steps that discard useful evidence, yet the final reward of one does not distinguish these flawed actions from genuinely informative ones.

ABC maps each training question (q, a*) to a set of clues C = c1, c2,..., cK, where each ck is a verifiable piece of intermediate evidence relevant to answering q, such as a specific entity, fact, attribute, or relationship connecting the query to the verified answer.

The recovery process is itself an active ReAct loop: "the recovery model conducts web searches and visits pages through the same tool-call protocol as the forward agent, tracing evidence from the answer back toward the query and anchoring each clue in actual web content. Clues that survive this verification serve as reliable, answer-backtracked reference points for subsequent step scoring."

The paper illustrates this with a concrete example: given a four-constraint query and the verified answer CeraVe, the recovery model produces six clues: Ceramides as the clinically supported ingredient, L'Oréal as the acquiring company, and Eugène Schueller as its founder who graduated in 1904, among others, forming a verified evidence chain.

Given the recovered clue set, each step in a trajectory is scored. Each step starts with a base score of 1.0, ensuring that reasonable exploration without an obvious error is not penalized. The scoring rubric (Table 1) is:

  • Discovers or verifies a correct clue: +0.8

  • Rules out an incorrect candidate: +0.4

  • Incorrectly dismisses a correct clue: −0.8

  • Submits the verified answer: +1.0

  • Submits an incorrect answer: −1.0

The score accumulates deltas and is clipped:


r t = clip(1.0 + Σ Δ j, 0, 2.0)

A key property: "A step that discovers a correct clue in a trajectory that ultimately fails still receives positive credit, whereas a step that incorrectly dismisses a correct clue in a trajectory that ultimately succeeds still receives a penalty."

In the worked example: Step 22 receives 1.8 for discovering Ceramides; Step 35 is clipped to 2.0 for verifying both L'Oréal acquisition and Schueller's 1904 graduation; Step 56 receives 0.2 because the agent abandons the accumulated evidence supporting CeraVe and returns to an incorrect candidate; and Step 64 receives 0.0 for submitting an incorrect final answer.

ABC-SFT: Standard SFT is reweighted so each turn's loss contribution depends on its reward:


L SFT(θ) = −Σ t w(r t) Σ j log p θ(x t,j x t,<j)

where the weight is a sigmoid function w(r t) = σ(α·(r t − β)). In practice, the paper uses w(r t) = 2σ(2(r t − 1)), so the neutral reward r t = 1.0 is mapped to weight 1.0. High-scoring steps contribute more to the gradient; low-scoring steps contribute little.

ABC-GRPO: Step-level rewards are used in GRPO. Rewards are normalized within each rollout group, and a discounted step-level advantage is computed:


A i,t = Σ k=t T i γ k−t R̂ i,k

with γ = 0.25. The resulting advantage A i,t is assigned to all policy-generated tokens at step t, while environment-provided tool responses are masked from optimization. The standard clipped GRPO objective is used with this step-specific advantage replacing the trajectory-level advantage.

  • Training data: OpenSeeker trajectories with both correct (5.5K) and incorrect (3.0K) final answers, totaling 8.5K trajectories; maximum 200 steps per trajectory.

  • Backbone: Qwen3.5-4B.

  • SFT: 3 epochs, global batch size 64, Adam optimizer, learning rate 5×10−5, cosine decay.

  • RL: 1,000 questions, 8 rollouts per question, batch size 16, up to 200 interaction turns, temperature 1.0, top-p 0.95, KL loss coefficient 0.001.

  • Recovery and scoring model: DeepSeek-V4-Flash.

  • Evaluation benchmarks: BrowseComp, BrowseComp-ZH, xbench-2505, xbench-2510, and GAIA-text; each evaluation run three times and averaged, with up to 200 tool calls allowed.

ABSeeker achieves the best performance among 4B search agents on every benchmark: 55.3% on BrowseComp, 52.9% on BrowseComp-ZH, 77.0% on xbench-2505, 46.0% on xbench-2510, and 81.6% on GAIA-text (with context management on the first two). "Despite its smaller model size, ABSeeker also remains competitive with substantially larger search agents. It outperforms all reported 30B agents on xbench-2505 and GAIA-text, while surpassing several 30B systems on both BrowseComp and BrowseComp-ZH. The authors note that although our method is trained exclusively on BrowseComp-style questions, it generalizes effectively to xbench and GAIA, indicating strong cross-benchmark generalization."

Analysis of the 8.5K SFT trajectories shows that "even successful trajectories contain approximately 4% low-quality steps with rewards below 1.0. More importantly, nearly 10% of the steps in failed trajectories receive rewards above 1.0, indicating that they still discover or verify useful clues despite ultimately producing an incorrect answer." This demonstrates the inadequacy of trajectory-level supervision and the value of step-level credit assignment.

ABC-GRPO achieves consistently stronger BrowseComp performance after training begins while producing longer search trajectories, showing that step-level credit assignment improves both search accuracy and exploratory behavior.

Following MiroThinker and LongSeeker, the authors apply a discard-all strategy with a maximum context of 256K tokens: ABSeeker improves from 37.3% to 55.3% on BrowseComp and from 39.1% to 52.9% on BrowseComp-ZH.

  • ABC-SFT vs. Standard SFT: ABC-SFT improves performance on BrowseComp, BrowseComp-ZH, xbench-2510, and GAIA-text, while remaining comparable on xbench-2505 (72.0 vs. 73.0).

  • ABC-GRPO vs. Standard GRPO: Building on the ABC-SFT initialization, ABC-GRPO consistently outperforms standard trajectory-level GRPO across all benchmarks (e.g., 37.3 vs. 33.5 on BrowseComp; 81.6 vs. 77.7 on GAIA-text).

These results demonstrate that fine-grained step-level credit assignment improves both SFT and RL by enabling the model to emphasize useful actions and suppress erroneous ones throughout training.

The paper's stated contributions are:

  1. We propose Answer-Backtracked Credit Assignment, a fine-grained credit assignment framework that rewards useful actions in failed trajectories while suppressing erroneous actions in successful ones.

  2. Based on ABC, we develop ABC-SFT, which reweights the loss of each turn according to its step reward, and ABC-GRPO, which incorporates step-level rewards into GRPO.

  3. We first train ABSeeker based on Qwen3.5-4B using ABC-SFT, and then further optimize it with ABC-GRPO, achieving 37.3% on BrowseComp and 39.1% on BrowseComp-ZH.

The paper concludes: "Instead of treating all steps within a trajectory uniformly, ABC recovers intermediate evidence clues from verified answers and uses them to evaluate each search step... These methods reward useful actions in failed trajectories while suppressing erroneous or redundant actions in successful ones. The authors state that answer-backtracked step-level supervision improves reward quality, training dynamics, and long-horizon exploration, highlighting the importance of explicit process supervision for scalable search-agent training."

For future work, the authors note: "Due to computational constraints, our experiments focus on a compact 4B model. A natural next step is to scale ABSeeker to larger backbone models and examine whether answer-backtracked credit assignment brings stronger gains under higher model capacity. Beyond web search, we also plan to extend this framework to other long-horizon agent tasks where final outcomes can be backtracked into intermediate evidence, subgoals, or decision points to provide fine-grained process supervision."

Improvements for AI systems

Here are the specific improvements to AI systems suggested by this paper, and what the resulting system can do:

  1. Dense, step-level credit assignment instead of sparse trajectory rewards.

The improved system scores every intermediate action (reasoning, tool call, evidence gathering) based on whether it discovers, verifies, rules out, or incorrectly dismisses answer-backtracked clues.

→ It can learn from failed trajectories: a step that finds correct evidence still receives positive training signal even if the final answer is wrong. It can also correct successful trajectories: a step that discards a correct clue is penalized despite the final answer being right.

  1. Automatic recovery of intermediate evidence clues from the verified answer.

A recovery model runs a reverse ReAct loop, searching the web from the answer back toward the query, producing verifiable clues (entities, facts, relations, constraints) anchored in actual web content.

→ The system can automatically generate high-quality process supervision labels without human annotation, making it feasible to train on large-scale search trajectories.

  1. Reward-weighted supervised fine-tuning (ABC-SFT).

Each token's loss is reweighted by a sigmoid of its step reward, where neutral steps get weight 1.0, high-value clue-discovery steps get higher weight, and erroneous/redundant steps get lower weight.

→ The base policy learns to emphasize useful reasoning and search behaviors while suppressing poor ones, improving downstream RL performance and final accuracy on benchmarks like BrowseComp and GAIA-text.

  1. Step-level GRPO with discounted advantages (ABC-GRPO).

Instead of using one trajectory-level reward, the system normalizes step rewards within a rollout group and computes discounted step-level advantages (γ=0.25), masking environment-provided tool responses from optimization.

→ RL training becomes more stable and effective, yielding longer, more exploratory search trajectories and consistently higher accuracy than trajectory-level GRPO (e.g., 37.3% vs. 33.5% on BrowseComp).

  1. Integration with long-context management.

Combined with a discard-all strategy capping context at 256K tokens, the improved system avoids context overflow and maintains critical evidence across long horizons.

→ It can sustain deep, long-horizon searches (up to 200 tool calls), improving BrowseComp from 37.3% to 55.3% and BrowseComp-ZH from 39.1% to 52.9%.

  1. Cross-benchmark generalization from a single training distribution.

Trained only on BrowseComp-style questions, the improved system transfers effectively to other benchmarks (xbench-2505, xbench-2510, GAIA-text), matching or exceeding much larger 30B-parameter agents.

→ It can handle diverse web-search and reasoning tasks without task-specific retraining.

  1. Scalable process supervision for any backtrackable long-horizon task.

The ABC framework generalizes beyond web search to any agent task whose final outcome can be backtracked into intermediate evidence, subgoals, or decision points.

→ The resulting system can be applied to tool-use agents, code-generation pipelines, robotic task planning, and multi-step reasoning systems, providing fine-grained step rewards without manual process annotations.

In short, the improved AI system is a long-horizon agent that learns reliably from mixed-quality trajectories by giving credit to partial progress, penalizing harmful steps, exploring longer and more effectively during RL, and generalizing across benchmarks and task types with only a compact backbone model.

Abstract

Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate evidence to reach a final answer. However, existing methods for training these agents typically treat all steps within a trajectory uniformly during both supervised fine-tuning (SFT) and reinforcement learning (RL), failing to distinguish useful actions from erroneous or redundant ones. In this paper, we propose Answer-Backtracked Credit Assignment (ABC), a fine-grained credit assignment framework for training long-horizon search agents by converting sparse trajectory-level outcomes into dense step-level supervision that rewards useful actions (even in failed trajectories) while suppressing erroneous or redundant actions. Specifically, given a potentially obscure query and its corresponding ground-truth answer, ABC first performs Answer-Backtracked Clue Recovery, which traces back from the answer to recover intermediate clues required to solve the question. It then applies Clue-Anchored Step Scoring to evaluate each search step against these clues, converting sparse binary outcome supervision into dense step-level rewards. Based on these rewards, we develop ABC-SFT, which reweights the loss of each turn, and ABC-GRPO, which uses the step-level scores as rewards in GRPO. Building on this framework, we train ABSeeker based on Qwen3.5-4B with only 8.5k examples. ABSeeker achieves 37.3% on BrowseComp and 39.1% on BrowseComp-ZH. With context management, the scores further improve to 55.3% and 52.9%, respectively, significantly outperforming same-scale (4B) agents and even matching the performance of larger ones (approximately 30B). These results demonstrate the effectiveness of answer-backtracked step-level credit assignment for training long-horizon search agents.

Sources

Related papers