ICA: Information-Aware Credit Assignment for Visually Grounded Long-Horizon Information-Seeking Agents

arXiv:2602.10863 · cs.LG, cs.AI · Submitted 2026-02-11 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "ICA: Information-Aware Credit Assignment for Visually Grounded Long-Horizon Information-Seeking Agents".

Jane: The paper was written by author1 and author2 from University1 and Company2.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary of Findings: Tom: So, the researchers at "ICA: Information-Aware Credit Assignment for Visually Grounded Long-Horizon Information-Seeking Agents" did something clever to solve that noise problem, right? They moved away from text parsing and using visual snapshots instead.

Jane: Exactly. Instead of just reading the words on a webpage, they capture the whole visual scene—the layout, the tables, everything. This is huge because often critical evidence is presented in charts or specific positions that standard text extraction misses entirely.

Lu: It’s like moving from reading a transcribed voicemail to looking at the actual video call; you see all the context and nuance that was lost in transcription. The visual structure provides that grounding.

Meng: From an engineering standpoint, this snapshot approach gives us such a consistent data unit to work with across multiple runs, which is critical for building reliable training pipelines. We aren't fighting the inconsistencies of poor parsing heuristics anymore.

Lalam: I love how this helps because it allows the AI to see information as a grounded object rather than just a sequence of words, giving us more context for its decision-making process.

Improvements Suggested by the Paper: Tom: The way they measure success is also really innovative, using something called Information-Aware Credit Assignment or ICA. It’s not just looking at the final answer; it's looking at every single piece of information acquired.

Jane: That's the biggest technical leap, Tom. Instead of just getting one big reward for success or failure, they assign a 'utility score' to each retrieved piece of evidence based on its likelihood P b(R=one I e=one) minus P b(R=one I e=zero).

Lu: It’s basically calculating the marginal contribution of every atomic piece of data to the final outcome. If that piece was necessary for success, it gets a high credit score.

Meng: And then this utility signal is fed back into the previous steps, which is what makes it so powerful. We're getting dense feedback at the turn level, not just sparse feedback at the end of a long chain of reasoning.

Lalam: It’s moving from treating reasoning like a black box to making every single piece of acquired knowledge accountable for its importance in AI decision-making.

Paper Discussion Segment 1 — Title and Authors: Tom: We've talked about the mechanics, but let's start by discussing the title itself—"ICA: Information-Aware Credit Assignment for Visually Grounded Long-Horizon Information-Seeking Agents." What does that even mean?

Jane: It’s a promise of accuracy. "Information-Aware" suggests that AI won't just look at the final answer; it will scrutinize *how* it got there. And "Visually Grounded" tells us that the agents aren't just reading text, they are seeing the actual webpage structure.

Lu: It sounds like a paradigm shift in agentic learning, moving away from merely hoping for a successful outcome to actively tracking the evidence acquisition path itself.

Meng: From an engineering standpoint, it promises much greater stability in training sets that will be chaotic and noisy right now. The paper is clearly addressing the bottleneck of sparse rewards.

Lalam: I hope this leads to more sophisticated and trustworthy AI systems that can understand not just what a page says, but how the layout conveys its truth.

Conclusion: Tom: We've covered so much ground today regarding "ICA: Information-Aware Credit Assignment for Visually Grounded Long-Horizon Information-Seeking Agents." It’s clear this paper is changing the way we train agents to be reliable web researchers.

Jane: It’s a massive win for data efficiency and accuracy in long, complicated tasks.

Lu: I see huge potential here for creating truly autonomous AI that can solve problems requiring deep, iterative research.

Meng: I'm excited to see how this translates into real-world systems that need to perform reliable information extraction at scale.

Lalam: This allows us to build an AI culture where we trust the answers because we know exactly where the evidence came from.

Tom: And so, as "ICA: Information-Aware Credit Assignment for Visually Grounded Long-Horizon Information-Seeking Agents" provides a strong foundation for a much smarter, more grounded future of AI, we're going to wrap up our show.

Jane: Thank you all for sharing your insights today. Goodbye everyone!

University1 · Company2

cs.LG, cs.AI

Submitted: 2026-02-11

Updated: 2026-08-25

Code: https://github.com/pc-inno/ICA_MM_deepsearch

Project page: https://x.ai/news

Importance score: 7/100

The gist: The paper "ICA: Information-Aware Credit Assignment for Visually Grounded Long-Horizon Information-Seeking Agents" addresses the inherent challenges in open-web information seeking, particularly

Key concepts

Visually Grounded Agents
These agents do not rely solely on reading transcribed text. Instead, they capture the entire visual scene of a webpage—including its layout and tables. This provides crucial context that standard text extraction misses, allowing the AI to see information as a grounded object.
Information-Aware Credit Assignment (ICA)
This technique measures success by evaluating every single piece of data acquired during research. Instead of only rewarding the final answer, it assigns a 'utility score' to each retrieved evidence based on its likelihood, providing dense feedback at the turn level.
Long-Horizon Information-Seeking Agents
These agents are designed to handle complex tasks that require deep and iterative research. They are built to solve problems that involve a long chain of reasoning, requiring them to acquire and process many pieces of information over time.

Terminology

Summary

The paper ICA: Information-Aware Credit Assignment for Visually Grounded Long-Horizon Information-Seeking Agents addresses the inherent challenges in open-web information seeking, particularly concerning long-horizon trajectories under noisy information sources. The core problem is that agents often struggle when redundant, low-quality, or misleading observations accumulate and ultimately obscure the evidence most relevant to the final decision, leading to a persistent signal-to-noise bottleneck.

To address this limitation, the authors propose a dual innovation: a visual-native search framework and Information-Aware Credit Assignment (ICA).

1. Visual Grounding Framework

The system replaces traditional text parsing with a representation that preserves the page's semantic organization through its visual layout. The agents are trained to view webpages as visual snapshots. This approach allows the agents to leverage stable structural cues... and spatial grouping while effectively suppressing distractors. Crucially, these snapshots retain visually grounded information that is frequently distorted or discarded by text extraction pipelines, including figures, charts, and other non-textual elements, providing a more faithful and consistent interface to external evidence than lossy text-based representations.

2. Information-Aware Credit Assignment (ICA)

The central technical contribution is the ICA framework, a post-hoc method designed to overcome the limitations of sparse terminal rewards in long-horizon learning. Instead of attributing success solely based on final task outcomes, ICA assigns credit according to how individual decisions affect the acquisition of external information that supports subsequent reasoning.

  • Atomic Evidence Unit (e): An atomic evidence unit is defined as the minimal identifiable, self-contained fragment of external information acquired during a tool interaction.

  • Posterior Success Association: ICA estimates the empirical success probability conditioned on acquiring this unit (P n(B(R=1 I e=1))) versus the probability of success when that unit is not acquired (P n(B (R = 1 I e = 0)).

  • Atomic-Evidence Counterfactual Contribution (e): This contribution is defined as the difference between these two conditional success probabilities: e = P n(B(R=1 I e=1) - P n(B (R = 1 I e = 0)).

  • Turn-Level Credit Aggregation: The raw turn-level credit (t) is calculated by aggregating the e values for all atomic evidence units (E t) introduced at a single turn. To account for the difference in how information is acquired, temporal decay only to fetch turns is applied, scaling the aggregated credit by a factor.

  • Decoupled Advantage: This dense signal is then integrated into a GRPO-based training pipeline. The final learning signal combines the task-level advantage (A(n)) and the local information-aware advantage (t), weighted by a hyperparameter lambda: t(n) = A(n) + lambda t.

3. Experimental Results

The researchers evaluated their approach across several long-horizon benchmarks, including BrowseComp, GAIA, Xbench-DS, and SealQA.

  • Snapshot vs. RAG: The study found that adopting snapshot-based fetching in Supervised Fine-Tuning (SFT) consistently yields improvements over RAG-style textual fetches, with gains being particularly pronounced on benchmarks sensitive to layout and presentation (e.g., GAIA and XDS).

  • ICA vs. GRPO: When using the snapshot retrieval policy during Reinforcement Learning (RL), ICA consistently outperforms vanilla GRPO on all benchmarks.

  • Performance: The full-scale model, Qwen3-VL-30B-A3B-ICA, demonstrated strong performance across challenging tasks. For instance, it achieved 83.5 on Xbench-DS and 86.0 on SealQA, surpassing the best text baseline by significant margins (e.g., +18.0 points).

In summary, the paper concludes that grounding reinforcement learning in visually structured external observations and performing credit assignment at the granularity of acquired evidence effectively alleviates the credit-assignment bottleneck in open-ended web environments.

Improvements for AI systems

The following points detail specific architectural and methodological improvements to AI systems based on the principles of Information-Aware Credit Assignment (ICA) and visual grounding.


  • The Improvement: Replace traditional text-based parsing pipelines (e.g., using BeautifulSoup or Trafilatura on raw HTML) with a visual snapshot retrieval mechanism. The system must capture and store the webpage as a high-fidelity, rendered image/snapshot (C(l)).

  • Specific Technical Detail: Utilize a headless browser environment (e.g., Playwright) configured with an auto-scrolling mechanism to ensure all lazy-loaded content is visible. Implement adaptive slicing strategies (e.g., 4,480px height with 112px overlap) and downsampling (Lanzos interpolation at 0.7x) to maintain visual semantic continuity across the entire webpage, bounding the maximum rendering height to prevent unbounded computational overhead.

  • Benefit: Eliminates information loss and distortion caused by text extraction heuristics. Preserves critical layout cues, structural elements (tables, charts), and multimodal data (figures), providing a stable interface that is cross-trajectory consistent.

  • The Improvement: Establish a rigorous definition for the smallest identifiable unit of external information (e). This unit must be precisely mapped to the source of acquisition:

  • Web Search: A single, ranked search result item (rt).

  • Fetch URL: A specific, rendered webpage snapshot (C).

  • Specific Technical Detail: The system must maintain a persistent index linking these atomic units to the precise turn (t) in which they were acquired. This allows for granular tracking of utility across all subsequent reasoning steps.

  • The Improvement: Instead of scoring an action's immediate utility, the system calculates the marginal contribution of each atomic evidence unit (e) to the final task success rate (R).

  • Specific Technical Detail: For every completed trajectory tau(n), calculate two conditional success probabilities: P b(R=1 I e=1) (success given e was acquired) and P b(R=1 I e=0) (success given e was not acquired). The counterfactual contribution is defined as the difference: e = P b(R = 1 I e = 1) - P b(R = 1 I e = 0).

  • Benefit: This quantifies the causal impact of specific information acquisition, providing a dense, informative signal that is far more actionable than a single final reward.

  • The Improvement: Aggregate the e values associated with all atomic units (E t) introduced at a specific turn t, providing a unified, scalar credit for that decision point.

  • Specific Technical Detail: The raw turn-level credit (rt) is the average of all associated e. To account for temporal relevance, apply a decay factor (in (0, 1]) specifically to credits derived from FETCH actions (since they provide richer content that can be referenced over multiple turns). The final aggregated signal is t = rt times Tn-t-1.

  • Benefit: Provides a granular, time-aware learning signal for every step, allowing the the agent to learn which retrieval actions are critical early on versus those that occur later in a long trajectory.

  • The Improvement: Integrate ICA into a Group-in-Group Policy Optimization (GRPO) framework, using t as the primary input for policy updates.

  • Specific Technical Detail: The final policy update uses a mixed advantage (t) that combines the overall task-level advantage (A(n)) with the local, information-aware turn-level advantage (t), weighted by lambda. This allows the agent to learn both global strategy and specific evidence utility simultaneously.

  • Benefit: Enables efficient, decoupled learning where the model is trained not just to achieve a final answer, but to maximize the acquisition of high-utility information at every single turn.

The improved system achieves a level of reliability and data efficiency that current state-of-the-art text-based agents cannot match:

  1. Eliminate Information Loss: It guarantees the preservation of critical visual evidence (e.g., finding a specific value in a table, or identifying a date next to an image) regardless of how poorly the HTML is structured, by using snapshots as its primary observation.

  2. Self-Diagnose Failure Modes: Because it assigns credit based on counterfactual success rates (e), the the system can precisely identify which retrieval action led to a successful outcome and which action introduced noise or distractors, allowing it to learn from failure modes at the atomic evidence level.

  3. Optimize Long-Horizon Trajectories: It learns an explicit strategy for prioritizing high-utility information. Instead of blindly searching, the agent learns that specific visual cues (e.g, finding a Support Lifetime table) are more valuable than generic text snippets, leading to significantly faster and more focused search paths in complex tasks.

  4. Achieve Superior Performance: It consistently outperforms baseline systems on long-horizon benchmarks by providing a denser, more targeted learning signal to the RL objective function.

Sources

Related papers