RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic Evaluation

arXiv:2607.24772 · cs.AI, cs.CL · Submitted 2026-08-22 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic Evaluation".

Jane: The paper was written by Bingxian Wu, Yu Zhang, Zonghao Guo, Tang Liu, Chen Qian et al. from Institute of Geographic Sciences and Natural Resources Research, Chinese Academy of Sciences and Tsinghua University and University of the Chinese Academy of Sciences and China University of Geosciences, Beijing and Aerospace Information Research Institute, Chinese Academy of Sciences and Shanghai Jiao Tong University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Summary: Tom: So, we just touched on the title and the authors; now that we know what they are aiming for, let’s look at the core summary of "RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic Evaluation."

Jane: The paper highlights that current agents, even those using powerful LLMs, lack domain expertise and their workflows are brittle and error-prone in complex remote sensing scenarios.

Lu: It's a huge limitation that they face—these generalist AI models don't naturally understand concepts like NDVI or BSI interactions, making them fragile when faced with real-world geospatial data.

Meng: The researchers are tackling this by providing two key components: Hierarchical Knowledge Grounding and Failure-Aware Experience Refinement, which is a clever way to structure the knowledge they are giving the agent.

Lalam: This means the AI isn's just guessing or following static templates; it's actively learning from its mistakes, building a more robust internal understanding of how certain Earth processes work.

Improvements: Tom: That summary tells us what the problem is and what they built; now let’s talk about the actual improvements that "RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic Evaluation" brings to the table.

Jane: The paper shows that by grounding the agent in a pre-distilled hierarchical domain corpus, we can guide its planning and tool selection much more effectively than before.

Lu: And it's not just about the initial data; they are iteratively integrating online experience, which means the system improves over time by distilling failure-annotated tool-use traces.

Meng: The engineering benefit here is that instead of relying on resource-intensive, expert-authored static workflow templates, we have a system that evolves and adapts to low marginal cost.

Lalam: This iterative learning capability ensures the AI isn't just a powerful calculator; it’s becoming an increasingly reliable scientific partner because its memory grows organically with every interaction.

Results: Tom: That's why the system is so smart; it learns from failures and uses structured knowledge. Let’s look at the hard data in "RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic Evaluation" and examine the results of the experiments.

Jane: The authors tested this against EarthBench, which is a challenging benchmark, and they showed that RSMeM consistently improves tool-use performance across various LLMs.

Lu: I was really struck by how efficient it is; we're talking about achieving a six percent accuracy improvement on DeepSeek-V3 point 2 while adding less than one percent more experience tokens.

Meng: That efficiency is impressive, but the results also show that stronger backbones benefit more from this system, which suggests the AI itself has to be capable of summarizing that domain knowledge effectively.

Lalam: The evidence shows that by combining structured guidance with instance-level execution experience, we're seeing a much higher level of success on tasks where other methods just fail.

Conclusion: Tom: Well, we’ve covered the mechanics and the data; so, what does this all mean for our listeners? Let’s wrap up our discussion of "RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic Evaluation."

Jane: This paper demonstrates a shift from static, expert-curated workflows to dynamic, self-improving AI systems that are fundamentally reliable in geospatial analysis.

Lu: It means the future of AI isn't just about processing data; it's about internalizing scientific knowledge and evolving its own execution capabilities.

Meng: Practically, this makes building complex Earth observation tools much more feasible because we’ve found a path toward low-cost, high-performance memory management.

Lalam: The ultimate impact is that this allows AI to bridge the gap between generalist models and specialized scientific requirements, enabling us to make better decisions about our planet based on sophisticated remote sensing data.

Tom: That's a powerful way to end the show; we hope you enjoyed hearing about "RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic Evaluation."

Lu: I can't wait to see how this approach is applied to other scientific fields.

Meng: I think this solves a real bottleneck in building robust AI tools today.

Lalam: It’s truly exciting work, allowing us all to look forward to the next paper on our list.

Bingxian Wu, Yu Zhang, Zonghao Guo, Tang Liu, Chen Qian, Yuxiang Lu, Xingbo Du, Yanghao Li, Yidan Zhang,5,23,54,623,5182, Chi Chen2104970896370466, Ling Yao3104970896370466, Maosong Sun21.55.525182

Institute of Geographic Sciences and Natural Resources Research, Chinese Academy of Sciences · Tsinghua University · University of the Chinese Academy of Sciences · China University of Geosciences, Beijing · Aerospace Information Research Institute, Chinese Academy of Sciences · Shanghai Jiao Tong University

cs.AI, cs.CL

Submitted: 2026-08-22

Updated: 2026-08-25

Code: https://github.com/AI9Stars/RSMeM

Importance score: 80/100

The gist: RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic Evaluation The paper introduces RSMeM, a "knowledge-enhanced memory evolution mechanism" designed to address the

Key concepts

Remote Sensing Agents
These are AI systems designed for complex geospatial analysis using remote sensing data. The paper focuses on improving these agents because generalist models often lack the specific domain expertise needed to interpret concepts like NDVI or BSI interactions.
Knowledge Grounding
This technique guides the agent's planning and tool selection by grounding it in a pre-distilled hierarchical domain corpus. This structured knowledge helps guide the AI more effectively than relying solely on general LLM capabilities.
Failure-Aware Experience Refinement
This is a core mechanism where the system improves over time by actively learning from its mistakes. It achieves this by distilling failure-annotated tool-use traces, making the agent's memory grow organically and building robustness.

Terminology

Summary

RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic Evaluation

The paper introduces RSMeM, a knowledge-enhanced memory evolution mechanism designed to address the limitations of existing Remote Sensing (RS) agents. These current agents are described as domain-agnostic, which results in brittle and error-prone workflows. To overcome this, RSMeM is an iterative framework that bootstraps RS agents with pre-distilled domain knowledge and iteratively integrates online experience for robust multi-step tool execution.

RSMeM is composed of two core components:

  1. Hierarchical Knowledge Grounding (HKG): This component performs taxonomy-aware retrieval over a hierarchical domain corpus to guide planning and tool selection. The knowledge base (KB) is meticulously constructed using a three level hierarchy consisting of 3 high-level Geoscience Domains, which are further decomposed into 12 Sub-domains and 64 Atomic Analytical Tasks. This structure allows the agent to navigate from broad application areas down to specific operational steps.

  2. Failure-Aware Experience Refinement (FAM): This component leverages a failure-critique environment to identify erroneous tool-use trajectories and distill them into reusable constraints that are stored as experience memory.

The methodology involves iteratively coupling these two processes: By iteratively employing these two processes, RS agents can evolve to absorb task-level domain knowledge and effectively translate it into instance-level execution experience. This iterative learning process is designed to move beyond the limitations of static expert templates, which are inflexible and costly to curate, by enabling the agent to internalize corrections from execution failures and improve its behavior over time.

Evaluation and Results:

Extensive experiments conducted on EarthBench demonstrate that RSMeM consistently improves tool-use performance and end-to-end answer across a diverse set of LLM backbones. The framework achieves significant gains in efficiency. Specifically, RSMeM achieves a 6.07% absolute accuracy improvement on DeepSeek-V3.2 with less than 1% additional experience tokens, demonstrating the strong knowledge density of our distilled experience.

Key Performance Metrics:

The performance is evaluated across several dimensions:

  • Tool Calling Accuracy: Measures success rates in Any-Order (AO), In-Order (IO), Exact-Match (EM, which measures the whether the entire model-generated tool-calling trajectory is an identical sequence to the reference steps), and Wise-Score (WS).

  • Outcome: End-to-end Accuracy (Acc).

  • Efficiency: Metrics include Trajectory Redundancy Index (TRI), Token Efficiency (Tok.), and Experience Density (ED) to quantify the cost-performance trade-off.

The results show that RSMeM provides a favorable effectiveness–efficiency trade-off, achieving consistent gains through superior experience density. The analysis of component ablation shows that while HKG optimizes global planning, FAR significantly enhances trajectory robustness.

Conclusion:

In summary, RSMeM is presented as a plug-and-play framework that can be applied to diverse LLM backbones and consistently improves both tool-use and end-to-end performance on challenging remote sensing tasks. The paper concludes that the integration of structured grounding and failure-aware refinement allows the agent to achieve consistent gains through superior experience density.

Improvements for AI systems

As a diligent researcher whose work demands absolute precision, I have analyzed the RSMeM (Knowledge-Enhanced Memory Evolution) framework. The core of RSMeM is not just about adding memory; it is about structured, purposeful memory that combines static domain knowledge with dynamic, failure-derived experience.

The following improvements outline specific architectural and functional enhancements for any AI system handling complex, multi-step Geoscience or Remote Sensing (RS) tasks.


We are moving beyond standard Retrieval-Augmented Generation (RAG) systems and replacing static workflow templates with a dynamic, dual-stream grounding mechanism. This requires three specific architectural upgrades:

The Improvement: Instead of relying solely on semantic similarity to retrieve relevant knowledge, the system must implement a structured, taxonomy-aware retrieval pipeline that grounds planning in a hierarchical corpus (e: 3 Domains to 12 Subdomains to 64 Atomic Tasks).

  • Mechanism: The retriever must utilize a hybrid scoring function that combines standard neural similarity ((v q, v i)) with specific lexical bonuses (I T, I K, I C) that fire when the query tokens overlap with task titles, curated domain keywords, or content.

  • Goal: This ensures the LLM is not just generally relevant, but is guided to a specific, expert-validated toolchain (S i) and domain definition (D i) appropriate for the task taxonomy.

The Improvement: The system must incorporate an iterative self-improvement loop that converts execution failures into compressed, reusable constraints stored in a Failure-Aware Memory (FAM). This avoids the brittle nature of static templates.

  • Mechanism:

  • Failure Detection: A lightweight symbolic critic (rho in R) must monitor the tool-use trajectory (a i) against defined failure modes (e.g., invalid tool calls, empty responses, parameter mismatch).

  • Compression (GeoTraceCompress): The verbose I/O of the failed trace is compressed using a type-aware strategy (e.g., retaining only the first three items in a list or two key-value pairs in a dictionary) while preserving the full reasoning text.

  • Distillation: A Reflection Generator (R) takes the query, HKG context, and compressed trace to produce a structured, 3-5 sentence diagnosis (r i). This r i is then appended to the FAM.

  • ** Goal:** This allows the the agent to learn from mistakes by storing specific guardrails (e.g., LST must be derived from emissivity) rather than re-executing the entire failed workflow.

The Improvement: The LLM's input prompt is not monolithic; it is dynamically constructed from two complementary streams:

  • Stream A (HKG Guidance): Provides high-level, static, taxonomy-based guidance (The What and How of the domain).

  • Stream B (FAM Retrieval): Provides dynamic, instance-specific constraints and historical fixes (The Why it failed before this specific).

By implementing these architectural changes, the AI system will transform from a brittle ReAct agent into a robust, expert-level problem solver.

The system will achieve significantly higher Tool Calling Accuracy (AO, IO, EM) because it is guided by two independent sources of truth:

  • HKG ensures: It selects the correct tools for the task based on established geoscientific taxonomy (e.g., choosing lst single channel instead of a generic temperature lookup).

  • FAR ensures: It avoids repeating past, specific errors (e.g., it will not use an invalid tool name or misinterpret BT10 as LST if the FAM contains a constraint against that specific semantic drift).

The system's performance gains are driven by Experience Density (ED), not by brute-force token consumption. It can:

  • Self-Correct with Minimal Overhead: By distilling failures into concise, actionable constraints (r i), the agent can correct its trajectory in subsequent rounds using minimal additional tokens (less than 1% additional experience tokens), making it highly cost-effective.

  • Avoid Redundancy: It will avoid unnecessary exploration or repetitive steps, concentrating computation on high-impact, corrected actions.

The system is capable of maintaining analytical rigor even when the static domain knowledge (HKG) is suboptimal:

  • Mismatched Guidance Handling: If HKG suggests a general Urban Trend Analysis (Stream A), but the user query demands LST retrieval, the system will leverage its specific historical memory (Stream B) to pivot and execute the correct LST pipeline, maintaining analytical consistency.

The architecture is designed to be plug-and-play across various LLM backbones (e.g., DeepSeek V3.2, Qwen3). It does not rely on proprietary or highly curated model strengths but provides a standardized layer of expertise that consistently elevates performance regardless of the underlying LLM' capacity for general reasoning.

Sources

Related papers