GeoForge: Non-Parametric Self-Evolving Agents for Earth-Observation Reasoning

arXiv:2608.10494 · cs.AI, cs.MA · Submitted 2026-08-11 · Read on arXiv

Xin Xiao, Jiang Zhong, Junnan Zhu, Yingchao Feng, Peijin Wang, Yidan Zhang, Kaiwen Wei

School of Computer Science, Chongqing University · MAIS, Institute of Automation, Chinese Academy of Sciences · Aerospace Information Research Institute, Chinese Academy of Sciences

cs.AI, cs.MA

Submitted: 2026-08-11

Updated: 2026-08-12

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: GeoForge is a training-free, self-evolving framework for Earth observation (EO) agents that transforms completed trajectories into a structured nonparametric execution state.

Terminology

Summary

GeoForge is a training-free, self-evolving framework for Earth observation (EO) agents that transforms completed trajectories into a structured nonparametric execution state. The paper states: GeoForge, a training-free, self-evolving framework that transforms completed trajectories into a structured nonparametric execution state. It addresses the challenge that EO workflows are constrained by sensing semantics, product dependencies, spatial and temporal compatibility, and parameter requirements, and that existing agents often search a broad operation space for each query, while recent self-evolving systems do not fully organize heterogeneous EO trajectories into reusable knowledge across different decision levels.

The framework constrains the operation space according to the sensing context, then retrieves a task-conditioned prior from three complementary memories. These memories are: Workflow Graph Memory captures global operation order, Action-Level Experiences provide local corrections, and the Adapted Skill Standard Operating Procedure preserves procedural and data constraints. The retrieved prior guides tool execution, while current observations remain the basis of the final answer. After each task, a safety-gated distillation process converts grounded trajectories into reusable execution knowledge for future retrieval. This execution, distillation, and reuse loop improves planning without updating the backbone LLM.

The methodology section describes the problem formulation: An EO task is x = (q, d, Y), comprising a scientific request, data collection, and answer space. A trajectory is defined as τ = (q, a1, o1,..., am, om, y), ot = Ti (θt). The paper notes that "A valid trajectory must respect remote-sensing dataflow: relevant observations and modalities are identified before product generation, returned artifacts are propagated to downstream tools, and aggregation is performed only after checking spatial and temporal compatibility. Final answers must be supported by the current EO data rather than retrieved memory."

The Workflow Graph Memory is a directed, reliability-aware graph G = (W, A, U) of workflow nodes, tool-transition edges, and usage statistics, encoding EO dependencies such as data discovery → modality-aware input organization → geospatial product generation → spatial or temporal aggregation → scientific interpretation. Each node is wi = (ci, ri, zi, pi, bi, Qi, n+ i, ni−), where ci is the task type, ri the workflow description, zi an order-preserving compressed tool sequence, pi parameter hints, bi cautions, Qi provenance, and n+ i, ni− success/failure counts. Retrieval scores combine tool sequence Jaccard and text semantics similarity via sG (q, wi) = α · Jtool (q, wi) + β · Jtext (q, wi) + λC 1[ci = ĉ(q)] − λL max(0, zi − L0).

The Action-Level Experiences store fine-grained execution knowledge as non-parametric corrections for local decisions under previously observed conditions, with items ei = (χi, αi, ai, ρi, mi), where χi is a trigger, αi a correction, ai the source operations, ρi provenance, and mi metadata. Retrieval uses a gated relevance model with scoring sE (q, ei) = V (q) ∩ V (χi ∥αi ∥ρi) / max(V (q), 1) + λE 1[ĉ(χi ∥ρi) = ĉ(q)]. The bank evolves through novelty-preserving consolidation, admitting only entries outside existing equivalence classes: Et+1 = TailNE (Et ∪ e ∈ ∆E: ν(e) ∈/ ν(Et)).

The Adapted Skill SOP specifies reusable task decompositions, intermediate artifacts, and evidence aggregation, with each skill sj = (µj, ωj, πj, βj, κj), encoding the EO task, processing workflow, data and argument constraints, sensor/modality semantics, and failure-prevention rules. The adaptation operator Sq = AΘ (S, q, Eq; LS) produces a compact task-local SOP, updated via replacement-style consolidation only when it is complete, compact, grounded in an executable trajectory, and passes the safety gate.

At inference, GeoForge assembles a compact execution context by concatenating retrieved graph context, action-level experiences, the adapted SOP, and deterministic EO task guidance: Cq = TruncLC [Gq ∥Eq ∥Sq ∥Hq]. The agent then executes a standard ReAct transition over the filtered MCP tool space: at ∼ πΘ (· q, Cq, ht, Tq), ht+1 = ht ∪ at, ot. Memory is treated only as a processing prior, and every data-grounded EO question requires a tool call, typically beginning with data inventory.

The safety gate for self-evolution is "ψ(τ, ŷ) =1[Tools(τ) > 0] 1[ϕ(ŷ) = 0] 1[Calls(τ) ≤ Lmax] 1[max nT (τ) ≤ Rmax] 1[∆S ∩ F = ∅], with memory transition Mt+1 = (UG (Gt, ∆W), UE (Et, ∆E), US (St, ∆S; ψ))."

Experiments evaluate GeoForge on Earth-Bench, ThinkGeo, and GeoPlan-Bench. On Earth-Bench, "GeoForge achieves the best accuracy on four backbones: GPT-5 at 74.33%, DeepSeek-V3.1 at 77.09%, Gemini-2.5-Flash at 67.91%, and Qwen3-Max at 69.72%, consistently outperforming Earth-Agent, OpenEarth-Agent, and GeoEvolver. It also improves trajectory-level metrics such as Tool-Any-Order, Tool-In-Order, and Tool-Exact-Match on most backbones. On ThinkGeo, GeoForge achieves the highest Instruction Alignment at 97.27%, Tool Accuracy at 80.31%, Answer Accuracy at 60.98%, and grounded Answer Accuracy at 62.43%. On GeoPlan-Bench, it obtains the best Recallkey of 1.00, F1key of 0.77, Structural score of 0.79, and Holistic score of 1100.65."

The ablation study shows "all three memory components contribute to GeoForge. Without memory, the model achieves only 52.23% accuracy. Removing Skill causes the largest accuracy drop, from 74.33% to 52.66%... removing the Workflow Graph leads to the largest degradation in trajectory quality, reducing Tool-Any-Order, Tool-In-Order, and Tool-Exact-Match from 85.89%, 70.01%, and 51.18% to 75.69%, 61.74%, and 45.09%, respectively. Against general agents, GeoForge consistently outperforms all general-purpose agents on Earth-Bench, raising average accuracy from 47.84% to 61.85% over the strongest baseline Earth-Agent, with the highest scores on Spectrum at 77.00% and Products at 71.26%, with gains of 27.00 and 29.15 percentage points respectively."

Error analysis reveals "substantial reductions in tool planning errors across most backbones. Reasoning errors are eliminated for the majority of models, and answer or format failures are also notably reduced. Remaining errors are mainly execution-trace and parameter-related issues. A case study shows the baseline follows a tangled trajectory with many redundant tool calls, ultimately failing to produce a valid answer, while GeoForge executes a concise workflow that retrieves the file list, computes the NBR index, detects hotspot regions, and performs a final directional analysis, producing the correct answer B with substantially fewer steps. Sensitivity analysis shows GeoForge is robust to both the number of retrieved skills and the retrieval score threshold, with Top-k = 3 and Min-Retrieve-Score = 0.6 providing the best trade-off between retrieval quality and planning performance."

The paper concludes: "GeoForge constrains the action space by sensing context and retrieves task-conditioned priors from workflow graphs, action-level experiences, and an adapted skill SOP. Retrieved priors guide execution while observations ground final answers. A safety-gated distillation process enables continuous memory evolution without backbone LLM updates. Experiments across multiple benchmarks show that GeoForge consistently improves task accuracy and tool-use trajectory quality while reducing planning and reasoning errors, offering a scalable paradigm for reliable scientific agents."

Improvements for AI systems

Improvements to AI Systems:

  1. Add a nonparametric, self-evolving execution memory – The AI system can store completed task trajectories as structured, reusable knowledge (workflow graphs, action-level corrections, and skill SOPs) without retraining or updating the backbone LLM. This enables continuous improvement from experience while keeping the model frozen.

  2. Implement sensing-context-aware action space filtering – Before planning, the system constrains its available tools/operations based on the current sensing context (e.g., data modality, spatial/temporal constraints). This reduces search space and prevents invalid tool calls, improving efficiency and accuracy in domain-specific tasks.

  3. Integrate a three-tier memory retrieval mechanism – The system can retrieve: (a) global workflow order from a reliability-aware graph, (b) local corrections from past action-level experiences, and (c) procedural/data constraints from an adapted skill SOP. This provides task-conditioned priors that guide execution while keeping final answers grounded in current observations, not stale memory.

  4. Add a safety-gated distillation loop – After each task, the system can distill successful trajectories into new memory entries only if they pass a safety gate (e.g., valid tool usage, no harmful outputs, bounded complexity). This prevents memory poisoning and ensures only reliable knowledge accumulates.

  5. Use provenance-aware memory consolidation with novelty preservation – The system can admit new experiences only if they are outside existing equivalence classes, avoiding redundant or conflicting memory. This keeps memory compact and high-quality over time.

  6. Enable hybrid retrieval scoring combining semantic and structural similarity – For workflow retrieval, the system can score candidates using both tool-sequence Jaccard similarity and text semantic similarity, plus task-type matching and length penalties. This improves retrieval precision for complex multi-step tasks.

  7. Incorporate deterministic domain guidance into the prompt context – The system can assemble a compact execution context by concatenating retrieved memory with hard-coded domain rules (e.g., must call a tool before answering data-grounded questions). This enforces data-grounded reasoning and prevents hallucination.

  8. Support adaptive skill SOP generation – The system can generate a task-local standard operating procedure on-the-fly by adapting a base skill to the current query and available experiences, ensuring procedural constraints (e.g., data dependencies, argument requirements) are respected.

What the Improved AI System Can Do:

  • Solve complex Earth observation (EO) tasks with higher accuracy (e.g., 74.33% on Earth-Bench with GPT-5, up to 77.09% with DeepSeek-V3.1) than general-purpose agents, by using domain-aware memory and constrained action spaces.

  • Produce more reliable tool-use trajectories – Achieve higher Tool-Exact-Match scores (e.g., 51.18% vs. 45.09% without workflow memory), meaning the system executes the correct sequence of tools in the correct order, reducing redundant or tangled calls.

  • Reduce planning and reasoning errors – The system eliminates most reasoning errors and significantly cuts tool-planning failures, leading to more grounded, executable answers.

  • Self-improve without retraining – After each task, it distills successful trajectories into memory, so performance improves over time on similar tasks without any gradient updates to the LLM.

  • Handle heterogeneous scientific workflows – Beyond EO, the system can generalize to any domain with tool dependencies, data compatibility checks, and multi-step pipelines (e.g., climate modeling, biomedical data analysis, robotic task planning) by adapting its memory structures.

  • Provide explainable and safe evolution – The safety gate ensures only verified, non-harmful trajectories enter memory, and provenance tracking allows auditing of why certain actions were taken.

  • Operate robustly under varying memory sizes and retrieval thresholds – The system remains stable and effective even when the number of retrieved skills or the retrieval score threshold changes, making it practical for deployment in dynamic environments.

Sources

Related papers