ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval

arXiv:2608.12720 · cs.CL, cs.AI · Submitted 2026-08-13 · Read on arXiv

Shenzhen International Center for Industrial and Applied Mathematics · Shenzhen Research Institute of Big Data · The Chinese University of Hong Kong, Shenzhen · Shenzhen Campus of Sun Yat-sen University · Shenzhen Loop Area Institute

cs.CL, cs.AI

Submitted: 2026-08-13

Updated: 2026-09-28

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

The gist: ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval Abstract While Large Language Model (LLM) agents increasingly rely on long-term memory for persistent interactions, the retrieval

Terminology

Summary

ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval

Abstract

While Large Language Model (LLM) agents increasingly rely on long-term memory for persistent interactions, the retrieval mechanisms governing this memory are rarely treated as evolvable components. This static approach limits performance on heterogeneous memory queries, which often demand diverse evidence construction strategies. To address this, we introduce ERSkill, a retrieval-centric framework for self-evolving, skill-guided memory access. ERSkill compiles interaction histories into a structured memory store and represents retrieval behaviors as executable skills composed of fundamental primitives. At inference time, a trained router dynamically matches each query to the optimal skill to construct tailored evidence for answer generation. To enable continuous improvement, ERSkill co-evolves the skill set and the router during training. It employs an experience trie to efficiently record explored retrieval paths, alongside a double-frontier mechanism that safely decouples the expansion of new skill capabilities from stable, router-facing deployment. Experiments across multiple agent memory benchmarks demonstrate that ERSkill substantially outperforms strong non-evolving and self-evolving baselines. Notably, it improves the overall average across F1, BLEU-1, and LLM-judge scores by 31.3% with Qwen3-Next-80B-A3B-Instruct and by 28.1% with GPT-5.4-nano.

Introduction

Large Language Model (LLM) agents are increasingly expected to function as persistent collaborators rather than one-shot assistants. Over extended interactions spanning weeks or months, an agent accumulates user preferences, tracks changing events, remembers prior decisions, and reuses experiences, motivating rapid progress in agent memory. Recent work has made progress toward persistent agents by building explicit external memory to store interaction history. These systems construct and maintain memory through operations such as information extraction, compression, updating, and forgetting, and retrieve stored memories as evidence for downstream tasks. More recent work further moves toward self-evolving agents: instead of only storing key information, agents reflect on past trajectories, distill reusable experiences or insights, store and apply them to guide future behavior.

Despite this progress, the retrieval side of agent memory remains underexplored as an object of evolution. Existing self-evolving methods mainly use past task experience to improve future reasoning; for example, ReasoningBank reflects on past task traces to distill reusable reasoning experiences. MemSkill evolves LLM-based memory extraction skills, but the resulting memories are still accessed through a predefined dense retrieval strategy at query time. Thus, while memory content and reasoning guidance become increasingly adaptive, retrieval behavior itself often remains predefined. This becomes limiting for agent memory question answering, where queries can differ substantially in their information demands. For example, What gift did Alice buy for Bob during her Hawaii trip? primarily requires retrieving a specific event, whereas Why did Alice later stop planning another Hawaii trip with Bob? requires connecting earlier events with later developments to uncover causal relations. Such queries require qualitatively different evidence construction behaviors, raising a central question about query-time memory access: How can an LLM agent evolve and learn to compose complex retrieval actions so that its retrieval behavior adapts to the heterogeneous information demands of different queries?

In this paper, we propose ERSkill (Evolving Retrieval Skill), a retrieval-centric self-evolving framework that models agent memory access as skill-guided evidence construction. To support diverse retrieval perspectives, ERSkill first builds a retrieval-oriented memory storage that exposes memory via a library of retrieval primitives, such as dense retrieval, BM25 retrieval, and query rewriting. Based on these primitives, ERSkill represents retrieval behavior as a set of executable retrieval skills. Each skill specifies a concrete retrieval schema by composing primitives into a sequence. This skill abstraction makes retrieval behavior reusable, interpretable, and refinable. To choose the appropriate retrieval behavior for each query, ERSkill trains a skill router that matches the query's information demand to skills. At inference time, the selected skill controls how memory atoms are gathered, expanded, and organized into a task-specific evidence view for answer generation. In this way, ERSkill realizes query-adaptive memory access while keeping the retrieval process executable and interpretable.

However, a predefined static set of retrieval skills fundamentally limits the expressiveness of memory access, as retrieval requirements vary across queries. We therefore design a skill-router co-evolution procedure that jointly updates the skill set and the router by analyzing previous task traces and identifying gaps between current skill behavior and desired evidence construction. To reduce inefficient exploration, ERSkill stores explored primitive paths in an experience trie, allowing the skill generator to reuse past rollout experience and avoid repeatedly proposing equivalent retrieval programs. Moreover, because the router may not always select the best skill during deployment, skill evolution should consider not only whether a skill is useful in principle, but also whether its utility can be reliably activated by the router. To this end, ERSkill uses a Pareto-style double-frontier mechanism: the capability frontier maintains a compact set of skills with the best retrieval capabilities, while the deploy frontier maintains a router-facing skill set validated under routed inference, allowing ERSkill to expand retrieval capability while keeping deployment stable.

Methodology

Overview. As shown in Figure 2, ERSkill adapts memory retrieval by selecting, for each query, a retrieval skill that matches its information demand. It first compiles the interaction history into memory storage, then uses a trained router to select the most suitable skill, and finally executes the selected skill to construct evidence for answering. During training, the skills and the router co-evolve.

Memory Storage. Let D be the interaction history. ERSkill splits D into atom-level records A = a1,..., an, where each ai stores the atom text, metadata, and timestamp. The compiled memory is M(D) = (A, I, G), where I is a collection of indexes for atom search and G is a collection of graphs for atom expansion. Indexes provide entry points into candidate atoms (e.g., embedding indexes for dense retrieval and entity-to-atom indexes for entity-based retrieval), while graphs connect atoms for expansion (e.g., similarity-based).

ERSkill's memory is not merely a storage, but an executable substrate for constructing query-specific evidence. Specifically, ERSkill exposes the memory storage through a primitive library P = p1,..., pm, where each primitive is a state transition p: (q, s, M(D)) → s'. Here, q is the query, s is the current evidence state, and s' is the updated state. Retrieval primitives use the indexes in I to access memory from different perspectives. Formally, we initialize: (1) entity search that identifies query-relevant entities and retrieves atoms through the entity–atom graph; (2) lexical search that performs BM25-style surface-form matching; and (3) dense search that retrieves atoms by query–atom embedding similarity. Besides, we define that expansion primitives use the graphs in G to grow the evidence state. Specifically, temporal focus expand adds atoms from a given temporal span, similarity expand propagates along similarity edges, and relation expand follows typed relation edges. We also include llm process for query rewriting, evidence filtering, and other dynamic control operations. The primitive library remains fixed throughout evolution; skills differ only in how these primitives are composed. Because all skills share the same primitives, their outcomes can be accumulated in a common experience structure.

Inference. A retrieval skill is an executable primitive program κ = (cκ, ρκ), where cκ is the skill description and the information preference, and ρκ = (pκ,1,..., pκ,Lκ) is the primitive sequence. Given query q, the skill executes its primitives sequentially by applying sj = pκ,j(q, sj−1, M(D)) and returns sLκ as the evidence view. The state maintains the current evidence set and execution context, enabling the skill to organize retrieved information according to the query's evidence needs. Each skill is stored as a markdown file.

ERSkill employs a query-conditioned routing model to select retrieval skills. Given a query q and a retrieval skill set K, the router assigns a relevance score to each skill based on its compatibility with the query. Specifically, we encode the query q and each skill κ into vector representations using a shared encoder Enc(·), resulting in hq = Enc(q) and hκ = Enc(κ), where κ ∈ K. A learnable scoring function uθ(q, κ) is then used to measure the compatibility between the query q and the skill κ, reflecting the performance of applying skill κ to the query q. The routing model Rθ(·) is defined by normalizing these scores into a probability distribution over skills: Rθ(κ q, K) = exp(uθ(q, κ)) / Σκ'∈K exp(uθ(q, κ')), which allows the router to score newly evolved skills from their textual descriptions, information preferences, and programs without requiring changes to the output space. We keep the text encoder frozen and optimize only the routing parameters.

Given a query q and skill set K, the router selects the skill that best matches the query's information demand. We denote this routed skill as πθ(q; K) = arg maxκ∈K Rθ(κ q, K), and write κ̂ = πθ(q; K) at inference time. ERSkill then executes κ̂ over M(D) to refine evidence state sLκ̂, and generates the answer as ŷ = fLLM(q, sLκ̂). In this way, ERSkill enables adaptive query-time memory access through skill selection and execution.

Skill-Router Co-Evolution. ERSkill evolves skills through two Pareto-style frontiers. A frontier is the active boundary of explored skills, retaining compact and non-redundant skills from the evolution history. The capability frontier Ct preserves skills with oracle-side value, i.e., skills that lead to the best attainable utility when the best skill can be chosen for each query without router errors. It therefore tracks the retrieval capability discovered so far. The deploy frontier Bt is the router-facing skill set used at inference time, containing only validation-gated skills that the router can select reliably. This separation allows ERSkill to explore new retrieval capabilities while exposing only stable skills for deployment.

At evolution step t, ERSkill has experience trie Tt, frontiers Ct and Bt, and router parameters θt. Evolution starts from three single-primitive seed skills: semantic-clue with dense search, entity-focus with entity search, and surface-fact with lexical search; both frontiers are initialized from them, so B0 = C0. Given a training batch Qt, ERSkill executes every skill in Ct on every query and records query–skill performance, execution traces, and ability overlaps in Tt. These experiences reveal weak or redundant regions of the current capability frontier, guiding skill candidate generation, capability frontier update, router update, and deploy frontier update.

ERSkill generates new skill candidates by editing paths from the current capability frontier, e.g., adding or replacing primitives. To reuse past search experience and avoid duplicate exploration, ERSkill maintains an experience trie T over primitive paths. Since skills are sequences over the fixed primitive library P, each root-to-node path represents a primitive prefix, and shared prefixes are stored only once. This allows ERSkill to detect duplicate candidates. Each explored path in T corresponds to a skill and stores its train-batch rollouts, validation rollouts, and frontier status. The skill generator proposes candidates from skill summaries, rollout statistics, failure/success patterns, and trie summaries, while excluding paths already recorded in Tt. Thus, both accepted and rejected candidates remain available to guide later evolution.

ERSkill updates the capability frontier by retaining skills with unique oracle-side value. For each query–skill pair, we measure performance by a score r(q, κ) ∈ [0, 1], instantiated as LLM-as-a-Judge accuracy in our experiments. For a batch Q, skill utility is Util(κ; Q) = (1/Q) Σq∈Q r(q, κ), and the oracle score profile of a skill set K is gK(q) = maxκ∈K r(q, κ). ERSkill uses a frontier recomputation operator Φ(K; Q) that sorts skills by utility and removes a skill only when doing so does not decrease gK(q) for any query. Thus, Φ preserves oracle performance while pruning redundant skills. Given generated candidates Ut, ERSkill first computes a train-side temporary frontier C̃t+1 = Φ(Ct ∪ Ut; Qt) and rejects candidates not retained in it. The retained candidates Vt are then evaluated on validation data, and the capability frontier is updated as Ct+1 = Φ(Ct ∪ Vt; Qval). This two-stage update uses training batches for efficient filtering and validation data for stable frontier recomputation.

Training uses a window Wt of rollout instances (q, Kq, r(q, κ) κ∈Kq), where Kq is the historical skill set used for that rollout. Let M be a mini-batch of instances sampled from Wt. For each instance, the rollout scores induce a soft target distribution over its associated skill set Kq, and the router is optimized with soft-label cross-entropy: p̃(κ q, Kq) = exp(r(q, κ)) / Σκ'∈Kq exp(r(q, κ')), Lrouter = −Σ(q,Kq,·)∈M Σκ∈Kq p̃(κ q, Kq) log Rθ(κ q, Kq).

After updating the router, ERSkill refreshes the deploy frontier only when the capability frontier changes, using the newly retained candidates Ht = Ut ∩ Ct+1; otherwise, it sets Bt+1 = Bt. For any skill set K, given available rollout scores for all q ∈ Q and κ ∈ K, its routed performance is Routed(K, θ; Q) = (1/Q) Σq∈Q r(q, πθ(q; K)). ERSkill forms a candidate deploy frontier Bt' = Φ(Bt ∪ Ht; Qval) and computes Δroute(Bt'; θt+1) = Routed(Bt', θt+1; Qval) − Routed(Bt, θt+1; Qval). It accepts Bt' and sets Bt+1 = Bt' if Δroute(Bt'; θt+1) ≥ γroute, or if Δroute(Bt'; θt+1) ≥ −ξdrop and Bt' ≤ Bt, where γroute is a routed-gain margin and ξdrop is a compactness tolerance that allows a small routed-performance drop to accept a stronger, more compact deploy skill set. If the candidate is rejected, ERSkill sets Bt+1 = Bt. Accept/reject outcomes are written back to the experience trie, allowing later skill generation to consider router-aware deployment performance. After training, the final deploy frontier is used for inference.

Proposition 2.1 (Oracle-safe two-level frontier update). Denote the oracle coverage as OCov(K; Q) = (1/Q) Σq∈Q gK(q). For every evolution step t, ERSkill satisfies OCov(Ct+1; Qval) ≥ OCov(Ct; Qval) and OCov(Bt+1; Qval) ≥ OCov(Bt; Qval). Proposition 2.1 shows that both frontiers maintain non-decreasing oracle coverage on the validation set, even though the deploy frontier may lag behind capability updates due to router-aware validation.

Experiments

Experimental Setup. We evaluate ERSkill on three agent memory benchmarks: LoCoMo, LongMemEval, and PerLTQA. LoCoMo and LongMemEval contain multi-session conversational histories, while PerLTQA further evaluates memory use over heterogeneous sources beyond dialogue history.

We compare ERSkill against strong baselines including: (1) Non-evolving baselines include A-Mem, MemoryOS, and LightMem, which focus on memory construction and maintenance without self-evolution from past task traces; (2) Self-evolving baselines include Dynamic Cheatsheet, ReasoningBank, GEPA, and MemSkill. Dynamic Cheatsheet, ReasoningBank, and GEPA reflect on past task traces to distill reusable experiences or insights for future tasks, while MemSkill evolves LLM-based memory construction skills for downstream QA. Dynamic Cheatsheet, ReasoningBank, and GEPA use the standard RAG memory storage.

We evaluate the performance on two LLM backbones, Qwen3-Next-80B-A3B-Instruct and GPT-5.4-nano, using GPT-4o-mini as the judge model. We report F1, BLEU-1 (B1), and LLM-judge score (L-J), where higher values represent better alignment with the ground-truth. For ERSkill, we set the evolution's train batch sizes as 20 and 40 for LoCoMo and PerLTQA, respectively. Queries and skills are encoded with Qwen3-Embedding-0.6B as the Enc(·). For the router, both embeddings are first projected into a representation space via a Linear Layer. The resulting representations are then concatenated and fed into a two-layer multilayer perceptron (MLP) to produce the scalar score uθ(q, k) for each query–skill pair. ERSkill is trained for one epoch. For dense memory retrieval, we use Contriever as the embedder for all methods.

Comparison Experiments. ERSkill achieves the strongest overall performance. Table 1 reports the main comparison results, where we observe that ERSkill achieves the best overall average under both backbones. Specifically, ERSkill improves the overall average by 31.3% across F1, BLEU-1, and L-J with qwen3-next-80b-a3b-instruct, and by 28.1% with gpt-5.4-nano against the strongest baseline. Compared with non-evolving methods, ERSkill shows that its memory storage is sufficiently informative, while skill-guided adaptive retrieval can extract suitable evidence. For self-evolving methods, the gains indicate that updating summaries, prompts, or reasoning traces alone is insufficient when query-time access remains fixed; ERSkill gains from exploiting retrieval-path experience through the experience trie and the skill-router co-evolution mechanism.

As shown in Figure 4, ERSkill's advantages are more pronounced in tasks that require accurately locating specific evidence points, such as Single Hop and Multi Hop, indicating that its effective regulation of retrieval behavior leads to improved evidence-searching capability.

ERSkill achieves a leading cost-performance trade-off. Figure 5 compares token costs for memory construction and inference. Dynamic Cheatsheet, ReasoningBank, and GEPA are omitted from the construction-cost plot because they use standard chunk-and-embedding memory storage rather than LLM-based construction. ERSkill achieves a favorable cost-performance trade-off: it is the lightest among LLM-based memory construction methods, as it uses the LLM only for relation extraction among memory atoms, and remains in the lower-cost tier at inference while achieving the highest L-J score. This suggests that ERSkill's gains come from targeted evidence construction rather than simply feeding more retrieved content.

ERSkill transfers effectively across datasets. ERSkill also shows strong transferability. On LongMemEval, we directly reuse the router and retrieval skills trained on LoCoMo without additional training. As shown in Table 1, in this transfer setting, ERSkill still achieves the best performance under both LLM backbones, which suggests that the learned retrieval skills capture reusable retrieval behaviors.

Ablation Study. Ablation results verify the effectiveness of each design component. We ablate four core components. w/o skill evolution keeps only the initial seed skills. w/o router replaces the learned router with LLM-based skill selection. w/o double frontier removes the capability/deploy frontier and accepts all generated skill candidates. w/o experience trie generates candidates only from the current frontier without historical path-level records. As shown in Figure 6a, all ablations underperform the full ERSkill across datasets and backbones. The largest drops come from removing skill evolution and the router, showing the importance of both discovering effective retrieval skills and learning query-dependent skill selection. The double frontier and experience trie also contribute: the former stabilizes skill acceptance, while the latter reduces repeated exploration and guides candidate generation with accumulated path-level experience.

Hyperparameter Study. We study the train batch size, which controls the granularity of skill evolution. ERSkill runs one pass over the training set. A larger train batch provides more rollout evidence for each evolution step, making frontier recomputation and router updates more stable, but it also reduces the number of evolution steps. Conversely, a smaller batch allows more frequent skill updates, but each update is based on less reliable rollout statistics. As shown in Figure 6b, a batch size of 20 achieves a good balance between stable per-step evolution and sufficient update frequency.

Case Study of Evolution. ERSkill expands retrieval capability while keeping evolution controlled. Figure 7 visualizes a representative run on LoCoMo with GPT-5.4-nano. The left figure shows that both capability- and deploy-frontier oracle accuracy increase steadily, indicating that frontier recomputation preserves and improves oracle-side skill coverage. Routed deploy accuracy may temporarily drop when the deploy frontier replaces weaker or redundant skills with higher-coverage ones before the router has fully adapted. Subsequent updates then recover the routed performance. This suggests that ERSkill can expand retrieval ability without persistent deployment instability, while avoiding an overly conservative deploy frontier. The right figure shows that frontier sizes remain controlled throughout evolution. The capability frontier grows only when new skills add non-redundant oracle value, while the deploy frontier is more conservative. ERSkill avoids accumulating all generated skills, pruning weaker or redundant ones once their utility is covered by stronger alternatives.

Evolution Stability Analysis. ERSkill evolves stably across runs. We evaluate the stability of ERSkill's skill evolution across independent runs. Table 2 summarizes five runs on LoCoMo with GPT-5.4-nano and compares them against the strongest baseline. The coefficients of variation (CV) are low across all metrics, with the largest value below 5.4%. It shows the robustness of ERSkill's evolution.

Related Work

Agent Memory. Memory systems enable agents to retrieve historical information and use it as evidence for subsequent reasoning and decision-making tasks. Recent work on agent memory studies how to externalize long interaction histories into persistent memory stores. A typical pipeline decomposes interactions into memory atoms, compresses or consolidates salient information, stores the resulting memories in external storage, and retrieves relevant evidence when a future query arrives. Some methods focus on designing more effective memory pipelines, such as A-MEM and LightMem, while others improve the agent's memory management ability through training, such as MemoryR-1, Mem-α, and MemAgent. Despite these advances, most methods still expose memory through a predefined retrieval strategy at query time. By contrast, we focus on adaptive memory retrieval: how an agent should access, expand, and organize memory evidence according to heterogeneous query demands.

Self-Evolving Agents and Skill Discovery. Recent work on self-evolving LLM agents studies how agents can improve from interaction experience by using LLMs to analyze past task traces. A typical pipeline is to collect trajectories from previous tasks, let the LLM identify key success or failure factors, and distill them into reusable experiences, rules, prompts, or memory items that can guide future behavior. For example, ReasoningBank accumulates historical reasoning paths and asks the LLM to analyze correct and erroneous reasoning steps, turning the resulting insights into memory items for later reasoning. MemSkill studies memory construction skills: it adjusts how an agent extracts memory items and improves these skills using task signals from downstream memory QA traces. By contrast, ERSkill focuses on constructing skills for memory retrieval rather than memory construction. It treats query-time memory access as an evolvable retrieval behavior and introduces an experience trie and a double-frontier mechanism to enable effective, stable self-evolution.

Conclusion

We proposed ERSkill, a retrieval-centric framework for self-evolving, skill-guided agent memory retrieval. ERSkill represents memory access as executable retrieval skills composed from a shared primitive library, and uses a trained router to select the skill matching each query's information demand. To build effective and deployable skills, ERSkill co-evolves the skill set and router with an experience trie and a double-frontier mechanism, separating capability expansion from router-facing deployment. Experiments on three agent memory benchmarks show that ERSkill outperforms strong non-evolving and self-evolving baselines across different backbone LLMs, while further analyses demonstrate cross-dataset transfer, favorable cost-performance trade-offs, and stable evolution. These results highlight retrieval-side evolution as a promising direction for long-term LLM agents.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems:

  • Improvement: Replace fixed retrieval strategies with a router that dynamically selects among multiple retrieval skills (dense search, entity search, lexical search, temporal expansion, similarity expansion, relation expansion) based on query type.

  • Capability: The AI system can distinguish between queries requiring single-event retrieval (What gift did Alice buy?) versus causal reasoning across events (Why did Alice stop planning another trip?) and apply the appropriate retrieval strategy automatically.

  • Improvement: Represent retrieval behaviors as composable primitive sequences (e.g., entity search → temporal focus expand → llm process) rather than monolithic retrieval functions.

  • Capability: The system can interpret, reuse, and refine retrieval strategies as modular programs, making memory access transparent and debuggable.

  • Improvement: Maintain a trie of explored primitive paths to avoid redundant exploration and reuse successful retrieval patterns across tasks.

  • Capability: The system learns from past retrieval attempts—both successes and failures—and proposes new retrieval programs that build on proven sub-sequences, reducing trial-and-error.

  • Improvement: Separate capability expansion (oracle-optimal skills) from deployment (router-validated skills) using Pareto-style frontiers.

  • Capability: The system can safely explore new retrieval capabilities without destabilizing current performance—new skills are only deployed after the router can reliably select them, preventing regression during online use.

  • Improvement: Jointly update the skill set and the router using soft-label cross-entropy on rollout performance scores.

  • Capability: The router continuously improves its query-to-skill matching as new skills emerge, and skills are selected based on demonstrated utility rather than static heuristics.

  • Improvement: Train retrieval skills on one dataset and directly reuse them on another without retraining.

  • Capability: The system can generalize learned retrieval behaviors to new domains or memory types, reducing the need for per-dataset training.

  • Improvement: Use targeted retrieval programs that construct only necessary evidence rather than retrieving large volumes of content.

  • Capability: The system achieves higher answer quality with lower token costs at inference, making it suitable for long-running agents with budget constraints.


  1. Handle heterogeneous memory queries intelligently: Automatically adapts retrieval strategy per query—locating specific facts, tracing causal chains, or aggregating temporal patterns.

  2. Continuously improve its memory access over time: Evolves its retrieval skills based on task performance, without manual intervention.

  3. Deploy safely in production: New retrieval capabilities are validated under routed inference before exposure, avoiding performance drops.

  4. Transfer learned retrieval behaviors across domains: Reuses effective retrieval patterns from one application (e.g., customer support) to another (e.g., personal assistant).

  5. Operate with lower computational cost: Achieves higher accuracy while using fewer tokens, enabling deployment on resource-constrained platforms.

  6. Provide interpretable retrieval: Each skill is an executable program with a textual description, allowing humans to audit and refine retrieval logic.

Abstract

While Large Language Model (LLM) agents increasingly rely on long-term memory for persistent interactions, the retrieval mechanisms governing this memory are rarely treated as evolvable components. This static approach limits performance on heterogeneous memory queries, which often demand diverse evidence construction strategies. To address this, we introduce ERSkill, a retrieval-centric framework for self-evolving, skill-guided memory access. ERSkill compiles interaction histories into a structured memory store and represents retrieval behaviors as executable skills composed of fundamental primitives. At inference time, a trained router dynamically matches each query to the optimal skill to construct tailored evidence for answer generation. To enable continuous improvement, ERSkill co-evolves the skill set and the router during training. It employs an experience trie to efficiently record explored retrieval paths, alongside a double-frontier mechanism that safely decouples the expansion of new skill capabilities from stable, router-facing deployment. Experiments across multiple agent memory benchmarks demonstrate that ERSkill substantially outperforms strong non-evolving and self-evolving baselines. Notably, it improves the overall average across F1, BLEU-1, and LLM-judge scores by 31.3% with Qwen3-Next-80B-A3B-Instruct and by 28.1% with GPT-5.4-nano.

Sources

Related papers