EvoIR-Agent: Self-Evolving Image Restoration Agentic System via Experience-Driven Learning

arXiv:2605.22208 · cs.CV · Submitted 2026-05-21 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "EvoIR-Agent: Self-Evolving Image Restoration Agentic System via Experience-Driven Learning".

Jane: EvoIR-Agent introduces a novel, self-evolving image restoration agentic system designed to overcome the limitations of existing methods that struggle with zero-shot planning and lack compatibility with new tools or degradations.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So Jane, I was just looking at the title of this paper we're discussing: "EvoIR-Agent: Self-Evolving Image Restoration Agentic System via Experience-Driven Learning." It sounds pretty technical, but what it really suggests is that they've built an agent that learns how to fix images by building its own experience as it goes.

Jane: That title makes sense, Tom; it points toward a system that doesn't just follow fixed rules but actively improves its decision-making based on what it actually sees during the restoration process. It suggests this agent can adapt its approach over time, which is a big deal for complex image restoration tasks.

Lu: I think the real excitement here is in that "Self-Evolving" part; it implies a learning loop where the agent gets smarter just by doing things, which opens up possibilities for creating truly autonomous restoration systems.

Meng: From an engineering standpoint, I'm thinking about how this self-evolution works; does it mean we have to constantly retrain the whole thing every time it learns something new? That kind of iterative learning sounds like it could be very computationally intensive to manage in a real-world deployment.

Lalam: I see this as a cultural shift, Meng; if an AI system can truly evolve its own restoration skills through experience, it means we move from simply using static tools to having systems that possess an emergent competence in handling novel visual problems.

Tom: Exactly, Lalam; it moves us beyond the current limitations where models either need endless training or they are too dumb to handle new degradation types without explicit guidance.

Jane: And looking at the authors, Kailin Zhuang and Jiawei Wu from Sun Yat-sen University, it shows this is coming from a strong academic foundation focusing on intelligent systems engineering.

Lu: Their focus on systematically formulating experience components seems very rigorous; they aren't just throwing data at a black box and hoping for the best.

The paper's summary: Tom: Okay, so what this paper is actually saying in its summary, "EvoIR-Agent: Self-Evolving Image Restoration Agentic System via Experience-Driven Learning," boils down to solving a specific problem where existing image restoration agents struggle with planning without prior experience.

Jane: They identify that current methods have a real dilemma: training-based ones are fast but can't handle new tools, and training-free ones can handle new things but waste time on "naive experience." So, the summary explains they want to keep the compatibility of the latter while getting the speed of the former.

Meng: It sounds like they are proposing a structured way for an agent to collect and use that experience so it doesn't just wander around trying random things every time it sees a new degradation. That sounds like practical problem-solving, not just theoretical modeling.

Lalam: The core idea is that the agent needs to systematically define what constitutes experience—things like selecting the right tool for a specific pattern and deciding the best order to remove degradations—and then use that experience to guide its planning efficiently.

Lu: I think their summary clearly lays out how they break down those components into Tool Selection, Degradation Removal Order, and Visual Quality, which are all tied together in a hierarchical structure.

Tom: Right, so they aren't just saying "it learns"; they are detailing *how* it learns by defining specific experience components and building a tiered pool to store that knowledge.

Jane: It sounds like the summary is really emphasizing the transition from blind trial-and-error to a guided, structured learning process where the agent knows which tool to try first based on what it's already learned.

The paper's improvements: Tom: Now that we know what they’re summarizing, let’s talk about the actual improvements they propose for this EvoIR-Agent system. They suggest a systematic way to break that planning bottleneck by formulating the experience components first.

Jane: They focus on identifying three specific elements: Tool Selection based on degradation patterns, the optimal Removal Order depending on those patterns, and Visual Quality which they treat as a preference requirement involving fidelity and perception.

Meng: That level of detail in defining what constitutes "experience" is interesting; it’s not just recording a win or loss, but capturing *why* a certain action was taken in relation to the specific degradation pattern it faced.

Lu: The hierarchical structure they build for this experience pool, going from Insight Level down to Fine-grained Pattern-oriented retrieval using CLIP and MLLM evaluation, is what really seems like their main technical contribution.

Lalam: That hierarchy sounds incredibly smart because it balances a high-level textual guidance from the LLM with very specific visual pattern matching for the finest level of control.

Tom: So they are essentially creating a multi-layered memory system that allows the agent to switch between broad strategic guidance and extremely specific pattern recognition when needed.

Jane: It seems like this structure is what allows them to achieve that balance they mentioned—retaining compatibility while boosting inference speed compared to the methods they compared it against.

Conclusion: Tom: We’ve covered how EvoIR-Agent tackles the planning overhead by defining those experience components and building a hierarchical pool. So, to wrap up, what is the main implication of this whole work for image restoration agents?

Jane: The main implication is that we can move toward image restoration agents that are both compatible with new tools and efficient enough for real-time use without constantly needing extensive retraining.

Lu: I think the biggest impact is demonstrating a viable path to end-to-end learning where an agent can effectively manage complex degradation coupling scenarios through this structured experience mechanism.

Meng: For practical application, it means we might see restoration tools deployed in less controlled environments because the agent won't need perfect prior training for every single edge case.

Lalam: This work suggests that future AI systems in creative or technical domains will be able to develop expertise through interaction, making the restoration process itself more adaptive and intelligent.

Tom: Fantastic stuff; so we’re looking at a system that learns robustly by structuring its own learning process, all thanks to this EvoIR-Agent paper. It’s been fascinating following this research.

Kailin Zhuang, Jiawei Wu, Zhi Jin

School of Intelligent Systems Engineering, Shenzhen Campus of Sun Yat-sen University

cs.CV

Submitted: 2026-05-21

Updated: 2026-09-30

Code: https://github.com/rightleft-123/EvoIRAgent

Importance score: 82/100

The gist: EvoIR-Agent introduces a novel, self-evolving image restoration agentic system designed to overcome the limitations of existing methods that struggle with zero-shot planning and lack compatibility

Key concepts

Experience Components
These are the specific elements the agent learns: which tool is best for a particular degradation pattern, the correct order to remove degradations in coupled scenarios, and different visual quality preferences. Learning these components allows the agent to make smarter decisions during restoration.
Hierarchical Experience Pool
The experience is organized into three levels: Insight (LLM guidance), Coarse-grained (degradation type mapping), and Fine-grained (specific pattern profiles). This structure allows the system to handle both general degradation knowledge and highly specific visual details efficiently.
Self-Evolving Mechanism
The agent updates its knowledge by acquiring new experiences, using a model like Bradley–Terry–Davidson to determine win/loss relationships, and evolving the LLM's guidance based on these records. This iterative process ensures the system continuously improves its planning strategy without constant retraining.

Terminology

Summary

EvoIR-Agent introduces a novel, self-evolving image restoration agentic system designed to overcome the limitations of existing methods that struggle with zero-shot planning and lack compatibility with new tools or degradations. By systematically formulating experience components and constructing a hierarchical experience pool, EvoIR-Agent bridges the gap between training-based methods' inference efficiency and training-free methods' compatibility, achieving a remarkable Pareto-optimal balance between performance and efficiency in degradation coupling scenarios.

The Core Problem Addressed

The paper addresses the dilemma faced by Multimodal Large Language Model (MLLM)-driven Image Restoration Agents (IRAs): zero-shot planning often fails without experience, necessitating severe trial-and-error overhead. Existing paradigms are limited: training-based methods embed experience into parameters, lacking compatibility with new tools or degradations; conversely, training-free methods use explicit experience storage but suffer from trial-and-error overhead due to naive experience. The authors aim to retain the compatibility of training-free methods while endowing the model with the inference efficiency of training-based methods by resolving this dilemma through a systematic approach.

System Formulation and Experience Components

EvoIR-Agent systematically formulates experience components for an IRA, identifying three core elements:

  1. Tool Selection: The optimal tool depends not merely on broad degradation type but on specific degradation pattern. Experience is designed to explicitly map these patterns to their optimal tool.

  2. Degradation Removal Order: In coupling settings, the optimal order is pattern-dependent.

  3. Visual Quality: Multiple paths can address degradation but yield significantly different visual outcomes, conceptualized as a Preference requirement (Fidelity and Perception).

Hierarchical Experience Pool Structure

To manage the complexity of coarse-to-fine guidance, EvoIR-Agent constructs a hierarchical experience pool structured across three granularities:

(i) Insight Level (Preference-oriented):

The highest level is instantiated by an LLM that leverages distilled textual experience to guide the planning on o.

(ii) Coarse-grained Level (Degradation Type-oriented):

This level instantiates a direct key-value mapping using Q and D as the query, where the value is considered as o.

(iii) Fine-grained Level (Degradation Pattern-oriented):

This level anchors specific patterns using a pattern profile P, which is constructed from images characterizing the pattern and its semantic description. Retrieval involves a two-stage cascade: first, using a CLIP encoder to recall Top-K relevant profiles based on cosine similarity, and second, having the MLLM evaluate these candidates to identify the most accurate profile P∗.

Self-Evolving Mechanism

The experience mechanism is self-evolving and operates in a batch-wise manner to update the priority function p from accumulated records. This process involves:

  1. Experience Acquisition: Exhaustive exploration of all potential outcomes generates an atomic experience record by evaluating restoration images against preference Q using IQA metrics, deriving win rates for each tool.

  2. Experience Evolution (Coarse-grained): The Bradley–Terry–Davidson (BTD) model is used to parameterize win/loss/tie relations, yielding the latest priority function p = argsort(ˆθ). Significance testing determines if fine-grained experience is necessary based on an ability gap threshold.

  3. Insight Level Evolution: Pairwise win/loss/tie relations are deduced from the BTD model and concatenated into a guidance prompt, which is then used by the LLM to distill textual experience, resulting in p = LLM(Guidance, Concatenate(P(i ≻ j), P(i = j))).

Empirical Evaluation and Results

Extensive experiments demonstrate that EvoIR-Agent achieves a significant lead in the full reference metrics and yields a remarkable Pareto-optimal balance between performance and efficiency. Ablation studies confirm that the hierarchical design provides an optimal equilibrium: coarse-grained experience ensures generalizable stability, while fine-grained experience drives peak inference efficiency. Furthermore, ablation on evolution times reveals that Times = 2 emerges as the optimal sweet spot between performance and efficiency, with a trajectory showing diminishing marginal returns for Times > 2. The system also demonstrates strong adaptability to Out-of-Distribution (OOD) data domains, with EvoIR-Agent-S consistently achieving the highest Unified Quality Index (UQI).

Implementation Details

The framework is implemented using Qwen3-VL-Flash for standard operations and Qwen3-Plus for experience evolution. The tool set configurations are detailed in two settings: EvoIR-Agent-I adheres to prior agent settings, while EvoIR-Agent-II utilizes tool settings from agents like 4KAgent. The system employs a batch size B=25 for the self-evolving mechanism and sets the Top-K recall to 3 for retrieval.

Improvements for AI systems

Based on the provided paper, here are specific improvements that can be made to existing Image Restoration Agent (IRA) systems by adopting EvoIR-Agent's architecture, along with what these improved systems will be able to achieve:


  1. The proposed system replaces static or naive experience mechanisms with a dynamic, hierarchical experience pool guided by a priority-driven mapping function.

  2. This mechanism enables the agent to select restoration tools and determine removal orders based on complex degradation patterns, not just broad degradation types.

  3. The system introduces a self-evolving mechanism that iteratively updates this experience pool using newly accumulated records (via pairwise comparison and BTD modeling).

  4. The improved AI system will achieve a significant lead in full-reference metrics (PSNR, SSIM) and yield a remarkable Pareto-optimal balance between restoration performance and inference efficiency compared to state-of-the-art methods.

  5. The system can handle degradation coupling scenarios more effectively by flexibly selecting tools and determining removal orders based on learned patterns, overcoming the trial-and-error overhead of zero-shot planning.

  6. The improved system can operate efficiently across different data domains (Out-of-Distribution/OOD) through three distinct variants:

  7. Direct application of prior experience (EvoIR-Agent-Z), Transfer learning using OOD data (EvoIR-Agent-T), and Scratch initialization from scratch with OOD data (EvoIR-Agent-S).

  8. The system can maintain its performance and efficiency trade-off even when confronted with significantly out-of-distribution degradation scenarios, ensuring robust generalization.

  9. The hierarchical experience pool provides an optimal equilibrium: the coarse-grained level ensures generalizable stability, while the fine-grained level drives peak inference efficiency by focusing only on specific degradation patterns that require deep visual comprehension.

  10. The system can mitigate risks associated with fine-grained experience (like semantic dependency or API overhead) by leveraging the robust mapping of coarse-grained experience as a stable fallback mechanism.

  11. The system can be optimized for specific user preferences by explicitly modeling them as preference requirements (Fidelity and Perception), allowing it to generate restoration outputs that align with human subjective quality standards.

Sources

Related papers