REFLEX: Reflective Evolution from LLM Experience
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "REFLEX: Reflective Evolution from LLM Experience".
Jane: Large language models (LLMs) have emerged as powerful tools for guiding evolutionary search toward interpretable programmatic policies,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So we're looking at this paper called "REFLEX: Reflective Evolution from LLM Experience," and it seems like they're tackling a real problem with how we use large language models for finding good programming policies.
Jane: Exactly, Tom. The main issue they point out is that right now, existing methods mix looking at pictures of failures with writing the code to fix them all at once, which makes it really hard to see why a specific change was made and stop those insights from carrying over between different tests.
Lu: What's interesting about this paper is their core idea: they argue that the visual diagnosis needs to be completely separate from the actual code generation part of the process so you get an auditable search.
Meng: So, instead of one giant step where the model looks at a picture and writes code, they propose splitting that into two distinct roles to keep things transparent.
Lalam: I see how it works: there's this vision-enabled Critic that just looks at the behavioral evidence image and spits out a structured JSON diagnosis detailing failure modes and root causes, without writing any code itself.
Tom: That decoupling is what they call the Multi-turn Diagnostic Chain or MDC, which ensures that every single mutation in the search process explicitly records that visual problem before any actual code edit gets applied.
Jane: And then you have this text-optimized Actor who takes that structured diagnosis and combines it with the parent code and some reusable snippets to synthesize a new child program, which is where they introduce this persistent Skill Memory.
Meng: This Skill Memory is a library of reusable code snippets pulled from successful individuals, and they use something called a UCB1 bandit to decide which operator—like diagnosing or text-optimizing—to use next based on past success.
Lu: What really stands out is how this skill memory allows them to abstract and transfer learned concepts, like energy-pumping heuristics, across completely different evolutionary campaigns without having to re-learn everything every time.
Tom: That cross-run knowledge transfer sounds pretty powerful for making the search process way more efficient than just running things one by one from scratch.
Jane: They test this separation on a bunch of domains, including classic control problems like Lunar Lander and Pendulum, and they show that this decoupling drastically speeds up the search.
Meng: Their results are pretty impressive; for instance, they say REFLEX reaches the Pendulum task ceiling, which is normalized at one point zero zero zero MLES score, in as few as two LLM calls <ref:2606.16496#pg1,in as few as 2 LLM calls>.
Tom: And they also show competitive performance on other tasks too; they match the Acrobot ceiling with a score of zero point nine eight five and Lunar Lander performance with a best test score of zero point eight two two NWS <ref:2606.16496#pg1>.
Lu: Beyond those benchmarks, they even apply this framework to a much harder problem called the Concentric Circular Antenna Array synthesis task, which is a thirty-six dimensional continuous design problem that’s quite different from the control tasks they started with.
Jane: And for that engineering design task, they achieved a mean best score of twenty-five point three one across different random seeds, which is significantly faster than the traditional black-box optimizers that took about one hundred and sixty evaluations to hit a similar target score.
Meng: The diagnostic feedback loop seems to be what really accelerates things here; they found that existing REFLEX campaigns can reach the twenty-five point two five region in just seven evaluations on average, which is much quicker than the hundreds it would take before.
Tom: So, what this paper suggests is that by separating the vision diagnosis from code generation and keeping a persistent library of skills, you get a framework that’s not just faster at finding solutions but also transparent about how it gets there.
Jane: They call this separation between the Critic and the Actor a way to make multimodal LLM-guided evolution auditable, which means you can actually inspect the logic behind every single mutation.
Lu: It reframes how we think about these LLMs in search; they aren't just one monolithic model making everything happen at once, but a two-role process where diagnosis comes first and then repair follows.
Tom: And they mention that this approach works across different types of control problems, from discrete control to continuous control, and even engineering design tasks.
Jane: The paper also flags a limitation: they're focused on the structure of the search itself, and while it shows great acceleration, they haven't explicitly mentioned how it handles all kinds of novel failure modes outside the specific benchmarks they tested.
Meng: So in simple terms, REFLEX is a new way to guide AI search by making sure we can see exactly why an AI changed its mind on a design or control problem and learn from that experience for future problems.
Tom: It’s about improving sample efficiency while still keeping the resulting policies interpretable, which is a big deal when you're building systems that need to be trusted.
Jane: That's the main point of "REFLEX: Reflective Evolution from LLM Experience," showing how structured diagnosis and persistent memory can make AI search much more efficient and understandable.
Conclusion: Tom: So we've been talking about REFLEX, which is this new way to use LLMs for finding good programming policies.
Jane: Right, it’s this paper called "REFLEX: Reflective Evolution from LLM Experience." The authors are looking at how to make that whole process more transparent.
Lu: It’s interesting because they're decoupling the visual diagnosis from the code generation part of the search.
Meng: Decoupling sounds good, but how does it actually work when you’re trying to evolve a program?
Lalam: Basically, you have this vision-enabled Critic that just looks at the image of failure and spits out a structured report on what went wrong.
Tom: So that report isn't code; it's just a diagnosis of the problem.
Jane: Exactly. And then another part, the Actor, takes that diagnosis and combines it with the existing code to make a new version.
Lu: The real kicker is this Skill Memory they build in, which lets them store useful code snippets from successful tests and transfer those ideas between different search runs.
Meng: That sounds like a big deal for practical engineering. So you can reuse solutions instead of reinventing the wheel every time?
Tom: It’s more than just reusing wheels; it’s about making the search process track exactly *why* a change happened, which helps with understanding the policy itself.
Jane: They show that this whole loop works across control problems and even complex engineering design tasks.
Lalam: And they found that this structured way of thinking leads to much faster results, cutting down on how many tests you need to run.
Tom: So if you’re listening and you want to know what REFLEX really does, it's about making AI search more efficient and giving us a better look at the reasoning behind the code it produces.
Jane: It’s a big step toward building programs that are not just functional, but also understandable.
Pan Wang
University of Science and Technology of China
cs.CL, cs.LG
Submitted: 2026-06-15
Updated: 2026-10-05
Comments: NeurIPS 2026
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 88/100
The gist: Large language models (LLMs) have emerged as powerful tools for guiding evolutionary search toward interpretable programmatic policies, but existing frameworks suffer from an opaque feedback loop
Key concepts
- Multi-turn Diagnostic Chain (MDC)
- This chain structurally separates the process into two distinct roles: a 'vision-enabled Critic' that diagnoses the visual evidence, and a 'text-optimized Actor' that synthesizes code. This separation ensures the evolutionary trace is fully inspectable because each code edit is preceded by an explicit diagnosis of the visual problem.
- Vision-enabled Critic
- This component analyzes behavioral evidence images to emit a structured JSON diagnosis detailing failure modes and root causes. It does not write code; its sole job is to provide standardized, actionable feedback on what went wrong visually, such as identifying an incorrect altitude threshold in a lunar lander scenario.
- Skill Memory
- This is a persistent library of reusable code snippets extracted from successful individuals in previous evolutionary campaigns. It allows the framework to abstract and transfer learned programmatic concepts across independent searches. A UCB1 bandit manages this memory, selecting operators based on historical success to maximize utility.
- Actor
- The Actor takes the parent code, the Critic's structured diagnosis, and relevant snippets from Skill Memory to synthesize a new child program. By relying on external diagnostic input rather than purely autonomous generation, it ensures that mutations are guided by explicit problem-solving rationale.
Terminology
Summary
Large language models (LLMs) have emerged as powerful tools for guiding evolutionary search toward interpretable programmatic policies, but existing frameworks suffer from an opaque feedback loop caused by entangling visual evidence interpretation with code synthesis, obscuring the rationale behind mutations and preventing cross-run knowledge transfer <ref:2606.16496#pg1>. This paper introduces REFLEX, a train-free evolutionary framework that structurally decouples visual diagnosis from code generation to achieve auditable and efficient policy search <ref:2606.16496#pg1>.
How it works
REFLEX operates through a Multi-turn Diagnostic Chain (MDC) that structurally decouples visual diagnosis from code generation, ensuring that the evolutionary trace is fully inspectable <ref:2606.16496#pg4>. First, a vision-enabled Critic
analyzes the task-specific Behavioral Evidence (BE) image to emit a structured, auditable JSON diagnosis outlining specific failure modes and root causes <ref:2606.16496#pg4>. Crucially, this Critic does not write code <ref:2606.16496#pg4>.
Subsequently, a text-optimized Actor
synthesizes a child program by combining the parent code, the Critic’s diagnosis, and reusable code snippets retrieved from a persistent Skill Memory <ref:2606.16496#pg4>. This structural decoupling makes the evolutionary trace fully inspectable because Each mutation explicitly records the diagnosed visual problem Dp before any code edit p' is applied
<ref:2606.16496#pg4>.
Skill Memory and Evolutionary Operators
REFLEX augments the Actor with a persistent, self-evolving Skill Memory—a library of reusable code snippets extracted from elite individuals—enabling the framework to abstract, store, and transfer learned programmatic concepts across independent evolutionary campaigns <ref:2606.16496#pg3>. This memory is managed via a UCB1 bandit that selects operators from a set O = 5 (m diag), 10 (m text), 15 (x div), and i1 (initialization) based on historical success, maximizing the utility Ut(o) <ref:2606.16496#pg5>. New skills are extracted only when a child policy yields a positive fitness improvement, and their utility ui is updated based on the fitness delta of the child policy, promoting active skills that contribute to successful children and evicting those that fail
<ref:2606.16496#pg5>.
Interpretability of Critic Diagnoses
The framework enforces a rigorous diagnostic thought process by prompting the Critic to output a structured JSON object containing specific keys: failure mode,
visual evidence,
root cause,
and suggested fix
<ref:2606.16496#pg2>. This taxonomy ensures that the Actor receives standardized, actionable feedback rather than unstructured text <ref:2606.16496#pg2>. Qualitative examples demonstrate this grounding, such as diagnosing a Lunar Lander failure by identifying The altitude threshold for the braking heuristic is set too high
as the root cause <ref:2606.16496#pg2>.
Cross-run and Cross-task Knowledge Transfer
A key advantage is the ability to transfer learned heuristics across independent evolutionary campaigns via Skill Memory <ref:2606.16496#pg5>. This is tested by seeding new runs with a pre-populated Skill Memory bank, which results in warm-start recipients converging significantly faster, with one case reaching NWS = 1.115 in just 6 LLM calls compared to 14 calls for a cold start <ref:2606.16496#pg9>. Furthermore, testing cross-task knowledge transfer showed that skills from Pendulum campaigns could lift the average second-step score from near failure to near solution by capturing control philosophies—such as phase-aligned energy pumping, velocity-gated stabilization—that transcend specific state and action spaces
<ref:2606.16496#pg8>.
Generalization to Engineering Design
REFLEX demonstrates generalization beyond control benchmarks by successfully applying the framework to a 36-dimensional continuous design problem, the Concentric Circular Antenna Array (CCAA) synthesis task <ref:2606.16496#pg6>. The method achieves a robust mean best score of 25.31 across independent seeds, significantly accelerating discovery compared to traditional black-box optimizers which required roughly 160 evaluations to reach the target score <ref:2606.16496#pg6>. The diagnostic feedback loop drastically accelerates the search process, showing that existing REFLEX campaign traces can reach the 25.25 region in a median of just 7 evaluations <ref:2606.16496#pg6>.
The complete evolutionary loop, integrating the Critic’s visual diagnosis, Actor’s synthesis, and Skill Memory updates, is formalized in Algorithm 1 <ref:2606.16496#pg4>. REFLEX reframes multimodal LLM-guided evolution as an auditable two-role process: visual diagnosis first, code repair second. Existing results suggest that this separation, combined with persistent Skill Memory and forced exploration bursts, can improve sample efficiency while preserving interpretable programmatic policies <ref:2606.16496#pg9>. The same loop applies across discrete control, continuous control, and engineering design, making REFLEX a promising direction for train-free, transparent search.
REFERENCES
Bouzenia, I., Devanbu, P., and Pradel, M. Repairagent: An autonomous, llm-based agent for program repair. In 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE), pp. 2188–2200. IEEE, 2025
Chen, X., Lin, M., Scharli, N., and Zhou, D. Teaching large language models to self-debug. arXiv preprint arXiv:2304.05128, 2023
Dib, N. and Sharaqa, A. Synthesis of thinned concentric circular antenna arrays using teaching-learning-based optimization. International Journal of RF and Microwave Computer-Aided Engineering, 24(4):443–450, 2014
Dolph, C. L. A current distribution for broadside arrays which optimizes the relationship between beam width and side-lobe level. Proceedings of the IRE, 34(6):335–348, 1946
Gou, Z., Shao, Z., Gong, Y., Shen, Y., Yang, Y., Duan, N., and Chen, W. Critic: Large language models can selfcorrect with tool-interactive critiquing. arXiv preprint arXiv:2305.11738, 2023
Guo, Q., Wang, R., Guo, J., Li, B., Song, K., Tan, X., Liu, G., Bian, J., and Yang, Y. Connecting large language models with evolutionary algorithms yields powerful prompt optimizers. arXiv preprint arXiv:2309.08532
Hu, Q., Tong, X., Yuan, M., Liu, F., Lu, Z., and Zhang, Q. Multimodal llm-assisted evolutionary search for programmatic control policies. arXiv preprint arXiv:2508.05433, 2025
Hu, Y., Wen, Z., Liu, X., Wang, P., Zhang, X., and Wu, W. Seal: Synergistic co-evolution of agents and learning environments. arXiv preprint arXiv:2605.24426
Jiang, N., Li, X., Wang, S., Zhou, Q., Hossain, S. B., Ray, B., Kumar, V., Ma, X., and Deoras, A. Ledex: Training llms to better self-debug and explain code. Advances in Neural Information Processing Systems, 37:35517–35543, 2024
Kohler, H., Delfosse, Q., Akrour, R., Kersting, K., and Preux, P. Interpretable and editable programmatic tree policies for reinforcement learning. arXiv preprint arXiv:2405.14956
Lehman, J., Gordon, J., Jain, S., Ndousse, K., Yeh, C., and Stanley, K. O. Evolution through large models. In Handbook of evolutionary machine learning, pp. 331–366. Springer, 2023
Lin, J., Zhu, C., Kneuertz, P. J., Bai, Y., and Xue, Y. Medcausalx: Adaptive causal reasoning with self-reflection for trustworthy medical vision-language models. arXiv preprint arXiv:2603.23085
Lin, Y.-A., Lee, C.
Improvements for AI systems
-
A train-free evolutionary framework that
structurally decouples visual diagnosis from code generation
enables auditable and targeted programmatic mutations, solving thediagnosis–repair entanglement bottleneck.
This allows for explicit logging of failure modes before any code edit is applied, ensuringtransparent mutation traces.
-
The introduction of a persistent Skill Memory managed via a UCB1 bandit enables
cross-run programmatic knowledge transfer across independent evolutionary campaigns,
allowing the system to abstract and store learned concepts likeenergy-pumping heuristics or altitude-gated braking
for reuse. -
The architecture allows the system to achieve
exceptional sample efficiency
across discrete and continuous control environments, reaching task ceilings in minimal calls, such as solving Acrobot and Pendulum inunder 10 LLM calls.
-
A vision-language Critic generates structured JSON diagnoses with specific keys like
failure mode,
visual evidence,
androot cause,
enabling human developers to audit the process by identifying whether a failure was due to aperception error (the Critic misread the BE) or a synthesis error (the Actor wrote buggy code).
-
The system can generalize its learned heuristics to complex, out-of-distribution engineering design tasks, such as CCAA synthesis, where it reaches high performance by leveraging visual evidence to ground diagnoses in
observable physical or geometric anomalies.
Sources
- Teaching Large Language Models to Self-Debug
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
- EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers
- Multimodal LLM-assisted Evolutionary Search for Programmatic Control Policies
- SEAL: Synergistic Co-Evolution of Agents and Learning Environments
- Interpretable and Editable Programmatic Tree Policies for Reinforcement Learning
- When Models Learn to Ask Why: Adaptive Causal Reasoning for Trustworthy Medical Vision-Language Models
- Evolution of Heuristics: Towards Efficient Automatic Algorithm Design Using Large Language Model
- Synthesizing Programmatic Reinforcement Learning Policies with Large Language Model Guided Search
- End-to-End Neuro-Symbolic Reinforcement Learning with Textual Explanations
- Eureka: Human-Level Reward Design via Coding Large Language Models
- BlendRL: A Framework for Merging Symbolic and Neural Policy Learning
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agents
- MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolution
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering