REMAC: Self-Reflective and Self-Evolving Multi-Agent Collaboration for Long-Horizon Robot Manipulation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "REMAC: Self-Reflective and Self-Evolving Multi-Agent Collaboration for Long-Horizon Robot Manipulation".
Jane: The paper was written by Puzhen Yuan, Angyuan Ma, Yunchao Yao, Huaxiu Yao, Masayoshi Tomizuka et al. from Tsinghua University and Carnegie Mellon University and University of North Carolina at Chapel Hill and University of California, Berkeley.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everybody. Today we are digging into a paper that just landed on arXiv, and the title alone got me excited: "REMAC: Self-Reflective and Self-Evolving Multi-Agent Collaboration for Long-Horizon Robot Manipulation." Jane, when you see a title that long, what jumps out at you?
Jane: Honestly, Tom, the phrase "self-reflective" and "self-evolving" is what grabbed me. We've seen robots follow plans before, but a robot that can think about its own plan and then improve it over time? That's a whole different ballgame. It's like the difference between someone reading a recipe and someone who's cooked that dish a hundred times and knows exactly where they're going to mess up.
Tom: Right, and it's not just one robot doing this. The "multi-agent" part means we've got multiple robots in the kitchen, potentially helping each other out. I mean, imagine two robots trying to cook a meal together without bumping into each other or duplicating work. That's the dream, right?
Jane: Exactly. And the authors are from Tsinghua, CMU, UNC, and Berkeley. So we've got a solid crew of researchers who clearly care about making robots useful in real homes, not just in controlled labs. They're tackling the problem of long-horizon tasks, which is a fancy way of saying tasks that take many steps and require remembering what you did ten minutes ago.
Tom: And that's where most robots fall apart. They're great at one specific move, like picking up a cup, but terrible at a sequence of moves that depend on each other. This paper is trying to fix that by giving the robot a way to check its own work as it goes.
Jane: So the big implication here is that we're moving from robots that execute commands to robots that understand the intent behind those commands. That's a huge leap for household robotics. I mean, who hasn't wanted a robot to just "clean up the kitchen" without having to spell out every single step?
Tom: And that's the hook for the rest of our discussion. We're going to break down how they actually built this system, what they tested it on, and whether it really works. Stick around, because this one has some numbers that are going to surprise you.
Summary: Tom: So, Jane, we've got the title, we've got the authors, and now let's talk about what this paper actually does. The summary in the abstract is pretty dense, but the core idea is that they built a framework called REMAC that lets robots plan and execute long tasks by constantly checking their own assumptions.
Jane: And the key word there is "constantly." Most planning systems generate a plan once and then just try to execute it. If something goes wrong, they're stuck. REMAC, on the other hand, has these two built-in modules: self-reflection and self-evolvement. The reflection part is like a robot looking in the mirror after each step and asking, "Did that actually work? Is the next step even possible right now?"
Tom: And the evolvement part is where it gets really clever. After the robot finishes a task, even if it succeeded, it looks back at all the reflections it recorded and asks, "Could I have done that better?" Then it uses that experience to plan a better sequence next time. It's literally learning from its own past, not from some massive dataset.
Jane: And they didn't just test this in a simple simulation. They built a whole benchmark based on RoboCasa, which is a realistic kitchen simulation. They have four task categories, twenty-seven task styles, and over fifty different objects. So we're talking about tasks like opening a microwave, putting food in it, defrosting meat in a bowl, heating vegetables on a stove. These are the kinds of chores that actually take humans a while to figure out.
Tom: And here's the kicker. They compared their system against some of the biggest reasoning models out there, like DeepSeek-R1, o3, QwQ, Qwen3, and Grok3. And REMAC didn't just edge them out; it boosted the average success rate by forty percent and execution efficiency by fifty-two point seven percent compared to a single robot baseline. Those are not small numbers, Jane.
Jane: No, they're not. And what's really impressive is that they did this without any task-specific prompting or fine-tuning. The robot gets a vague instruction like "heat the vegetables," and it has to figure out the rest on its own. That's the kind of generalization that makes this feel like a real step toward robots that can actually live in our homes.
Tom: So the summary is: robots that can check themselves, learn from their own mistakes, and even call in a second robot for help. We're going to dig into how they made that happen next, so don't go anywhere.
Improvements: Tom: Alright, Jane, so we've established that REMAC is a framework that helps robots plan and execute long tasks. But what's the actual improvement over what we had before? Because it's not like robots couldn't plan at all. They just did it badly.
Jane: Right, and that's the crux of it. The improvement here is in the checking mechanism. They introduced something called pre-condition and post-condition checks. Before the robot does a subtask, it looks at the scene and asks, "Is this even possible right now?" And after the subtask, it asks, "Did that actually work?"
Tom: And that sounds simple, but it's incredibly powerful. Let me give you an example from the paper. Imagine a robot is told to heat a carrot in the microwave. A naive plan might say: pick up the carrot, open the microwave, put the carrot inside, close the door. But if the robot is holding the carrot, it can't open the microwave door. That's a pre-condition failure.
Jane: And without the check, the robot would just try to open the door, fail, and then maybe drop the carrot or get stuck. With the check, the robot realizes, "Oh, I need to put the carrot down first, then open the door, then pick it up again." That's a simple fix, but it's the difference between a robot that works and a robot that frustrates you.
Tom: But the real improvement is the self-evolvement part. After the robot completes the task, it stores all those reflections in a database. Then, on the next attempt, it uses that database to generate a better initial plan. So instead of making the same mistake twice, it learns from the first attempt and skips the redundant steps entirely.
Jane: And that's where the efficiency gains come from. In their experiments, the initial plan length dropped by thirty-five percent to sixty-two percent after a few iterations. That means the robot is doing fewer unnecessary moves, which saves time and reduces the chance of failure.
Tom: And then there's the multi-agent twist. Once the robot has a solid plan, it can look at the pre-conditions and think, "Hey, I need to open the microwave door, but I'm busy with the carrot. Why don't I call the other robot to open the door for me?" That's parallel execution, and it's a huge efficiency boost.
Jane: And that's a big deal because most multi-robot systems are just splitting a task into chunks and assigning them. But here, the robots are actively coordinating based on what they've learned. It's not just parallel; it's intelligent parallel.
Tom: So the improvements are: better checking, learning from mistakes, and smarter collaboration. We're going to look at the actual first page of the paper next and see how they set all this up, so stay with us.
First Page: Tom: So Jane, we've talked about the framework and the improvements, but let's actually look at the first page of the paper, because that's where they set the stage. And the first thing they do is point out a fundamental problem with current vision-language models.
Jane: And that problem is that these models, as powerful as they are, often ignore the physical constraints of the environment. They'll generate a plan that sounds logical in text but is completely impossible in the real world. Like, they'll tell a robot to open a door while it's holding a heavy object, which just doesn't work.
Tom: Exactly. And the paper gives a great example. If you ask a model to "heat the vegetables," it might plan to open the microwave, put the vegetables in, and start it. But if the microwave door is closed and the robot's hands are full, that plan is dead on arrival. The model didn't check the scene first.
Jane: So the authors propose REMAC as a solution, and the first page really emphasizes that this is a zero-shot framework. That means you don't need to train it on your specific task. You just give it a high-level instruction, and it figures out the rest using its own reasoning and the checking mechanisms we talked about.
Tom: And that's a big deal for real-world deployment. If you want a robot to do a new task, you don't want to spend weeks collecting data and fine-tuning. You want to just tell it what to do and let it figure it out. REMAC is designed for exactly that.
Jane: And the first page also introduces the idea of the reflection database. Every time the robot checks a pre-condition and finds a problem, it records that problem. Over time, that database becomes a kind of memory that the robot can use to plan better in the future.
Tom: And that's what makes it "self-evolving." The robot isn't just reacting to failures; it's actively building a better plan for the next time. It's like a chef who writes down notes after every service and then uses those notes to improve the menu.
Jane: So the first page sets up the problem, introduces the solution, and gives us a roadmap for the rest of the paper. And the roadmap leads to some pretty impressive results. We're going to wrap up our discussion in a moment, but I think it's worth noting that this paper is not just about making robots better at kitchen tasks. It's about making robots better at adapting to new situations.
Tom: And that's the future of robotics. We're going to summarize everything we've learned and say goodbye to this paper in just a moment.
Conclusion: Tom: Alright, Jane, we've spent a good chunk of time with "REMAC: Self-Reflective and Self-Evolving Multi-Agent Collaboration for Long-Horizon Robot Manipulation," and I think it's time to wrap up. So let's recap the big takeaways.
Jane: Sure. The core idea is that robots can plan and execute long tasks much better if they check their own work as they go. The pre-condition and post-condition checks catch mistakes before they happen and verify success after each step.
Tom: And the self-evolvement part means the robot learns from those checks. It stores reflections and uses them to generate better plans on the next attempt. That's how they got that forty percent boost in success rate and the fifty-two point seven percent improvement in efficiency.
Jane: And we can't forget the multi-agent collaboration. Once the robot has a solid plan, it can call another robot to handle pre-conditions in parallel. That's how you get two robots working together to open a microwave and place food inside without stepping on each other's toes.
Tom: And the fact that they tested this across four task categories, twenty-seven task styles, and fifty objects in a realistic kitchen simulation makes the results feel solid. This isn't a toy example; it's a real step toward robots that can handle the messiness of daily life.
Jane: And the implications are huge. If robots can plan, check, and evolve on their own, then we're moving toward a future where you can just tell a robot to "clean up the kitchen" and trust that it will figure out the details. That's the kind of technology that could actually change how we live.
Tom: So, as we say goodbye to this paper, I want to thank the authors for giving us a framework that feels both practical and visionary. And to our listeners, if you're excited about robots that can think on their feet, this is one to watch.
Jane: And with that, we're ready to move on to the next paper. Thanks for tuning in, and we'll see you on the next episode.
Puzhen Yuan, Angyuan Ma, Yunchao Yao, Huaxiu Yao, Masayoshi Tomizuka, Mingyu Ding
Tsinghua University · Carnegie Mellon University · University of North Carolina at Chapel Hill · University of California, Berkeley
cs.RO, cs.AI, cs.CL, cs.CV
Submitted: 2026-08-10
Updated: 2026-08-11
Comments: 24 pages, 8 figures
Project page: https://reamac-repo.github.io/remac-project-page
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 55/100
The gist: "We propose a novel framework for long-horizon, multi-robot task planning called REMAC, which enables the agent to continuously interact with the environment and improve planning through
Key concepts
- Self-Reflective
- This capability allows a robot to think about its own plan by constantly checking its assumptions during execution. It asks questions like, 'Did that actually work?' or 'Is the next step even possible right now?' This helps the robot correct errors in real-time.
- Self-Evolving
- This involves the robot learning from its own past experiences. After completing a task, it stores reflections and uses this data to generate a better initial plan for future attempts, rather than relying on external training data.
- Pre-condition and Post-condition Checks
- These are mechanisms where the robot checks the environment before performing a subtask (pre-condition) to ensure it is possible. After the subtask, it checks if that action was successful (post-condition), preventing failures like trying to open a door while holding an object.
- Multi-Agent Collaboration
- This involves multiple robots working together on a task. Instead of just splitting work, the robots actively coordinate based on their learned knowledge. For example, one robot might call another to handle a pre-condition in parallel.
Terminology
Summary
Summary
The paper introduces REMAC, a framework for long-horizon, multi-robot task planning that integrates self-reflection and self-evolvement mechanisms. The authors state: "We propose a novel framework for long-horizon, multi-robot task planning called REMAC, which enables the agent to continuously interact with the environment and improve planning through self-reflection and self-evolvement. REMAC is described as
a zero-shot framework that generates efficient and feasible multi-robot collaborative task plans based on concise task descriptions provided by humans across diverse scenarios."
The framework addresses two critical limitations in existing approaches: "1) The tendency of VLM to neglect logical and spatial constraints within the environment during task planning under ambiguous instructions and 2) The inability of prevailing multi-robot task planning methods to handle tasks beyond short-duration horizons effectively."
The method consists of two interconnected stages. First, a task planning and check mechanism: Given scene information and task goal, a large language model decomposes the high-level task description into subtasks that can be executed by low-level policies.
For each subtask, the framework performs 1) pre-conditions check to assess the feasibility of the plan and 2) post-condition checks to evaluate the successful execution of subtasks.
The pre-conditions check verifies whether current conditions meet planning requirements. If met, the robot executes the subtask; otherwise, the VLM identifies failure reasons for replanning the upcoming subtask sequence.
The post-conditions check evaluates whether success criteria are met. If satisfied, the robot proceeds to the next subtask; otherwise, it retries.
Second, an iterative self-evolving framework: "Long-horizon tasks may achieve success after undergoing pre-conditions checks, reflection, post-conditions checks, and retries, but there may still be redundant steps in the execution; alternatively, the planning may remain a failure. Nevertheless, the reflections generated throughout the process are stored in a long-term memory database to guide the planning in the subsequent iteration. The framework uses multiple iterations
to allow the planning to evolve through reflection. After a single robot can execute the task successfully and efficiently,
we prompt the planning LLM to call the other robot to assist, further improving efficiency without sacrificing task success rate."
The authors built a multi-agent benchmark based on RoboCasa with 4 task categories, 27 task styles and 50+ different objects.
The four tasks are OpenCabinetPnP, OpenMicrowavePnP, DefrostInBowl, and HeatOnStove, each incorporating 6–8 spatial layouts, 5–12 dynamically configured environmental styles, and over 50 graspable objects.
The tasks use scene-agnostic high-level instructions
requiring the robot to locate necessary items and decompose tasks autonomously.
Experiments were conducted under four settings: "1. single-robot planning (BASE): A single-robot system without pre- and post-condition checking or the reflection-based iterative evolvement mechanism. 2. + condition checking (CC): A single-robot system incorporating pre- and post-condition checking. 3. + reflective evolvement (RE): A single-robot system with both pre- and post-condition checking and the reflection-based iterative evolvement mechanism. 4. + multi-agent collaboration (REMAC): A multi-robot system integrating both pre- and post-condition checking and the reflection-based iterative evolvement mechanism."
Results show REMAC improve[s] the average success rate by 40.0% and efficiency by 52.7% compared to a single robot baseline.
Specifically, task success rates improved from near-zero baseline to 20-60% across tasks, and self-evolvement enhances planning efficiency by reducing initial plan length by 35-62% through pruning redundant subtasks.
Transitioning from single-robot (RE) to multi-robot (REMAC) coordination reduced average subtask duration by approximately 20% through optimized task decomposition.
The authors also benchmarked reasoning models including DeepSeek-R1, o3, QwQ, Qwen3, and Grok3
on their reflective ability. Results show Grok3 outperforms Qwen3, o3, DeepSeek-R1, and QwQ, showing superior reflection ability,
with reflection success rates of 82.5% for Grok3, 66.25% for Qwen3, 62.5% for o3, 37.5% for DeepSeek-R1, and 26.25% for QwQ.
Real-world experiments were conducted on a dual-arm mobile manipulation robot based on HEXMOVE and equipped it with two PiPER robotic arms,
reproducing the OpenMicrowavePnP task. Results show BASE fails entirely (0% success), CC achieves 10% overall success with 128.5s average time, and REMAC achieves 25% overall success with 57.2s average time.
The authors conclude: "Our work resolves two critical limitations in existing approaches: 1) The tendency of VLM to neglect logical and spatial constraints within the environment during task planning under ambiguous instructions and 2) The inability of prevailing multi-robot task planning methods to handle tasks beyond short-duration horizons effectively."
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to an AI system, along with what the improved system can do:
-
Improvement: Before executing any subtask, the AI system checks whether the current state (from visual observations) satisfies the requirements for that subtask. If not, it flags a planning error and triggers replanning.
-
What it can do: Prevents failures caused by ignoring spatial/logical constraints (e.g., attempting to open a microwave door while holding an object).
-
Improvement: After executing a subtask, the AI system verifies whether the subtask actually succeeded by comparing the post-execution scene with the expected outcome. If it failed, it retries or rechecks pre-conditions.
-
What it can do: Detects low-level policy execution failures and avoids cascading errors in long-horizon tasks.
-
Improvement: Store reflections from failed or suboptimal plans in a long-term memory database. On subsequent task attempts, feed these reflections back into the planner to generate a more efficient and feasible initial plan.
-
What it can do: Reduces redundant steps (e.g., eliminating the need to put down an object before opening a door) and improves task success rate over iterations without task-specific prompting.
-
Improvement: After a single robot successfully completes a task, the system prompts the planner to identify pre-conditions that can be handled by a second robot in parallel (e.g., opening a door or fetching a container). The main robot then focuses on the primary object-centric plan.
-
What it can do: Improves execution efficiency by 20% on average (up to 52.7% vs. single-robot baseline) without sacrificing success rate, by parallelizing independent subtasks.
-
Improvement: Instead of relying on task-specific prompts with detailed scene information, the system uses only high-level instructions (e.g.,
heat the vegetables
) and discovers necessary objects (microwave, pan) from the environment via visual grounding. -
What it can do: Generalizes to unseen layouts and object configurations, avoiding brittle plans that fail when the scene changes.
-
Improvement: Use a reasoning-capable LLM (e.g., DeepSeek-R1, o3, Grok3) specifically for generating and refining plans, while a faster VLM (e.g., GPT-4o) handles real-time condition checks. This splits cognitive load and improves both speed and accuracy.
-
What it can do: Achieves up to 82.5% reflection success rate (with Grok3) and significantly outperforms non-reasoning models in iterative planning.
-
Execute long-horizon, multi-stage tasks (e.g.,
prepare a meal with coffee and vegetables
) with high success rates (up to 60% vs. near-zero baseline) in simulated kitchen environments. -
Self-correct in real time by detecting infeasible actions before execution and verifying completion after each step.
-
Learn from past mistakes across task attempts, reducing plan length by 35–62% and improving efficiency without manual intervention.
-
Coordinate multiple robots to work in parallel on independent subtasks, cutting execution time by 20% while maintaining robustness.
-
Operate in unseen environments using only high-level language instructions, without needing task-specific prompts or fine-tuning.
-
Benchmark and select the best reasoning model for reflection tasks, enabling adaptive use of models like Grok3 or Qwen3 for optimal planning.
These improvements directly address the paper’s core contributions: self-reflection, self-evolvement, and multi-agent collaboration for robust, efficient long-horizon manipulation.
Sources
- ViLPAct: A Benchmark for Compositional Generalization on Multimodal Human Activities
- Towards Human Awareness in Robot Task Planning with Large Language Models
- Multi-Agent Collaboration: Harnessing the Power of Intelligent LLM Agents
- LaMMA-P: Generalizable Multi-Agent Long-Horizon Task Allocation and Planning with LM-Driven PDDL Planner
- Guiding Long-Horizon Task and Motion Planning with Vision Language Models
- RePLan: Robotic Replanning with Perception and Language Models
- RoboMatrix: A Skill-centric Hierarchical Framework for Scalable Robot Task Planning and Execution in Open-World
- Generalizable Long-Horizon Manipulations with Large Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Qwen2 Technical Report
- Qwen2.5 Technical Report
- Qwen3 Technical Report
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving