CLEA: Closed-Loop Embodied Agent for Enhancing Task Execution in Dynamic Environments
cs.RO, cs.AI
Submitted: 2025-03-02
Updated: 2025-03-02
Journal ref: 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp. 21048-21054, 2025
DOI: 10.1109/IROS60139.2025.11246638
Project page: https://sp4595.github.io/CLEA
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Large Language Models (LLMs) exhibit remarkable capabilities in the hierarchical decomposition of complex tasks through semantic reasoning.
Terminology
Abstract
Large Language Models (LLMs) exhibit remarkable capabilities in the hierarchical decomposition of complex tasks through semantic reasoning. However, their application in embodied systems faces challenges in ensuring reliable execution of subtask sequences and achieving one-shot success in long-term task completion. To address these limitations in dynamic environments, we propose Closed-Loop Embodied Agent (CLEA) -- a novel architecture incorporating four specialized open-source LLMs with functional decoupling for closed-loop task management. The framework features two core innovations: (1) Interactive task planner that dynamically generates executable subtasks based on the environmental memory, and (2) Multimodal execution critic employing an evaluation framework to conduct a probabilistic assessment of action feasibility, triggering hierarchical re-planning mechanisms when environmental perturbations exceed preset thresholds. To validate CLEA's effectiveness, we conduct experiments in a real environment with manipulable objects, using two heterogeneous robots for object search, manipulation, and search-manipulation integration tasks. Across 12 task trials, CLEA outperforms the baseline model, achieving a 67.3% improvement in success rate and a 52.8% increase in task completion rate. These results demonstrate that CLEA significantly enhances the robustness of task planning and execution in dynamic environments.
Sources
- Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language
- Inner Monologue: Embodied Reasoning through Planning with Language Models
- VIMA: General Robot Manipulation with Multimodal Prompts
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- LLM-as-BT-Planner: Leveraging LLMs for Behavior Tree Generation in Robot Task Planning
- Guiding Long-Horizon Task and Motion Planning with Vision Language Models
- Embodied Task Planning with Large Language Models
- Large Language Models for Multi-Robot Systems: A Survey
- TwoStep: Multi-agent Task Planning using Classical Planners and Large Language Models
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- Qwen2.5-VL Technical Report
- Reducing the Barrier to Entry of Complex Robotic Software: a MoveIt! Case Study
- YOLOv11: An Overview of the Key Architectural Enhancements
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving