Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)
cs.RO, cs.AI, cs.LG
Submitted: 2026-06-25
Updated: 2026-09-20
Comments: Solution of the LeHome Challenge at ICRA 2026
Code: https://github.com/huggingface/lerobot
Project page: https://lehome-challenge.com
License: http://creativecommons.org/licenses/by/4.0/
The gist: I describe my solution to the LeHome Challenge 2026, an ICRA 2026 competition on bimanual garment folding.
Terminology
Abstract
I describe my solution to the LeHome Challenge 2026, an ICRA 2026 competition on bimanual garment folding. The system placed 1st of 62 teams in the online (simulation) round and 2nd in the real-world final. It improves a vision-language-action (VLA) policy with a reinforcement-learning loop. The policy is its own value function: the same network that predicts actions also predicts success, progress, and a few task-relevant future quantities, and those predictions drive advantage estimation, live failure detection, and candidate selection. The work mostly recombines existing RL ideas with engineering and optimization contributions that can be used together as one recipe or individually: AWR + RECAP combined for flow-matching VLA; an asynchronous distributed training / rollout pipeline through HuggingFace Hub; inference-time hyperparameters optimization via Thompson sampling; a sim-to-real recipe with camera-alignment tooling, heavy augmentation and DAgger-like HIL data collection.
Sources
- LeHome: A Simulation Environment for Deformable Object Manipulation in Household Scenarios
- NVIDIA Isaac Sim: Enabling Scalable, GPU-Accelerated Simulation for Robotics
- Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
- $\pi^{*}_{0.6}$: a VLA That Learns From Experience
- Task adaptation of Vision-Language-Action model: 1st Place Solution for the 2025 BEHAVIOR Challenge
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- AWAC: Accelerating Online Reinforcement Learning with Offline Datasets
- A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning
- Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout Replay
- Sigmoid Loss for Language Image Pre-Training
- Gemma: Open Models Based on Gemini Research and Technology
- Flow Matching for Generative Modeling
- FAST: Efficient Action Tokenization for Vision-Language-Action Models
- Exclusive Self Attention
- A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation
- High-Dimensional Continuous Control Using Generalized Advantage Estimation
- Real-Time Execution of Action Chunking Flow Policies
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving