Real2Gym: Building Gyms from Videos, Bringing Skills to Robots
cs.CV
Submitted: 2026-09-29
Updated: 2026-09-29
Project page: https://real2gym.github.io
Terminology
Sources
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents
- Challenges of Real-World Reinforcement Learning
- Leave no Trace: Learning to Reset for Safe and Autonomous Reinforcement Learning
- CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation
- Inner Monologue: Embodied Reasoning through Planning with Language Models
- VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
- Agent as Policy for Robotic Manipulation
- DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset
- MoGe-3: Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric Refinement
- RoboGSim: A Real2Sim2Real Robotic Gaussian Splatting Simulator
- ASPIRE: Agentic /Skills Discovery for Robotics
- MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations
- Splatting Physical Scenes: End-to-End Real-to-Sim from Imperfect Robot Data
- RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots
- ProgPrompt: Generating Situated Robot Task Plans using Large Language Models
- Sim-to-Real: Learning Agile Locomotion For Quadruped Robots
- ManiSkill3: GPU Parallelized Robotics Simulation and Rendering for Generalizable Embodied AI
- Reconciling Reality through Simulation: A Real-to-Sim-to-Real Approach for Robust Manipulation
- Voyager: An Open-Ended Embodied Agent with Large Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models