Leveraging Human-In-The-Loop Demonstrations in Reinforcement Learning for Digital Twin-Driven Robot Flexibility
Yuzhu Sun, Mien Van, Nguyen Minh Nhat, Stephen McIlvanna, Sean McLoone
Queen's University Belfast
cs.RO, cs.SY, eess.SY
Submitted: 2026-10-08
Updated: 2026-10-08
Code: https://github.com/ultralytics/ultralytics
The gist: The gist The proposed framework combines a digital twin, reinforcement learning, and human demonstrations to create an online training system for collaborative robots that can adapt to changes in
Terminology
Summary
The gist The proposed framework combines a digital twin, reinforcement learning, and human demonstrations to create an online training system for collaborative robots that can adapt to changes in their physical workspace.
Digital Twin and Hardware-in-the-Loop System
The framework utilizes a hardware-in-the-loop DT that stays synchronized with the physical workspace and resumes RL training when a change causes the current policy to fail The digital twin is synchronized with the physical system in real time through camera feeds, allowing the virtual robot to update its observations and policy from real-world feedback If the twin detects that the current policy fails an obstacle-avoidance task after a workspace change, the RL agent resumes training in simulation The physical setup includes the Ufactory Xarm5 robot, a ZED 2i depth camera, physical obstacles and goal objects, and an HTC VIVE system Camera images update the DT in two stages: we detect goals and obstacles, then map their positions into PyBullet coordinates The linear form of Eq. (1) rests on the following assumptions: (i) the camera pose is fixed with respect to the workspace, and the perspective transformation removes perspective distortion, so pixel coordinates in the transformed view are proportional to metric coordinates on the table plane; (ii) all objects rest on the table plane, so the height hz fully determines the z coordinate of an object’s center; and (iii) residual lens distortion after calibration is negligible over the taped workspace area.
Dual Actor Framework for Demonstrations
The framework employs a dual actor framework in which demonstrations train a separate behaviour-cloning actor and guide the RL actor only through the critic target Unlike methods that add an imitation loss to the RL actor, this approach allows it to explore beyond poor demonstrations. The second actor learns only from VR demonstrations through behaviour cloning (BC) with objective JBC (ξ): JBC (ξ) = N X D i=1 πBC (si; ξ) − ai 2. The value function is then updated via: Vˆ (st+1) ← min i=1,2 Qθi(st+1, at+1). This approach helps the agent exploit the BC actor’s path more while exploring further when the agent’s performance surpasses that of the human.
Integration of Human Demonstrations
Human demonstrations are captured in virtual reality (VR) and transferred to the DT, where they guide the virtual robot. The DT evaluates each demonstrated step using the task reward and labels the resulting transitions automatically This allows an actor-critic learner to use human demonstrations without manual reward annotation. The dual actor framework adds no direct imitation loss to the RL actor. Instead, a separate behaviour-cloning (BC) actor proposes an alternative action.
Performance with Non-Optimal Demonstrations
With fixed non-optimal demonstrations, the proposed framework reached over 80% mean final training success, while two imitation-loss baselines rarely completed the task The results show that SACBC and SACBC + Q-Filter converge in none of the six seeds and end with training success rates of 7.3% and 3.5%, whereas DA-SACBC reaches 80.5% and 85.0% for the two target forms. With non-optimal demonstrations, the DASACBC policies reach a mean deterministic evaluation success of 83.3% and 100.0%, against 0.0% for SACBC and 16.7% for SACBC + Q-Filter.
Generalization to More Complex Scenarios
The framework was tested in three additional scenes derived from the nominal scene by one fixed geometric change in simulation. In all three scenes the existing policies fail, motivating retraining of the tested policies With the pillar (S1), DA-SACBC was the only method to converge in any seed and reached a mean deterministic evaluation success of 83%, against 50% for SACBC and 0% for SAC. With the moved goal (S3) and complete demonstrations, DA-SACBC and SAC without demonstrations reached the same mean deterministic evaluation success (67%), both above SACBC (33%), and DA-SACBC had the lowest mean number of sampled-contact episodes of the three.
Human Effort Measurement
Operator effort and workload were measured for one person in one session. For complete demonstrations, effort was highest (90), followed by mental demand and performance (75 each). The partial set took fewer attempts and less time, while unsuccessful trials in the DT posed no collision risk to the physical robot.
The framework currently assumes that relevant workspace changes appear in the RL observations, such as obstacle size; changes outside that representation and gaps in camera coverage may limit adaptation. Future work will address obstacles that move during task execution The state must include velocity estimates or short-horizon motion predictions, and the policy must be trained across a range of motions rather than retrained after every change. Because a trajectory judged safe in the DT may be unsafe by the time the physical robot executes it, contact-penalty rewards should be complemented by safe-RL mechanisms such as control barrier functions. Lidar could improve scene sensing, while force/torque sensing could help detect contact with coworkers. Moving obstacles also make demonstrations harder to time; we will integrate VR headsets into the DT so operators can see the virtual robot and obstacles from their own viewpoint.
The paper is supported by the European Commission through the Horizon Europe project HARMONICA under grant agreement No 101294331 The remainder of this paper is organized as follows. Section II reviews related work. Section III describes the DT, and Section IV details the dual actor framework. Section V presents the experiments and discussion; Section VI concludes the paper
The references section lists numerous prior works related to reinforcement learning, digital twins, and imitation learning in robotics. The authors compare their proposed framework against various baselines, including those that add an imitation loss to the actor objective. The comparison shows that the dual actor framework achieves a much higher final success rate than two methods that add an imitation loss to the actor. This work is supported by the European Commission through the Horizon Europe project HARMONICA under grant agreement No 101294331. The authors are with the School of Electronics, Electrical Engineering and Computer Science, Queen’s University Belfast, Northern Ireland, UK.
The paper is supported by the European Commission through the Horizon Europe project HARMONICA under grant agreement No 101294331. The authors are with the School of Electronics, Electrical Engineering and Computer Science, Queen’s University Belfast, Northern Ireland, UK.
The paper is supported by the European Commission through the Horizon Europe project HARMONICA under grant agreement No 101294331.
Improvements for AI systems
-
Bold header: Hardware-in-the-loop DT for Resumed RL Training. This system can
resume training after a change in the physical workspace
by synchronizing a PyBullet digital twin with real-time camera feeds, allowing the RL agent toupdate its observations and policy from real-world feedback.
-
Bold header: Dual Actor Framework for Robust Learning. The proposed framework utilizes
a dual actor framework in which demonstrations train a separate behaviour-cloning actor and guide the RL actor only through the critic target,
enabling the RL actorto explore beyond poor demonstrations
without adding a direct imitation loss to its objective. -
Bold header: Real Human VR Demonstration Exploitation. The system can utilize
real human VR demonstrations collected in virtual reality (VR)
where the DTlabels each demonstrated step using the task reward and labels the resulting transitions automatically,
achieving high success rates even withdemonstrations that never reach the goal.
-
Bold header: Adaptive Policy Resumption During Workspace Change. The framework achieves adaptation when a change occurs by allowing RL training to resume, as it addresses the gap where traditional DTs
may no longer reflect the physical scene
andthe model or policy may need to be updated.
-
Bold header: Efficient Demonstration Guidance Without Direct Loss. By employing a dual actor structure, the system exploits demonstrations without adding a direct imitation loss to the RL actor, contrasting with methods that add an
imitation loss on the RL actor,
which can limit exploration of better solutions when usingnoisy, non-optimal demonstrations.
Sources
- AWAC: Accelerating Online Reinforcement Learning with Offline Datasets
- AW-Opt: Learning Robotic Skills with Imitation and Reinforcement at Scale
- Robust Multi-Modal Policies for Industrial Assembly via Reinforcement Learning and Demonstrations: A Large-Scale Study
- Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations
- Demonstration-Guided Reinforcement Learning with Learned Skills
- Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards
- Soft Actor-Critic Algorithms and Applications
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving