Experience-Based Feasibility-Aware Generative Adversarial Imitation from Observation under Embodiment Mismatch

arXiv:2610.01171 · cs.RO · Submitted 2026-10-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Experience-Based Feasibility-Aware Generative Adversarial Imitation from Observation under Embodiment Mismatch".

Dev: Imitation from observation (IfO) learning robot behaviors from state-only demonstrations is enhanced by introducing Experience-Based Feasibility-Aware Generative Adversarial Imitation from Observation (EF-GAIfO),

Rosa: First, who's behind it and why it matters.

Paper summary: Dev: So, to summarize this paper on Experience-Based Feasibility-Aware Generative Adversarial Imitation from Observation under Embodiment Mismatch, the main idea is that imitation from observation can be improved by estimating whether demonstrated state transitions are feasible for the robot based on its own accumulated experience.

Rosa: That's right; they propose EF-GAIfO, which avoids needing explicit dynamics models or large prior exploration datasets by instead using the robot’s experience to judge feasibility. The key claim is that this notion of feasibility evolves alongside policy learning, meaning as the robot experiences more state transitions, the set of feasible motions it can successfully follow gets progressively larger.

Taro: What makes this significant for autonomy is that it allows the policy to learn from demonstrations while simultaneously ensuring those demonstrations are physically realizable given the robot's current capabilities and past interactions.

Dev: It matters because traditionally, if a demonstration shows a motion the robot simply can't do, it can degrade performance; EF-GAIfO claims that by weighting the discriminator based on this estimated feasibility, the policy learns to prioritize demonstrations that are actually supported by the robot’s experience.

Rosa: So it directly addresses the embodiment mismatch problem by making sure we only learn from feasible trajectories, which is a big step forward from just trying to imitate raw state sequences without regard for physical possibility.

Taro: It also implies that the robot's own interaction with the environment becomes an active part of validating what kind of behaviors are appropriate to adopt during imitation learning.

Dev: If we look at the mechanism, they train a state-transition reconstruction model using experience data to evaluate demonstrations via their reconstruction error, and this error is then converted into a feasibility weight between zero and one.

Rosa: That sounds like a very clever way to quantify feasibility without needing a perfect physics engine or explicit dynamics model upfront.

Taro: The paper suggests that this approach allows the system to incorporate more demonstrations as the policy improves, creating a self-improving feedback loop regarding what is feasible within the robot's operational envelope.

Conclusion: Rosa: Looking at the title, Experience-Based Feasibility-Aware Generative Adversarial Imitation from Observation under Embodiment Mismatch, it really highlights that the novelty here is tying feasibility estimation directly into the imitation learning process using robot experience. The authors are Yoshiki Takebayashi and Giovanni Perantoni among others.

Dev: I think the implication boils down to making imitation from observation much more robust when you're dealing with real-world robots that have different physical constraints than the creators of the demonstrations, because it dynamically filters out infeasible paths.

Taro: From a broader autonomy view, this suggests that we can build imitation systems that are inherently more conservative about what they adopt unless their own operational history supports those actions, which is essential for safety in unpredictable environments.

Rosa: So in simple terms, the paper shows how to use a robot's past actions to judge if a new demonstration step is actually possible for the robot right now, leading to a policy that only learns from physically viable examples.

Dev: It moves beyond simply learning what humans did; it starts learning what the robot *can* do based on its own accumulated knowledge of state transitions, which is crucial for deployment longevity.

Taro: If this approach scales well—and I'm hoping it does—it opens up possibilities for deploying imitation learning in hardware that has significant embodiment mismatches, like different sized or shaped robots, where traditional methods struggle.

Rosa: That’s the big picture: we are moving toward imitation systems that are not just good at mimicking behavior but are also intrinsically aware of their own physical limitations based on what they have actually experienced.

Yoshiki Takebayashi, Giovanni Perantoni, Hikaru Sasaki, Matteo Saveriano, Takamitsu Matsubara

Nara Institute of Science and Technology · University of Trento

cs.RO

Submitted: 2026-10-01

Updated: 2026-10-01

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 87/100

The gist: Imitation from observation (IfO) learning robot behaviors from state-only demonstrations is enhanced by introducing Experience-Based Feasibility-Aware Generative Adversarial Imitation from

Key concepts

Imitation from Observation (IfO)
A learning method where a robot learns behaviors by observing demonstrations that only show states and observations, without needing explicit knowledge of the underlying physics. It tries to mimic expert actions directly from sensory inputs.
State-Transition Reconstruction Model
This model is trained using the robot's own collected experience data during policy learning. Its job is to learn how to accurately recreate the sequences of state transitions that actually occurred when the robot was following its current policy.

Terminology

Summary

Imitation from observation (IfO) learning robot behaviors from state-only demonstrations is enhanced by introducing Experience-Based Feasibility-Aware Generative Adversarial Imitation from Observation (EF-GAIfO), a framework that estimates whether demonstrated state transitions are feasible for the robot based on its own accumulated experience, thereby reducing the influence of robot-infeasible motions.

How it works

The core idea is to estimate feasibility from what the robot has actually experienced rather than relying on explicit dynamics models or large prior exploration datasets. This mechanism ensures that as the policy improves and experiences a broader range of state transitions, the feasible region is progressively expanded, allowing more demonstrations to be incorporated into learning.

  1. The feasibility estimation is based on training a state-transition reconstruction model using the robot’s experience data in GAIfO. This model learns to reconstruct state transition sequences generated by the robot during policy learning, minimizing a reconstruction loss (Equation 3).

  2. The feasibility of demonstrated motions is then evaluated based on their reconstruction error. The error, calculated as Equation 4, measures the difference between the original and reconstructed states across all dimensions and time steps.

  3. This error is converted into an estimated feasibility weight, denoted as a function of the reconstruction error: fw(zt, R(zt)) = clip τhigh − e(zt, R(zt)) τhigh − τlow, 0, 1 (Equation 5). This weight is clipped to the range [0, 1] using thresholds that are updated based on reconstruction errors.

How it works (Continued)

The estimated feasibility is then incorporated into the Generative Adversarial Imitation from Observation (GAIfO) objective to control the contribution of each demonstration. The policy is trained using a discriminator D, which is optimized to distinguish between state transitions generated by the policy and those from expert demonstrations.

  1. The overall objective function for EF-GAIfO modifies the standard GAIfO objective by weighting the discriminator's training based on feasibility: max D Eπe [fw(zt, R(zt)) log D(st, st+1)] + Eπ [log (1 − D(st, st+1))].

  2. By incorporating this feasibility weight into the discriminator’s objective for demonstration data, the policy is trained to prioritize demonstrations with high feasibility, effectively reducing the influence of demonstrations with low estimated feasibility.

How it works (Continued)

The learning process involves a joint update of three components: the policy, the reconstruction model, and the discriminator. The algorithm iteratively collects experience transitions using the current policy, updates it via PPO (lines 3–7), updates the reconstruction model R using collected experience data (line 8), computes feasibility weights fw (lines 9–11), and finally updates the discriminator D using these weighted demonstrations and robot experience (line 12).

How it works (Continued)

The thresholds τhigh and τlow for the feasibility weight are updated based on the reconstruction errors of demonstration samples, with a percentile parameter q gradually increased during training. This dynamic adjustment allows the system to adapt as policy experience grows, leading to experience-based feasibility estimation that becomes more discriminative over time.

How it works (Continued)

The paper validates EF-GAIfO on two primary tasks: controlled point-mass experiments and a real quadruped robot performing an object-reaching-and-grasping task. In the point-mass environment, EF-GAIfO achieved the highest success rate and the smallest Hausdorff distance, outperforming GAIfO and WGAIfO by selectively imitating feasible trajectory modes. In the quadruped robot experiment, EF-GAIfO consistently achieved higher success rates across all settings compared to GAIfO and WGAIfO, successfully executing object-picking behavior on the real robot.

How it works (Continued)

The results demonstrate that the feasibility estimates become more discriminative as robot experience accumulates, where weights assigned to feasible demonstration transitions increase substantially while those for infeasible transitions remain low. This outcome confirms that EF-GAIfO can reduce the influence of robot-infeasible demonstrations without requiring explicit dynamics models or separate prior exploration datasets. Furthermore, the study shows that the learned policy can be successfully executed on a real robot, proving its practical applicability in handling embodiment mismatch.

How it works (Continued)

Future work suggests extending this estimator to condition feasibility not only on robot motion but also on environmental context, such as terrain and obstacles, to address environment-dependent feasibility. This involves jointly representing robot state transitions and environmental observations, including visual and depth information.

The gist

EF-GAIfO estimates the feasibility of state-only demonstration transitions using experience data collected during GAIfO-based policy learning and reduces the influence of demonstrations with low estimated feasibility.

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed the proposed framework, Experience-Based Feasibility-Aware Generative Adversarial Imitation from Observation (EF-GAIfO). This method fundamentally addresses the critical embodiment mismatch problem in imitation learning from observation (IfO) by replacing explicit dynamics models with an experience-based feasibility estimator.

Here are the specific improvements and capabilities this system enables:


) Experience-Based Feasibility-Aware Generative Adversarial Imitation from Observation (EF-GAIfO) System Enhancements:

  1. The core improvement lies in the integration of a learned, experience-driven reconstruction model as a proxy for robot feasibility.

  2. The discriminator's objective function is dynamically weighted by this feasibility score, allowing the policy to learn selectively and adaptively from demonstrations.

) Specific Improvements and Enabled Capabilities:

  1. The system learns a state-transition reconstruction model (e.g., a Variational Autoencoder) on-the-fly using the robot's own generated experience during policy training.

  2. It computes an explicit feasibility weight, which is a function of the reconstruction error between demonstration transitions and the robot's experienced transitions, clipped between learned thresholds.

  3. This feasibility weight is directly incorporated into the Generative Adversarial Imitation Learning (GAIfO) objective function:

  4. The discriminator is trained to prioritize imitation of demonstrations that exhibit low reconstruction error (i.e., are physically feasible for the robot's current embodiment and dynamics).

) What the Improved AI System Can Do (Specific Outcomes):

  1. The system can learn complex manipulation or locomotion behaviors from human demonstrations even when those behaviors involve motions that are physically impossible or highly unstable for a specific robot (e.g., high pitch angles, fast movements).

  2. It enables Adaptive Imitation Learning: As the robot accumulates more experience (i.e., as its state-transition reconstruction model becomes more accurate), it progressively learns to recognize which parts of the demonstration dataset are feasible, allowing it to incorporate previously infeasible demonstrations later in training without degradation.

  3. The resulting policy will exhibit significantly higher success rates and lower trajectory discrepancies (Hausdorff distance) compared to standard IfO methods (like GAIfO or WGAIfO), especially under conditions of high feasibility mismatch (e.g., when the demonstration involves motions exceeding the robot's physical limits).

  4. For real-world deployment on a quadruped robot, the system can reliably execute object-reaching and grasping tasks with higher success rates (up to 80% reported in experiments) because it actively suppresses robot-infeasible actions that would otherwise lead to catastrophic failures (like excessive pitch causing ground contact).

Sources

Related papers