Does Continual Imitation Learning Remain Grounded? A Language-Perturbed Benchmark for Robotic Task Retention

arXiv:2610.00542 · cs.RO · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Does Continual Imitation Learning Remain Grounded? A Language-Perturbed Benchmark for Robotic Task Retention".

Dev: Continual imitation learning evaluates whether a robot can learn new knowledge without forgetting previously learned skills,

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: So, we're looking at a paper titled "Does Continual Imitation Learning Remain Grounded? A Language-Perturbed Benchmark for Robotic Task Retention." It seems like they are really digging into the core problem of whether robots can keep learning new stuff without forgetting the old stuff, specifically when those instructions are given in language.

Dev: I agree with Rosa; I'm interested in how they frame this issue. The title suggests they’re testing if just achieving a successful action is enough, or if that action is actually tied to the language instruction itself across different learning stages.

Taro: From my perspective, it sounds like they are trying to find the boundary where a policy relies on actual knowledge versus relying on learned scene cues or memorized patterns when it gets new language input.

Rosa: Exactly, and the authors set up this benchmark protocol to see how language-guided behavior shifts as the robot learns more tasks. It’s about finding that gap between competence and true language grounding.

Dev: So, essentially they are building a way to measure if a policy is just guessing based on scene context or if it actually understands the meaning of what we're saying.

Taro: That’s the crux of it; when things get tricky or the world doesn't cooperate, does the robot just keep doing what it used to do because that’s easier, or can it adapt its behavior based on a new instruction?

Rosa: And this paper seems to be testing that adaptation directly by changing instructions in different ways. It’s not just about success; it’s about the robustness of the language connection itself.

The paper's summary: Dev: The summary of "Does Continual Imitation Learning Remain Grounded? A Language-Perturbed Benchmark for Robotic Task Retention" focuses on introducing a temporal evaluation protocol to see how language-guided behavior evolves when a policy moves from one task to the next during continual learning.

Rosa: They construct several instruction variants, like paraphrases that keep the meaning but change the wording, minimal changes that alter just one part of the goal, and even collisions where the request is completely impossible for the current scene.

Taro: That’s interesting because they are explicitly setting up tests to see if a policy preserves its behavior when the meaning stays exactly the same versus when it's subtly altered.

Dev: They evaluate these policies against original instructions, paraphrases, minimal contrasts, and incompatible collisions after training at successive checkpoints. This allows them to track changes over time without having to retrain every time they test something new.

Rosa: The key finding they highlight is that strong continual-learning performance in the standard sense doesn't automatically mean the behavior remains reliably guided by language instructions as learning progresses.

Taro: So, if a policy gets really good at switching tasks but loses its connection to what we are saying, that’s what this paper is showing us is possible. It suggests that relying on scene cues or memorized structure can replace direct language processing when things get tough.

Dev: That points toward a real challenge in deploying these systems reliably in dynamic environments where the underlying scene might change unexpectedly.

The paper's improvements: Rosa: The paper proposes several diagnostic metrics to move beyond simple task success, focusing on semantic robustness and goal adaptation rather than just whether the robot succeeded or failed a single step.

Dev: They use goal-based metrics like Original Goal Persistence and Goal Switch Accuracy, which are designed to check if the policy actually achieved the modified goal instead of just executing some other action.

Taro: I like that they also have behavioral diagnostics for when instructions are incompatible, things like Action Initiation Rate, which tells us how much the robot tries to move or act when it gets a weird request.

Rosa: On top of those, they look at action-level metrics such as Action Divergence and Language Margin to see if the expert action is more likely under the correct instruction than under a perturbed one.

Dev: It seems like their improvement is in creating a suite of complementary metrics that diagnose different aspects of language grounding—semantics, goal execution fidelity, and behavioral sensitivity—all evaluated across the continual learning checkpoints.

Taro: That comprehensive diagnostic approach is important because it lets us pinpoint exactly *why* a policy might be failing to follow instructions in a specific situation, whether it's a semantic slip or just a bad execution path.

Conclusion: Rosa: So, to wrap up this discussion on "Does Continual Imitation Learning Remain Grounded? A Language-Perturbed Benchmark for Robotic Task Retention," the central conclusion is that conventional task retention doesn't guarantee semantic invariance or reliable goal switching when language is involved.

Dev: They’ve shown that policies can maintain high continual learning performance while still relying on scene cues or memorized structures instead of actually grounding their actions in the language instruction.

Taro: That implies a significant risk for autonomous systems operating in complex, evolving environments where they might default to old habits even when asked to do something new.

Rosa: It’s a necessary caution because it tells us we need more than just tracking task success; we need measures that test the actual connection between what's said and what the robot does.

Dev: This benchmark protocol gives us a much clearer lens through which to study these issues over time, allowing us to see degradation in language robustness as tasks pile up.

Taro: It opens the door for developing stronger goal adaptation mechanisms and scene-compatible instruction handling, moving beyond just simple retention toward true linguistic understanding in robotics.

Siddeshwar Raghavan, Ziqin Yuan, Fengqing Zhu, Byung-Cheol Min

Department of Electrical and Computer Engineering, Purdue University · Department of Computer and Information Technology, Purdue University · Computer Science and Intelligent Systems Engineering, Indiana University Bloomington

cs.RO

Submitted: 2026-09-30

Updated: 2026-09-30

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 83/100

The gist: Continual imitation learning evaluates whether a robot can learn new knowledge without forgetting previously learned skills, but this paper introduces a benchmark protocol to study how

Key concepts

Continual Imitation Learning
This is a method where a robot learns new tasks sequentially without forgetting what it learned previously. The goal is to maintain a skill set while acquiring more abilities, ensuring the robot can handle many different manipulation tasks over time.
Language Grounding
This refers to how well the robot's actions match the meaning of the language instruction it receives. The paper investigates whether a robot's behavior stays correct when instructions are slightly rephrased or become contradictory, testing if it truly understands the language.
Paraphrase Evaluation
This involves testing a policy against a new instruction that has the same meaning but different wording. The researchers check if the robot still achieves the original goal despite this change, which tests whether it relies on literal language or actual task understanding.
Scene-Incompatible Collision
This is a test where the new instruction conflicts with what is currently happening in the robot's environment. Instead of checking for success, researchers look at how the robot responds to this conflict, probing its robustness when faced with impossible requests.

Terminology

Summary

Continual imitation learning evaluates whether a robot can learn new knowledge without forgetting previously learned skills, but this paper introduces a benchmark protocol to study how language-guided behavior changes as robotic policies learn successive tasks, revealing that strong continual-learning performance does not always translate to reliable language grounding.

The gist: Strong continual-learning performance does not always translate to reliable language grounding.

Introduction and Motivation

Robot policies conditioned on language must acquire new behaviors while retaining previously learned skills through streams of incoming manipulation tasks, a challenge studied by continual imitation learning methods like LIBERO. However, metrics such as task success do not establish whether retained behavior remains appropriately guided by language because policies might rely on scene cues, object associations, or memorized task structure rather than the instruction itself. The paper posits that Success under a paraphrase does not necessarily indicate language use, and a behavioral change does not guarantee completion of the modified goal. This distinction is crucial because it raises the question: as new tasks are learned, does the connection between language and behavior remain stable?

Benchmark Protocol Design

The researchers introduce a temporal evaluation protocol to study how language-guided behavior changes across successive continual learning checkpoints. The protocol evaluates policies saved at successive stages against four instruction conditions:

  1. The original instruction.

  2. A paraphrase with the same meaning, which is evaluated against the original goal, testing for semantic invariance.

  3. A minimal semantic change that requests a different goal, evaluated to see if it induces an appropriate response.

  4. A collision that makes the request incompatible with the current scene, where behavioral outcomes are examined instead of success targets.

Instruction Construction and Validation

The construction of instruction variants is controlled using an LLM (gpt-4.1-mini) to propose natural-language rewrites and candidate interventions for every semantic slot. The protocol enforces strict validation criteria:

(5) Paraphrase:

The paraphrase is evaluated against the original goal gj.

(6) Validated minimal contrast:

A minimal contrast is retained only if its required entities occur in Ej, its modified goal uses predicates supported by Pj, and the goal annotation can be parsed and registered by the simulator evaluator. This check ensures the modified goal is scene-compatible.

(8) Incompatible collision:

A collision l×(k)j also changes one slot, but references an absent entity or unsupported predicate. These are treated as sensitivity probes because no correct alternate rollout exists.

Continual Grounding Metrics

The protocol employs specific metrics to diagnose language grounding:

  1. For executable goals, they use goal-based metrics like Original Goal Persistence (OGP) and Goal Switch Accuracy (GSA). For instance, GSA requires the policy to achieve the modified goal rather than merely change its actions.

  2. For incompatible instructions, they use behavioral diagnostics such as Original-Goal Persistence (OGP), Canonical Goal Activation Rate (CGAR), and Action Initiation Rate (AIR).

  3. Action-level metrics include Action Divergence, which measures the normalized Euclidean distance between the mean actions predicted for the same observation under the two instructions, and Language Margin, which indicates if the expert action is more likely under the correct instruction than under the perturbed one.

Experimental Results Summary

The experiments compare eight representative continual imitation learning methods across LIBERO-Goal. Key findings include:

(Table I)

Methods with stronger continual-learning performance generally perform better under paraphrasing. DMPEL achieved the highest Original and Paraphrase AUC, but still exhibited a nonzero paraphrase gap.

(Table II)

Strong benchmark performance does not necessarily imply reliable goal switching. DMPEL achieves the highest Target SR (30.91) and GSA (21.68), followed by TAIL, IsCiL, and ER, while L2M and SeqFT have GSA below 4%.

(Figure 4)

Incompatible instructions rarely cause policies to stop acting, with AIR remaining between 83 to 92% across methods. Stronger methods are also more likely to fall back to a previously learned goal.

Temporal Evaluation Analysis

The study examines the effect of evaluation timing. Figure 6 shows that final-only and temporal evaluation can tell different stories. For instance, PackNet has the smallest Final Gap (1.16) but a larger AUC Gap (6.41), suggesting that final-only evaluation can miss earlier instability, while temporal averages can hide poor final-stage performance. This demonstrates that temporal averages can conceal pronounced degradation at the final checkpoint.

Conclusion and Limitations

The paper concludes that conventional task retention does not guarantee semantic invariance, reliable goal switching, or appropriate behavior under scene-incompatible instructions.

Improvements for AI systems

Here are the specific improvements to AI systems derived from this research, and what those improved systems can achieve:


The core improvement is shifting evaluation metrics from simple task success to comprehensive language grounding diagnostics across continual learning stages.

  1. Improve Continual Learning Evaluation for Language Grounding:

  2. Implement Semantic Invariance Testing via Paraphrase Evaluation:

  3. Develop Robust Goal Adaptation Mechanisms:

  4. Establish Scene-Incompatible Instruction Handling Capabilities:

  5. Integrate Multi-Modal Diagnostic Feedback Loops (Action and Outcome):

Specific Improvements and Capabilities:

The improved AI system, utilizing these diagnostics, can achieve the following specific capabilities:

The improved AI system can perform the following specific actions:

  1. Perform Continual Learning Evaluation for Language Grounding using a temporal diagnostic protocol, allowing researchers to track how language guidance changes across successive learning stages without modifying the training process.

  2. Test and quantify semantic invariance by comparing policy behavior under original instructions versus meaning-preserving paraphrases (measuring AUC Gap and signed paraphrase gap), providing a rigorous measure of whether retained skills remain grounded in the instruction's semantics.

  3. Develop policies with superior Goal Adaptation capabilities, ensuring that when the requested goal changes, the policy correctly executes the new objective (measured by Goal Switch Accuracy, GSA), rather than merely executing a default or original action.

  4. Establish reliable and sensible behavior under scene-incompatible or contradictory instructions (e.g., Pick up an object that does not exist in the scene), ensuring policies respond appropriately—either by continuing toward a known goal, initiating motion, or changing action distribution—instead of failing silently.

  5. Integrate multi-modal diagnostic feedback loops by reporting complementary metrics like Action Divergence and Language Margin alongside outcome metrics (like Original Goal Persistence). This allows for a deeper understanding of why a policy might succeed or fail, distinguishing between failures in language interpretation versus limitations in task execution capability.

Abstract

Continual imitation learning evaluates whether a robot can learn new knowledge without forgetting previously learned skills. However, retaining task performance does not ensure the behavior remains grounded in language because policies may rely on scene cues, object associations, or memorized task structure. We introduce a benchmark protocol to study how language-guided behavior changes as robotic policies learn successive tasks. We construct meaning-preserving and meaning-changing instruction variants for the Goal, Spatial, Object, and Long suites of LIBERO. Policy experiments focus on LIBERO-Goal, evaluating Original and Paraphrase instructions after each continual-learning stage. We compare representative continual imitation learning methods under their original assumptions while separating task competence from language sensitivity. The proposed diagnostics complement standard learning and forgetting metrics by measuring semantic robustness, goal adaptation, and language sensitivity. Results show that strong continual-learning performance does not always translate to reliable language grounding, and our diagnostics help determine whether retained skills remain correctly guided by their instructions. Additional materials are available at https://sites.google.com/view/stillgrounded

Sources

Related papers