Memorize, Adapt, Ignore: Diagnosing Robot Learning Mechanisms under Training Data Variation

arXiv:2609.38401 · cs.RO · Submitted 2026-09-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "Memorize, Adapt, Ignore".

Rosa: Training data variation, whether through designing a domain randomization (DR) scheme in simulation or curating demonstrations for imitation learning,

Dev: First, who's behind it and why it matters.

Paper summary: Rosa: So to recap what we’ve covered in this first part of our discussion about "Memorize, Adapt, Ignore: Diagnosing Robot Learning Mechanisms under Training Data Variation," the central thesis is that understanding how training data variation shapes a model's internal representations is crucial for improving the robustness of robotic manipulation policies.

Dev: They use a diagnostic framework based on representation kernels, specifically the empirical neural tangent kernel or eNTK, to isolate and track how this data variation influences what the network learns during training.

Taro: The paper claims that we can distinguish between three learning regimes: memorizing different situations with insufficient variation, adapting to the situation at hand with sufficient variation, or ignoring task-irrelevant factors.

Rosa: They achieved this by systematically varying factors like object size, color, and type in simulation and using structured randomization in imitation learning to structure the training data distributions.

Dev: The methodology involves three stages: controlled training where they structure joint distributions as either confounded or independent, creating probe sets that sweep through variations of target factors while keeping others static, and then computing representation kernel diagnostics on model checkpoints and probes.

Taro: They compare the eNTK with last-hidden-layer and output kernels to show why the eNTK is a reasonable default choice because it captures parameter-update effects which are key for predicting cross-input effects during training.

Rosa: The core claim is that these diagnostic metrics, like phase transition heatmaps and the Effective rank, allow researchers to move past simple success rates and gain insights into whether a policy is adapting or just memorizing.

Dev: They also define metrics like the Factor Sensitivity Ratio to quantify how strongly different factors influence each other within the model's representation space.

Taro: Essentially, they provide a toolkit for practitioners to diagnose *how* the learning mechanism is operating, rather than just observing what the policy achieves in terms of final performance metrics.

Rosa: This matters because it shifts the focus from just tweaking randomization parameters by trial and error to understanding exactly what kind of variation is needed for genuine generalization.

Dev: It’s a framework that helps us understand the underlying mechanics of how data composition directly affects the learned behavior.

Taro: If this framework works as intended, it gives us a structured way to approach improving autonomy when the environment misbehaves because we can pinpoint where the model's generalization is failing.

Conclusion: Rosa: Thinking about the title "Memorize, Adapt, Ignore: Diagnosing Robot Learning Mechanisms under Training Data Variation," it really encapsulates the paper’s main contribution—using representation kernels to dissect those three distinct learning modes based on data variation.

Dev: That distinction is important because it moves beyond just saying a policy works or doesn't work; it tells us *why* it’s succeeding or failing under different conditions.

Taro: The authors, Ke Zhang, Danica J. Sutherland, and Chao Liu, have provided a method for practitioners to analyze the internal learning process without needing to run massive numbers of new experiments every time they want to test a hypothesis.

Rosa: Their implication is that we can start designing training protocols that are explicitly aimed at promoting adaptation instead of just brute-force memorization when deploying policies on real hardware.

Dev: It means we move away from blind trial and error toward a more informed approach to building policies that are inherently more robust against the kinds of unseen situations we encounter in the field.

Taro: For autonomy, this suggests that future AI development shouldn't just focus on making models perform better on known test sets, but on understanding the mechanisms of how they generalize when those tests change.

Rosa: It means we need to treat data variation not as a nuisance to be randomly added, but as a carefully managed lever to steer the model toward building genuine generalization capabilities.

Dev: In short, this paper offers a way to diagnose the learning regime so we can predict policy behavior before it hits the real world.

Taro: It gives us a diagnostic lens for understanding complex AI systems under uncertainty, which is something we desperately need when dealing with unpredictable real-world interactions and misbehavior.

Ke Zhang, Danica J. Sutherland, Chao Liu

PRIME Robotics Lab · Department of Mechanical Engineering, The University of British Columbia · Department of Computer Science, University of British Columbia

cs.RO

Submitted: 2026-09-29

Updated: 2026-09-29

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 82/100

The gist: Training data variation, whether through designing a domain randomization (DR) scheme in simulation or curating demonstrations for imitation learning, is a primary lever for improving the robustness

Key concepts

Representation Kernels
These are mathematical tools used to measure how similar a network considers two different inputs during training. They help researchers see exactly how the model's internal understanding of the world changes as the training data is varied, revealing whether the model is memorizing or adapting.
Phase Transition Heatmap
This visual tool shows how the kernel structure changes as training variation increases. A 'clear block structure' suggests a memorization regime where situations are treated separately. As variation grows, this structure flattens into an 'adapt regime,' indicating the model is learning to generalize.
Factor Sensitivity Ratio (FSR)
FSR measures how much the representation changes when one training factor is varied compared to another. A high FSR for a task-relevant factor shows that the model's internal understanding is highly sensitive to that specific variation, which is crucial for determining if it's focusing on the right information.

Terminology

Summary

Training data variation, whether through designing a domain randomization (DR) scheme in simulation or curating demonstrations for imitation learning, is a primary lever for improving the robustness of robotic manipulation policies.

How it works

The core diagnostic framework introduces representation kernels to isolate and track how training data variation shapes a model’s internal representations during training. The default choice is the empirical Neural Tangent Kernel (eNTK), which measures how similar a network considers two inputs under some feature map, such as a layer’s activations or the network’s parameter gradients. This geometric toolkit allows researchers to explicitly distinguish whether a network is memorizing different situations with insufficient variation to adapting to the situation at hand with sufficient variation, or if it is ignoring task-irrelevant factors.

The methodology involves three stages:

  1. Controlled Training: Systematically varying one or two target factors while holding others fixed, structuring joint distributions as either confounded (co-varying) or independent (factorial design).

  2. Probe Construction: Creating specialized probe sets that sweep through a range of values for a target factor while keeping all other input variables static, using expert trajectories and validation frames.

  3. Representation Kernel Diagnostics: Computing the eNTK on model checkpoints and probes to extract statistics that offer actionable insights into the learning mechanism.

Key Diagnostic Metrics

The paper defines several statistics derived from the representation kernel to provide actionable insights:

(a) Phase transition heatmap:

(b) Effective rank:

(c) Success rate:

(d) PCA adapt vs ignore:

These metrics are used to characterize the learning regime. For instance, when training on few colors, the kernel matrix exhibits a clear block structure, indicating a memorize regime, where the model treats different situations as isolated cases. As variation increases, this structure flattens into an adapt regime.

The Factor Sensitivity Ratio (FSR) is defined as:

(a) FSRa:b = ¯da/¯db:

This ratio asks how strongly the representation changes when varying one factor relative to another. When a factor is task-relevant and another is task-irrelevant, this is called the Signal-to-Noise Ratio (SNR):

(b) SNR = FSRs:n = ¯ds/¯dn:

A larger SNR indicates that task-relevant variation produces larger changes in the representation relative to task-irrelevant variation.

Findings Across Model Classes

The diagnostics are applied across various model types, including Reinforcement Learning (RL), Imitation Learning, and Vision-Language-Action (VLA) models.

(a) RL Policies:

For pick-and-place tasks in ManiSkill, the eNTK successfully distinguishes between memorizing and adapting regimes. Furthermore, sufficient DR smooths out the kernel structure into an adapt style regime. Analysis of training performance under changing colors and backgrounds shows that sufficient DR leads to an ignore regime, which collapses irrelevant visual variance and leads to out-of-distribution (OOD) policy robustness.

(b) Transformer Visuomotor Policies:

When trained on increasing table-texture variation, the SNR diagnostics correctly track OOD success, from 0.05 to 0.43, showing that sufficient DR improves generalization. Joint variation of two factors (e.g., texture and color) is needed for joint shifts, consistent with prior work on data composition effects.

(c) Vision-Language-Action Policies:

At this scale, the FSR statistic supports two practical uses: ranking models by robustness and detecting shortcut learning without evaluation rollouts. For example, lighting is treated as a task-irrelevant nuisance and task state as the relevant signal; a higher SNR means the representation is more sensitive to state than to lighting. Shortcut learning is detected by comparing scene-to-language variation (FSRscene:language); a decrease in FSR after introducing a third task indicates reduced relative sensitivity to scene variation and increased sensitivity to language variation.

Representation Kernel Comparison

The paper compares the eNTK with last-hidden-layer and output kernels. The results indicate that the eNTK is a reasonable default diagnostic representation because it captures parameter-update effects, which are crucial for predicting cross-input effects of training. While activation and output representations are substantially cheaper to compute, they have characteristic blind spots: the output kernel can miss distinctions that do not yet affect the action, and the last-hidden layer representation cannot capture arm-specific decoupling introduced by task-irrelevant factors. The eNTK is therefore chosen as a primary diagnostic tool due to its ability to characterize how training data variation shapes a model’s internal representations.

Hardware Validation

The factor-conditioned kernel diagnostics are validated on hardware using an ACT policy.

Improvements for AI systems

Here are specific improvements for AI systems based on the findings in this paper, categorized by application:


)Improvement 1: Robustness Tuning via Data Randomization (DR) Optimization

Based on Figure 1 and Section 4, the system should move away from brute-force DR tuning toward smart randomization.

  • The system should dynamically adjust the intensity of domain randomization (e.g., object size, color, lighting) based on real-time diagnostic signals derived from the Neural Tangent Kernel (NTK).

  • Specifically, if the NTK indicates a memorize regime (high effective rank), the system should increase variation in task-relevant factors to force adaptation.

  • If the NTK suggests a shift toward an ignore regime (low effective rank/high Signal-to-Noise Ratio), the system should prioritize varying task-irrelevant factors (like color or background) to explicitly promote invariance, thereby maximizing Out-of-Distribution (OOD) success rates on new, unseen scenarios.

)Improvement 2: Shortcut Learning Detection and Mitigation

The system can be equipped with a dedicated Shortcut Learning Monitor based on the Factor Sensitivity Ratio (FSR).

  • The system should continuously compute FSRs between task-relevant factors (state) and task-irrelevant factors (lighting, language, texture).

  • If the FSR for a specific factor drops significantly during training or deployment, it signals that the model is relying too heavily on that factor.

  • In Vision-Language-Action (VLA) models, if the scene-to-language ratio FSR decreases after introducing a third task, this indicates reduced reliance on visual cues and increased reliance on language for disambiguation—a signal to reinforce learning in that direction.

)Improvement 3: Representation Selection via Kernel Comparison

Instead of defaulting to the empirical NTK (eNTK), the system should employ a dynamic representation selection module.

  • The system should evaluate the performance of different kernels (eNTK, last-hidden-layer, output kernel) during training.

  • If a task requires distinguishing between subtle state variations (like in bimanual tasks where arm structure matters), the system should select the Last Hidden Layer representation to capture fine structural details missed by other kernels.

  • If the primary goal is robust generalization across broad, varied inputs, it should rely on the eNTK for its aggregate information across layers.

)Improvement 4: Zero-Shot Model Selection and Ranking

For deploying VLA models (like OpenVLA), the system can use kernel metrics for pre-deployment auditing.

  • The system should compute a robustness score based on the SNR (FSR state:lighting). A high SNR indicates that state information is strongly encoded relative to lighting noise.

  • This metric allows for zero-shot model selection by ranking models based on their predicted OOD robustness, bypassing expensive real-world evaluation rollouts.

)Improvement 5: Bimanual/Modular Task Decoupling

For complex manipulation tasks involving multiple agents (e.g., bimanual peg picking), the system should explicitly monitor for head-specific decoupling.

  • The system should use a metric derived from the Last Hidden Layer representation to detect if different modules (arms) are learning task-relevant factors specific to their own pose, which is crucial for coordinated action.

  • This allows the system to ensure that joint policies learn arm-specific nuances rather than collapsing into identical representations shared across all arms.

Abstract

Training data variation, whether through designing a domain randomization (DR) scheme in simulation or curating demonstrations for imitation learning, is a primary lever for improving the robustness of robotic manipulation policies. Yet its underlying mechanisms remain poorly understood, and practitioners typically select randomization parameters through expensive trial and error. We investigate these mechanisms through a series of case studies, randomizing object size, color, and type as well as scene lighting and linguistic prompts across settings including pick-and-place RL in ManiSkill and fine-tuning of vision-language-action (VLA) models on LIBERO and RoboTwin. We examine both model behavior and internal representations, using the empirical neural tangent kernel (NTK) as our primary diagnostic tool. We show that the NTK distinguishes a shift in the internal learning mechanism from memorizing different situations with insufficient variation (e.g. learning what to do for a large cube, and what to do for a small cube) to adapting to the situation at hand with sufficient variation. An NTK-based signal-to-noise ratio also helps distinguish when policies have learned to ignore task-irrelevant factors (e.g. treating blue and red cubes identically, instead of learning a blue sub-policy and a red sub-policy). We use these diagnostics to develop practical guidance for designing DR schemes, selecting models, and detecting shortcut learning. We further compare different kinds of representations and validate our findings with real-world hardware experiments using ACT-based imitation learning.

Sources

Related papers