DiSA-IQL: Offline Reinforcement Learning for Robust Soft Robot Control under Distribution Shifts
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "DiSA-IQL: Offline Reinforcement Learning for Robust Soft Robot Control under Distribution Shifts".
Jane: The paper was written by Linjin He, Xinda Qi, Dong Chen, Zhaojian Li and Xiaobo Tan from Georgetown University and Michigan State University and Mississippi State University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Hey everyone, welcome back to the show! Today we're diving into a fresh arXiv paper that's got me genuinely pumped — it's called "DiSA-IQL: Offline Reinforcement Learning for Robust Soft Robot Control under Distribution Shifts." Jane, have you had a chance to look at this one?
Jane: I have, Tom, and honestly, the title alone tells a great story. You've got soft robots — think squishy, flexible machines like snake robots — and you're trying to control them using reinforcement learning, but without letting the robot interact with the real world during training. That's the "offline" part, and it's a big deal.
Tom: Right, and the team behind it is from Georgetown, Michigan State, and Mississippi State. Linjin He, Xinda Qi, Dong Chen, Zhaojian Li, and Xiaobo Tan. They're really tackling a practical problem — soft snake robots are amazing in complex environments, but their dynamics are so nonlinear that traditional control methods just don't cut it.
Jane: Exactly. And that's where reinforcement learning comes in. But here's the catch — training an RL agent online means letting it stumble around and make mistakes, which is expensive and potentially damaging for a real soft robot. So offline RL uses pre-collected data instead, which is safer. But then you hit this wall called distribution shift.
Tom: Distribution shift — that's when the policy learns to exploit actions that weren't really represented in the training data. The robot thinks it's found a great move, but that move was never actually tested, so the Q-values are inflated and unreliable. It's like studying for a test using only practice problems from chapter one, then getting questions from chapter five.
Jane: That's a perfect analogy, Tom. And the authors are saying, look, existing offline RL methods like Behavior Cloning, Conservative Q-Learning, and even Implicit Q-Learning all struggle with this in soft robot control. So they built an extension of IQL that adds a robustness mechanism to penalize those unreliable state-action pairs.
Tom: And the results? They're pretty impressive. In their in-distribution tests, DiSA-IQL hit a one hundred percent success rate on goal-reaching tasks. But the real test is out-of-distribution — training in one region, testing in another — and there they got ninety-one point two percent success, while vanilla IQL only managed seventy-five point eight percent. That's a massive jump.
Jane: It really is. And I love that they open-sourced everything — the code and the simulator are on GitHub. That's how you make research reproducible and actually useful for the community.
Tom: Totally agree. So we've got the setup — soft snake robots, offline RL, distribution shift. But how did they actually make this work? That's what we're digging into next.
Summary: Jane: So we've established the problem — soft snake robots are hard to control, and offline RL has this distribution shift issue. Now let's talk about what the paper actually does. Tom, you want to walk us through the robot itself?
Tom: Love to. So this is a modular soft snake robot made of pneumatic bending actuators. It moves using serpentine locomotion — basically a traveling wave that propagates down its body, and the friction between the snake skin and the ground does the rest. The shape is modeled as an inextensible curve, and the local curvature is controlled by pressure differences in the air chambers.
Jane: And the control task is goal-reaching. The robot's center of mass needs to get within a small neighborhood of a target point. The state includes relative distances to the target, the deflection angle, and the pressure biases. The action is the chamber biases plus the wave propagation direction — forward or backward.
Tom: Right. And the reward function is a mix of sparse success reward plus penalties for distance and deflection. Now, the key innovation here is DiSA-IQL itself. It builds on Implicit Q-Learning, which uses expectile regression to estimate the value function and advantage-weighted regression for the policy. But vanilla IQL doesn't explicitly punish out-of-distribution actions.
Jane: So what they did is add a penalty term into the Q-function update. The penalty depends on how often a state-action pair appears in the dataset. If something is rare, the penalty is large, which pushes down the estimated Q-value. That way, the policy won't be tempted to exploit actions that the data doesn't really support.
Tom: And they tested a few different penalty functions — Wasserstein, KL divergence, chi-square, total variation. They settled on the KL divergence penalty because it performed best in their simulations. They also added a decay schedule for the robustness coefficient, so early training is more conservative and later training relaxes a bit.
Jane: That's clever. It's like being extra cautious when you're not sure about the terrain, then gradually trusting your model as you gather more confidence. And the results speak for themselves — we already mentioned the success rates. But let's talk about the baselines. They compared against Behavior Cloning, CQL, and vanilla IQL.
Tom: And the gap is stark. In the out-of-distribution setting, CQL only got thirty-three point two percent success — it's too conservative, it basically refuses to try anything unseen. BC got fifty-eight point eight percent because it just imitates the data without any reasoning about returns. IQL got seventy-five point eight percent. DiSA-IQL? ninety-one point two percent. And it also had the fewest average steps per episode, meaning it's not just more successful, it's more efficient.
Jane: Efficiency matters a lot for a physical robot. Fewer steps means less energy, less wear and tear. And the trajectories were smoother too — less oscillation, more direct paths. That's a big deal for real-world deployment.
Tom: Absolutely. So the algorithm works, but I'm curious about the broader implications. What does this mean for soft robotics beyond snake robots? Let's bring in Lu and Meng to get their take.
Lu: Thanks, Tom. I think this is a significant step for the whole field of soft robot control. The fact that you can train a policy entirely offline and still generalize to unseen goal regions opens up a lot of possibilities. Think about surgical robots — you can't let them experiment on patients. Offline RL with robustness against distribution shift is exactly what you need there.
Meng: And from an engineering standpoint, the open-source code is huge. I can actually take this framework, plug in my own dataset, and see if it works for my robot. The fact that they normalized actions and rewards, formatted the data in the D4RL style — that's practical. It lowers the barrier to entry significantly.
Jane: Great points from both of you. So we've got the algorithm, we've got the results. But what are the limitations? What's the next step? That's coming up.
Improvements and Implications: Tom: Alright, we've covered the what and the how. Now let's talk about the improvements this paper suggests — not just for the algorithm, but for the field. Jane, what stood out to you?
Jane: For me, it's the way they framed the problem. They didn't just say "here's a new algorithm." They systematically benchmarked existing offline RL methods on soft snake robot control, identified where they fail, and then designed a targeted fix. That's good research methodology. And the fix itself — penalizing unreliable state-action pairs — is conceptually simple but effective.
Lu: I'd push back slightly on the simplicity, Jane. The choice of penalty function and the decay schedule — those details matter. They tested four different penalty designs and found that KL divergence worked best. That's empirical tuning, and it's not obvious a priori which one would win. But the framework they built allows for that kind of exploration, which is valuable.
Meng: From my side, the practical improvement is the robustness. In real-world deployment, you're almost always facing distribution shift — different terrain, different wear on the robot, different environmental conditions. A policy that can handle unseen scenarios without retraining is incredibly valuable. The fact that they got ninety-one point two percent success in the OOD setting, compared to seventy-five point eight percent for IQL, that's not a marginal gain.
Tom: And let's not forget the trajectory quality. They showed that DiSA-IQL produces smoother, more direct paths. For a physical robot, that means less stress on the actuators, less energy consumption, and potentially longer lifespan. It's not just about reaching the goal — it's about how you get there.
Jane: Right. And the paper also acknowledges its limitations. It's all simulation-based so far. The dynamics model, the friction model — those are approximations. Real-world validation on physical robots is the obvious next step. They also mention that the dataset is task-specific, and extreme distribution shifts could still be a problem.
Lu: That's where I think the future work gets exciting. They suggest exploring generative models and adaptive curriculum strategies. Imagine training on a diverse set of terrains and goal distributions, then having the policy adapt online. That could make soft robots truly versatile in unstructured environments.
Meng: And the sim-to-real transfer angle — they mention integrating simulated and real data in the offline dataset. That's a practical path forward. You can collect some real trajectories, augment with simulation, and train a robust policy without excessive real-world experimentation.
Tom: So the improvements here are threefold — the algorithm itself, the benchmarking methodology, and the open-source framework that lets others build on it. But what's the bigger picture? What does this mean for the world? Let's bring in Lalam to weigh in.
Conclusion: Jane: We've had a great discussion about "DiSA-IQL: Offline Reinforcement Learning for Robust Soft Robot Control under Distribution Shifts." Tom, before we wrap up, let's do a quick recap.
Tom: Sure. The paper tackles offline RL for soft snake robots, addressing the distribution shift problem by adding a robustness penalty to Implicit Q-Learning. They showed it outperforms BC, CQL, and vanilla IQL in both in-distribution and out-of-distribution goal-reaching tasks, with higher success rates and smoother trajectories. And they open-sourced everything.
Jane: And the implications go beyond snake robots. This approach could apply to any soft robot where online training is impractical or unsafe — surgical robots, search-and-rescue, agricultural harvesting. The ability to train robustly from static datasets is a game-changer.
Lu: I'd add that the theoretical framework — the distributionally robust perspective — could influence other areas of offline RL, not just robotics. Any domain with high-stakes decisions and limited data could benefit.
Meng: And from a practical standpoint, the open-source code means we can start experimenting immediately. That's how the field advances — building on each other's work.
Lalam: I'd like to emphasize the cultural impact. As we develop robots that can safely learn from past data without risky exploration, we're making automation more accessible and more reliable. That has implications for how we design human-robot collaboration, how we deploy robots in homes and hospitals, and how we think about safety in autonomous systems. This paper is a step toward that future.
Tom: Beautifully said. So we're saying goodbye to this paper — thank you for the insights, the robust framework, and the open-source spirit. We're ready to move on to the next one. Stay tuned, folks.
Jane: And as always, if you want to dig into the details, the paper is on arXiv, and the code is on GitHub. We'll link everything in the show notes. Thanks for listening, everyone!
Linjin He, Xinda Qi, Dong Chen, Zhaojian Li, Xiaobo Tan
Georgetown University · Michigan State University · Mississippi State University
cs.RO, cs.AI
Submitted: 2026-08-15
Updated: 2026-08-18
Code: https://github.com/hlj0908/DiSA-IQL-for-Soft-Robot-Control
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 48/100
Key concepts
- Offline Reinforcement Learning (RL)
- A type of reinforcement learning where an agent learns a policy entirely from pre-collected, static datasets rather than interacting with the environment online. This is used here because training real robots online is expensive and potentially damaging.
- Distribution Shift
- This occurs when an RL policy learns to take actions that were not well-represented in the training data. The robot might find a seemingly good move in a new situation, but since it wasn't trained on that scenario, its performance becomes unreliable.
- DiSA-IQL
- An extension of Implicit Q-Learning that adds a robustness mechanism to penalize state-action pairs that appear infrequently in the training data. This penalty prevents the policy from exploiting actions based on unverified or rare experiences.
Terminology
Summary
Summary
This paper addresses the challenge of controlling soft snake robots, which offer flexibility and adaptability in complex environments but are difficult to control due to highly nonlinear dynamics. Traditional model-based and bio-inspired controllers rely on simplified assumptions that limit performance. While deep reinforcement learning (DRL) has emerged as a promising alternative, online training is often impractical due to costly and potentially damaging real-world interactions. Offline RL provides a safer option by leveraging pre-collected datasets, but it suffers from distribution shift, which degrades generalization to unseen scenarios.
To overcome this challenge, the authors propose DiSA-IQL (Distribution-Shift-Aware Implicit Q-Learning), an extension of Implicit Q-Learning (IQL) that incorporates robustness modulation by penalizing unreliable state–action pairs to mitigate distribution shift. The main contributions are: (1) systematically evaluating and benchmarking classical offline RL algorithms—Behavior Cloning (BC), Conservative Q-Learning (CQL), and Implicit Q-Learning (IQL)—on locomotion and navigation tasks of a soft snake robot; (2) enhancing IQL with an out-of-distribution action suppression mechanism; (3) validating the proposed algorithm across two environments with varying complexity and open-sourcing the implementation.
The soft snake robot is modeled as an inextensible curve in 2D space, with shape described by the center of mass and local curvature determined by pressure differences in pneumatic chambers. The dynamics include anisotropic friction between the snake skin and ground. The task is goal-reaching: an RL agent drives the robot's center of mass from a start point to a random target point, with success defined as reaching within a small neighborhood (radius ϵ) of the target. The MDP formulation includes a state space with relative distances, deflection angle, and pressure biases; an action space with chamber biases and wave propagation direction; and a reward function combining sparse position reward with penalties for distance and deflection.
The paper details the offline RL baselines: BC reduces the problem to supervised learning by imitating dataset actions; CQL adds a pessimistic regularizer penalizing Q-values of out-of-distribution actions; IQL uses expectile regression and advantage-weighted regression to bias toward high-return trajectories without explicit pessimism. DiSA-IQL extends IQL by adding a penalty term P(s,a) to the Bellman regression and advantage function, with several penalty designs based on visitation counts (Wasserstein, KL, Chi-square, Total Variation). The KL divergence penalty is selected after performance comparison, with a decay schedule for the robustness coefficient α to balance robustness and convergence.
Simulation results are presented for two settings. In the in-distribution setting (training and testing in the same region), DiSA-IQL achieves a 100% success rate with an average reward of 42.5 and average 37 steps per episode, outperforming BC (99.2% success, 42.3 reward, 36 steps), IQL (93.3% success, 34.0 reward, 50 steps), and CQL (65.8% success, 3.7 reward, 91 steps). In the out-of-distribution (OOD) setting (training in one half-annulus, testing in both training and unseen half-annulus), DiSA-IQL achieves a 91.2% success rate with average reward 27.9 and 59 steps, significantly outperforming IQL (75.8% success, 17.1 reward, 76 steps), BC (58.8% success, 6.5 reward, 86 steps), and CQL (33.2% success, -22.5 reward, 118 steps). Visualizations confirm DiSA-IQL produces smoother, more stable trajectories with fewer failures, especially in unseen regions, while CQL exhibits severe instability and BC fails to generalize.
The paper concludes that DiSA-IQL achieves stable learning in in-distribution tasks and better generalization to out-of-distribution scenarios, consistently outperforming baselines. Limitations include reliance on simulations, task-specific datasets, and sensitivity to extreme distribution shifts. Future work will validate DiSA-IQL on physical robots, expand datasets to more diverse terrains, and explore generative models and adaptive curriculum strategies to further enhance robustness.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to an AI system, along with what the improved system can do:
1. Distribution-Shift-Aware Robustness Modulation
-
Implement the DiSA-IQL algorithm, which extends Implicit Q-Learning (IQL) by adding a penalty term to the Bellman target that scales inversely with state-action visitation counts in the offline dataset.
-
Use the KL-divergence-based penalty function:
P(s,a) = α√(2/(N(s,a)+1)), where N(s,a) is the visitation count, with a linear decay schedule on α over training. -
Modify the advantage function to subtract this penalty before policy weighting, preventing the policy from exploiting unreliable Q-value estimates in out-of-distribution regions.
2. Offline RL with Pessimistic Value Estimation
-
Train the value function using expectile regression with τ < 0.5 to bias toward lower returns, ensuring conservative value estimates.
-
Use the robust Bellman regression for Q-function updates, incorporating the penalty term to suppress overestimation of rarely-seen actions.
-
Apply advantage-weighted regression for policy updates, weighting actions by their robust advantage to favor high-reward, well-supported behaviors.
3. Goal-Conditioned Control with State-Action Reliability Awareness
-
Design the observation space to include relative goal distances (Δx, Δy), heading deflection (Δθ), and actuation biases (bi, bi,pre), enabling the agent to track goal-relative states.
-
Incorporate a reliability score for each state-action pair based on dataset coverage, allowing the system to dynamically adjust its confidence in Q-value predictions.
1. Robust Soft Robot Control Under Distribution Shift
-
Achieve 91.2% success rate on out-of-distribution goal-reaching tasks, compared to 75.8% for vanilla IQL, 58.8% for BC, and 33.2% for CQL.
-
Maintain 100% success rate on in-distribution tasks while producing smoother, more direct trajectories than baselines.
-
Generalize to unseen target regions (right half-annulus) when trained only on left half-annulus data, demonstrating transferability across structurally related goal spaces.
2. Sample-Efficient and Safe Policy Learning
-
Learn effective policies entirely from pre-collected static datasets without any online interaction, eliminating the need for costly and potentially damaging real-world robot trials.
-
Avoid catastrophic failures from overestimating out-of-distribution actions, a critical safety improvement for physical soft robots.
-
Reduce average episode length to 59 steps in OOD settings (vs. 76 for IQL, 86 for BC, 118 for CQL), indicating faster and more efficient goal-reaching.
3. Adaptive Confidence-Based Decision Making
-
Automatically down-weight actions that are infrequently observed in the training data, preventing the system from relying on unreliable predictions.
-
Provide a natural robustness-accuracy trade-off via the α decay schedule, allowing the system to be more conservative early in training and more aggressive as confidence grows.
-
Produce stable, oscillation-free trajectories even in unfamiliar environments, as visualized in the paper's trajectory plots.
4. Transferable Control Framework for Deformable Systems
-
Apply the same algorithm to any soft robot with continuous state and action spaces, given a pre-collected dataset of trajectories.
-
Support both discrete (wave propagation direction) and continuous (pressure biases) action components within a unified MDP formulation.
-
Enable sim-to-real transfer by integrating simulated and real data in the offline dataset, as the framework is agnostic to data source.
5. Reproducible and Extensible Research Infrastructure
-
Provide open-source code and simulator (github.com/hlj0908/DiSA-IQL-for-Soft-Robot-Control) that allows direct replication and extension of the results.
-
Offer a D4RL-style dataset format, making it compatible with existing offline RL evaluation pipelines and enabling fair benchmarking against future methods.
Sources
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- Offline Goal-Conditioned Reinforcement Learning for Shape Control of Deformable Linear Objects
- Offline Reinforcement Learning with Implicit Q-Learning
- Distributionally Robust Offline Reinforcement Learning with Linear Function Approximation
- Bridging Distributionally Robust Learning and Offline RL: An Approach to Mitigate Distribution Shift and Partial Data Coverage
- Conservative Q-Learning for Offline Reinforcement Learning
- D4RL: Datasets for Deep Data-Driven Reinforcement Learning
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving