Yes, Q-learning Helps Offline In-Context RL

arXiv:2502.17666 · cs.LG, cs.AI · Submitted 2026-08-13 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Yes, Q-learning Helps Offline In-Context RL".

Jane: The paper was written by Denis Tarasov, Alexander Nikulin, Ilya Zisman, Albina Klepach, Andrei Polubarov et al. from ETH Zürich and Moscow Institute of Physics and Technology and Skolkovo Institute of Science and Technology and Innopolis University and HSE University and Accenture.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Alright, welcome back to the show, everyone. Today we're digging into a paper that's been making the rounds, and the title alone got me hooked: "Yes, Q-learning Helps Offline In-Context RL."

Jane: Tom, that title is practically a mic drop. It's like the authors are answering a question that a lot of people in the field have been asking, and they're answering it with a resounding "yes." I love it.

Tom: Exactly. And for our listeners who might be new to this, let's break down what we're even talking about. In-context reinforcement learning, or ICRL, is this idea of training a big model so it can adapt to a new task just by seeing a few examples, kind of like how a large language model can learn from a prompt.

Jane: Right. So instead of retraining a robot every time you give it a new maze, you want it to just... figure it out as it goes. The old way of doing this, a method called Algorithm Distillation, was basically supervised learning. It just watched what a bunch of other algorithms did and tried to copy their actions.

Tom: And that's where the problem lies, right? It's copying, not learning. The paper argues that this misses the whole point of reinforcement learning, which is to maximize reward, not just mimic a teacher.

Jane: So the authors, a big team from places like ETH Zürich and MIPT, they basically said, "What if we just... use actual reinforcement learning objectives inside this in-context framework?" And that's the core of "Yes, Q-learning Helps Offline In-Context RL."

Tom: And the results are pretty wild. They tested this across over one hundred fifty different datasets, and just by swapping out the learning objective, they got about a thirty percent improvement on average over the old Algorithm Distillation method.

Jane: Thirty percent is huge. That's not a small bump. That's a fundamental change in how well these agents can adapt.

Tom: It really is. And I think the title is also a bit of a jab at the community, you know? Like, "Hey, we've been overcomplicating this, maybe we should just go back to basics."

Jane: A friendly jab, but a jab nonetheless. It's a bold statement, and the data seems to back it up. I'm really curious to see how they actually implemented this, because it can't be as simple as just plugging in a Q-learning loss function.

Tom: Oh, for sure. There's always a catch. But before we get into the nitty-gritty of the methodology, I want to know what this means for the bigger picture. If this holds up, it could change how we train general-purpose agents for everything from robotics to game playing.

Jane: Absolutely. It suggests that the path to more intelligent, adaptable AI might not be through more complex supervised learning tricks, but through embracing the core principles of reinforcement learning. Let's get into the details of how they pulled this off.

Summary: Tom: So, Jane, we've established that this paper, "Yes, Q-learning Helps Offline In-Context RL," is a big deal. But how did they actually go from that idea to a thirty percent performance boost? Let's get into the summary.

Jane: Right. The key was to take the standard Transformer backbone used in Algorithm Distillation and just swap out the head. Instead of a head that predicts the next action, they put on heads that predict value functions, which are the expected future rewards.

Tom: And they didn't just use one flavor of reinforcement learning. They tested a bunch. They had a plain vanilla DQN, which is like the classic online RL algorithm. Then they had two offline-specific ones, CQL and IQL, which are designed to be more careful when learning from a fixed dataset.

Jane: And that distinction is crucial. In offline RL, you can't go out and explore. You have to learn from a static pile of data. So if your algorithm is too optimistic, it might think an action is great because it never saw the bad consequences in the data.

Tom: So it's like learning to drive by only watching a video of a perfect lap. You might think you can just floor it, but you don't know what happens when you hit a curb.

Jane: Exactly. That's why the offline-specific methods, CQL and IQL, they add a bit of "conservatism." They're pessimistic about actions they haven't seen before. And the paper found that this conservatism really matters.

Tom: Right. The plain DQN did better than Algorithm Distillation on average, but the offline-regularized ones, IC-CQL and IC-IQL, were the real stars. They were more robust, especially when the data was messy or limited.

Jane: And they didn't just test this in one simple environment. They went through a whole gauntlet. They used simple grid worlds, more complex partially observable ones, and even continuous control tasks like a half-cheetah running.

Tom: Yeah, the MuJoCo stuff. That's where you have to control a simulated robot with continuous joint movements. It's a much harder problem than just moving up, down, left, and right on a grid.

Jane: And the results held up there too. The offline RL methods, like TD3+BC and IQL, they consistently beat Algorithm Distillation. It really shows that this isn't a fluke of one specific environment.

Tom: The other thing I found interesting was how they handled data. Algorithm Distillation needs these "learning histories," which are like a full log of an agent's training from dumb to smart. But in the real world, you often just have a random collection of trajectories.

Jane: That's a great point. The paper tested that scenario too, where the data was just shuffled. And again, the RL-based methods were much more resilient. They could still learn to adapt, while AD really struggled.

Tom: So the summary is pretty clear: if you want an agent that can learn in-context from offline data, you should be optimizing for reward, not just mimicking actions. And you should probably use an offline-RL algorithm to do it.

Jane: And the improvements are so consistent that it's hard to argue with. I'm really curious about the specific experiments they ran, especially that XLand-MiniGrid environment. That's where they said the RL methods doubled the performance of AD.

Improvements: Tom: We're back, and Jane just brought up XLand-MiniGrid. That's the perfect segue into the improvements section of "Yes, Q-learning Helps Offline In-Context RL." Because that's where the results got really dramatic.

Jane: Right, Tom. XLand-MiniGrid is this incredibly complex environment with a huge variety of tasks. It's designed to be a real test for generalization. And the authors used a "tiny" version of the dataset, just one percent of the original size.

Tom: And even with that tiny amount of data, the RL-based methods absolutely crushed it. They got NAUC scores around zero point four, while Algorithm Distillation could only manage zero point two two. That's literally double the performance.

Jane: It's a massive jump. And it shows that when you're in a data-sparse regime, the ability to actually learn from experience, rather than just imitate, becomes even more critical. The RL methods are extracting more signal from less data.

Tom: But the improvements aren't just about raw performance. The paper also digs into something called "mixture of dynamics." This is where the training data comes from environments that behave differently.

Jane: Oh, that's the Janus experiment. They trained agents on a grid where the same action could mean different things, like one environment where "up" moves you up, and another where "up" moves you down.

Tom: And then they deployed the agent into a single environment that had both of these dynamics at the same time. It's a really clever test for out-of-distribution generalization.

Jane: And again, the RL methods were much better at handling this confusing situation. They were able to adapt their behavior to the specific part of the environment they were in, while AD just fell apart.

Tom: It's not perfect, though. The paper is honest that none of the methods fully solve the out-of-distribution problem. But the RL methods are clearly more robust.

Jane: And that's a huge practical improvement. It means these agents are less brittle when you deploy them in the real world, where things don't always match your training data.

Tom: I also want to mention the model size experiments they ran. They found that bigger isn't always better. There's a sweet spot for the Transformer's size, and going too big can actually hurt performance.

Jane: That's a really important finding for people who are trying to scale these models. It suggests that we don't just need more parameters; we need the right learning objective first. The architecture is secondary to the training signal.

Tom: Exactly. It's like saying you can't just build a bigger engine and expect to win the race if you're still driving in the wrong direction. You need to point the car at the finish line first.

Jane: And that's the core message of this whole paper. The improvement isn't a new architecture or a new trick. It's a fundamental realignment of the goal. They're saying, "Let's stop trying to predict the past and start trying to optimize the future."

Tom: And the future looks a lot brighter for in-context RL because of it. Let's bring in Lu and Meng to get their take on the practical side of this.

Lu: From my perspective, the most exciting part is that this validates a whole line of research. It says that the principles we've learned in offline RL, like conservatism, are directly transferable to this new, more general setting of in-context learning.

Meng: And from an engineering standpoint, the fact that they used a simple, shared architecture and just changed the loss function is huge. It means we can take existing, well-optimized codebases and adapt them without having to build everything from scratch.

Tom: So we have a clear win in performance, a clear win in robustness, and a clear win in simplicity. This is a pretty compelling package.

Conclusion: Tom: Well, we've covered a lot of ground today on "Yes, Q-learning Helps Offline In-Context RL." Let's try to wrap this up. Jane, what's the one thing you want our listeners to remember?

Jane: I think it's that the learning objective matters more than we thought. For a long time, the field was focused on clever architectures and data processing. This paper shows that just going back to the fundamental RL goal of maximizing reward gives you a massive boost.

Tom: It's a strong argument for simplicity. They took a proven baseline, Algorithm Distillation, and just swapped out the supervised loss for a reinforcement learning loss. And that simple change led to a thirty percent average improvement across a huge battery of tests.

Jane: And it's not just about being better on average. It's about being more reliable. The offline-regularized methods, like CQL and IQL, were much more consistent across different data qualities and coverage levels. That's what you need for real-world deployment.

Tom: We also saw that this approach is more data-efficient, which is a huge deal. It can work with random, unstructured data, not just carefully curated learning histories. That makes it much more practical.

Jane: And the implications are broad. This isn't just a win for game-playing agents. It points toward a better way to train adaptable systems in robotics, autonomous driving, and any field where an agent has to learn on the job from a fixed set of experiences.

Tom: The authors are also clear that this is just the beginning. They're pointing toward testing in even more complex environments, like NetHack, and exploring visual observations. The future work section is full of exciting possibilities.

Jane: So, as we say goodbye to this paper, I think the message is one of cautious optimism. The path to more general AI might not require a completely new paradigm. Sometimes, it just requires remembering the basics and applying them well.

Tom: Well said, Jane. That's a wrap on "Yes, Q-learning Helps Offline In-Context RL." A big thank you to Lu and Meng for joining us and sharing their insights.

Lu: Thanks for having me. It was a great discussion.

Meng: Yeah, really enjoyed it. Looking forward to the next one.

Tom: And to our listeners, thanks for tuning in. We'll be back soon with another paper to break down. Until then, keep exploring.

Denis Tarasov, Alexander Nikulin, Ilya Zisman, Albina Klepach, Andrei Polubarov, Nikita Lyubaykin, Alexander Derevyagin, Igor Kiselev, Vladislav Kurenkov

ETH Zürich · Moscow Institute of Physics and Technology · Skolkovo Institute of Science and Technology · Innopolis University · HSE University · Accenture

cs.LG, cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: https://github.com/dunnolab/yesq

Code: https://github.com/dunnolab/yesq

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 69/100

The gist: Existing offline in-context reinforcement learning (ICRL) methods have predominantly relied on supervised training objectives, which are known to have limitations in offline RL settings.

Key concepts

In-context Reinforcement Learning (ICRL)
The idea of training a large model to adapt to a new task by observing only a few examples, similar to how an LLM learns from a prompt. It aims for agents to figure out tasks as they go, rather than needing full retraining.
Algorithm Distillation (AD)
An older method of training agents where the model is supervised and learns by watching and copying the actions of other algorithms. The paper argues this misses the core goal of maximizing reward, only focusing on imitation.
Offline RL
A type of reinforcement learning where an agent must learn from a fixed, static dataset (trajectories) without being able to explore or interact with the environment. This requires 'conservatism' to prevent over-optimism.
Q-learning
A classic reinforcement learning algorithm used here as a learning objective. Instead of predicting the next action, the model predicts value functions (expected future rewards), which helps agents maximize their cumulative reward.

Terminology

Summary

Existing offline in-context reinforcement learning (ICRL) methods have predominantly relied on supervised training objectives, which are known to have limitations in offline RL settings. In this study, we explore the integration of RL objectives within an offline ICRL framework. Through experiments on more than 150 GridWorld and MuJoCo environment-derived datasets, we demonstrate that optimizing RL objectives directly improves performance by approximately 30% on average compared to widely adopted Algorithm Distillation (AD), across various dataset coverages, structures, expertise levels, and environmental complexities. Furthermore, in the challenging XLand-MiniGrid environment, RL objectives doubled the performance of AD. Our results also reveal that the addition of conservatism during value learning brings additional improvements in almost all settings tested. Our findings emphasize the importance of aligning ICRL learning objectives with the RL reward-maximization goal, and demonstrate that offline RL is a promising direction for advancing ICRL.

Improvements for AI systems

Based on the paper Yes, Q-learning Helps Offline In-Context RL, here are specific improvements to AI systems and what the improved systems can do:

Improvement: Instead of training in-context RL agents with supervised learning (predicting the next action from history), train them with explicit RL objectives like Q-learning, CQL, or IQL.

What the improved system can do:

  • Achieve 30% higher average performance (NAUC) compared to Algorithm Distillation across diverse offline datasets

  • Double performance in complex environments like XLand-MiniGrid (NAUC from 0.22 to 0.46)

  • Adapt to unseen tasks more effectively during deployment without parameter updates

These improvements collectively enable AI systems to learn adaptive policies from static, imperfect, and unstructured offline data—a critical capability for real-world deployment in robotics, healthcare, and autonomous systems where online interaction is costly or unsafe.

Abstract

Existing offline in-context reinforcement learning (ICRL) methods have predominantly relied on supervised training objectives, which are known to have limitations in offline RL settings. In this study, we explore the integration of RL objectives within an offline ICRL framework. Through experiments on more than 150 GridWorld and MuJoCo environment-derived datasets, we demonstrate that optimizing RL objectives directly improves performance by approximately 30% on average compared to widely adopted Algorithm Distillation (AD), across various dataset coverages, structures, expertise levels, and environmental complexities. Furthermore, in the challenging XLand-MiniGrid environment, RL objectives doubled the performance of AD. Our results also reveal that the addition of conservatism during value learning brings additional improvements in almost all settings tested. Our findings emphasize the importance of aligning ICRL learning objectives with the RL reward-maximization goal, and demonstrate that offline RL is a promising direction for advancing ICRL.

Sources

Related papers