Representation and Invariance in Reinforcement Learning

arXiv:2112.07752 · cs.AI, cs.GT, cs.LG · Submitted 2026-08-11 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Representation and Invariance in Reinforcement Learning".

Jane: The paper was written by Samuel Allen Alexander and Arthur Paul Pedersen from Independent Researcher and The City University of New York.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back, everyone. Today we’re digging into a paper that’s been making the rounds on arXiv, and it’s called “Representation and Invariance in Reinforcement Learning.” Jane, I’ve got to say, the title alone got me excited.

Jane: Oh, absolutely, Tom. And I think the title is actually doing a lot of work here. It’s not just about reinforcement learning in the abstract. It’s asking whether the way we *represent* RL — the formal rules we pick — actually changes what the agents can do.

Tom: Right, and that’s a big deal, because in the field, people talk about RL like it’s one thing. You have an agent, an environment, rewards, actions. But the authors, Samuel Allen Alexander and Arthur Paul Pedersen, they’re saying, hold on, the details matter.

Jane: Exactly. And the word “invariance” is the key. They want to know: if you change the rules of the game, does the relative intelligence of agents stay the same? If agent A is smarter than agent B in one framework, is that still true after you convert them to another framework?

Tom: So it’s like, if you take a chess grandmaster and make them play checkers, are they still a grandmaster?

Jane: That’s the spirit, but even more subtle. Because in RL, you’re not just changing the game. You’re changing the language the game is written in. And the authors show that sometimes, that language change actually breaks the comparison.

Tom: And that’s the provocative part. They’re not just saying, “be careful.” They’re showing concrete cases where two perfectly standard RL frameworks are not equivalent. No transformation exists that preserves relative intelligence.

Jane: Which, if true, means the field has been a bit sloppy. People publish results in one framework and assume they apply to another. This paper says, maybe not.

Tom: And that’s why I’m excited. This is foundational work. It’s not a new algorithm that gets you two percent better on a benchmark. It’s asking whether the whole edifice is built on sand.

Jane: Right. And the authors are careful to say this is a first step, a modest stab, as they put it. But it’s a step that could change how we interpret a huge body of research.

Tom: So, Jane, before we get into the guts of the paper, what’s the one thing you want listeners to take away from the title itself?

Jane: That “representation” isn’t a neutral choice. How you set up the rules of RL — deterministic or stochastic, integer or real rewards — that choice has consequences. And “invariance” is the property we’d like to have, but the paper shows we don’t always get it.

Tom: And that’s the hook. Next segment, we’re going to look at the actual summary and the big claims they make. Stick around.

Summary: Jane: Welcome back. We’re still on “Representation and Invariance in Reinforcement Learning,” and now we’re looking at the summary of what the paper actually claims.

Tom: And Jane, the summary is punchy. They introduce this idea of a “transformation” between RL frameworks. It’s a way to convert agents and environments from one setup to another, and they say if such a transformation exists, then relative intelligence is preserved.

Jane: Right. So imagine you have a lab that builds agents for one version of RL, but you have environments built for a different version. You can’t just plug them together. You need a translator. And the paper says, if you can build a good translator, then the smart agents stay smart.

Tom: But here’s the kicker. They don’t just define this. They actually test it on four concrete RL frameworks. And these aren’t exotic. These are the standard ones: deterministic agents, stochastic agents, deterministic environments, stochastic environments.

Jane: And the results are, frankly, a little alarming. Out of twelve possible pairs, only five have a transformation in one direction. And no two frameworks are mutually transformable. So none of the four are equivalent to each other.

Tom: That’s huge. I mean, in practice, people use pseudo-random number generators to make deterministic agents act stochastic. They assume it’s the same thing. This paper says, mathematically, it’s not.

Jane: And the key technical tool they use is something called an ultrafilter. Now, that sounds scary, but the paper explains it in terms of elections. Imagine environments are voters, and they’re voting on which agent is smarter.

Tom: And the ultrafilter is the rule for deciding who wins the election, even when there are infinitely many voters.

Jane: Exactly. And it turns out, this election-based approach to comparing intelligence was proposed by one of the authors, Alexander, in an earlier paper. So this is a continuation of that line of work.

Tom: So the summary is: they define a rigorous notion of reducibility between RL frameworks, they prove that this notion preserves intelligence comparisons, and then they show that the four most common frameworks fail to be reducible to each other.

Jane: Which is a strong statement. It’s not saying one framework is better. It’s saying they’re genuinely different, and we can’t pretend otherwise.

Tom: And I love that they’re honest about the limitations. They say, look, we only looked at integer rewards. We don’t know if this holds for rational rewards. That’s the kind of humility you don’t always see.

Jane: Right. And that’s actually a perfect segue, because next we’re going to look at the first page of the paper, where they lay out the motivation and the questions they’re asking. That’s where the real philosophical meat is.

Tom: And I can’t wait. Because the questions they ask on that first page are the ones that keep me up at night.

Improvements: Tom: We’re back, still on “Representation and Invariance in Reinforcement Learning.” And Jane, we’ve talked about the title and the summary. Now I want to get into what this paper actually improves on. What does it add to the field?

Jane: Great question, Tom. And I think the biggest improvement is that it gives us a *language* for talking about equivalence. Before this paper, if someone said “these two RL frameworks are basically the same,” there was no formal way to check that claim.

Tom: Right. It was all vibes. Someone would say, “Oh, deterministic agents are just a special case of stochastic agents,” and everyone would nod along.

Jane: And this paper says, okay, let’s make that precise. Let’s define what it means to convert one framework into another, and let’s prove whether that conversion preserves the thing we care about — relative intelligence.

Tom: And that’s a real improvement, because it turns a philosophical debate into a mathematical one. You can’t just assert equivalence anymore. You have to exhibit a transformation, or prove none exists.

Jane: Exactly. And they also improve on the existing intelligence measurement literature. The Legg-Hutter intelligence measure, which is famous, is mathematically unwieldy. It involves infinite sums and Kolmogorov complexity, which is noncomputable.

Tom: So it’s beautiful but impractical.

Jane: Right. And the ultrafilter approach they use is more tractable. It lets them actually prove preservation theorems, which they couldn’t do with the Legg-Hutter measure.

Tom: And there’s another improvement I want to highlight. They introduce this notion of “well-behaved environments.” Because the naive definition of expected total reward doesn’t always converge. You can get infinite sums that blow up.

Jane: Oh, that’s a good point. So they restrict attention to environments where the value function is guaranteed to converge. That’s a practical fix that makes the theory usable.

Tom: And it’s not just a technical trick. It’s a real modeling decision. In practice, you don’t want agents in environments where rewards diverge to infinity. You want well-behaved ones.

Jane: So the improvements are: a formal notion of reducibility, a tractable way to compare intelligence, and a clean way to handle convergence issues. That’s a solid contribution.

Tom: And the payoff is that they can now prove something surprising: that the four standard frameworks are not equivalent. That’s the result that’s going to get people talking.

Jane: And it’s going to get people arguing, too. Because some of those negative results are counterintuitive. Like, why can’t you map stochastic agents into deterministic ones? That seems like it should be easy.

Tom: Right, and that’s exactly what we’re going to dig into next, when we look at the first page of the paper and the specific questions they raise. That’s where the real fireworks are.

First Page: Jane: Welcome back to the show. We’re still on “Representation and Invariance in Reinforcement Learning,” and now we’re looking at the first page, where the authors set the stage.

Tom: And Jane, the first page is basically a manifesto. They start by asking, “If we changed the rules, would the wise become fools?” That’s the opening line, and it’s fantastic.

Jane: It really sets the tone. And they immediately get concrete. They ask: what if rewards are only zero? Clearly weaker. What if rewards are only zero and one? What about −one zero one? What about all integers? What about rationals? What about reals?

Tom: And they admit they don’t know the answers. They say this is a “modest first stab” at the problem. But the questions themselves are the contribution.

Jane: Right. Because those questions force you to realize that the details of RL aren’t arbitrary. They’re design choices, and those choices have consequences.

Tom: And then they introduce the four frameworks we’ve been talking about. Deterministic agents, stochastic agents, deterministic environments, stochastic environments. Four combinations.

Jane: And they say, look, all four are used in practice. Deterministic agents often masquerade as stochastic through pseudo-random number generators. So if these frameworks aren’t equivalent, that’s a real problem.

Tom: And the first page also mentions the broader context. They cite Silver et al., who wrote “Reward is enough,” arguing that RL will lead to AGI. And the authors here are saying, well, which RL?

Jane: That’s the key question. If the formal details matter, then you can’t just say “RL will lead to AGI.” You have to say *which* RL, with *which* rewards, *which* action spaces, *which* convergence criteria.

Tom: And that’s a profound point. Because the AGI debate is often conducted at a high level of abstraction, and this paper drags it down to the gritty details.

Jane: It does. And I think that’s healthy. Because if we’re going to make grand claims about intelligence, we need to be precise about what we mean.

Tom: And the first page also sets up the structure of the paper. They’re going to define transformations, they’re going to introduce ultrafilters, and then they’re going to prove the main theorem about the four frameworks.

Jane: Which we’ve already spoiled, but that’s okay. The journey is the fun part.

Tom: And the journey includes some genuinely surprising results. Like, you can map deterministic agents into stochastic agents, and you can map stochastic environments into deterministic environments. But you can’t do the reverse.

Jane: And that asymmetry is the kind of thing that makes you go, “Wait, really?” And that’s the hook for our final segment, where we wrap up and talk about what this all means.

Conclusion: Tom: And we’re back for the final segment on “Representation and Invariance in Reinforcement Learning.” Jane, let’s pull it all together.

Jane: Let’s do it. So the paper gives us a formal way to ask whether one RL framework can be reduced to another. It defines transformations, proves that those transformations preserve relative intelligence, and then shows that the four standard frameworks are not mutually reducible.

Tom: And the key result is that asymmetry. You can go from deterministic agents to stochastic agents, and from stochastic environments to deterministic environments. But not the other way around. And that’s a genuinely new finding.

Jane: It is. And it suggests that the choice of whether agents are deterministic or stochastic isn’t just a convenience. It’s a fundamental design decision that changes the mathematical landscape.

Tom: And the implications for the field are big. Researchers need to be more careful when they say “reinforcement learning” as if it’s one thing. They need to specify which framework they mean.

Jane: And that’s not just pedantry. It affects how we interpret results, how we compare agents, and even how we think about the path to AGI.

Tom: Right. Because if the frameworks aren’t equivalent, then a result proven in one framework might not transfer to another. And that’s a warning shot across the bow.

Jane: But the authors are also humble. They say this is an initial step. They only looked at integer rewards. They don’t know if the results hold for rational rewards. So there’s a lot of open work.

Tom: And that’s exciting. This paper opens up a whole research program. What about real-valued rewards? What about multi-agent RL? What about non-Archimedean rewards?

Jane: Exactly. And that’s why I’m glad we covered this paper. It’s not a flashy result. It’s a foundational one. It’s the kind of paper that might not get a thousand citations, but it changes how the people who do cite it think.

Tom: Well said, Jane. So that’s “Representation and Invariance in Reinforcement Learning” by Samuel Allen Alexander and Arthur Paul Pedersen. We’re going to say goodbye to this paper and get ready for the next one.

Jane: And we hope you enjoyed the discussion as much as we did. Thanks for listening, and we’ll see you next time.

Samuel Allen Alexander, Arthur Paul Pedersen

Independent Researcher · The City University of New York

cs.AI, cs.GT, cs.LG

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: 16 pages, 1 figure

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 38/100

The gist: This paper lays foundations for studying relative-intelligence-preserving mappability between reinforcement learning (RL) frameworks.

Key concepts

Invariance and Reducibility
The paper defines a rigorous notion of reducibility between RL frameworks. If a transformation exists that converts one framework into another, then relative intelligence is preserved. The authors prove that this preservation does not always hold for the standard frameworks.
RL Framework Types
These are the standard combinations of agents and environments used in practice: deterministic agents/environments and stochastic (random) versions thereof. The study tests these four types, finding they are generally not transformable into one another.
Ultrafilter
A technical tool used by the authors to rigorously compare agent intelligence. It functions as a rule for deciding who wins an 'election' among infinitely many voters (the environment), allowing for a formal comparison of intelligence.

Terminology

Summary

This paper lays foundations for studying relative-intelligence-preserving mappability between reinforcement learning (RL) frameworks. The authors introduce a criterion which is sufficient for relative intelligence to be preserved according to one particular method of measuring intelligence, and show that this criterion cannot be met when mapping between certain deterministic and stochastic RL frameworks, suggesting inherent fundamental differences between these different versions of RL.

The paper begins by noting that researchers have formalized RL in different ways, and that if an agent in one RL framework is to run within another RL framework’s environments, the agent must first be converted, or mapped, into that other framework. The authors state: "Implicit in treatments of RL is that answers to these questions are inconsequential. The problem addressed in this paper is whether this is really so. If answers to such questions are inconsequential to problems for reinforcement learning, then evaluation of agent performance — measures of their relative intelligence — would be expected to be invariant with respect to transformations between different RL frameworks."

They introduce a formal definition of an RL framework: "By a reinforcement learning framework (or RL framework) we mean a triple (A, E, V) where: 1. A is a set whose members are called agents; 2. E is a set whose members are called environments; 3. V: A×E → R is a function assigning to every agent π ∈ A and environment µ ∈ E a total expected reward Vµπ ∈ R representing how well π performs in µ."

They then define a transformation between frameworks: "Suppose F = (A, E, V) and F = (A, E, V) are RL frameworks. A transformation from F to F is a pair (•∗: A → A, •∗: E → E) of functions such that: 1. (Faithfulness) For all π, ρ ∈ A and µ ∈ E, V µ < V µ iff Vµπ∗ < Vµρ∗. 2. (Nontriviality 1) There exist π, ρ ∈ A, µ ∈ E such that V µ < V µ. 3. (Nontriviality 2) There exist π ∈ A, µ, ν ∈ E such that V µ < V ν."

For comparing intelligence, the paper uses an approach from [1] based on ultrafilters. The idea is that environments act as voters in an election comparing agents: "For any particular environment µ, if Vµπ > Vµρ, then µ votes that π is more intelligent than ρ. If Vµπ < Vµρ, then µ votes that π is less intelligent than ρ. If Vµπ = Vµρ, then µ votes that π and ρ are equally intelligent. They define an ultrafilter as a set of subsets (majorities) satisfying Properness, Monotonicity, Maximality, and ∩-closure. They define the intelligence comparator ≤U by: π ≤U ρ iff µ ∈ E: Vµπ ≤ Vµρ ∈ U."

The main preservation theorem states: "Theorem 1. (Preservation Theorem) Suppose F = (A, E, V), F = (A, E, V) are RL frameworks and (•∗: A → A, •∗: E → E) is a transformation from F to F. For any ultrafilter U on E, the transformation (•∗, •∗) preserves relative intelligence in the following sense: for all π, ρ ∈ A, we have π ≤U∗ ρ iff π ∗ ≤U ρ∗."

The paper then introduces four concrete RL frameworks differing only in whether agents and environments are deterministic or stochastic. They fix finite action set A and percept set E, with a reward function R: E → Z (integer-valued, including 0 and 1). They define deterministic agents as functions from agent histories to actions, stochastic agents as functions from agent histories to probability distributions over actions, and similarly for environments. They define expected total reward Vµπ as the limit of expected rewards over n steps, and restrict to well-behaved environments where this limit always converges.

The four frameworks are: Fdet det (deterministic agents, deterministic environments), Frnd det (stochastic agents, deterministic environments), Fdet rnd (deterministic agents, stochastic environments), and Frnd rnd (stochastic agents, stochastic environments).

The main result is: "Theorem 2. For all G, H ∈ Frnd rnd, Fdet rnd, Frnd det, Fdet det with G ̸= H, there is a transformation from G to H iff G = Fdet rnd or H = Frnd det. In other words: there is a transformation from G to H if and only if there is an arrow from G to H in Figure 1." Figure 1 shows arrows: Fdet det → Frnd det, Fdet det → Fdet rnd, Frnd det → Frnd rnd, Fdet rnd → Frnd rnd, and Fdet det → Frnd rnd (the latter via composition).

The positive parts are proved by embedding deterministic agents among stochastic agents (via a function that assigns probability 1 to the deterministic action) and deterministic environments among stochastic environments (similarly). The negative parts are proved using a Mixing Lemma that allows constructing mixtures of agents or environments. For example, Theorem 3 shows no transformation from Fdet det to Frnd rnd by constructing an infinite strictly increasing sequence of integer rewards that would be impossible. Theorem 4 similarly shows no transformation from Frnd rnd to Fdet det.

The paper concludes: "Theorem 2 suggests that, at least if rewards are limited to integers, the nature of reinforcement learning may be inherently different depending whether agents be deterministic or stochastic, and whether environments be deterministic or stochastic. We do not currently know whether Theorem 2 would remain true if arbitrary rational-number rewards were allowed. For lack of any better evidence, though, the analysis here at least urges that researchers should exercise caution before speaking about RL as if these decisions don’t matter."

The authors also note: Our high-level hope is that these results will encourage authors to be more specific, when talking about reinforcement learning, about which version of RL they mean. They acknowledge that the paper is a tentative initial step toward the difficult problem of comparing different reinforcement learning frameworks in general.

Improvements for AI systems

Based on the paper, here are specific improvements I can make to AI systems, particularly those using reinforcement learning:

  • Improvement: Implement a formal verification step that checks whether an RL agent's design framework (deterministic/stochastic agent, deterministic/stochastic environment) matches the deployment framework before training or deployment.

  • What the improved system can do: Automatically detect when an agent trained in one RL framework (e.g., deterministic agent, stochastic environment) is being deployed in an incompatible framework (e.g., stochastic agent, deterministic environment), and either reject deployment or apply a formal transformation (as defined in Definition 2) to convert the agent.

  • Improvement: Build a module that, given a source RL framework and a target RL framework, checks whether a transformation exists (per Theorem 2) and, if so, automatically applies it to convert agents.

  • What the improved system can do:

  • Convert a deterministic agent to a stochastic agent (via the embedding in Lemma 6) without loss of expected reward.

  • Convert a stochastic environment to a deterministic environment (via the embedding in Lemma 6) when needed.

  • Refuse to convert between frameworks where no transformation exists (e.g., from deterministic-agent/deterministic-environment to stochastic-agent/stochastic-environment), preventing silent performance degradation.

  • Improvement: Add a runtime check using ultrafilter-based intelligence comparators (Definition 4) to verify that after any agent conversion, relative intelligence ordering is preserved (Theorem 1).

  • What the improved system can do: Before deploying a converted agent, run a small set of benchmark environments and verify that for any pair of agents, if agent A was more intelligent than agent B in the source framework, then the converted A remains more intelligent than converted B in the target framework. If this fails, flag a framework mismatch.

  • Improvement: Add a diagnostic that detects when the reward function's codomain (e.g., integers vs. rationals vs. reals) might cause framework non-equivalence (as the paper shows for integer rewards).

  • What the improved system can do: Warn users when they switch from integer rewards to rational rewards, because the paper's Theorem 2 (which shows non-equivalence) may not hold for rational rewards, and the system cannot guarantee transformation existence. This prevents users from assuming equivalence when none is proven.

  • Improvement: Implement the Mixing Lemma (Lemma 7) as a formal operation, but with a safety check that the mixed agent's performance is exactly the weighted average of the component agents' performances.

  • What the improved system can do: When an RL system needs to combine multiple agents (e.g., for exploration vs. exploitation), it can create a mixture agent with provable performance bounds, rather than relying on ad-hoc blending that might produce unpredictable results.

  • Improvement: Build a tool that, given two RL frameworks, automatically checks whether transformations exist in both directions (making them equivalent) or only one direction, or neither.

  • What the improved system can do: Before migrating an RL system from one codebase to another (e.g., from a deterministic-agent library to a stochastic-agent library), the tool can tell the developer whether the migration is theoretically sound or whether it will introduce fundamental behavioral differences that cannot be corrected.

Suppose you have a robot controller trained as a deterministic agent in a deterministic environment (e.g., a simulated factory floor with fixed sensor readings). You want to deploy it in a real factory where sensors are noisy (stochastic environment).

Without the paper's improvements: You'd naively add noise to the sensor inputs and hope the controller still works.

With the improvements: The system would:

  1. Detect the framework mismatch (deterministic agent → stochastic environment).

  2. Check Theorem 2: transformation exists from deterministic-agent/deterministic-environment to deterministic-agent/stochastic-environment (the arrow in Figure 1).

  3. Apply the formal transformation (embedding the deterministic environment into a stochastic one via Lemma 6).

  4. Verify that relative intelligence is preserved using the ultrafilter comparator.

  5. Deploy with a guarantee that performance ordering between any two candidate controllers is unchanged.

This prevents the common failure mode where a controller that was better than another in simulation becomes worse in the real world due to framework mismatch, potentially saving millions in failed deployments.

Related papers