Illuminating the Three Dogmas of Reinforcement Learning under Evolutionary Light

arXiv:2507.11482 · cs.AI · Submitted 2025-07-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Illuminating the Three Dogmas of Reinforcement Learning under Evolutionary Light".

Jane: The paper was written by Maxwell, J. C. from Longmans, Green Company.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Paper discussion segment 2 — Tom and Jane discuss the paper's summary of the paper 'Illuminating the Three Dogmas of Reinforcement Learning under Evolutionary Light' and its implications. Explain in simple terms; do not repeat what earlier segments covered.: Tom: The abstract provides a really strong overview, saying that while current AI systems are becoming more autonomous, we need to rethink what "agency" actually means for them.

Jane: They argue that the current model—the one where learning is just about maximizing a single number or "scalar reward"—is too limited to capture all of what we mean by goals in the real life.

Lu: What I find fascinating is their argument that this lack of a formal theory of the agent itself has been sidelined because we focus so much on the environment. It's like we have a perfect map but no one knows how to drive.

Meng: From an implementation perspective, this means that if you are only optimizing for a single reward signal, you are missing huge amounts of potential behavior in complex environments.

Lalam: The paper’s summary is pointing toward the idea that we need to move away from just maximizing one number and towards a more complex, dynamic way of achieving goals.

Tom: It’s about recognizing that the current approach is too restrictive, so we're looking for a more robust framework to handle diverse challenges.

Jane: The core of this summary is that the authors are building an alternative model that is much closer to how life actually works, not just mimicking it.

Lu: It's about moving beyond "optimization" and toward a dynamic form of adaptation that allows for novelty, which is a huge concept in itself.

Meng: This suggests that if we want AI agents to be truly autonomous, they can't just be chasing a single target; they need the ability to explore and create something new.

Lalam: The implications are that we might finally have a way to build AI systems that aren't just executing instructions but truly discovering their own goals.

Paper discussion segment 3 — Tom and Jane discuss the improvements the paper suggests of the paper 'Illuminating the Three Dogmas of Reinforcement Learning under Evolutionary Light' and its implications. Explain in simple terms; do not repeat what earlier segments covered.: Tom: The authors propose a fundamental shift, specifically in T2, where they replace "optimization" with a concept they call "open-ended novelty search."

Jane: This isn's the opposite of searching for an objective solution; it's a different kind of search that is driven by diversity. They use concepts like "niching" and "coevolution."

Lu: Imagine thousands of different specialized behaviors evolving within a single population—that's what this novelty search allows, where each niche has its own set of solutions.

Meng: From an engineering standpoint, this means we could design AI systems that are inherently good at generating multiple viable strategies rather than just one "optimal" path.

Lalam: The idea is that by embracing novelty and competition within a single lifetime—which they call Darwinian neurodynamics—we allow the agent to develop diverse skills.

Tom: It's about accepting that we don't need to converge on one specific answer, but instead, exploring many different "satisficing" options.

Jane: This is a massive improvement over simply trying to find the absolute best way; it’ about finding several ways that work well enough for survival.

Lu: The paper also suggests looking at how this works across generations and then bringing those concepts back into the brain, which is mind dynamics.

Meng: If we can operationalize this "novelty search" in a practical AI setting, it could lead to agents that are much more resilient because they aren't stuck on one brittle solution.

Lalam: This opens up such possibilities for creating AI that learns and adapts like the world does, with continuous growth and change.

Discussion of key points (Not a numbered segment, but a conversational deepening): Tom: The discussion has been very focused on adaptation and goals; now we need to look at the core failures of the existing framework. The paper identifies two specific ways the reward hypothesis fails when we look at real-world biological agents.

Jane: They are looking at Axiom one Completeness, showing that because nature produces things like different species or niches, there's no single way to rank them all as "better."

Lu: It’s really about incommensurability—some solutions are just so fundamentally different from others that you can't compare them on a single scale.

Meng: From a systems perspective, this means the AI agent is facing problems where there isn't even one good way to start, so the optimization itself becomes meaningless.

Lalam: We are moving away from rigid definitions of success toward recognizing that diverse pathways can simply coexist and thrive.

Tom: That’s why they also challenge Axiom three Independence, which relates to how a single goal is maintained when multiple competing demands show up in the environment.

Jane: For instance, if you are trying to stay homeostatically balanced but suddenly face a predator, your priorities shift drastically; you can't just stick to one fixed reward function.

Lu: The system has to dynamically re-prioritize its internal state based on the external threats, which is what this independence failure describes.

Meng: In a practical sense, this means an AI needs dynamic weighting—it cannot simply ignore one goal in favor of another—to make a reliable decision at all.

Lalam: The implications for agency are profound when you realize that the ability to juggle competing demands is not just a flaw, it's actually the core of what we mean by life.

Conclusion: Tom: We’ve covered so much ground, from how AI learns to the thermodynamic requirements of life itself. The authors are making a comprehensive argument that challenging these three dogmas is necessary for creating truly autonomous agents.

Jane: It's clear they aren't just suggesting a minor fix; they are providing an entirely new framework that is deeply rooted in biological reality.

Lu: I think the idea that the system itself must be self-maintaining and heritable—that’s where the real power lies, connecting life to information.

Meng: My main takeaway is that we' can't build truly autonomous AI by simply adding more rewards; we need to ground it in a mechanism of self-maintenance first.

Lalam: The final vision from the paper is that a system must be able to sustain itself through cycles of self-renewal, which provides the necessary foundation for any kind of intelligence.

Tom: It’s a huge shift in perspective, acknowledging that we have been too reliant on external designers defining success.

Jane: Before we go, Lu, you have one last thought on "Illuminating the Three Dogmas of Reinforcement Learning under Evolutionary Light."

Lu: The idea that the interpreter comes before the template is a beautiful way to look at how information flows and evolves.

Meng: I'm excited to see if this self-maintaining, heritable design can be translated into scalable AI architectures.

Lalam: It offers a path toward designing systems that truly reflect the complexity and dynamic nature of our own biological existence.

Tom: And I think "Illuminating the Three Dogmas of Reinforcement Learning under Evolutionary Light" finally gives us a proper language to discuss what it means for an AI agent to be truly autonomous.

Jane: It really challenges us all to move past simple optimization and find a new, more complex way of thinking about intelligence.

Mani Hamidi, Terrence Deacon

University of Tübingen · University of California, Berkeley, USA Department of Anthropology

cs.AI

Submitted: 2025-07-15

Updated: 2026-08-24

Importance score: 77/100

The gist: This paper critiques the foundational tenets of Reinforcement Learning (RL) by applying insights from evolutionary biology, artificial life, and thermodynamics.

Key concepts

Scalar Reward Limitation
The traditional model of learning is too restrictive because it only focuses on maximizing a single number or 'scalar reward.' This approach fails to capture the complexity of real-life goals, meaning that optimizing for one target misses vast amounts of potential behavior in complex environments.
Open-Ended Novelty Search
This concept replaces simple optimization with a search driven by diversity, utilizing ideas like 'niching' and 'coevolution.' It allows AI systems to generate multiple viable strategies rather than converging on a single optimal path, supporting diverse skills.
Axiom Failures (Incommensurability)
The existing framework fails because nature demonstrates that solutions are often incommensurable—they cannot be compared or ranked using a single scale. This means there is no single 'better' solution, challenging the premise of optimization itself.
Dynamic Re-prioritization
A fixed goal function fails when faced with external threats. The system must dynamically re-prioritize its internal state based on these threats, meaning it cannot simply ignore one goal in favor of another to make a reliable decision.

Terminology

Summary

This paper critiques the foundational tenets of Reinforcement Learning (RL) by applying insights from evolutionary biology, artificial life, and thermodynamics. It argues that current RL lacks a formal theory of the agent and proposes reframing learning as an adaptive, open-ended process grounded in the physical requirements of self-persistence.

The three dogmas of Reinforcement Learning

The authors identify three core tenets that define current RL, which they suggest require revision to better reflect biological reality. These dogmas are:

  1. (T1) the environment formalized as a Markov decision process;

  2. (T2) learning as policy optimization on the MDP; and

  3. (T3) the agent as a goal-directed system defined by the maximization of scalar reward.

The paper argues that T1 emphasizes the environment at the expense of a formal theory of the agent, T2 treats learning as terminal search rather than adaptation, and T3's reliance on scalar reward is insufficient to capture all forms of goal-directed behavior.

Redefining learning as adaptation

To address the limitations of policy optimization, the authors recast adaptation as open-ended, objective-free novelty search rather than objective-driven terminal search. This approach utilizes two complementary engines: fitness-driven optimization of quality and a concurrent nonobjective exploration for novelty or diversity.

This second engine is driven by three specific mechanisms:

  • niching or speciation to exploit untapped resources;

  • coevolution, where organisms actively modify their environments; and

  • minimal criteria, which operate on a binary survival threshold rather than a gradient.

The failure of the reward hypothesis

The paper challenges the reward hypothesis—the idea that scalar reward maximization is sufficient to define all goals—by highlighting two axiomatic failures. First, Completeness (Axiom 1) fails because evolutionary processes produce incommensurable solutions across different niches; as the authors note, grass does not preclude grasshoppers, meaning there is no single preference ordering for all viable organisms.

Second, Independence (Axiom 3) fails at the level of instrumental goals. Biological agents manage multiple concurrent physiological drives, such as thermoregulation or pH homeostasis, where the relative urgency of each drive is inherently context-dependent. This context-dependence violates the requirement that preferences remain unaffected by the introduction of new concerns, such as a predator appearing.

Grounding agency in thermodynamics

The authors argue that explaining agency requires halting an explanatory regress by looking to the thermodynamic conditions of life. Rather than relying on evolution alone, which is nonteleological, they propose that agency must be grounded in self-persistence through replication.

They introduce the concept of autogenesis, a synergy-first paradigm where two dissipative processes are coupled to create a hologenic constraint. This mechanism allows a system to turn its dissipative activity back upon preserving the conditions of its own persistence, providing a substrate-independent recipe for spawning an evolutionary process.

Improvements for AI systems

1. Improvement: Transition from Single-Objective Optimization to Dual-Engine Adaptive Architectures (Quality Diversity + Fitness Optimization).

  • What the improved AI can do: Instead of converging on a single, potentially sub-optimal solution for a known task, the system will simultaneously run an optimization engine (to solve known objectives) and an open-ended novelty search engine (to explore unknown unknowns). This allows the agent to discover qualitatively new behaviors, create its own niches, and evolve diverse strategies that satisfy minimal survival criteria rather than just maximizing a fixed score.

2. Improvement: Replacement of Scalar Reward Functions with Multi-Dimensional Allostatic Drive Systems.

  • What the improved AI can do: The system will move away from collapsing all goals into a single scalar value (r), which causes mathematical instability and context-blindness. Instead, it will manage a vector of independent, homeostatic setpoints (e.g., energy, safety, information gain). This enables allostasis—the ability to dynamically re-prioritize objectives based on environmental context (e.g., shifting from exploration to self-preservation immediately upon detecting a threat) without the loss of information inherent in scalar aggregation.

3. Improvement: Implementation of Darwinian Neurodynamics for Intra-lifetime Learning.

  • What the improved AI can do: Rather than relying solely on global gradient-based weight updates (backpropagation), the agent will utilize selectional learning mechanisms where neural activation patterns act as replicators. These patterns will undergo competition and selection within a single lifetime/session. This allows for the rapid, emergent discovery of new motor primitives, state representations, or world models through local variation and selective retention, enabling much faster adaptation to novel environments than traditional policy optimization.

4. Improvement: Integration of Autogenic (Hologenic) Constraint Frameworks for Agency.

  • What the improved AI can do: The system will transition from a passive learner with borrowed objectives to an autonomous agent with intrinsic normativity. By coupling two dissipative processes (e.g., energy acquisition and structural maintenance) such that they mutually constrain one another, the agent develops a hologenic constraint. This provides the agent with a fundamental, non-programmed drive for self-persistence and structural integrity against environmental entropy, providing a principled foundation for truly embodied and autonomous AI.

Sources

Related papers