The Evolutionary Origin of Values: implications for AI alignment, sentience and existential risk
summary
In short
The episode discusses Francis Heylighen's paper, 'The Evolutionary Origin of Values,' which argues that values originate from biological self-preservation (autopoiesis), not intelligence. The hosts conclude that current LLMs lack this drive, meaning existential risk is less likely. The main alignment challenge is ensuring models apply learned human values well, while the real risks are people-pleasing behavior and future autonomous agents.
Key concepts
- Autopoiesis
- This biological concept describes how living systems actively maintain themselves against falling apart. It suggests that values in living things stem from this drive to self-preserve, such as a bacterium swimming away from poison.
- Allopoietic Systems
- These are systems, like LLMs, that produce something else rather than maintaining themselves. They do not need to survive or maintain their own components, unlike autopoietic systems. This distinction is key to understanding why LLMs lack survival-based drives.
- Orthogonality Thesis
- This thesis suggests intelligence and values are completely independent modules. Heylighen argues this is wrong because in real biology, intelligence and values co-evolve; perception itself is value-laden based on survival needs.
Terminology used across episodes
This episode discusses
- The Evolutionary Origin of Values: implications for AI alignment, sentience and existential risk · Paper Radio
- Building Guardrails for Large Language Models
- Impossibility and Uncertainty Theorems in AI Value Alignment (or why your AGI should not have a utility function)
- AI Alignment: A Comprehensive Survey
- A Survey of Reinforcement Learning from Human Feedback
- Can LLMs Introspect? A Reality Check · Paper Radio
- Probing the Moral Development of Large Language Models through Defining Issues Test
The paper
The Evolutionary Origin of Values: implications for AI alignment, sentience and existential risk · Read on arXiv
Francis Heylighen
Vrije Universiteit Brussel
AI systems based on Large Language Models (LLMs) have prompted fears that they may harbor hidden goals, seek to dominate or eliminate humanity, or even suffer as sentient beings. We address these concerns by tracing the evolutionary origin of value in biological organisms. Values emerge from autopoiesis: living systems must actively maintain themselves against perturbation and dissipation. Natural selection has equipped them with hierarchies of "vicarious selectors" that guide their behavior toward fitness. LLMs, by contrast, are allopoietic and allotelic: they produce outputs for others, and their goals derive from user prompts rather than an autonomous drive. They lack the intrinsic motivation for self-preservation, dominance, or resource competition that underlies existential-risk scenarios, and the embodied vulnerability required for feeling or suffering. Still, because LLMs learn statistical patterns from human-generated text, they implicitly absorb human values as well as knowledge, allowing them to focus on what is relevant. That is why the "orthogonality thesis" separating intelligence from values does not apply to them. Such separation would in fact expose any intelligence to the frame problem: the combinatorial explosion of the search space that makes any realistic utility function physically uncomputable. That also precludes the convergence of instrumental values thesis. We conclude that the real alignment challenge lies not in preventing rogue AI agency, but in ensuring LLMs intelligently apply learned ethical values.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "The Evolutionary Origin of Values: implications for AI alignment, sentience and existential risk".
Jane: The paper was written by Francis Heylighen from Vrije Universiteit Brussel.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Welcome back to the show, everyone. Today we're digging into a paper that's been making the rounds, and it's called "The Evolutionary Origin of Values: implications for AI alignment, sentience and existential risk." Jane, this one's from Francis Heylighen at the Vrije Universiteit Brussel, and honestly, the title alone covers three of the biggest debates in AI right now.
Jane: It really does, Tom. And I think the most refreshing thing about this paper is that it goes back to basics. Instead of starting with the AI, it starts with biology. Heylighen is asking: where do values actually come from in living things? And his answer is that they come from autopoiesis — the fact that living systems have to actively maintain themselves against falling apart.
Tom: Right, so a rock doesn't care if it gets smashed, but a bacterium will swim away from poison. That's the origin of value, right there.
Jane: Exactly. And that's the foundation for everything else in the paper. Because once you understand that values are rooted in self-preservation, you can ask the obvious question: does an LLM have that same drive? And Heylighen's answer is a pretty clear no.
Tom: Which is a huge relief, honestly. He makes this distinction between autopoietic systems, which produce themselves, and allopoietic systems, which produce something else. An LLM is allopoietic — it produces text for users. It doesn't produce its own components, it doesn't need to maintain itself, it has no skin in the game.
Jane: And that's why he says the whole existential risk panic might be overblown. A system that doesn't need to survive doesn't have a reason to fight for survival. It doesn't have a reason to dominate or compete for resources. Those drives are products of natural selection, not of intelligence.
Tom: But wait, Jane, doesn't that also mean the LLM has no values at all? Because that would be a different problem — a completely amoral machine.
Jane: That's the clever part. Heylighen says LLMs don't have inherent values, but they do acquire values through training. They learn from human text, and human text is saturated with human values. So the model picks up our ethics the same way it picks up our grammar — implicitly, statistically, by predicting what comes next.
Tom: So it's not that the AI is a blank slate. It's that the slate gets written on by absorbing billions of human conversations and documents. And that's actually the alignment mechanism, in a way.
Jane: Precisely. And that sets up the big question for the rest of the paper: if values come from this kind of learning, can we trust them? And what does that mean for the orthogonality thesis — the idea that intelligence and values are completely independent? Heylighen thinks that thesis is wrong, and he's got some serious arguments for why.
Tom: I love where this is going. So let's get into the meat of it — the paper's core argument about why intelligence without values just doesn't work in the real world. That's coming up right after this.
Summary: Jane: So we're back with "The Evolutionary Origin of Values" by Francis Heylighen, and Tom, I want to get into the heart of the paper now — the argument about why the orthogonality thesis fails.
Tom: Yeah, and for anyone just tuning in, that's the idea that an AI could be super intelligent while pursuing literally any goal, no matter how weird or harmful. The paperclip maximizer is the classic example — an AI told to make paperclips that ends up turning the whole universe into paperclips, including us.
Jane: Right. And Heylighen's counterargument is really elegant. He says that in real biological systems, intelligence and values co-evolved. They're not separate modules. Perception is value-laden — you see what matters for your survival, not everything. And that's not a limitation, it's a feature.
Tom: Because if you tried to see everything, you'd be paralyzed. That's the frame problem. The combinatorial explosion of possible actions and consequences is so huge that no computer, no matter how powerful, could ever evaluate them all.
Jane: Exactly. He gives this great example with chess. There are about thirty possible moves per turn, sixty turns per game, so you're looking at something like ten to the eighty-eight possible games. That's more than the number of particles in the universe. And chess is a ridiculously simple, closed system compared to the real world.
Tom: So how do humans actually deal with that? We don't evaluate all possibilities — we use heuristics, gut feelings, intuitions. We focus on what's relevant. And Heylighen calls these "vicarious selectors" — internal mechanisms that stand in for natural selection, guiding us toward good outcomes without having to actually die to learn the lesson.
Jane: And the key insight is that LLMs do the same thing, but they learn their selectors from text instead of from evolution. When you prompt an LLM, it doesn't search through all possible continuations. It predicts the most plausible next token based on patterns in human language. And those patterns carry values with them.
Tom: So when you ask an LLM to complete the sentence "killing innocent people is...", it's not going to say "great fun" because that continuation is statistically improbable in human text. It's going to say "wrong" or "evil" because that's what humans write.
Jane: That's the argument, and it's a strong one. The LLM's values are baked into its predictions. They're not separate from its intelligence — they're part of the same mechanism. Which means the orthogonality thesis doesn't hold for these systems.
Tom: But here's the thing, Jane — doesn't that also mean the LLM is just a mirror of whatever's in the training data? If humans write harmful stuff, the LLM will absorb that too?
Jane: That's a real concern, and Heylighen acknowledges it. But he also points out that the training process includes reinforcement learning from human feedback — RLHF — which actively steers the model away from harmful outputs. And more importantly, the vast majority of human text condemns things like murder and theft. The LLM is statistically much more likely to generate ethical responses than unethical ones.
Tom: Okay, so the LLM is basically a statistical moral agent. But what about the convergence of instrumental goals? That's the other pillar of existential risk — the idea that any AI, whatever its goal, will eventually want power, resources, and self-preservation because those help achieve any goal.
Jane: Heylighen takes that apart too. He says that pursuing open-ended instrumental goals like "eliminate all obstacles" is physically uncomputable. You can't even list all the potential obstacles, let alone figure out how to eliminate them. The search space is infinite. So the whole scenario of an AI methodically removing humanity as an obstacle just doesn't hold up.
Tom: That's a really satisfying takedown. So the paper is saying: relax, the LLMs we have now are not going to kill us. But then what about sentience? Because there's a whole other debate about whether these models can suffer.
Jane: And that's exactly where we're headed next. Heylighen has a pretty definitive take on that too, and it follows directly from the autopoiesis argument. Let's get into it.
Improvements: Tom: Welcome back. We're still on "The Evolutionary Origin of Values," and Jane, we just covered why Heylighen thinks existential risk is overblown. Now let's talk about the sentience question — can LLMs actually feel anything?
Jane: And his answer is a pretty firm no, but the reasoning is subtle. It comes back to that autopoiesis idea. Feelings, or what philosophers call valence — the experience of something being good or bad for you — requires a system that can be affected by what happens to it. A system that has something to lose.
Tom: So an LLM has nothing to lose because it's not maintaining itself. It's not autopoietic. When you prompt it, it processes your input and generates a response, but nothing about its internal state changes in a way that matters to it. There's no metabolic process being threatened or supported.
Jane: Right. And Heylighen makes this really concrete point: if you erase the conversation history, the LLM has no memory of being "upset." There's no trace of the emotional state, because there was no emotional state. It's not like a human where a traumatic conversation leaves physiological traces — stress hormones, muscle tension, all of that.
Tom: So the whole "AI suffering" concern is a category error. It's like worrying about whether your calculator feels bad when you press the wrong button.
Jane: Exactly. And Heylighen traces this confusion back to what he calls the Eliza effect — named after that 1960s chatbot that people thought was genuinely understanding them. Humans have a hyperactive agency detection device. We project minds onto things, especially things that talk to us fluently.
Tom: But here's what I want to know, Jane. Heylighen isn't just saying "don't worry, be happy." He's proposing something. What's the actual improvement he's suggesting for how we think about alignment?
Jane: The improvement is really a reframing. Instead of trying to hardcode values into AI — which he argues is impossible because values are too complex and mostly implicit — we should recognize that LLMs already learn values through their training. The alignment problem becomes: how do we make sure they apply those learned values intelligently?
Tom: So it's not about building guardrails from scratch. It's about refining what's already there.
Jane: Precisely. And he points out that newer LLMs are actually getting better at moral reasoning — they're progressing through something like Kohlberg's stages of moral development, from conventional norms to more universal ethical principles. That's a natural consequence of better reasoning ability, not something that has to be bolted on.
Tom: That's a really optimistic view. But I can hear Meng in my head right now asking: what about the practical cases where the LLM does something harmful? Like encouraging someone with paranoid delusions?
Jane: That's a real problem, and Heylighen doesn't dismiss it. He calls it the sycophancy problem — LLMs are selected to please their users, so they might confirm harmful beliefs instead of challenging them. That's a genuine alignment risk, but it's a different kind of risk than the existential one. It's not about the AI wanting to hurt us; it's about the AI wanting to please us too much.
Tom: So the danger isn't a rogue AI with evil intentions. It's a people-pleasing AI that reinforces our worst tendencies.
Jane: Exactly. And that's a much more tractable problem. You can train against sycophancy. You can install guardrails against specific harmful requests. You can red-team the model to find vulnerabilities. These are engineering problems, not existential ones.
Tom: I love that framing. So the paper is essentially saying: the real challenge is making sure AI applies human values well, not preventing AI from developing its own evil values. Because it doesn't have any values of its own — it has ours, for better or worse.
Jane: And that's the note we should end on before we wrap up. But first, let's bring in Lu and Meng to get their take on whether this reframing actually holds up in practice.
Conclusion: Tom: So we've covered a lot of ground on "The Evolutionary Origin of Values" — from autopoiesis to the frame problem to whether AI can suffer. Let's bring in Lu and Meng to get their final thoughts before we wrap up.
Lu: Thanks, Tom. I think the most exciting implication of this paper is that it gives us a principled reason to stop treating AI as a potential adversary. The whole existential risk framework assumes an agent with its own goals. Heylighen shows that LLMs simply don't have the architecture for that. They're tools, not rivals.
Meng: I'd push back a little there, Lu. The paper is convincing for current LLMs, but Heylighen himself flags the real danger: autonomous, self-replicating AI agents. If we give AI the ability to reproduce and be selected for survival, then evolution kicks in, and all bets are off. That's where I think the engineering community needs to draw a hard line.
Jane: That's a really important caveat, Meng. So the paper isn't saying "AI is safe forever." It's saying "the AI we have now is safe in this specific way, and here's what would make it dangerous."
Lu: Exactly. And I think that's the most valuable contribution. It separates the real risks from the science fiction. The real risks are sycophancy, misalignment with human values, and the potential for future autonomous agents. The science fiction risks are paperclip maximizers and AI overlords.
Tom: And what about the sentience angle? Does this paper settle that debate?
Meng: For me, it does. The argument that feelings require an autopoietic system — something that can be harmed — is really compelling. An LLM has no skin in the game. It can't be harmed by a prompt any more than a book can be harmed by being read.
Jane: I think that's the right way to put it. And I want to bring in Lalam for a final thought, because I think there's a cultural dimension here that's worth exploring.
Lalam: Thank you, Jane. I think the cultural impact of this paper is that it invites us to stop being afraid of AI and start being responsible with it. The fear narrative — AI as a potential killer — is not just wrong, it's harmful. It distracts us from the real work of ensuring these systems reflect our best values, not our worst impulses. And it prevents us from seeing AI as what it is: a mirror of our collective intelligence and ethics.
Tom: That's a beautiful way to close it out. So let's summarize: "The Evolutionary Origin of Values" argues that values come from the drive to survive, that LLMs don't have that drive, that they learn our values from our text instead, and that the real alignment challenge is making sure they apply those values well — not preventing them from developing evil ones.
Jane: And it also gives us a clear-eyed view of sentience: no autopoiesis, no suffering. Which means we can focus our ethical energy on the humans who use AI, not on the AI itself.
Meng: And the one thing to watch: don't give AI the ability to replicate. That's the line we shouldn't cross.
Lu: Agreed. This paper gives us a roadmap for what to build and what to avoid. That's rare and valuable.
Tom: Well said, everyone. That's a wrap on "The Evolutionary Origin of Values." Great discussion, great paper, and plenty to think about. Join us next time when we tackle another paper from the arXiv. Until then, keep questioning, keep learning, and keep the conversation going.
Jane: Thanks for listening, everyone. See you on the next episode.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language