An evolutionary model of animats with VLM-based subjective evaluation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "An evolutionary model of animats with VLM-based subjective evaluation".
Jane: The paper was written by Shota Miyazaki, Takaya Arita and Reiji Suzuki from Graduate School of Informatics, Nagoya University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Alright, welcome back to the show, everyone. Today we’re digging into a paper with a title that’s a mouthful: “An evolutionary model of animats with VLM-based subjective evaluation.” Jane, I’m going to need you to break that down for me, because I’ve got about three buzzwords stuck in my throat.
Jane: Happy to, Tom. So, an “animat” is basically a virtual animal — a soft robot made of blocks that can move and deform. And “VLM” stands for Vision-Language Model, which is an AI that can look at images and understand them using language. So the paper is about using one of those models to decide which virtual creatures are “better” in an evolutionary sense.
Tom: So instead of a human sitting there judging which creature looks cool, they’re using an AI to do the judging?
Jane: Exactly. And that’s a big deal, because traditionally, if you want to evolve things based on subjective feelings — like “this one looks adorable” — you need a human in the loop. And humans get tired. They get bored. They make inconsistent choices after the fiftieth round.
Lu: Right, and that fatigue problem is real. The paper cites Interactive Evolutionary Computation, or IEC, as the old approach. Dawkins’ Biomorphs, where people picked shapes they liked, was a classic example. But you can’t scale that up because human evaluators are a bottleneck.
Tom: So they’re replacing the human judge with an AI judge. But wait, can an AI really understand “adorable” or “weird”? Isn’t that a stretch?
Jane: That’s exactly the question the paper is trying to answer. They ran experiments where the AI had to pick between two animats based on terms like “adorably,” “weirdly,” and “solemnly.” And the results show that the AI does push evolution in specific directions — the creatures end up looking different depending on the word used.
Meng: But hold on, Jane. From an engineering standpoint, I want to know how they actually got the AI to do this. You can’t just show it a video and ask “which one is weirder?” — that’s computationally expensive.
Jane: Good point, Meng. They used still images — four snapshots of each creature in a row, showing its motion over time. Then they stacked two of those rows, one on top of the other, and asked the model to compare them. That way, one image can represent the motion of two individuals, and the AI can make a judgment.
Tom: So it’s like a comic strip for robots. And the AI reads the comic and says, “Yeah, the top one moves more adorably.” That’s wild.
Lu: It is, and it gets even more interesting when you look at what the AI says it’s basing its decision on. They asked the model to output a one-word criterion for its judgment. For “adorably,” the top criterion was “fluidity.” For “weirdly,” it was “disjunction.” So the model is decomposing these abstract words into concrete visual features.
Tom: So the AI isn’t just guessing — it’s actually building an internal checklist for what “weird” looks like in motion. That’s a whole new way of thinking about how these models understand language.
Jane: And that’s what we’re going to dig into next — how the evolution actually played out, and what kinds of creatures emerged from each word. Stick around, because the results are genuinely surprising.
Summary: Tom: Welcome back. We’re still on “An evolutionary model of animats with VLM-based subjective evaluation,” and Jane, you promised me surprising results. Let’s hear them.
Jane: So the core experiment was simple: they evolved populations of thirty virtual creatures for fifty generations. Each generation, the AI compared pairs of creatures and picked the winner based on a subjective term. Then the winners “reproduced” with crossover and mutation, just like in natural evolution.
Tom: And the key question is whether the AI’s judgment actually steers the evolution, right? Because if it’s just random noise, the creatures won’t converge on anything.
Jane: Right. And they compared their results against a control where the winner was chosen randomly. The random selection showed a slow, steady decrease in genetic diversity. But with the AI’s subjective selection, the diversity dropped much faster — meaning the population converged on specific traits much more quickly.
Meng: So the AI is applying real selection pressure. But faster convergence isn’t necessarily good, is it? You could converge on something that doesn’t match the term at all.
Jane: That’s why they did a second test. They took creatures from the first generation and the final generation, paired them up, and asked the AI which one better matched the term. In every condition, the final-generation creatures won more often. So the evolution was actually moving toward the target, not just collapsing into a random blob.
Tom: But here’s the thing — what does “adorable” actually look like to this AI? Because I bet it’s not what I’d pick.
Jane: That’s the fascinating part. Under “adorably,” they got two distinct types. Some trials produced creatures with leg-like structures that walked like small animals. Others produced blob-like creatures that bounced and deformed softly. So the AI had two different internal notions of adorable, and different trials went down different paths.
Lu: And that’s actually a really important finding. The paper shows that the VLM doesn’t have a single fixed definition of these words. It decomposes them into multiple criteria, and which one dominates depends on the initial random population.
Tom: So it’s like asking five different people what “cute” means — you’ll get different answers, but they’re all valid.
Jane: Exactly. And for “weirdly,” the creatures evolved slit-like structures that opened and closed, like the body was splitting apart. For “solemnly,” they got near-square shapes with hollow interiors that barely moved. The AI was mapping these abstract words onto concrete visual features.
Meng: I’m curious about the practical side, though. How many times did the AI have to make a judgment in one run? Because if it’s thousands of inferences, that’s a real computational cost.
Jane: They used a quantized version of Gemma three which is a smaller model that can run locally. And by using still images instead of videos, they kept the cost down. But it’s still a lot of inference — every generation, thirty pairwise comparisons, each done twice to check consistency.
Tom: So it’s not cheap, but it’s way cheaper than hiring a human to sit through fifty generations of robot evolution.
Jane: Exactly. And that’s the whole point — automating subjective judgment in a way that scales. But the really interesting part is what happens when you compare the AI’s choices to a human’s. That’s coming up next.
Improvements: Tom: We’re back with “An evolutionary model of animats with VLM-based subjective evaluation,” and Jane, you mentioned the paper suggests improvements over the old way of doing things. Let’s talk about that.
Jane: Right. The old way was Interactive Evolutionary Computation, where a human sits and makes choices. The paper’s improvement is replacing that human with a VLM. But they didn’t just swap the evaluator — they had to design the whole evaluation process to make it work reliably.
Meng: And that’s where I want to dig in. Because a VLM isn’t a deterministic function. It can be inconsistent. How did they handle that?
Jane: They used a clever trick. For each comparison, they showed the AI the same two creatures twice — but with the top and bottom rows swapped. The answer was only accepted if the AI picked the same creature both times. If it flip-flopped, that was counted as a “draw” and no selection pressure was applied.
Tom: So they’re basically testing the AI’s consistency before trusting its judgment. That’s smart.
Jane: It is. And they also designed the prompt to force the AI to reason step by step. First describe the shapes, then describe how they change, then define what the subjective term means in this context, and only then make a choice. That reduces the chance of the AI just guessing.
Lu: And this is a genuine improvement over naive approaches. If you just ask a VLM “which is weirder?” it might latch onto irrelevant features like color or position. By forcing it to focus on shape transformation over time, they’re steering it toward the actual motion characteristics.
Tom: So the improvement isn’t just “use an AI instead of a human” — it’s a whole protocol for getting reliable subjective judgments out of the AI.
Jane: Exactly. And the results show it works. The evolution converged faster than random selection, and the final creatures were judged to better match the target terms. But there’s a caveat — the paper also ran a human experiment to see how well the AI agrees with real people.
Meng: Oh, this is where it gets interesting. How much did they agree?
Jane: Not as much as you’d hope. For “adorably,” the agreement was below chance — forty-three percent — meaning the AI and humans were picking different creatures more often than not. For “weirdly” it was sixty-two percent, and for “solemnly” it was fifty-six percent. So there’s some overlap, but it’s far from perfect.
Tom: So the AI isn’t a perfect substitute for human taste. But is that actually a problem?
Lu: That depends on what you want. If you’re trying to replicate human subjectivity, then yes, it’s a problem. But the paper argues that’s not the goal. The goal is to explore what these subjective terms mean to the AI itself — to understand how a VLM maps language onto visual features.
Jane: And that’s a different research question entirely. Instead of asking “can the AI judge like a human?”, they’re asking “what does the AI think these words mean?” And that’s what we’ll dig into next — the actual evaluation criteria the AI used.
First Page: Tom: Welcome back. We’re deep into “An evolutionary model of animats with VLM-based subjective evaluation,” and Jane, you just teased that the AI’s internal criteria are the real story. Let’s unpack that.
Jane: So the paper’s first page sets up the whole problem: how do you evolve things when the fitness function is subjective? And the answer they propose is to use a VLM as the judge. But the fascinating part is what the AI actually said it was looking for.
Tom: Give me an example. What did the AI say it was judging on?
Jane: For “adorably,” the top criterion was “fluidity,” appearing in forty-six percent of judgments. Then “whimsy” at nineteen percent and “gentleness” at fifteen percent. So the AI was mostly looking for smooth, flowing motion. For “weirdly,” the top criterion was “disjunction” at seventeen point five percent, followed by “abruptness” and “discontinuity.”
Lu: And that maps directly onto the evolved morphologies. The “weird” creatures had slit-like structures that opened and closed — that’s disjunction, the body appearing to split apart. The “adorable” creatures either walked with leg-like structures or bounced softly — both of which involve fluid, continuous motion.
Tom: So the AI’s words and the creatures’ bodies line up. It’s not just making random choices — it’s consistently applying its internal criteria.
Jane: Exactly. And they even computed Jaccard indices to measure how similar the criteria were between different subjective terms. The highest similarity was between “solemnly” and “motionlessly” at zero point three one, which makes sense — both involve minimal movement. And “weirdly” and “dynamically” were also similar, both involving large, unpredictable shape changes.
Meng: So the AI is grouping these words by underlying visual features, not by their literal dictionary meanings. “Solemnly” and “motionlessly” aren’t synonyms, but the AI treats them similarly because both lead to similar motion patterns.
Jane: Right. And that’s a really interesting finding about how VLMs understand language. They’re not just matching words to definitions — they’re mapping words onto visual and motion concepts.
Tom: So this paper is as much about understanding the AI as it is about evolving robots.
Jane: That’s the deeper implication. The authors say this framework lets them visualize how subjective linguistic expressions get mapped onto embodied phenotypes. You’re literally watching the AI’s interpretation of “weird” take physical form.
Lu: And there’s a cultural angle here too. Different VLMs, trained on different data, might have different internal criteria for the same word. So this could become a tool for studying how language models from different backgrounds understand subjective concepts.
Tom: That’s a big idea. But before we wrap up, I want to hear about the human experiment — how did real people react to doing this task?
Conclusion: Tom: Alright, we’ve covered a lot of ground on “An evolutionary model of animats with VLM-based subjective evaluation.” Let’s pull it all together before we say goodbye.
Jane: The core finding is that a Vision-Language Model can act as a subjective judge in evolutionary computation. It steers populations of virtual creatures toward morphologies and motions that match abstract terms like “adorably” and “weirdly,” and it does so consistently enough to produce clear convergence.
Tom: And the really interesting part is that the AI’s internal criteria — the one-word summaries it gives for its judgments — align with the actual evolved bodies. “Weird” creatures have splitting motions, “solemn” creatures barely move, “adorable” creatures flow.
Lu: The paper also shows that this isn’t about replacing human judgment. The agreement between the AI and human participants was modest — around forty-three percent to sixty-two percent depending on the term. But the evolved morphologies were qualitatively similar, which suggests the AI and humans are picking up on overlapping visual features even when they disagree on individual choices.
Meng: And from a practical standpoint, the human experiment confirmed that doing this manually is exhausting. Participants reported fatigue and difficulty, especially for abstract terms like “solemnly.” So automating subjective evaluation with a VLM isn’t just convenient — it opens up experiments that would be impractical with human judges.
Jane: The limitations are clear too. They used a single VLM, a quantized Gemma three with a fixed prompt. Different models or different prompts could give different results. And they used still images instead of video, which compresses the temporal information.
Tom: So this is really a first step toward a new way of doing evolutionary computation — one where the fitness function is subjective and the judge is an AI.
Jane: And it’s a step toward understanding how VLMs think about abstract language. When the AI says “weird” means “disjunction,” that tells us something about how the model has learned to connect words to visual concepts.
Tom: Well, we’ve had a great time with this one. Thanks to Lu and Meng for joining us, and to our listeners for sticking around. Next time we’ll be looking at another paper from the arXiv, so stay tuned. Goodbye, everyone.
Jane: Goodbye!
Shota Miyazaki, Takaya Arita, Reiji Suzuki
Graduate School of Informatics, Nagoya University
cs.NE, cs.AI, cs.CL, cs.HC, cs.MA
Submitted: 2026-07-29
Updated: 2026-08-11
Comments: 18 pages, 12 figures, 2 tables. This manuscript has been accepted for publication in Artificial Life and Robotics following peer review
Journal ref: Shota Miyazaki, Takaya Arita and Reiji Suzuki: An evolutionary model of animats with VLM-based subjective evaluation, Artificial Life and Robotics (2026). https://link.springer.com/article/10.1007/s10015-026-01135-4
DOI: 10.1007/s10015-026-01135-4
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 68/100
Key concepts
- Animat
- An animat is a virtual animal made of soft robot blocks that can move and deform. The paper uses these virtual creatures to test an evolutionary model driven by AI judgment.
- VLM (Vision-Language Model)
- A VLM is an AI capable of looking at images and understanding them using language. In this study, the VLM acts as the judge to decide which animat is 'better' based on subjective terms.
- Subjective Evaluation
- This refers to judging something based on personal feelings or opinions, like whether a creature looks 'adorable.' The paper investigates how an AI can be trained to make these subjective judgments.
- Criterion Decomposition
- The VLM does not have a single definition for words; it decomposes abstract terms into multiple concrete visual features. For example, 'adorably' was linked to 'fluidity' and 'gentleness,' showing how the AI maps language onto motion concepts.
Terminology
Summary
Summary
This paper proposes a framework that incorporates subjective evaluations provided by a Vision–Language Model (VLM) into the fitness evaluation and selection processes of a genetic algorithm, enabling the evolution of animats whose morphologies and behaviors reflect given subjective linguistic expressions. The authors employ virtual soft robots with flexible morphologies and locomotion, implemented using Evolution Gym, and present the VLM with sequence images representing the locomotion of two individuals. Selection is performed via pairwise comparisons based on subjective evaluation terms such as adorably and weirdly, with the outcomes used as selection pressure within the genetic algorithm, enabling the simultaneous evolution of morphology and locomotion.
The animats are composed of square blocks (voxels) arranged within a W times W grid, with four types of voxels: rigid, soft, horizontal actuator, and vertical actuator. The genotype is a one-dimensional list of integers corresponding to voxel types and, for actuator voxels, real-valued phase parameters determining the timing of expansion and contraction. For evaluation, four snapshot images of an animat at predetermined time points are arranged horizontally in chronological order, and two such sequences are stacked vertically to form a single image provided to the VLM along with a prompt. The prompt instructs the model to act as an impartial motion analysis specialist, describe the shape in each frame, describe how the shape transforms between frames, define the subjective evaluation term in the context of motion, and then perform the subjective evaluation. To prevent order-dependent decisions, the same evaluation is performed again with the vertical positions swapped, and the selection is accepted only when the model consistently selects the same individual.
Experiments were conducted with a population size of N = 30, a total of G = 50 generations, a crossover probability of p cross = 0.05, and a mutation probability of p mut = 0.01. The VLM used was a 4-bit quantized version of gemma-3-12b-it with temperature set to 0. Five conditions were tested using evaluation terms: adorably, weirdly, solemnly, dynamically, and motionlessly. The results showed that VLM-based subjective selection accelerated population convergence compared to random selection, with the mean Hamming distance between genotypes falling below 20 in the final generation for all conditions, whereas random selection reached 33.9. Post-hoc pairwise comparisons between initial and final generation individuals showed that final-generation individuals were selected more frequently under all conditions, indicating adaptation toward the given subjective terms. However, the magnitude of this effect varied: for adorably and solemnly, the number of ties was close to the number of selections favoring final-generation individuals, while for weirdly, dynamically, and motionlessly, the number of ties was relatively small and the selection advantage was larger.
The evolved animats exhibited distinctive morphologies and motions corresponding to each evaluation term. Under adorably, two types of individuals emerged: those with leg-like structures exhibiting walking-like motions, and those with blob-like morphologies deforming in a bouncing manner. Under weirdly, slit-like structures and motions resembling opening, closing, and expansion of slits were observed. Under solemnly, individuals with near-square outlines and internal hollow structures emerged. The dynamically condition produced large body-expanding movements similar to weirdly, while motionlessly produced near-square morphologies similar to solemnly. Voxel composition analysis showed that weirdly and dynamically conditions exhibited the highest average number of vertical actuator voxels, while solemnly exhibited the highest average number of soft voxels and motionlessly the highest average number of rigid voxels.
The evaluation criteria terms output by the VLM during comparisons were aggregated. For adorably, the criteria were strongly biased toward fluidity (0.461), with the top three criteria accounting for more than 80% of occurrences. For weirdly, the most frequent criterion was disjunction (0.175), and the Jaccard index with dynamically was the highest at 0.192, with both sharing abruptness and unpredictability in their top five criteria. For solemnly, the most frequent criterion was deliberation (0.241), and the Jaccard index with motionlessly was the highest at 0.312, with both sharing regularity and continuity. The authors note that "the VLM does not apply evaluation terms in a purely literal manner, but instead decomposes them into internal evaluation criteria, and that similarities among these criteria determine similarities in evolutionary outcomes."
An auxiliary human-subjective IEC experiment was conducted with nine undergraduate and graduate students who were native speakers of Japanese, focusing on three conditions: adorably, weirdly, and solemnly. Each trial was evaluated by a different participant, with 20 selection generations. The agreement rates between VLM and human choices in initial-generation tournament selections were 43.3% for adorably (below chance), 62.2% for weirdly, and 55.6% for solemnly. Population-level convergence was comparable between human-based and VLM-based evaluation by the 20th generation. Under adorably, humans evolved animats resembling animal eyes or bodies, with free-description responses mentioning an appearance of joy,
shape transformations resembling an animal face,
and small-animal-like movements.
Under weirdly, humans evolved animats with widely spreading bodies, with participants focusing on geometric changes in shape,
changes in motion,
and large changes in angle.
Under solemnly, humans evolved near-square morphologies with high density and slight deformation, with participants focusing on the small amount of motion
and the lack of empty space.
The questionnaire results indicated that repeated pairwise selection imposed a clear burden on participants, with fatigue increasing and concentration decreasing, and that perceived difficulty varied by term, being relatively low for adorably and higher for solemnly.
The authors identify several limitations. First, motion was evaluated using four static snapshots rather than full video input, which compresses temporal information and may not fully capture rhythm, velocity changes, and the distinction between abrupt and gradual deformation. Second, the number of participants and evaluation terms in the human experiment was limited. Third, the study depends on the particular VLM used, and the evolved phenotypes "should not be interpreted as reflecting universal human subjectivity or general notions of adorableness, weirdness, solemnity, dynamic motion, or motionlessness. Rather, they should be understood as phenotypes favored under the subjective selection pressure induced by this particular VLM in the present experimental setting." Future work should compare evolutionary outcomes across multiple VLMs and prompting conditions, and the framework has potential applicability to cultural evolutionary computation by introducing multiple evaluators with different personas.
Improvements for AI systems
Based on this paper, here are the specific improvements I can make to AI systems:
Improvement: Implement the paper's pairwise comparison protocol with position-swap verification. The system evaluates two candidates, then re-evaluates with positions swapped, and only accepts the judgment if the model consistently selects the same candidate. This eliminates order bias and ambiguous responses.
What the improved system can do: Provide reliable subjective evaluations for tasks where absolute scoring is impossible (e.g., which animation is more elegant?
), with built-in rejection of uncertain judgments rather than forcing arbitrary decisions.
Improvement: Adopt the paper's multi-stage prompt: (1) objective frame description, (2) change pattern analysis, (3) explicit definition of the subjective term in context, (4) identification of relevant characteristics, (5) final judgment with justification. This forces the VLM to decompose abstract terms into observable features before deciding.
Improvement: Replace video input with 4 chronological static snapshots arranged in a grid. This reduces computational cost while preserving temporal dynamics. The paper cites evidence that discrete snapshots outperform continuous video for motion understanding.
Improvement: Add a Key Criterion
output field that forces the VLM to name its primary evaluation dimension (e.g., fluidity,
disjunction,
deliberation
). Aggregate these criteria across evaluations to reveal the model's internal decision structure.
Improvement: Integrate VLM-based pairwise comparison as the selection operator in a genetic algorithm, replacing explicit fitness functions. Use tournament selection with winners as parents, and apply crossover/mutation with constraint repair.
Improvement: After evolution, compare initial vs. final generation individuals via the same pairwise protocol. Count selections favoring final-generation individuals to verify that evolution actually progressed toward the subjective goal, rather than merely converging.
Improvement: Encode evolved individuals using CLIP embeddings of their sequence images, then project into 2D with UMAP. This reveals clustering patterns and trial-to-trial divergence in phenotype space.
Improvement: Implement the paper's methodology for comparing VLM judgments with human judgments: same initial populations, same tournament pairs, compute agreement rates per term, and compare final phenotypes via embedding-based similarity.
Improvement: Use the paper's finding that human evaluators show fatigue and term-dependent difficulty to design adaptive evaluation systems that limit repeated judgments, alternate terms, and provide breaks—while VLM-based evaluation eliminates this bottleneck entirely.
Improvement: Apply the paper's framework to test whether different VLMs or different personas (occupation, age, cultural background) produce similar or divergent evolutionary outcomes for the same subjective term. Use Jaccard indices of extracted evaluation criteria to quantify overlap.
Sources
- Automating the Search for Artificial Life with Foundation Models
- Gemma 3 Technical Report
- UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
Related papers
- Evolutionary Ensemble of Agents
- Encoding and Decoding Temporal Signals with Spiking Bandpass Wavelets
- Large Language Models and Evolutionary Computation: A Critical Review of Bidirectional Interaction, Automated Algorithm Design, and Co-Adaptive Systems
- Learning Alzheimer's Disease Signatures by bridging EEG with Spiking Neural Networks and Biophysical Simulations
- Investigating Hyperparameter Optimization and Transferability for ES-HyperNEAT: A TPE Approach
- S-AI-Recursive: A Bio-Inspired and Temporal Sparse AI Architecture for Iterative, Introspective, and Energy-Frugal Reasoning