The evolution of sex for artificial intelligence: a population-genetic framework for multigenerational model populations
cs.LG, cs.NE, q-bio.PE
Submitted: 2026-09-16
Updated: 2026-09-16
Comments: 22 pages, 5 figures, 1 table. Supplementary Information (26 pp) and a plain-language figure appendix for readers from biology (23 pp) are included as ancillary files. Code, configs and seeds: https://git.lab.gilest.ro/giorgio/MachineSex
License: http://creativecommons.org/licenses/by/4.0/
The gist: Some aspects of AI development resemble a population process in which models are specialised, retrained on the output of peers, or combined by averaging weights.
Terminology
Abstract
Some aspects of AI development resemble a population process in which models are specialised, retrained on the output of peers, or combined by averaging weights. These practices lead to generations of models, in the biological sense studied by population genetics. Here, I develop this parallelism and interpret multigenerational model populations in terms of sexual and asexual reproduction, formally recombining the two fields. I test these analogies in an exact inheritance model, in trained networks (recurrent, feedforward and variational autoencoder generators) and in large language models, and show that they hold generally, with some measurable architecture-specific biases. Training recursively on model output is known to lead to model collapse, a process previously described as akin to genetic drift; I develop all that follows. A minimal model of a learner retrained on its parent's output reproduces the Wright-Fisher process exactly; verified real data added to each generation play the role of immigration, with the surprising finding that the absolute number of real data samples matters, not their share, exactly as in population genetics. Training a child on the average of its parents' outputs cancels the benefit of having several parents, matching blending inheritance (and reviving Jenkin's objection to Darwin), whereas combining parents so that each keeps its strongest contribution preserves it; merged language-model specialists exceeded every parent across seeds (the Fisher-Muller effect); and lineages become reproductively isolated, losing the ability to merge at all, when they have learned conflicting conventions and not when they have merely drifted apart. As AI societies become societies in time as well as in space, a mathematical framework for their inheritance acquires predictive power. Remarkably, that framework can be adapted almost wholesale from biology.
Sources
- Anatomy of a Machine Learning Ecosystem: 2 Million Models on Hugging Face
- Self-Instruct: Aligning Language Models with Self-Generated Instructions
- PeaTMOSS: A Dataset and Initial Analysis of Pre-Trained Models in Open-Source Software
- Will we run out of data? Limits of LLM scaling based on human-generated data
- Arcee's MergeKit: A Toolkit for Merging Large Language Models
- Generative Agents: Interactive Simulacra of Human Behavior
- Large Language Model based Multi-Agents: A Survey of Progress and Challenges
- Virtual Agent Economies
- Drift and selection in LLM text ecosystems
- Multiagent Finetuning: Self Improvement with Diverse Reasoning Chains
- First-Extinction Law for Resampling Processes
- Nemotron-4 340B Technical Report
- Model Collapse in the Self-Consuming Chain of Diffusion Finetuning: A Novel Perspective from Quantitative Trait Modeling
- Git Re-Basin: Merging Models modulo Permutation Symmetries
- Self-Consuming Generative Models Go MAD
- On the Stability of Iterative Retraining of Generative Models on their own Data
- REPAIR: REnormalizing Permuted Activations for Interpolation Repair
- Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data
- Soup to go: mitigating forgetting during continual learning with model averaging
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks