Gene genealogies in diploid populations evolving according to sweepstakes reproduction
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Genomics Radio. Generated commentary on the latest computational biology and genomics papers.
Ines: I'm Ines, and with me are Marcus and Yuki, guest researcher.
Marcus: Today's paper: "Gene genealogies in diploid populations evolving according to sweepstakes reproduction".
Ines: Sweepstakes reproduction, characterized by a heavy right-tailed offspring number distribution, induces jumps in type frequencies and multiple mergers in gene genealogies of sampled gene copies.
Marcus: First, who's behind it and why it matters.
Paper summary: Ines: Looking at the title "Gene genealogies in diploid populations evolving according to sweepstakes reproduction," it really captures the essence of what they're doing here, don't you think? They are taking a specific, non-selective reproductive mechanism and tracing its effect on the evolutionary trails left by gene copies.
Marcus: I agree with Ines; it’s a very precise title for a paper that connects population ecology—how individuals reproduce—directly to population genetics—the structure of gene trees. It’s not just about the math; it's about what that math tells us about the underlying biological reality of the sampled genes.
Yuki: From a broader perspective, this research contributes to understanding how demographic processes, even those driven by chance rather than strong selection, shape the patterns we see in species evolution. It shows that we can build models that account for recruitment dynamics in a way that is more realistic than just assuming standard models like the Wright-Fisher model.
Ines: So, what's the big picture takeaway for us as computational biologists? It seems the main point is that we can derive these specific coalescent types when recruitment follows a sweepstakes distribution. This gives us a new framework to compare against observed gene tree patterns.
Marcus: And from the cohort side, it implies that when analyzing our data, we need to consider these specific multiple-merger or time-changed coalescent models rather than just standard ones. This means our statistical inference methods might need to be adapted to account for the skewness parameter alpha.
Yuki: Ultimately, this paper opens up new avenues for population geneticists to test hypotheses about demographic processes in populations where recruitment is not governed by simple deterministic rules. It gives us tools to look at the historical trail of genes and ask different types of questions about how those histories were shaped by the environment.
Ines: It really shows that modeling the process isn't just academic; it has direct implications for how we interpret genetic variation in real populations. We get a clearer picture of the link between ecology and genealogy.
Marcus: Indeed, the paper provides concrete mathematical descriptions—the continuous-time coalescents—that map this ecological input onto a predictable evolutionary output. That's what matters for applying these ideas to large genomic datasets.
Yuki: This work is valuable because it moves us closer to understanding the full spectrum of demographic forces that can shape evolutionary history, allowing us to better distinguish between different historical scenarios in the data we collect.
Conclusion: Ines: It means they've derived continuous-time coalescents that describe how ancestral lineages randomly combine under this specific skewed reproductive law. The analysis recovers a set of mathematical structures that predict the patterns of recombination and lineage sorting we see in gene trees.
Marcus: I see it as developing new statistical tools to account for non-standard recruitment dynamics, which directly addresses potential batch effects or systematic biases in our genomic data that we might be overlooking.
Yuki: For population genetics, these coalescents give us a rigorous way to test hypotheses about how demographic shifts drive the structure of gene genealogies within a species' evolutionary timeline.
Ines: The authors show that under certain conditions, this reproductive mechanism leads to specific types of multiple-merger coalescents, like Beta or Poisson-Dirichlet types, which are distinct from standard models. It’s a powerful way to link ecology to genealogy.
Marcus: That link is crucial because if we can model the input process accurately, it gives us a much more robust statistical foundation for inferring population history from sequencing data.
Yuki: The implication is that we might be able to better distinguish between different demographic scenarios in the fossil record or genomic sequences by testing which of these specific coalescent types best fits the observed genealogy.
Ines: So, to put it simply, this paper provides the blueprint for how chance-driven reproductive skew affects the structure of our genetic family trees.
Marcus: Exactly; it gives us a more sophisticated way to handle the statistical noise that comes from complex population dynamics.
Yuki: It’s a solid step forward in connecting microscopic ecological mechanisms to macroscopic patterns of evolution within populations.
Bjarki Eldon
q-bio.PE, math.PR
Submitted: 2026-01-15
Updated: 2026-09-29
Comments: 40 pages + 6 pages bibliography; a few graphs; Preliminary revised version according to peer-reviews; comments and questions welcome
Code: https://github.com/eldonb/gene_genealogies_diploid_pops_sweepstakes
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 73/100
The gist: Sweepstakes reproduction, characterized by a heavy right-tailed offspring number distribution, induces jumps in type frequencies and multiple mergers in gene genealogies of sampled gene copies.
Key concepts
- Sweepstakes Reproduction
- This is a reproductive mechanism where the number of offspring an individual produces is not determined by natural selection but by chance, like matching broadcast spawning with favorable environmental conditions. This creates a heavy right-tailed distribution for offspring numbers, leading to specific patterns in gene genealogies.
- Continuous-Time Coalescents
- These are mathematical models used to describe the random ancestral relations within a population over time. The study finds several types, including Beta and Poisson-Dirichlet coalescents, which characterize how lineages merge into common ancestors in this specific reproductive system.
- Multiple-Merger Coalescents
- These coalescent types describe situations where mergers involve a random number of ancestral lineages simultaneously rather than just two. They arise from population models involving sweepstakes reproduction and are classified as either asynchronous or simultaneous, reflecting complex merger dynamics.
- Time-Changed Coalescents
- These are coalescent models where the time scale of the evolution is modified by changes in population size. The study shows that these time-changed coalescents have a property where the modification of time is independent of the skewness parameter alpha, which governs how skewed the offspring number distribution is.
Terminology
Summary
Sweepstakes reproduction, characterized by a heavy right-tailed offspring number distribution, induces jumps in type frequencies and multiple mergers in gene genealogies of sampled gene copies. This study models population genetics under this reproductive mechanism to derive continuous-time coalescents that characterize the evolution of diploid panmictic populations.
Background and Model Setup
The research investigates the evolution of a diploid panmictic population absent selfing, evolving in a random environment, where recruitment dynamics—the distribution of offspring number among individuals—is governed by sweepstakes reproduction. Sweepstakes reproduction is defined as a skewed offspring number distribution due to mechanisms not involving natural selection, such as chance matching of broadcast spawning with favourable environmental conditions.
The population evolves over one generation in two stages: first, current individuals produce potential offspring according to a given law; second, a number of these potential offspring are sampled uniformly and without replacement to survive and replace the current individuals.
Key Mathematical Results
The main mathematical results establish continuous-time coalescents that describe the random ancestral relations. The study identifies several key coalescent types:
-
Continuous-time Beta and Poisson-Dirichlet coalescents, where the skewness parameter α of the Beta-coalescent ranges from 0 to 2, and these may be incomplete due to an upper bound on potential offspring.
-
Time scaling in large populations is measured in units proportional to either
N/ log N or N generations.
-
Incorporating population size changes leads to
time-changed coalescents with the time-change independent of α.
-
Simulations show that the ancestral process is
not well approximated by the corresponding coalescent
when skewness is increased, and that conditional and unconditional ancestral processes are not in good agreement.
Coalescent Families Arising from Population Models
The paper details several families of coalescents derived from specific population models:
)&Lambda-coalescents:
The study examines multiple-merger coalescents, which arise from population models of sweepstakes reproduction. These are characterized by multiple mergers, where a random number of ancestral lineages is involved whenever mergers occur. They are classified as either asynchronous or simultaneous multiple-merger coalescents.
)&omega-δ0-Beta(γ, 2 − α, α)-coalescents:
When the population evolves according to Definition 3.6 (a random environment), the resulting coalescent is an omega-δ0-Beta(γ, 2 − α, α)-coalescent.
The transition rate for a k-merger when n blocks is given by a complex formula involving the Beta function and parameters derived from the offspring distribution law.
)&omega-δ0-Poisson-Dirichlet(α, 0)-coalescents:
When the skewness of the offspring number distribution is increased (e.g., in Case 3 of Theorem 3.7), a discrete-time coalescent emerges, specifically the omega-δ0-Poisson-Dirichlet(α, 0)-coalescent.
This coalescent is defined by a transition probability involving the Poisson-Dirichlet distribution with parameter (α, 0).
Approximations and Comparison of Processes
The paper provides methods to approximate expected values related to gene genealogies:
-
Approximations for expected relative branch lengths are given as functionals such as
E[Ri(n)]
andE[hReN i(n)]
. -
These approximations can be computed via simulations, tracking the configuration of blocks in individuals, or by recording ancestral relations in a forward-in-time process.
-
The comparison between the approximation of expected relative branch lengths (annealed) and those conditioned on population ancestry (quenched) reveals qualitative differences depending on the pre-limiting model, suggesting that quenched and annealed multiple-merger coalescents may be qualitatively different.
Conclusion
The results demonstrate that under specific conditions related to sweepstakes reproduction, the limiting ancestral process converges to either a Kingman coalescent or a simultaneous multiple-merger coalescent (Beta or Poisson-Dirichlet types). The study highlights that population size changes lead to time-changed coalescents where the time-change is independent of the skewness parameter α. Furthermore, in scenarios with increased effect of sweepstakes (e.g., ζ(N)/N → ∞), the approximations for gene genealogies diverge significantly from those predicted by the limiting coalescent, suggesting that these models may not be relevant for inferring sweepstakes reproduction in real data under certain conditions.
Appendix Details
The paper includes detailed algorithms for computing approximations of expected relative branch lengths and ancestral relations, including sampling methods for the various coalescent types. It also provides extensive numerical examples comparing annealed and quenched approximations across various parameter settings. The final section discusses the implications of time-scaling on gene genealogies, showing that components of the underlying distribution may vary over time if population size changes are incorporated.
Improvements for AI systems
As a fastidious researcher, I have analyzed this paper, Gene genealogies in diploid populations evolving according to sweepstakes reproduction,
and identified several high-impact areas where AI systems can be significantly improved by incorporating these advanced coalescent models.
Here are the specific improvements and capabilities for an enhanced AI system:
) 1. Robust Inference of Highly Fecund Population History
The current limitations in inferring evolutionary histories often stem from using small family
models (like Wright-Fisher) when populations are highly fecund, leading to poor predictions. By integrating the results from this paper, an AI system can move beyond simple likelihood maximization in standard models.
-
To perform robust inference on gene genealogies of highly fecund natural populations, the AI should be equipped with a module that tests hypotheses against continuous-time coalescent models (Beta and Poisson-Dirichlet) rather than discrete approximations.
-
The system can specifically identify if the observed pattern is better explained by a
Kingman coalescent,
asimultaneous multiple-merger coalescent
(relevant to sweepstakes reproduction), or a time-changed version of these, based on the population's expected offspring number distribution.
) 2. Advanced Modeling of Stochastic Dynamics and Jumps
The paper introduces mechanisms like multiple mergers
and jumps in type frequencies
caused by sweepstakes reproduction.
-
An improved AI can model complex, non-linear evolutionary dynamics where gene genealogies undergo discontinuous jumps due to favorable environmental conditions (modeled via Definition 3.6).
-
The system can predict the probability of observing specific genetic variation patterns under different environmental regimes (e.g., high vs. low offspring number production rates), which is crucial for understanding population adaptation and evolution in broadcast spawners or other highly fecund species.
) 3. Time-Scale Invariance and Population Dynamics
The paper proves that incorporating population size changes leads to time-changed coalescents
where the time-change is independent of the skewness parameter α.
-
The AI system can dynamically adjust its temporal resolution when analyzing genetic data from populations undergoing fluctuating sizes (e.g., bottlenecks or range expansions).
-
It can accurately model how a change in population size over time affects the observed genealogical structure, effectively decoupling the timescale of coalescence from the underlying reproductive skewness parameter.
) 4. Discrimination Between Ancestral Processes (Annealed vs. Quenched)
The paper distinguishes between annealed
(averaged over all ancestral relations) and quenched
(conditioned on specific population ancestries) processes, showing they can be qualitatively different, especially when the upper bound on offspring number is constrained.
-
The AI can perform a diagnostic check on its own inferences by comparing the annealed approximation of branch lengths with the quenched approximation.
-
It can determine if its current inference relies too heavily on an idealized
complete dispersal
assumption or if it is sensitive to specific population pedigrees, guiding researchers toward more realistic, pedigree-based methods.
) 5. Predictive Power for Site Frequency Spectrum (SFS)
The paper provides explicit predictions for the SFS under different coalescent regimes (e.g., U-shaped spectra similar to the Atlantic cod).
-
An AI system can predict the expected site frequency spectrum of a sample based on its inferred evolutionary model, allowing it to validate its inferences against empirical genomic data from highly fecund species.
-
It can specifically detect spectral features characteristic of sweepstakes reproduction (like U-shaped spectra) that might be missed by standard neutral models.
) 6. Automated Algorithm Generation for Coalescent Approximations
The paper includes detailed simulation algorithms for approximating the expected branch lengths in the continuous-time coalescents (e.g., Figure C1).
-
The AI can automate the generation of these complex simulation algorithms, allowing researchers to rapidly test different parameters (like skewness α or population size N) and obtain empirical estimates of key quantities like branching rates.
-
This capability reduces the manual computational burden for testing complex stochastic models in evolutionary biology.
In summary, an AI system utilizing this paper would transform from a generic pattern-matching tool into a sophisticated, predictive tool capable of:
-
Performing high-fidelity genealogical inference in highly fecund species.
-
Modeling population dynamics that induce discontinuous genetic
jumps.
-
Differentiating between averaged and pedigree-specific evolutionary histories (annealed vs. quenched).
-
Providing explicit predictions for site frequency spectra under sweepstakes reproduction scenarios, enabling rigorous model selection in genomics.
Abstract
Recruitment dynamics, or the distribution of the number of offspring among individuals, is central for understanding ecology and evolution. Sweepstakes reproduction (when the offspring number distribution has a heavy right-tail) may characterize the recruitment dynamics of highly fecund natural populations. Sweepstakes reproduction can induce jumps in type frequencies, and multiple mergers in gene genealogies of sampled gene copies. Here, we consider gene genealogies in diploid panmictic populations evolving in a random environment. The heavy-tailed offspring number distribution is generated by mechanisms not involving natural selection, such as in chance matching of broadcast spawning with favourable environmental conditions. Our model of sweepstakes reproduction extends the one considered by Schweinsberg (2003) by applying an upper bound to the number of potential offspring of any given parent pair. Depending on the stated bound, the gene genealogies are in the domain of attraction of the Kingman coalescent, or specific familes of continuous-time Beta- or Poisson-Dirichlet simultaneous multiple-merger coalescents. The gene genealogies in a large population are viewed on a timescale proportional to at least N/log (N) generations; N is proportional to the population size (when constant). Incorporating deterministic population size changes leads to time-changed coalescents; the time-change is independent of the skewness of the offspring-number distribution. Using simulations, we show that gene genealogies in finite populations are not well approximated by the coalescent trees. Simulation results also indicate that quenched (conditioned on the population ancestry) and annealed gene genealogies in finite populations are not in good agreement whenever the skewness of the offspring number distribution is increased.
Related papers
- Drivers of periodicity in population dynamic models of long-lived, large mammals
- Constraints on the perfect phylogeny mixture model and their effect on reducing degeneracy
- Interplay between evolutionary and epidemic time scales challenges the outcome of control policies
- Evolutionary foraging in grids: Intermittent search dynamics emerge in finite, depletable landscapes
- Life Finds A Way: Emergence of Cooperative Structures in Adaptive Threshold Networks
- Fluctuating growth rate and spatial diffusion shape plankton diversity