Evaluating multi-season occupancy models with autocorrelation fitted to heterogeneous datasets

arXiv:2510.08151 · stat.AP, q-bio.PE, q-bio.QM · Submitted 2025-10-09 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Genomics Radio. Generated commentary on the latest computational biology and genomics papers.

Ines: Today's paper: "Evaluating multi-season occupancy models with autocorrelation fitted to heterogeneous datasets".

Marcus: Predicting species distributions using occupancy models that account for spatial and temporal autocorrelation is crucial for ecological inference, especially when dealing with heterogeneous datasets where replication is limited.

Ines: First, who's behind it and why it matters.

Paper summary: Ines: So, we're looking at this paper titled "Evaluating multi-season occupancy models with autocorrelation fitted to heterogeneous datasets," and it seems like they're tackling a real headache in ecological inference by using these advanced models to handle the lack of replication. What's the core idea behind what they’re doing here?

Marcus: Exactly, Ines. The main thesis of this paper is that modeling spatial and temporal autocorrelation can help us make occupancy models work better when we have incomplete or heterogeneous data, which is a big problem when you're trying to predict species distributions. They are specifically looking at how these models perform under conditions where replication isn't perfectly controlled.

Yuki: From a population genetics standpoint, I see this as an attempt to bridge the gap between traditional occupancy modeling and the reality of patchy species distribution across a landscape, especially when individual site data points are sparse or missing. This research suggests that incorporating autocorrelation allows us to account for how sites influence each other both geographically and temporally, which is vital for understanding how populations persist over seasons.

Ines: That makes sense; so they're trying to use these autocorrelations—spatial and temporal random effects—to provide a kind of "fractional replication" effect, as Doser and Stoudt (two thousand twenty-four) call it, to compensate for the lack of strict repeated visits in their datasets. The abstract says they evaluate this model against challenging conditions like covariate overlap and spatiotemporal gaps.

Marcus: Right, and what I find interesting from the summary is that they address these challenges directly by using a hierarchical structure involving Gaussian processes for spatial effects and an AR(one) process for temporal effects, linking them to the latent occupancy state as described in Equation one. This is a sophisticated statistical setup designed to handle the inherent noise in real-world datasets.

Yuki: And that sophistication is what matters because it moves us closer to a model that can actually capture underlying ecological processes rather than just fitting noise into the data, which is something I've been thinking about regarding how movement patterns affect site selection.

Ines: They are testing this framework on a specific dataset—a fine-scale butterfly occupancy dataset from the French Southwest—which has some pretty messy characteristics, including a skewed distribution of survey numbers per site that starts at zero, which is tough for traditional models.

Marcus: The data design they used is quite rigorous, incorporating ecologically motivated constraints like phenology grouping observations and observer behavior grouping observations into specific places and times to make the simulation more realistic. This shows they aren't just running abstract simulations; they're trying to mirror real-world field collection patterns.

Paper summary: Yuki: When you look at how they set up those constraints, it connects back to the broader ecological understanding of species behavior across seasons and locations, which is a key area for population geneticists studying dispersal and range shifts.

Ines: Now, moving into the results section, the paper does a rigorous identifiability assessment using Bayesian posterior distribution means across fourteen thousand four hundred simulated datasets derived from sixteen different scenarios of spatial and temporal autocorrelation within three study designs. This is where they really test if their model structure can even produce stable parameter estimates.

Marcus: And what they found regarding identifiability is quite telling; they diagnosed it by looking at scatter plots of the estimated occupancy probability relative to the true value. They showed that in scenarios with high spatial decay or high temporal correlation, there were "elongated (flat) shape" densities for parameter combinations, suggesting weak identifiability.

Yuki: That finding is significant because it hints that when autocorrelation is very strong, the model struggles to distinguish between different parameter settings; it essentially indicates "little to no spatial and temporal autocorrelation when they truly exist," which points toward a limitation in how we interpret those high-autocorrelation scenarios.

Ines: And they also noted specific biases, like the spatial decay estimator phi being biased high when it should actually be low, and temporal parameters rho and sigma 2T generally being biased low across most autocorrelation scenarios. These are concrete statistical findings that explain why simply adding autocorrelation isn't a magic fix.

Marcus: It reinforces my view that the model framework itself has specific sensitivities; it seems it tends to under-estimate the strength of spatial and temporal structure when those structures are present in the data. This is crucial for anyone trying to build robust models from noisy data, even with sophisticated techniques like this one.

Yuki: This ties back to the broader ecological context because if a model consistently underestimates spatial structure, our predictions about species connectivity or metapopulation dynamics could be skewed by that underestimation.

Ines: Let's turn to the empirical analysis for a moment; they fit this model to the encounter history of two specific butterfly species: *Polyommatus icarus* and *Lycaena dispar*. For the common blue, they got an average yearly site occupancy estimate of fifty point nine three percent, which is quite different from a naive average of twenty-four point one two percent.

Marcus: That discrepancy for the common blue is pretty striking, as it shows the model successfully identified a fluctuating occupancy trend over time that wasn't captured by the simpler naive trend calculation, which speaks to the value of including temporal autocorrelation here. They also noted detection probability peaked during aural summer months in mid-July.

Yuki: For me, seeing that fluctuation is important because it suggests seasonal dynamics are playing a stronger role than we initially thought in this species' distribution pattern, which aligns with historical observations of their life cycles.

Paper summary: Ines: Then there’s the spatial aspect for the common blue where they found the estimated spatial decay parameter phi was large at thirty-one point six three, suggesting a relatively short spatial autocorrelation range, and temporal parameters showed high prior-posterior overlap close to or above thirty percent.

Marcus: That high PPO of over thirty percent for the temporal parameters is another indicator that the model is struggling with precise estimation there, which lines up with their earlier finding that those parameters were generally biased low. This suggests uncertainty in how strongly time influences occupancy across different scenarios.

Yuki: The results for the large copper butterfly offer a different picture; their averaged yearly site occupancy estimate was twelve point eight five percent, compared to a naive average of eight point five four percent. Occupancy was low overall, but they found high occupancy only in sites where the species was detected, which is an interesting spatial pattern to note.

Ines: And for the large copper, the analysis showed that coefficients for latitude, longitude, marsh cover, spatial variance sigma squared, temporal variance sigma 2T, and correlation rho all had high PPO close to or above thirty percent. This suggests that for this species, the model is very uncertain about how much each of those factors actually drives the occupancy.

Marcus: So, when we put it all together from this paper, we see a model that handles heterogeneity well but still shows significant identifiability issues and sensitivity to data gaps. It manages to provide estimates despite skewed survey distributions by leveraging the inclusion of spatial and temporal autocorrelation.

Yuki: The implication for us in population ecology is that if we rely too heavily on models that assume perfect replication or perfect covariate separation, we might be misinterpreting the true ecological signals coming from patchy environments.

Ines: Ultimately, this paper with its title "Evaluating multi-season occupancy models with autocorrelation fitted to heterogeneous datasets" shows us that while incorporating these autocorrelations is a powerful tool for modeling complex spatial and temporal dependencies, we still have to be cautious about identifiability problems and the influence of severe data gaps on the final predictions.

Marcus: We're seeing a model that can provide estimates even with skewed survey distributions, but it also reveals that under certain conditions regarding autocorrelation strength, the parameters themselves become unstable or biased. That level of statistical nuance is exactly what we need when dealing with complex genomic and observational datasets like ours.

Yuki: This work helps frame the future direction by showing that next steps should focus on developing ways to stabilize these parameter estimates when autocorrelation is high, perhaps through more structured priors informed by species history, which I think is where the population genetics aspect of this research can really help guide us.

Conclusion: Ines: So we're wrapping up this discussion on "Evaluating multi-season occupancy models with autocorrelation fitted to heterogeneous datasets," which essentially looks at how adding spatial and temporal dependence helps species distribution models when data is messy.

Marcus: I think the authors are really focusing on the statistical mechanics here, looking at how those autocorrelation parameters actually behave in practice across different scenarios.

Yuki: From a population genetics viewpoint, this research is important because it suggests we can better model how species respond to fluctuating environmental conditions over time and space when our sampling isn't perfectly uniform.

Ines: I'm thinking about the title itself; "Evaluating multi-season occupancy models with autocorrelation fitted to heterogeneous datasets" sounds like a very thorough test of a complex statistical framework.

Marcus: Yeah, and I see how the authors are scrutinizing those results to see if the model actually recovers biologically meaningful information or just fits noise in these tricky situations.

Yuki: The implication for us is that if we use models that ignore spatial structure, we risk getting skewed pictures of species connectivity, which is a huge issue when thinking about metapopulations.

Ines: Exactly, and I'm curious how this work on butterflies can translate to other complex ecological systems where replication is naturally limited or highly variable.

Marcus: It’s a solid piece of work because it moves beyond just applying existing models and actually tests the limits of the Doser and Stoudt framework under real-world constraints.

Yuki: It really highlights that understanding these subtle spatio-temporal patterns is key to accurately reconstructing species' historical ranges and dynamics.

André Luís Luza, Didier Alard, Frédéric Barraquand

UMR Biodiversité Gènes et Communautés, University of Bordeaux · Institute of Mathematics of Bordeaux, University of Bordeaux, CNRS · US Fauna, University of Bordeaux

stat.AP, q-bio.PE, q-bio.QM

Submitted: 2025-10-09

Updated: 2026-07-17

DOI: 10.1007/s13253-026-00753-6

Code: https://github.com/andreluza/butterfly_occupancy

Project page: https://observatoire-fauna

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 76/100

The gist: Predicting species distributions using occupancy models that account for spatial and temporal autocorrelation is crucial for ecological inference, especially when dealing with heterogeneous datasets

Key concepts

Occupancy Model
A statistical framework used to estimate whether a species is present or absent at specific locations over time. It links detection probabilities to underlying true occupancy states, allowing ecologists to infer distribution patterns even when not all sites are surveyed.
Spatial Autocorrelation (NNGP)
This accounts for the idea that species presence at one location is related to its neighbors. The study used Nearest Neighbor Gaussian Processes (NNGP) to model this in continuous space, helping the model understand how nearby sites influence each other's occupancy.
Identifiability
This refers to whether a statistical model can uniquely determine the true values of its parameters from the observed data. The analysis showed that certain combinations of spatial and temporal autocorrelation settings make it very difficult for the model to accurately estimate those underlying processes.

Terminology

Summary

Predicting species distributions using occupancy models that account for spatial and temporal autocorrelation is crucial for ecological inference, especially when dealing with heterogeneous datasets where replication is limited. This research evaluates the performance of multi-season occupancy models incorporating these autocorrelations against challenging data conditions, revealing that while the models are robust to certain issues like covariate overlap, they remain susceptible to identifiability problems and severe contamination from spatiotemporal gaps in real-world scenarios.

Model Framework

The core model is a hierarchical structure based on Doser & Stoudt (2024), consisting of an occupancy state process for a single species at sites and primary occasions, where the latent occupancy state is modeled as:

logit(ψit) = XTitβ + ωi + ηt (Equation 1).

The spatial random effects are defined through a Gaussian process: ω(s) ∼ N (0, Σ(D, θ)) (Equation 2), utilizing Nearest Neighbor Gaussian Processes (NNGP) to account for spatial autocorrelation in continuous space. Temporal random effects follow a zero-mean AR(1) process: Cov(ηt, ηt′) = σ2T × ρt−t′ (Equation 3). The observation process links the latent occupancy to detection probability: logit(pitj) = vTitjα (Equation 4).

Data and Simulation Design

The study utilized a heterogeneous fine-scale butterfly occupancy dataset from the French Southwest, consisting of 298,389 valid records across 24 years and approximately 90,290 grid cells. Key challenges addressed in simulations included:

  1. A skewed distribution of survey numbers per site exhibiting a Poisson-like distribution starting at zero (including NAs).

  2. Covariates affecting occupancy and detection models being fully random (uncorrelated) in space and time, as used in the baseline D&S design.

  3. The incorporation of ecologically-motivated constraints, such as phenology grouping observations at specific times and observers’ behavior grouping observations at specific places and times.

Identifiability Assessment

The paper rigorously assessed model identifiability using Bayesian posterior distribution means across 14,400 simulated datasets derived from 16 sub-scenarios of spatial and temporal autocorrelation within three study designs. Identifiability was diagnosed by examining scatter plots of estimated occupancy probability (ψˆit) relative to the true value (ψit).

- Global identifiability requires a one-to-one correspondence between parameters and the model. The analysis showed that in scenarios with high spatial decay ϕ and low spatial variance σ2, or high temporal correlation ρ and variance σ2T, there were elongated (flat) shape densities for parameter combinations, indicating weak identifiability. Specifically, the spatial decay estimator ϕˆ was biased high when it should be low. Temporal autocorrelation parameters ρ and σ2T were generally biased low across all autocorrelation scenarios. This suggests that the model tends to indicate little to no spatial and temporal autocorrelation when they truly exist.

Empirical Data Analysis

The model was fitted to the encounter history of two species: the common blue butterfly (Polyommatus icarus) and the large copper butterfly (Lycaena dispar).

- For the common blue, the average yearly site occupancy estimate was 50.93%, compared to a naive average yearly occupancy of 24.12%. The model identified a fluctuating occupancy trend over time that differed from the naive trend. Detection probability peaked during aural summer months (mid-July). The estimated spatial decay parameter ϕˆ was large (ϕˆ = 31.63), depicting a short spatial autocorrelation range, and the temporal parameters showed high prior-posterior overlap (PPO) close to or above 30%.

- For the large copper, the averaged yearly site occupancy estimate was 12.85%, against a naive average of 8.54%. Occupancy was low overall, with high occupancy only found in sites where the species was detected. The analysis indicated that coefficients for latitude, longitude, marsh cover, spatial variance σ2, temporal variance σ2T, and correlation ρ had high PPO (close or above 30%).

Conclusion on Robustness and Limitations

The study concludes that multi-season occupancy models with autocorrelation are robust to heterogeneous data and covariate overlap, but they still present identifiability issues and are challenged by severe data gaps, which contaminate predictions even in data-rich areas. The authors found that the inclusion of spatial and temporal autocorrelation in the model, alongside replication level contained in the heterogeneous simulated data, enabled parameter estimation despite skewed survey distributions.

Improvements for AI systems

Here are specific improvements for AI systems based on the insights from this scientific paper, categorized by capability:


) Specific Improvements for AI Systems

  1. A. Occupancy Model Robustness and Heterogeneity Handling:

Improve the core occupancy modeling algorithms (especially those incorporating spatial and temporal random effects) to explicitly account for the skewed distribution of survey numbers per grid cell (Poisson-like distribution starting at zero). The improved system should be able to maintain robust parameter estimation even when a significant portion of sites have zero visits, moving beyond assumptions made by models designed around a two-group mixture.

  1. B. Covariate Overlap Management:

Develop a mechanism within the occupancy model framework that can explicitly diagnose and mitigate the impact of overlap of covariates between occupancy and detection submodels (e.g., when latitude affects both). The improved AI should be able to distinguish whether this overlap is benign or detrimental to parameter identifiability, as the paper suggests it is generally less damaging when spatial/temporal autocorrelation is incorporated.

  1. C. Spatiotemporal Gap Mitigation:

Enhance predictive capabilities by integrating mechanisms that account for spatiotemporal gaps (missing data) more effectively than current methods. The improved AI should be trained to recognize patterns where gaps are caused by clustered observations in space and time (e.g., phenology or observer preferences), allowing it to make more informed out-of-sample predictions for unsampled areas, rather than defaulting to biased averages.

  1. D. Parameter Identifiability Diagnosis:

Implement a diagnostic layer within the model that specifically monitors the weak identifiability issues inherent in spatial models, particularly concerning the spatial decay parameter (ϕ) and variance parameter (σ2). The improved system should be able to flag when these parameters are estimated as non-existent or show evidence of being biased high when they should be low, providing a confidence score regarding the reliability of the inferred autocorrelation structure.

  1. E. Trend Estimation under Autocorrelation:

Upgrade time-series forecasting modules for species distribution trends. Instead of relying solely on simple AR(1) processes, the improved AI should utilize models that incorporate temporal random effects (AR(1) process with specified autocorrelation ρ and variance σ2T). This allows the system to differentiate between actual temporal occupancy trends and spurious temporal clustering introduced by gaps or sampling patterns.

) What the Improved AI System Can Do

The improved AI system, equipped with these enhancements, will be capable of:

  1. Predicting Species Distribution Maps with High Reliability: It will generate spatial maps of species occupancy probability that are significantly more accurate than current models, especially in areas with sparse or zero survey effort (e.g., predicting occupancy for unvisited cells).

  2. Quantifying Uncertainty in Autocorrelation Structure: It will provide a rigorous assessment of the spatial and temporal autocorrelation parameters (ϕ and ρ) it estimates, indicating whether these structures are truly present in the underlying ecological process or merely artifacts of the data sampling design or model misspecification.

  3. Robust Species Trend Forecasting: It will provide reliable long-term occupancy trend forecasts for species, correctly accounting for seasonal variations (phenology) and spatial clustering of fieldwork effort, leading to more accurate predictions of population stability versus decline.

  4. Optimizing Data Collection Strategy: By identifying areas where the model is most sensitive to data gaps or covariate overlap, the AI can provide actionable recommendations to researchers on where future sampling efforts should be concentrated for maximum information gain and minimal bias.

  5. Differentiating True Ecological Signals from Sampling Artifacts: The system will be capable of distinguishing between true ecological patterns (e.g., latitudinal gradients, habitat preferences) and artifacts caused by imperfect detection or non-random observation schedules, leading to more biologically meaningful inferences about species distribution drivers.

Abstract

Predicting species distributions using occupancy models accounting for imperfect detection is now commonplace in ecology. Recently, modeling spatial and temporal autocorrelation was proposed to alleviate the lack of replication in occupancy data, which often prevents model identifiability. However, how such models perform in highly heterogeneous datasets where missing or single-visit data dominates remains an open question. Motivated by a heterogeneous fine-scale butterfly occupancy dataset, we evaluate the performance of a multi-season occupancy model with spatial and temporal random effects to a skewed (Poisson) distribution of the number of surveys per site, overlap of covariates between occupancy and detection submodels, and spatiotemporal clustering of observations. Results showed that the model is robust to heterogeneous data and covariate overlap. However, when spatiotemporal gaps were added, site occupancy was biased towards the average occupancy, itself overestimated. Random effects did not correct the influence of gaps, due to identifiability issues of variance and autocorrelation parameters. Occupancy analysis of two butterfly species further confirmed these results. Overall, multi-season occupancy models with autocorrelation are robust to heterogeneous data and covariate overlap, but still present identifiability issues and are challenged by severe data gaps, which contaminate predictions even in data-rich areas.

Related papers