Modelling Geographic Atrophy Progression using Implicit Neural Representations

arXiv:2608.10807 · cs.CV, cs.AI · Submitted 2026-08-11 · Read on arXiv

Simone Sarrocco, Paul Friedrich, Florentin Bieder, Christina Bornberg, Philippe Valmaggia, Peter Maloca, Philippe Cattin

University of Basel · University Hospital Basel · Moorfields Eye Hospital NHS Foundation Trust · Vienna University of Natural Resources and Life Sciences

cs.CV, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: Accepted at MICCAI 2026 Off-Grid Workshop

Code: https://github.com/SimoneSarrocco/ga-progression-with-inrs

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 50/100

The gist: Age-related Macular Degeneration (AMD) is the major cause of blindness in the Western world.

Terminology

Summary

Age-related Macular Degeneration (AMD) is the major cause of blindness in the Western world. Its late dry phase is characterised by irreversible atrophic areas, namely Geographic Atrophy (GA). Longitudinal Fundus Autofluorescence (FAF) image acquisitions are currently the main tool for assessing lesion growth over time at the image level. However, due to its highly individualised progression, the evolution of late AMD remains poorly understood. In this work, we propose using Implicit Neural Representations (INRs) to model GA progression at the individual level in a low-data setting. Our approach generates both FAF and GA segmentation at both past and future time points. Among the comparison models, our method achieves competitive segmentation quality across different scenarios, yielding the lowest Mean Absolute Error (MAE) for the GA lesion area and the highest DICE score, without sacrificing FAF image quality. The code is available at https://github.com/SimoneSarrocco/ga-progression-with-inrs.

In this work, we model subject-specific GA progression in patients with late dry AMD despite limited longitudinal data. Our modelling assumption is that the GA progression follows a continuous, patient-specific trajectory, with FAF images representing observations at discrete time points. Using Implicit Neural Representations (INRs), we jointly learn lesion growth trajectories from FAF images and corresponding GA segmentation masks. By adapting a population-level disease trajectory to subject-specific data, we model the disease progression as a continuous process, enabling interpolation of intermediate states and prediction of future GA progression at both image and segmentation levels.

To the best of our knowledge, this work presents the first application of INRs to model individual lesion trajectories in GA secondary to AMD. Our method predicts FAF images and GA segmentation masks over time while providing accurate estimates of lesion size at each time point. In a low-data setting, the model generates past and future segmentations from a single combination of FAF image and corresponding GA lesion mask for previously unseen patients. By employing eye-specific latent vectors, we can generate accurate, individual predictions of lesion growth in shape. These predictions can be used to visualise how the disease is expected to evolve, potentially helping clinicians communicate future visual outcomes and treatment benefits to patients.

Our framework uses an INR that models the disease trajectory as a continuous function fθ parameterised by θ. The INR consists of an auto-decoder, which is a multi-layer perceptron (MLP) with SIREN as the activation function. Following Dannecker et al., to capture individual changes in a setting with multiple images per patient, we assign one latent vector to each individual eye shared across all its images. In this way, we force the model to learn eye-specific features rather than visit-specific features. The model is conditioned on the latent vectors through modulation layers, where each vector gets linearly mapped to scale ϕ ∈ RH and shift ψ ∈ RH parameters, where H is the hidden size of the network. Similar to Dannecker et al., ω0 is applied only to the scale parameter to let the shift represent temporal dynamics as low-level signals. To give the model temporal information, we condition the INR on the number of weeks elapsed since the baseline visit and the patient’s age at that visit. The two conditioning variables are concatenated to the latent vector before modulation. Following Vyas et al., the INR is split into a reconstruction head fθfaf and a segmentation head fθseg, which are jointly modelled by the same MLP up to the penultimate layer L − 1. Different from Vyas et al., which feeds the segmentation head only from the penultimate layer, we concatenate the features of the last two SIREN layers and pass them through a linear layer to predict two-class probabilities (GA vs. background). As modulation layers, they are conditioned on the eye-specific latent vector and the time variables, as are all layers in our architecture.

During training, a subset of coordinates x = (x, y) ∈ Xi ⊂ R2 is drawn from the domain Ii: Xi → R of a FAF image at time t (measured in weeks from baseline) of eye i and is given as input to the INR. Then, a latent vector zi sampled from N (0, 10−2) gets assigned to eye i. Following Dannecker et al., we make use of 3D spatial latent vectors zi ∈ RC×X1×X2, where C is the number of channels, and X1 and X2 the spatial dimensions. The value of the latent vector at x is obtained through bilinear interpolation, denoted as BiInterp(zi, x). The value of t and the age of the patient at time t, aget, are concatenated to the latent vector, resulting in zi(x) = [BiInterp(zi, x) ⊕ t ⊕ aget] which is then linearly mapped into scale and shift parameters in each modulation layer of the INR. The model parameters θ and the set of latent vectors zi Ni=1 are optimised to maximise the joint log posterior distribution over the N eyes. This is achieved by minimising a combination of reconstruction loss LMSE, weighted by α, and segmentation loss LSEG: L = αLMSE(fθfaf(zi(x)), Ii(x)) + LSEG(fθseg(zi(x)), Si(x)), where Ii(x) is the true pixel intensity value of the FAF image at x, Si(x) is the true segmentation label at x, and fθfaf and fθseg are the corresponding predicted intensity value and segmentation label, respectively. LSEG is the sum of Binary Cross Entropy loss and DICE loss.

During test-time adaptation, for a previously unseen eye k a new latent vector zk is randomly sampled from N (0, 10−2) and assigned to it. While keeping θ frozen, the latent vector z is optimised on the image modalities from N − 1 of the N available visits, using the same loss as for training. To predict the held-out images, the concatenation of optimised latent vector and conditions, z̃k = [zk ⊕ t ⊕ aget], is then fed into the frozen INR, and a single forward pass is performed as fθ(xz̃k) = (fθfaf(x, z̃k(x)), fθseg(x, z̃k(x))).

The method was evaluated on the OMEGA dataset, an in-house longitudinal cohort of patients with GA secondary to AMD. The OMEGA dataset comprises scans of 37 eyes (19 right, 18 left) from 30 patients, monitored over up to 4 visits (approximately 12 weeks apart). The patients’ ages range from 66 to 90 years, with an average of 78.4 years. At each visit, short-wavelength FAF images were performed using a Heidelberg Spectralis device, generating 768 × 768 grayscale images centred on the macula with a 30◦ field of view and a scale of ∼ 0.01 mm/px. GA lesions were semiautomatically annotated in the FAF using the RegionFinder software by three graders individually. For each visit, we then aggregated the three segmentation masks into a final mask using majority voting. The FAF images were pre-registered intra-patient by aligning each follow-up visit of an eye to its corresponding baseline with an adapted version of RIFT. The pixel intensities were normalised between [0,1], whereas spatial coordinates x and conditioning variables were normalised between [-1,1]. We cropped out the black frames introduced by image registration and resized our images to 512 × 512 pixels. We further resized to 256 × 256 pixels when comparing our method to all other DL models. The data were split into training (26 eyes), validation (5 eyes), and test (6 eyes) sets, patient-wise; that is, both eyes of a single patient were always included in the same split.

The best-configured INR consisted of 8 modulation layers. The hidden layer size was set to 384, and the 3D latent vectors had shape [256, 32, 32]. Similar to Stolt-Ansó et al., the value of α for the reconstruction loss was set to 10. During training, the learning rate was set to 10−4 for the INR and 5·10−3 for the latent vector. Following Sitzmann et al., ω0 was set to 30 for all layers. The network parameters and the latent vectors are jointly optimised using AdamW, with a weight decay of 0.1. The number of coordinates sampled at each iteration was set to 10’000. We trained and validated our model on a single NVIDIA RTX 2080 Ti 12GB for 50 epochs, taking on average 3.5 hours and using 1.3 GB of GPU memory. Test-time adaptation on a new eye took on average 1 minute with 25 epochs.

For future FAF image prediction, we compare our method with ImageFlowNet. ImageFlowNet operates on pairs of images from the same patient at two different time points ti and tk, with i < k. By integrating a learned flow field, it predicts the FAF image at time tk using the image at time ti as input, along with the time shift tk − ti. To replicate this setup, which we refer to as Scenario 1, we created all pairs of visits within each eye, optimised the latent vectors on the older visit, and tested them on the newer visit of the same pair. It also allows test-time optimisation by fine-tuning the flow on the full patient history to extrapolate the last image, which we refer to as Scenario 2 and represents the default test-time optimisation of our model. Similar to Liu et al., we include in the comparison a time-aware diffusion model (T-I2SBUNet) based on Liu et al., and a time-conditional UNet (T-UNet) architecture derived from Ho et al. We also include classical methods for comparison, such as linear interpolation, cubic B-spline interpolation, and copy-forward, which simply copies the last available image as a prediction (no progression scenario).

For a quantitative evaluation, we used Peak Signal-to-Noise Ratio (PSNR, in dB), Structural Similarity Index Measure (SSIM), and Learned Perceptual Image Patch Similarity (LPIPS) for assessing image reconstruction quality, and Hausdorff Distance (HD, in pixels), DICE score, and Mean Absolute Error (MAE, in mm2) for segmentation quality. For a qualitative evaluation, we visually inspected the reconstructed FAF image quality and the segmented GA regions to assess their clinical accuracy.

In Scenario 1, our model achieved performance comparable to both classical and DL approaches. Specifically, our method achieved among the best results in both segmentation quality (DICE: 0.85; lesion-area MAE: 0.31) and structural similarity (SSIM: 0.61). In Scenario 2, our method again achieved the best results across all segmentation metrics (DICE: 0.91, HD: 6.78; lesion-area MAE: 0.20 mm2), demonstrating strong performance in learning disease progression in terms of lesion shape, localisation, and estimated area. We also achieved comparable performance on the FAF reconstruction task, demonstrating the model’s ability to capture meaningful anatomical changes across both modalities. However, simply copying the last available ground truth yielded the best FAF reconstruction quality, which might be due to the short time shift between visits in our dataset. In Scenario 1, our model achieved performance competitive with methods that were specifically trained on this setting. Our framework achieved the best results for lesion area prediction and GA segmentation, demonstrating the strength of having a single latent vector per eye, but at the cost of FAF reconstruction quality. In Scenario 2, performance on both tasks improved significantly, yielding a much better FAF image and very accurate segmentation estimates. Notably, some methods (e.g. T-UNet, ImageFlowNet) achieved low lesion-area MAE yet predicted atrophic regions that visually mismatch the ground truth. We also tested our model’s ability to predict missing data by excluding data from one intermediate visit during test-time optimisation, and using the optimised latent vector to predict it. Here, our method achieved the best PSNR (17.88), SSIM (0.61), DICE (0.87), and MAE for the diseased area (0.15) compared to classical approaches such as linear and cubic B-spline interpolation.

Using the whole history of the subject, the model generated faithful predictions at new intermediate and future time points. The GA lesion areas extracted from the predicted interpolated and extrapolated segmentations are realistic relative to the ground truth values of nearby measurements and appear to follow the underlying disease progression curve. This is confirmed for each of the six eyes included in the test set, showing the subject-specific prediction capabilities of the method. Classical methods like linear extrapolation are still quite good at approximating the trajectory over time, mainly due to the limited sample size of the test set and the quite short observational period of the study, but looking at both interpolated and extrapolated values, on average, our method is able to yield the lowest lesion-area MAE even in such a scenario of limited data availability.

We performed different stages of experiments. First, we varied how temporal and clinical information are combined in the decoder. We noticed that encoding time as an input coordinate, similar to Bieder et al., consistently underperformed FiLM modulation. We therefore adopted raw-scalar modulation of (age, weeks) for all subsequent ablations. We then performed a grid search over spatial dimensions and the number of latent vector channels C. DICE was largely insensitive to these choices. For both FAF reconstruction and segmentation, a 1 × 1 latent vector consistently underperformed larger latent grids, regardless of C. Larger grids led to better image reconstruction quality but at the cost of worse segmentation performance. In terms of lesion-area MAE, most configurations with spatial 2D latent vectors performed well, with intermediate latent grids yielding the best scores. Since no combination outperformed the others, we opted for a 32 × 32 grid and C = 256 as a good trade-off between image reconstruction and segmentation quality.

We have introduced the first application of INRs for modelling the progression of GA secondary to AMD. We achieve competitive results in segmentation quality and in predicted changes in lesion shape and area. We demonstrate the usefulness of spatial latent vectors for learning the anatomical structure of each individual eye and for injecting time information via modulation layers. Our model also performs well at FAF prediction, both for past and future time points, and at predicting missing visits. The image quality of the predictions, however, especially during extrapolation, remains a challenge, and reconstructing the high-frequency details of FAF images will require further experiments and ablation studies. In some scenarios, simply copying the input image as the prediction yields the best image reconstruction quality across all comparison methods, primarily due to the study’s short follow-up period. To this end, an evaluation criterion that incorporates additional anatomical information, such as GA lesion segmentation masks, as well as extracted biomarkers, is essential for assessing the clinical usefulness of the predictions and the individual disease evolution. In future work, we would like to apply the method to a larger cohort of patients, including subjects at different stages of AMD, to generate risk prediction curves at the individual level and to analyse the latent vector space for more interpretable outcomes. We would also like to extend the method to multiple modalities, such as scanner laser ophthalmoscopy, and to apply it to 3D OCT volumes.

Improvements for AI systems

Improvements to AI Systems:

  1. Continuous Disease Progression Modeling via Implicit Neural Representations (INRs):
  • Implement INRs with auto-decoder architectures and FiLM modulation to model disease trajectories as continuous functions of time, rather than discrete snapshots.

  • Use eye-specific latent vectors (spatial grids) to capture individualized anatomical structure, enabling interpolation of intermediate disease states and extrapolation to future time points.

  • The improved system can predict both imaging (e.g., FAF) and segmentation (e.g., GA lesion masks) at arbitrary time points, even with sparse longitudinal data (e.g., 2–4 visits).

  1. Joint Multi-Task Learning with Shared Representations:
  • Use a shared MLP backbone with separate reconstruction and segmentation heads, concatenating features from the last two layers for segmentation to improve boundary delineation.

  • Optimize a combined loss (MSE for image reconstruction + BCE + DICE for segmentation) with a weighted α to balance tasks.

  • The improved system can simultaneously generate high-quality anatomical images and accurate lesion segmentations, reducing the need for separate models and improving clinical interpretability.

  1. Test-Time Adaptation for Personalized Predictions:
  • Freeze the population-level INR after training and optimize only the latent vector for a new patient using a subset of their available visits.

  • This enables rapid (≈1 minute) personalization to unseen patients, allowing the system to predict missing intermediate visits or future progression from just one baseline image + segmentation.

  • The improved system can generate patient-specific risk curves and visual forecasts, aiding clinicians in communicating disease evolution and treatment benefits.

  1. Handling Missing Data and Irregular Time Intervals:
  • Condition the model on continuous time variables (weeks since baseline) and patient age, allowing predictions at any time point, not just predefined visit schedules.

  • The improved system can impute missing visits (e.g., interpolate a skipped check-up) with higher accuracy (e.g., DICE 0.87, MAE 0.15 mm2) than classical interpolation methods, and can extrapolate beyond the last visit to forecast lesion growth.

  1. Spatial Latent Vector Design for Anatomical Consistency:
  • Use 3D spatial latent vectors (e.g., [256, 32, 32]) with bilinear interpolation to condition the INR, preserving local anatomical details and avoiding overfitting to visit-specific features.

  • The improved system can generate lesion shape changes that are spatially coherent and clinically plausible, as evidenced by lower Hausdorff Distance (6.78) and higher DICE (0.91) compared to baselines.

  1. Robustness to Low-Data Regimes:
  • Train on small cohorts (e.g., 26 eyes) with few visits per patient, leveraging the inductive bias of continuous functions and shared latent structure.

  • The improved system can achieve competitive or superior performance in segmentation and lesion-area prediction (MAE 0.20 mm2) even when data is scarce, outperforming diffusion models and UNets that require larger datasets.

  1. Evaluation with Clinically Relevant Metrics:
  • Incorporate anatomical metrics (DICE, Hausdorff Distance, lesion-area MAE) alongside image quality metrics (PSNR, SSIM, LPIPS) to assess clinical utility.

  • The improved system can flag when image reconstruction alone is misleading (e.g., copy-forward may have high PSNR but no progression), guiding users to rely on segmentation-based outcomes for disease monitoring.

  1. Interpretable Latent Space for Disease Phenotyping:
  • Analyze the learned eye-specific latent vectors to cluster patients by progression patterns, potentially revealing subtypes of GA growth.

  • The improved system can provide clinicians with latent-space visualizations or risk scores, enabling more personalized prognosis and treatment planning.

What the Improved AI System Can Do:

  • Given a single FAF image and its GA segmentation at baseline, predict the patient’s FAF and lesion mask at any future time point (e.g., 6, 12, 24 months) with high accuracy in lesion area and shape.

  • Interpolate missing visits from sparse follow-ups, reducing patient burden and enabling retrospective analysis of incomplete datasets.

  • Generate patient-specific visual forecasts of disease evolution, which can be used in shared decision-making to illustrate potential outcomes of no treatment vs. therapy.

  • Operate in low-data clinical settings (e.g., rare diseases) where large annotated datasets are unavailable, by leveraging continuous trajectory learning and test-time adaptation.

  • Provide both image-level and segmentation-level outputs, allowing clinicians to verify predictions against anatomical boundaries and quantify growth rates (e.g., mm2/year) for regulatory or research endpoints.

Abstract

Age-related Macular Degeneration (AMD) is the major cause of blindness in the Western world. Its late dry phase is characterised by irreversible atrophic areas, namely Geographic Atrophy (GA). Longitudinal Fundus Autofluorescence (FAF) image acquisitions are currently the main tool for assessing lesion growth over time at the image level. However, due to its highly individualised progression, the evolution of late AMD remains poorly understood. In this work, we propose using Implicit Neural Representations (INRs) to model GA progression at the individual level in a low-data setting. Our approach generates both FAF and GA segmentation at both past and future time points. Among the comparison models, our method achieves competitive segmentation quality across different scenarios, yielding the lowest Mean Absolute Error (MAE) for the GA lesion area and the highest DICE score, without sacrificing FAF image quality. The code is available at https://github.com/SimoneSarrocco/ga-progression-with-inrs.

Related papers