Stochastic Siamese MAE Pretraining for Longitudinal Medical Images

arXiv:2512.23441 · cs.LG, cs.CV · Submitted 2025-12-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Stochastic Siamese MAE Pretraining for Longitudinal Medical Images".

Jane: Temporally aware image representations are crucial for capturing disease progression in 3D volumes of longitudinal medical datasets, and this work proposes STAMP,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So we're diving into "Stochastic Siamese MAE Pretraining for Longitudinal Medical Images" today. The big idea here is tackling disease progression in three dee medical scans that are collected over time, and they propose STAMP as a way to get temporal awareness into those existing Masked Autoencoding models <ref:2512.23441#pg0>.

Jane: That sounds really cool, Tom. So what's the core problem they're trying to solve with this paper? Is it just getting better at looking at images, or is there something deeper about how disease actually changes over time that they are focusing on?

Lu: They are specifically addressing the fact that current state-of-the-art self-supervised learning approaches like MAE don't inherently capture temporal information, which is a big issue when dealing with longitudinal medical data <ref:2512.23441#pg0>.

Meng: From an engineering standpoint, that lack of inherent temporal awareness means they probably have to build in complex mechanisms if you want to use these models for progression tracking, which sounds computationally heavy.

Lalam: I think the core thesis is that they propose STAMP as a novel SSL approach that extends masked auto-encoding specifically for longitudinal medical imaging data within a Siamese framework <ref:2512.23441#pg1>.

Tom: Exactly, and what makes their proposal interesting is how they try to fix that by conditioning the encoding on the time difference between the two input volumes. They hypothesize that latent factors driving temporal evolution can be inferred without needing explicit supervision by contrasting a patient’s two visits through a variational process <ref:2512.23441#pg0>.

Jane: Conditioning on time difference sounds like a clever way to introduce that temporal context, but how do they actually make the model aware of this difference in practice? I need to understand the mechanism behind that time awareness.

Lu: They introduce a learnable encoding of the time difference between input images, which they achieve by generating a Temporal Encoding using a learnable two-layer MLP that takes a one-dimensional sin-cos wave from the discretized time difference as input <ref:2512.23441#pg0>.

Meng: A learnable encoding based on the sine and cosine of the time difference, that sounds like they're trying to map temporal distance into something the network can use directly. How does that actually feed into the reconstruction?

Lalam: That Temporal Encoding is used as an additive bias to the CLS token both before and after the encoder, which lets them "time-prompt" during inference <ref:2512.23441#pg0>.

Tom: That sounds like a very direct way to inject time information into the representation, giving them a prompt that guides how the model sees the image features. Jane, what about their handling of uncertainty? I noticed they seem concerned about disease progression not being smooth <ref:2512.23441#pg1>.

Paper summary: Jane: Right, and that leads to their second major contribution regarding time-conditioned stochasticity, where they introduce a stochastic component into future image representations by sampling from a learned time-conditioned probability distribution <ref:2512.23441#pg0>.

Lu: They learn two separate two-layer MLPs to generate the prior and the posterior distribution logits, where the prior is conditioned on the past visit volume and the time difference, while the posterior is aware of both by accounting for masking in partially visible future volumes <ref:2512.23441#pg0>.

Meng: So they are essentially sampling from a distribution that tells them what plausible future progression might look like, which addresses that variation in disease speed across patients. That sounds like a sophisticated way to model uncertainty rather than just predicting one single outcome <ref:2512.23441#pg1>.

Lalam: This stochastic embedding is obtained by sampling from the posterior and then concatenating it with the encoder's output token before it goes into the decoder <ref:2512.23441#pg0>.

Tom: And that feeds into their decoding process where they use cross-self-attention, using tokens from both visits as Key and Value while masking future tokens as Query <ref:2512.23441#pg0>.

Jane: So the decoder is comparing what it knows from the past with what it can see in the future, guided by that learned stochastic embedding, to reconstruct the next step in time. That's a complex interplay of information flowing through the network.

Lu: They’ve done this fully three dee approach to learn representations of volumetric scans at each time point, which is an adaptation of the Siamese MAE framework specifically for three dee medical scans <ref:2512.23441#pg0>.

Meng: That three dee volumetric extension is significant because traditional methods often struggle with the complexity of whole-volume data, so learning representations at each time point in three dee seems like a substantial step forward for practical application <ref:2512.23441#pg0>.

Lalam: This work helps improve culture by allowing us to build models that can forecast individual disease trajectories from just one baseline scan, which really facilitates tailored follow-up schedules <ref:2512.23441#pg0>.

Tom: It sounds like they've done a lot of heavy lifting here by framing forecasting as conditional variational inference to handle those tricky aspects of progression prediction, specifically long visit intervals and varying progression speeds <ref:2512.23441#pg1>.

Jane: So when we look at the title, "Stochastic Siamese MAE Pretraining for Longitudinal Medical Images," it really encapsulates that they are using a stochastic process within a Siamese MAE framework to handle longitudinal three dee medical scans <ref:2512.23441#pg0>.

Paper summary: Lu: They are proposing this because stacking temporal images or frames into three dee inputs increases computation substantially and limits inference to time-series input, which makes their approach more efficient in a certain way <ref:2512.23441#pg2>.

Meng: I'm interested in the computational overhead mentioned; they noted that STAMP has insignificant computational overhead compared to SiamMAE, while RSP is more expensive with three forward passes from its encoder Table V. That’s important for real-world deployment considerations <ref:2512.23441#pg0>.

Lalam: And in terms of performance, they achieved the highest PRAUC in the three-year AD progression task on the ADNI dataset, which shows their superior ability to model multi-outcome temporal features compared to other MAE based methods Table VI <ref:2512.23441#pg0>.

Tom: So, when we look at these results from pretraining on HARBOR data for wet-AMD and PINNACLE data for GA progression, STAMP consistently outperformed existing baselines and foundation models on six and twelve months progression <ref:2512.23441#pg1>.

Jane: It seems like the paper is really showing how this combination of temporal encodings and stochasticity yields superior performance when compared to methods that don't include those specific components <ref:2512.23441#pg0>.

Lu: The implication here, from a creative AI perspective, is that we can move beyond deterministic sequence completion by allowing the model to sample from a distribution of possibilities for the future progression <ref:2512.23441#pg1>.

Meng: Practically speaking, this means we can potentially use a single baseline scan to forecast different plausible disease trajectories, which could really help in planning personalized treatment schedules and reducing hospital load by catching high-risk patients earlier <ref:2512.23441#pg0>.

Lalam: This capability allows for the forecasting of individual disease trajectories from a single baseline scan, which really facilitates tailored follow-up schedules <ref:2512.23441#pg0>.

Tom: So, to wrap up this section on the STAMP framework, it successfully formulates forecasting as conditional variational inference to deal with long visit intervals and multiple potential outcomes in disease progression <ref:2512.23441#pg1>.

Jane: It really highlights how important it is to introduce mechanisms that account for the inherent uncertainty in medical data when modeling long-term changes <ref:2512.23441#pg0>.

Lu: The future work mentioned suggests they are looking at applying this to different modalities and irregular visit schedules, which points toward a much broader applicability of this core concept <ref:2512.23441#pg0>.

Meng: From an engineer's view, the promise is that these self-supervised frameworks can be applied across different medical imaging types, which opens up avenues for more generalized pretraining pipelines in the AI space <ref:2512.23441#pg0>.

Lalam: I think this work has a significant impact on our culture because it shows how we can leverage large unlabeled datasets to improve personalized monitoring, which is a huge win for precision medicine <ref:2512.23441#pg0>.

Conclusion: Tom: So, we’ve been deep in the weeds of how STAMP uses stochastic processes to capture those long-term temporal dependencies in medical scans today. Jane, looking at that title and the authors, what do you make of this work?

Jane: Well, I see that the core idea is taking a standard image reconstruction method, MAE, and making it smarter about time by injecting randomness into how it predicts what happens next in a sequence. The authors are using a Siamese framework to compare two visits over time.

Lu: What I find really compelling is their approach of reframing the MAE loss as a conditional variational inference objective; that means they aren't just deterministically trying to fill in the missing pixels, they’re trying to infer the underlying temporal dynamics stochastically. That’s a big conceptual leap for how we train these models.

Meng: From an engineering standpoint, I like that they focus on using only two visits; it keeps the complexity manageable compared to models that try to process every single frame in a long video. However, I wonder if sampling from a learned distribution adds too much overhead when we start deploying this on massive datasets.

Lalam: The most impactful vision here is that this method allows us to use just one baseline scan to forecast different plausible disease trajectories, which could really help improve culture by enabling personalized monitoring before a problem gets severe. That’s where the real value lies for precision medicine.

Tom: That’s exactly it, Lalam—moving from simply reconstructing an image to actually predicting the future path of a patient's health based on just one starting point. Jane, how do you simplify that for our listeners?

Jane: Think of it like having a movie where the first scene is visible and the last scene is blurry; STAMP helps you sample multiple possible endings based on what you see in between, instead of just guessing one single path forward. It teaches the model to be uncertain about the future, which is very realistic in medicine.

Lu: And that uncertainty handling via a time-conditioned posterior distribution is what makes it so powerful; it accounts for the fact that disease progression isn't a perfectly smooth line, but rather has inherent variability. That’s where you can capture real biological complexity.

Meng: I agree on the complexity aspect, though I do see a practical hurdle in making sure those learned distributions actually map cleanly onto real-world clinical outcomes without needing massive amounts of labeled data to train them properly.

Lalam: But that’s precisely the point; this framework shows how we can leverage large unlabeled datasets to improve personalized monitoring, which is a huge win for precision medicine. It's about using what we already have to anticipate what comes next.

Tom: So, in short, STAMP takes a basic image reconstruction task and uses clever temporal conditioning and randomness to help us forecast individual disease trajectories from just one scan. Jane, where should we go next?

PINNACLE Consortium

cs.LG, cs.CV

Submitted: 2025-12-29

Updated: 2026-10-06

Comments: Provisional Accept at IEEE TMI. Code is available in https://github.com/EmreTaha/STAMP

Code: https://github.com/EmreTaha/STAMP

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 83/100

The gist: Temporally aware image representations are crucial for capturing disease progression in 3D volumes of longitudinal medical datasets, and this work proposes STAMP, a novel Siamese MAE framework that

Key concepts

Time Awareness
This involves creating a learnable encoding of the time elapsed between two input images. A simple mathematical function (sin-cos wave) is used to generate this encoding, which is then added to the main image features. This helps the model understand temporal differences over long periods.
Time-Conditioned Stochasticity
To handle uncertainty in future predictions, STAMP introduces randomness. It learns two separate models to generate a probability distribution that depends on both the past image and the time difference. This sampling process allows the network to account for inherent unpredictability in disease evolution.
Decoding (Reconstruction)
The decoder reconstructs the future visit by comparing what is fully visible in the past with what is partially visible in the future. It uses cross-self-attention, using past tokens as keys and values, while masked future tokens act as queries. A sampled stochastic embedding is added to this process.

Terminology

Summary

Temporally aware image representations are crucial for capturing disease progression in 3D volumes of longitudinal medical datasets, and this work proposes STAMP, a novel Siamese MAE framework that encodes temporal information through a stochastic process by conditioning on the time difference between two input volumes.

The gist

STAMP learns temporal dynamics stochastically by reframing the MAE reconstruction loss as a conditional variational inference objective to capture long-term nondeterministic temporal dependencies using only two visits.

How it works

STAMP is a Siamese MAE framework that encodes temporal information through a stochastic process conditioned on the time difference between the two input volumes. Unlike deterministic Siamese approaches, STAMP learns temporal dynamics stochastically by reframing the MAE reconstruction loss as a conditional variational inference objective. It hypothesizes that latent factors driving temporal evolution can be inferred without explicit supervision by contrasting a patient’s two visits via a variational process.

  1. Time Awareness: The framework introduces a learnable encoding of time difference between input images, improving temporal image features over longer intervals. This is achieved by generating the Temporal Encoding (TE) using a learnable 2-layer MLP that takes a 1D sin-cos wave from the discretized time difference as input. TE is then used as an additive bias to the CLS token before the encoder and again before the decoder, allowing for time-prompting during inference.

  2. Time-Conditioned Stochasticity: To account for inherent uncertainty, STAMP introduces a stochastic component into future image representations by sampling from a learned time-conditioned probability distribution. This involves learning two separate 2-layer MLPs to generate the prior distribution and the posterior distribution logits. The prior is conditioned on the past visit volume and the time difference, while the posterior is aware of both the past and partially visible future volumes due to masking.

  3. Decoding (Reconstruction): The transformer decoder uses cross-self-attention where tokens from both visits are utilized as Key and Value, while masked future tokens are used as Query. This allows the network to reconstruct the future visit by comparing a fully visible past with a partially visible future. The stochastic embedding (SE) is obtained by sampling from the posterior, which is concatenated with the encoder's output token before being used in the decoder.

Key Contributions and Evaluation

The key contributions of STAMP include:

Time Awareness: We introduce a learnable encoding of time difference between input images, improving temporal image features over longer intervals.

"Time-Conditioned Stochasticity: We introduce a stochastic component into future image representations by sampling from a learned time-conditioned probability distribution, taking the uncertainty of the future into account."

3D Volumetric Extension: We employ a fully 3D approach to learn representations of volumetric scans at each time point, marking the first adaptation of the Siamese MAE framework to 3D medical scans.

STAMP was evaluated on two different modalities (OCT and MRI) across various prognostic tasks. For instance, when pretraining on HARBOR data, STAMP consistently outperformed existing baselines and foundation models on 6 and 12 months progression for wet-AMD (HARBOR) and 12 months GA progression (PINNACLE). Furthermore, in the ADNI MRI experiments, STAMP surpassed both MAE based and age-aware pretraining models, demonstrating robustness under irregular visit intervals and extended prediction windows. The results showed that the combination of temporal encodings and stochasticity yielded superior performance compared to methods lacking these components.

Conclusion

STAMP successfully formulates forecasting as conditional variational inference to address the challenging aspects of disease progression prediction: long visit intervals, varying progression speeds, and multiple potential outcomes. This approach enables the model to prompt the network for possible futures at inference time, using a single scan and a specified time window, providing a tool for personalized monitoring by forecasting individual disease trajectories from a single baseline scan. The work demonstrates that this self-supervised framework can be applied to different modalities and irregular visit schedules, highlighting its promise for advancing precision medicine in progressive diseases.

Table V

STAMP has insignificant computational overhead compared to SiamMAE, while RSP is more computationally expensive with three forward passes from its encoder. The model was pretrained on ViT-Base and evaluated using Cross-Entropy loss for downstream tasks, reporting metrics like AUROC, PRAUC, and BACC.

Table VI

STAMP achieved the highest PRAUC in the 3-year AD progression task on the ADNI dataset. This indicates its superior ability to model multi-outcome temporal features compared to other MAE based methods. The paper concludes that STAMP helps with utilizing large unlabeled datasets to improve personalized monitoring by forecasting individual disease trajectories from a single baseline scan, facilitating tailored follow-up schedules, earlier intervention for high-risk patients, and reducing hospital load.

Improvements for AI systems

Here are the specific improvements that can be made to existing AI systems based on the STAMP framework, and what those improved systems can achieve:


  1. Improve disease progression forecasting in longitudinal medical imaging (e.g., AMD, Alzheimer's Disease) by enabling prediction from a single baseline visit scan using only two visits as input.

  2. Enable models to learn non-deterministic temporal dynamics by modeling multiple possible future trajectories as a probability distribution rather than a single deterministic outcome.

  3. Introduce time awareness into image representations by incorporating the time difference between two scans as an additive bias (Time Encoding, TE) to the model's learned latent features.

  4. Account for inherent uncertainty in disease evolution by sampling temporal feature tokens from a learned, time-conditioned probability distribution (stochasticity), allowing predictions to reflect plausible future scenarios.

  5. Enhance the robustness of models against inter-subject variability and varying progression rates by learning a prior distribution that is conditioned on both the past visit embeddings and the time interval between visits.

  6. Improve feature localization and focus in visual representations by using time-conditioned temporal encodings (TE) to guide attention mechanisms, leading to more focused attention maps on relevant anatomical regions (e.g., the retina).

  7. Provide a computationally efficient pretraining pipeline for 3D volumetric medical scans by extending the Siamese MAE framework specifically for 3D inputs, avoiding the high computational cost of stacking multiple time points into 4D volumes.

  8. Improve downstream prognostic tasks (e.g., conversion prediction from iAMD to wet-AMD, or MCI to AD) by leveraging a single scan and a specified future time window for inference, facilitating timely treatment planning and risk prioritization.

  9. Increase the generalizability of medical image models across different modalities (OCT, MRI) and irregular visit schedules by demonstrating superior performance compared to existing MAE-based methods (like SiamMAE or RSP).

Abstract

Temporally aware image representations are crucial for capturing disease progression in 3D volumes of longitudinal medical datasets. However, recent state-of-the-art self-supervised learning approaches like Masked Autoencoding (MAE), despite their strong representation learning capabilities, lack temporal awareness. In this paper, we propose STAMP (Stochastic Temporal Autoencoder with Masked Pretraining), a Siamese MAE framework that encodes temporal information through a stochastic process by conditioning on the time difference between the 2 input volumes. Unlike deterministic Siamese approaches, which compare scans from different time points but fail to account for the inherent uncertainty in disease evolution, STAMP learns temporal dynamics stochastically by reframing the MAE reconstruction loss as a conditional variational inference objective. We evaluated STAMP on two OCT and one MRI datasets with multiple visits per patient. STAMP pretrained ViT models outperformed both existing temporal MAE methods and foundation models on different late stage Age-Related Macular Degeneration and Alzheimer's Disease progression prediction which require models to learn the underlying non-deterministic temporal dynamics of the diseases.

Sources

Related papers