Stochastic Siamese MAE Pretraining for Longitudinal Medical Images

summary

Video file (mp4)

The gist

Temporally aware image representations are crucial for capturing disease progression in 3D volumes of longitudinal medical datasets, and this work proposes STAMP, a novel Siamese MAE framework that

In short

STAMP is a novel Siamese MAE framework that learns temporal disease progression in 3D medical scans from just two visits. It captures long-term, unpredictable changes by using a stochastic process conditioned on the time difference between images. This allows the model to predict future disease states from a single baseline scan, improving personalized monitoring.

Key concepts

Time Awareness
This involves creating a learnable encoding of the time elapsed between two input images. A simple mathematical function (sin-cos wave) is used to generate this encoding, which is then added to the main image features. This helps the model understand temporal differences over long periods.
Time-Conditioned Stochasticity
To handle uncertainty in future predictions, STAMP introduces randomness. It learns two separate models to generate a probability distribution that depends on both the past image and the time difference. This sampling process allows the network to account for inherent unpredictability in disease evolution.
Decoding (Reconstruction)
The decoder reconstructs the future visit by comparing what is fully visible in the past with what is partially visible in the future. It uses cross-self-attention, using past tokens as keys and values, while masked future tokens act as queries. A sampled stochastic embedding is added to this process.

Terminology used across episodes

This episode discusses

The paper

Stochastic Siamese MAE Pretraining for Longitudinal Medical Images · Read on arXiv

PINNACLE Consortium

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Stochastic Siamese MAE Pretraining for Longitudinal Medical Images".

Jane: Temporally aware image representations are crucial for capturing disease progression in 3D volumes of longitudinal medical datasets, and this work proposes STAMP,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So we're diving into "Stochastic Siamese MAE Pretraining for Longitudinal Medical Images" today. The big idea here is tackling disease progression in three dee medical scans that are collected over time, and they propose STAMP as a way to get temporal awareness into those existing Masked Autoencoding models <ref:2512.23441#pg0>.

Jane: That sounds really cool, Tom. So what's the core problem they're trying to solve with this paper? Is it just getting better at looking at images, or is there something deeper about how disease actually changes over time that they are focusing on?

Lu: They are specifically addressing the fact that current state-of-the-art self-supervised learning approaches like MAE don't inherently capture temporal information, which is a big issue when dealing with longitudinal medical data <ref:2512.23441#pg0>.

Meng: From an engineering standpoint, that lack of inherent temporal awareness means they probably have to build in complex mechanisms if you want to use these models for progression tracking, which sounds computationally heavy.

Lalam: I think the core thesis is that they propose STAMP as a novel SSL approach that extends masked auto-encoding specifically for longitudinal medical imaging data within a Siamese framework <ref:2512.23441#pg1>.

Tom: Exactly, and what makes their proposal interesting is how they try to fix that by conditioning the encoding on the time difference between the two input volumes. They hypothesize that latent factors driving temporal evolution can be inferred without needing explicit supervision by contrasting a patient’s two visits through a variational process <ref:2512.23441#pg0>.

Jane: Conditioning on time difference sounds like a clever way to introduce that temporal context, but how do they actually make the model aware of this difference in practice? I need to understand the mechanism behind that time awareness.

Lu: They introduce a learnable encoding of the time difference between input images, which they achieve by generating a Temporal Encoding using a learnable two-layer MLP that takes a one-dimensional sin-cos wave from the discretized time difference as input <ref:2512.23441#pg0>.

Meng: A learnable encoding based on the sine and cosine of the time difference, that sounds like they're trying to map temporal distance into something the network can use directly. How does that actually feed into the reconstruction?

Lalam: That Temporal Encoding is used as an additive bias to the CLS token both before and after the encoder, which lets them "time-prompt" during inference <ref:2512.23441#pg0>.

Tom: That sounds like a very direct way to inject time information into the representation, giving them a prompt that guides how the model sees the image features. Jane, what about their handling of uncertainty? I noticed they seem concerned about disease progression not being smooth <ref:2512.23441#pg1>.

Paper summary: Jane: Right, and that leads to their second major contribution regarding time-conditioned stochasticity, where they introduce a stochastic component into future image representations by sampling from a learned time-conditioned probability distribution <ref:2512.23441#pg0>.

Lu: They learn two separate two-layer MLPs to generate the prior and the posterior distribution logits, where the prior is conditioned on the past visit volume and the time difference, while the posterior is aware of both by accounting for masking in partially visible future volumes <ref:2512.23441#pg0>.

Meng: So they are essentially sampling from a distribution that tells them what plausible future progression might look like, which addresses that variation in disease speed across patients. That sounds like a sophisticated way to model uncertainty rather than just predicting one single outcome <ref:2512.23441#pg1>.

Lalam: This stochastic embedding is obtained by sampling from the posterior and then concatenating it with the encoder's output token before it goes into the decoder <ref:2512.23441#pg0>.

Tom: And that feeds into their decoding process where they use cross-self-attention, using tokens from both visits as Key and Value while masking future tokens as Query <ref:2512.23441#pg0>.

Jane: So the decoder is comparing what it knows from the past with what it can see in the future, guided by that learned stochastic embedding, to reconstruct the next step in time. That's a complex interplay of information flowing through the network.

Lu: They’ve done this fully three dee approach to learn representations of volumetric scans at each time point, which is an adaptation of the Siamese MAE framework specifically for three dee medical scans <ref:2512.23441#pg0>.

Meng: That three dee volumetric extension is significant because traditional methods often struggle with the complexity of whole-volume data, so learning representations at each time point in three dee seems like a substantial step forward for practical application <ref:2512.23441#pg0>.

Lalam: This work helps improve culture by allowing us to build models that can forecast individual disease trajectories from just one baseline scan, which really facilitates tailored follow-up schedules <ref:2512.23441#pg0>.

Tom: It sounds like they've done a lot of heavy lifting here by framing forecasting as conditional variational inference to handle those tricky aspects of progression prediction, specifically long visit intervals and varying progression speeds <ref:2512.23441#pg1>.

Jane: So when we look at the title, "Stochastic Siamese MAE Pretraining for Longitudinal Medical Images," it really encapsulates that they are using a stochastic process within a Siamese MAE framework to handle longitudinal three dee medical scans <ref:2512.23441#pg0>.

Paper summary: Lu: They are proposing this because stacking temporal images or frames into three dee inputs increases computation substantially and limits inference to time-series input, which makes their approach more efficient in a certain way <ref:2512.23441#pg2>.

Meng: I'm interested in the computational overhead mentioned; they noted that STAMP has insignificant computational overhead compared to SiamMAE, while RSP is more expensive with three forward passes from its encoder Table V. That’s important for real-world deployment considerations <ref:2512.23441#pg0>.

Lalam: And in terms of performance, they achieved the highest PRAUC in the three-year AD progression task on the ADNI dataset, which shows their superior ability to model multi-outcome temporal features compared to other MAE based methods Table VI <ref:2512.23441#pg0>.

Tom: So, when we look at these results from pretraining on HARBOR data for wet-AMD and PINNACLE data for GA progression, STAMP consistently outperformed existing baselines and foundation models on six and twelve months progression <ref:2512.23441#pg1>.

Jane: It seems like the paper is really showing how this combination of temporal encodings and stochasticity yields superior performance when compared to methods that don't include those specific components <ref:2512.23441#pg0>.

Lu: The implication here, from a creative AI perspective, is that we can move beyond deterministic sequence completion by allowing the model to sample from a distribution of possibilities for the future progression <ref:2512.23441#pg1>.

Meng: Practically speaking, this means we can potentially use a single baseline scan to forecast different plausible disease trajectories, which could really help in planning personalized treatment schedules and reducing hospital load by catching high-risk patients earlier <ref:2512.23441#pg0>.

Lalam: This capability allows for the forecasting of individual disease trajectories from a single baseline scan, which really facilitates tailored follow-up schedules <ref:2512.23441#pg0>.

Tom: So, to wrap up this section on the STAMP framework, it successfully formulates forecasting as conditional variational inference to deal with long visit intervals and multiple potential outcomes in disease progression <ref:2512.23441#pg1>.

Jane: It really highlights how important it is to introduce mechanisms that account for the inherent uncertainty in medical data when modeling long-term changes <ref:2512.23441#pg0>.

Lu: The future work mentioned suggests they are looking at applying this to different modalities and irregular visit schedules, which points toward a much broader applicability of this core concept <ref:2512.23441#pg0>.

Meng: From an engineer's view, the promise is that these self-supervised frameworks can be applied across different medical imaging types, which opens up avenues for more generalized pretraining pipelines in the AI space <ref:2512.23441#pg0>.

Lalam: I think this work has a significant impact on our culture because it shows how we can leverage large unlabeled datasets to improve personalized monitoring, which is a huge win for precision medicine <ref:2512.23441#pg0>.

Conclusion: Tom: So, we’ve been deep in the weeds of how STAMP uses stochastic processes to capture those long-term temporal dependencies in medical scans today. Jane, looking at that title and the authors, what do you make of this work?

Jane: Well, I see that the core idea is taking a standard image reconstruction method, MAE, and making it smarter about time by injecting randomness into how it predicts what happens next in a sequence. The authors are using a Siamese framework to compare two visits over time.

Lu: What I find really compelling is their approach of reframing the MAE loss as a conditional variational inference objective; that means they aren't just deterministically trying to fill in the missing pixels, they’re trying to infer the underlying temporal dynamics stochastically. That’s a big conceptual leap for how we train these models.

Meng: From an engineering standpoint, I like that they focus on using only two visits; it keeps the complexity manageable compared to models that try to process every single frame in a long video. However, I wonder if sampling from a learned distribution adds too much overhead when we start deploying this on massive datasets.

Lalam: The most impactful vision here is that this method allows us to use just one baseline scan to forecast different plausible disease trajectories, which could really help improve culture by enabling personalized monitoring before a problem gets severe. That’s where the real value lies for precision medicine.

Tom: That’s exactly it, Lalam—moving from simply reconstructing an image to actually predicting the future path of a patient's health based on just one starting point. Jane, how do you simplify that for our listeners?

Jane: Think of it like having a movie where the first scene is visible and the last scene is blurry; STAMP helps you sample multiple possible endings based on what you see in between, instead of just guessing one single path forward. It teaches the model to be uncertain about the future, which is very realistic in medicine.

Lu: And that uncertainty handling via a time-conditioned posterior distribution is what makes it so powerful; it accounts for the fact that disease progression isn't a perfectly smooth line, but rather has inherent variability. That’s where you can capture real biological complexity.

Meng: I agree on the complexity aspect, though I do see a practical hurdle in making sure those learned distributions actually map cleanly onto real-world clinical outcomes without needing massive amounts of labeled data to train them properly.

Lalam: But that’s precisely the point; this framework shows how we can leverage large unlabeled datasets to improve personalized monitoring, which is a huge win for precision medicine. It's about using what we already have to anticipate what comes next.

Tom: So, in short, STAMP takes a basic image reconstruction task and uses clever temporal conditioning and randomness to help us forecast individual disease trajectories from just one scan. Jane, where should we go next?

More episodes

← Home