Investigation into using stochastic embedding representations for evaluating the trustworthiness of the Fr' e chet Inception Distance

arXiv:2601.21979 · cs.LG · Submitted 2026-08-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Evaluating the trustworthiness of the Fréchet Inception Distance with stochastic embedding representations".

Jane: The paper was written by Ciaran Bench, Vivek Desai, Carlijn Roozemond, Ruben van Engen and Spencer A. Thomas from National Physical Laboratory and Dutch Expert Centre for Screening.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Hey everyone, we are checking out a fascinating new paper titled "Evaluating the trustworthiness of the Fréchet Inception Distance with stochastic embedding representations."

Jane: That title sounds incredibly technical, Tom, but the concept is something we can all grasp.

Tom: Can you simplify that for our listeners, Jane?

Jane: Think of it this way: we use a specific mathematical ruler called FID to check if AI-generated images look real.

Tom: But that ruler might be a bit wonky if we use it on the wrong things, doesn't it?

Jane: Exactly, because that ruler was designed to measure natural images like cats and dogs, not medical X-rays.

Lu: This research from the National Physical Laboratory and the Dutch Expert Centre for Screening hits on a massive problem in AI reliability.

Meng: I see this all the time in production, where a metric works perfectly in the lab but fails when you hit real-world data.

Tom: The authors, including Ciaran Bench and Vivek Desai, are essentially asking if we can detect when that ruler is lying to us.

Jane: They want to know if we can measure our own uncertainty about the measurement itself.

Lu: It's a beautiful idea to build a self-aware evaluation system.

Lalam: If we can quantify when an AI's judgment is shaky, we can build much safer cultural tools for healthcare.

Tom: We should look at how they actually went about testing this ruler's reliability.

Summary: Tom: We've established that the FID metric might not be reliable for things like medical images, so let's look at the methodology.

Jane: The researchers used a technique called Monte Carlo dropout to see how much the model's internal representations wobble.

Tom: Can you explain how "shaking" the model works, Jane?

Jane: They basically run the same image through the model many times, but they randomly turn off different parts of the neural network each time.

Lu: If the model's answer stays the same, it's confident, but if the answer jumps around, the model is likely confused by the input.

Meng: I'm curious about the cost of doing that, especially if they have to run it twenty times for every single image.

Tom: They did exactly that, evaluating each batch twenty times to get a sense of the variance.

Jane: They focused on two specific things: the variance in the embeddings themselves, called pVar, and the variance in the FID score itself, called sigma FID.

Lu: Using randomness to probe the boundaries of what a model knows is a brilliant way to map out its knowledge.

Meng: I'd want to know if those two numbers actually tell us different stories.

Tom: That's exactly what they found, and the results are quite surprising.

Improvements: Tom: We just talked about how they "shook" the model, and now we have the results from those experiments.

Jane: The results in Table I show a very clear pattern for sigma FID, which is the uncertainty in the FID score.

Tom: Tell them about the mammography data, Jane, because those numbers are huge.

Jane: When they tested on standard ImageNet images, the sigma FID was tiny, only zero point zero zero nine.

Tom: But when they switched to mammography images, which are very different from natural photos, that number shot up to zero point three five zero.

Jane: That jump suggests that sigma FID is a great warning light for when the FID metric is becoming untrustworthy.

Lu: It's like a sensor that screams when you're driving into a fog that the car wasn't built for.

Meng: I noticed the paper mentions that pVar, the other metric, didn't actually follow a consistent trend.

Tom: You're right, Meng, pVar was all over the place and didn't reliably tell us how out-of-distribution the data was.

Jane: They also did a sensitivity test where they added noise to images to see how the metrics reacted.

Tom: As the images became more distorted and similar to each other, the FID and the uncertainty both dropped.

Jane: It makes sense because if every image is just a blurry mess, the model doesn't have much to be uncertain about.

Lu: This provides a way to create a "reliability score" for every single AI evaluation we perform.

Lalam: Imagine a future where every medical AI diagnostic comes with a built-in confidence interval that we can actually trust.

Conclusion: Tom: We've spent a lot of time on "Evaluating the trustworthiness of the Fréchet Inception Distance with stochastic embedding representations," and it's been eye-opening.

Jane: It really shows that we shouldn't just take a single quality score at face value.

Tom: The big takeaway is that sigma FID can act as a heuristic to tell us when our image quality metrics are losing their grip.

Jane: It's a vital step toward making generative models safer for high-stakes areas like medicine.

Lu: This is just the beginning of making AI evaluation as rigorous as traditional science.

Meng: I'm definitely going to be looking for these uncertainty signals in my next deployment.

Lalam: This moves us from blind faith in AI numbers to a culture of verifiable, measured truth.

Tom: Thanks for joining us, everyone, we'll catch you at the next paper.

Ciaran Bench, Vivek Desai, Carlijn Roozemond, Ruben van Engen, Spencer A. Thomas

National Physical Laboratory · Dutch Expert Centre for Screening

cs.LG

Submitted: 2026-08-17

Updated: 2026-08-18

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 79/100

The gist: This paper investigates whether uncertainty quantification (UQ) techniques, specifically Monte Carlo dropout (MCD), can be used to evaluate the trustworthiness of the Fréchet Inception Distance

Key concepts

Fréchet Inception Distance (FID)
A mathematical ruler used to check if AI-generated images look real. It measures the quality of generated content but is designed for natural images like cats and dogs, not medical X-rays.
σ_FID
The measure of uncertainty in the FID score itself. Researchers found this metric provides a clear warning light when AI evaluation is unreliable, particularly when comparing standard images to unusual data.
Monte Carlo dropout
A technique used to test model reliability. The researchers ran the same image through the neural network many times while randomly turning off different parts of the network.

Terminology

Summary

This paper investigates whether uncertainty quantification (UQ) techniques, specifically Monte Carlo dropout (MCD), can be used to evaluate the trustworthiness of the Fréchet Inception Distance (FID) when applied to out-of-distribution (OOD) data, particularly medical images.

The authors note that "Feature embeddings acquired from pretrained models are widely used in medical applications of deep learning to assess the characteristics of datasets; e.g. to determine the quality of synthetic, generated medical images. The Fréchet Inception Distance (FID) is one popular synthetic image quality metric that relies on the assumption that the characteristic features of the data can be detected and encoded by an InceptionV3 model pretrained on ImageNet1K (natural images). They state that It is widely known that FID is less suitable for medical image datasets (among other types of non-natural images) given they are out of distribution relative to the training set of the feature embedding model."

The aim is stated as: "We examine whether the predictive variance of embedding representations estimated with MCD, and the corresponding predictive variance in the FID, may provide a heuristic indication of whether the model encodes characteristic features representations of the data (given its reported sensitivity to out of distribution (OOD) inputs and data quality), and therefore, the effectiveness of the FID."

Methods: The authors trained an InceptionV3 architecture on ImageNet1K with dropout regularisation applied to every convolutional layer, initialised with pretrained weights. Each test batch was evaluated J=20 times. They computed two metrics:

  • pVar: the average of the normalised trace of the covariance in the test embeddings (Equation 2).

  • σFID (or vFID): the variance of the resultant FIDs computed across the J evaluations (Equation 4).

They performed two main experiments:

  1. Equal augmentation: Adding progressively larger amounts of additive Gaussian noise equally to two halves of ImageNet1K, to see how σFID and pVar change as image contents homogenise.

  2. OOD datasets and sensitivity analysis: Testing on CelebA, a mammography dataset, and lightly augmented versions of ImageNet1K (with overlaid images and noise). They used k-NN distance (k=5) to verify the extent to which datasets are OOD, and compared changes in FID with MAE, MS-SSIM, and top-5 accuracy.

Results:

  • Equal augmentation: "Fig. 2 shows that the FID decreases with increasing augmentation strength applied to both halves of ImageNet1K. This coarsely validates its effectiveness on noise augmented data when both the test and reference sets are augmented, as we would expect the metric to decrease and trend to zero as contents homogenise. Correspondingly, we find that σFID decreases, and critically, exhibits low magnitude relative to the sensitivity study shown in Fig. 3d where image contents become less similar. Given the FID appears accurate and is low magnitude, this suggests that σFID could be a reliable indicator of the effectiveness of the FID."

  • OOD datasets: "Table I shows how σFID increases with the extent to which the test data was OOD of the training data. If one can assume that effectiveness of the FID for data that is increasingly OOD will decrease, then this suggests σFID may indicate the trustworthiness of the FID to some extent. Though, the validity of this assumption needs to be verified to draw stronger conclusions. pVar on the other hand does not exhibit a consistent trend, which could suggest poor suitability as an indicator for OOD detection, or as a proxy measure of the effectiveness of the FID."

  • Sensitivity analysis: "The sensitivity study in Fig. 3 shows that both σFID and pVar initially increase and then decrease with augmentation strength. pVar may decrease at higher augmentation strengths due to the arguments outlined in [11]; indeed we find that the norm of the embeddings decreases with higher augmentation strengths (Fig. 3g) where the decrease in pVar corresponds to the decrease in the mean norm of the latents. They also note that An initial increase in σFID was not observed for the equal augmentation case (see Fig. 2) despite the same base sets of images being used. This result suggests that differences in the fidelity of image features/content between reference and test sets (along with the features themselves) can significantly affect uncertainty magnitude for OOD data."

They performed a variance decomposition of vFID (Equation 5) and found that all terms gradually trend to zero with increasing augmentation strength. They also ran an experiment with a fixed test set and increasingly augmented reference set (Fig. 4), finding that "σFID increases and then plateaus as the reference set becomes increasingly augmented. This suggests that, akin to term a in the variance decomposition, σFID scales with the extent of the difference between the test and reference sets, as well as the magnitude of pVar."

Conclusion: "Our results indicate that both pVar and σFID depend strongly on the properties of the input data. σFID generally correlates with the extent to which unaugmented datasets are OOD of ImageNet1K, and therefore, likely provides some indication for when the FID may be less informative/effective. pVar did not exhibit the same correlation, which suggests it is less reliable as a proxy measure a) of the extent to which data is OOD (despite being used as such in the literature [11]) and b) the effectiveness of the FID."

They note that the decrease in σFID at higher strengths requires a closer analysis of the complex interplay of the various terms of the FID and that a quantitative 'gold standard' metric for assessing the effectiveness of the FID is needed to robustly validate the suitability of the expressions of the trustworthiness of the FID presented here.

Finally, they state: "Nonetheless, our results provide useful insight into how our heuristic metrics for the trustworthiness of the FID change with respect to varying datasets, giving some indication of how uncertainty quantification techniques can be used to evaluate the trustworthiness of the FID. This work establishes a framework for uncertainty-aware evaluation of generative models, enabling trust assessment in high-stakes applications such as healthcare and other safety-critical domains."

Improvements for AI systems

Based on the paper, here are specific improvements I can implement in AI systems:

Improvement: Build a wrapper system around FID that automatically computes and reports σFID (standard deviation of FID across Monte Carlo dropout samples) alongside the FID score.

Capabilities:

  • Automatically flag FID scores as unreliable when σFID exceeds a threshold (e.g., >0.1 based on the paper's results)

  • Provide confidence intervals for FID comparisons between different generative models

  • Enable model selection based on both FID accuracy and reliability, preventing costly deployment of models with misleadingly good FID scores

These improvements directly leverage the paper's findings that σFID correlates with OOD extent and that pVar behaves differently across augmentation types, enabling more reliable deployment of generative models in safety-critical domains.

Abstract

Feature embeddings acquired from pretrained models are widely used in medical applications of deep learning to assess the characteristics of datasets; e.g. to determine the quality of synthetic, generated medical images. The Fr' e chet Inception Distance (FID) is one popular synthetic image quality metric that relies on the assumption that the characteristic features of the data can be detected and encoded by an InceptionV3 model pretrained on ImageNet1K (natural images). While it is widely known that this makes it less effective for applications involving medical images, the extent to which the metric fails to capture meaningful differences in image characteristics is not obviously known. Here, we use Monte Carlo dropout to compute the predictive variance in the FID as well as a supplemental estimate of the predictive variance in the feature embedding model's latent representations. We show that the magnitudes of the predictive variances considered exhibit varying degrees of correlation with the extent to which test inputs (ImageNet1K validation set augmented at various strengths, and other external datasets) are out-of-distribution relative to its training data, providing some insight into the effectiveness of their use as indicators of the trustworthiness of the FID.

Sources

Related papers