APEX: Assumption-free Projection-based Embedding eXamination Metric for Image Quality Assessment

arXiv:2605.07786 · cs.CV, cs.AI · Submitted 2026-05-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "APEX: Assumption-free Projection-based Embedding eXamination Metric for Image Quality Assessment".

Jane: The gist The APEX framework introduces an assumption-free projection-based embedding examination metric for image quality assessment by leveraging the Sliced Wasserstein Distance with CLIP and DINOv2 embeddings to overcome limitations…

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So we’ve talked about APEX as a projection-based framework that uses Sliced Wasserstein Distance to compare image embeddings from models like CLIP and DINOv2, aiming to be assumption-free.

Jane: The core thesis is that traditional metrics often fail because they have rigid parametric formulations or rely on outdated features, so APEX offers a way around those limitations.

Lu: They introduce the Sliced Wasserstein Distance as a mathematically grounded similarity measure for comparing these distributions of image embeddings directly.

Meng: It’s trying to move beyond just looking at pixels or simple feature counts and instead focus on the actual distribution of what the AI models are producing in their embedding space.

Tom: Right. The framework inherits scalability to high-dimensional spaces, which is a big deal when you’re dealing with these complex foundation models.

Jane: They showed that APEX can be used to evaluate both global semantic fidelity using APEX-CLIP and diagnostic utility by looking at which level of abstraction is affected with APEX-DINO.

Lu: The metrics can work in synergy, meaning you get a robust check on overall quality from one and a deeper look into the model’s behavior from the other.

Meng: So for someone who just wants to know if an image looks good overall, APEX-CLIP gives them that robust assessment of semantic fidelity.

Tom: And APEX-DINO gives you diagnostic utility, which lets you isolate whether a specific distribution shift is happening at a lower or higher level of visual abstraction.

Jane: This approach tackles the problem by focusing on the distribution itself rather than forcing it into a pre-defined statistical shape, which is what many current methods struggle with.

Lu: It shifts the focus from rigid assumptions about how features should look to measuring actual similarity between the embeddings generated by state-of-the-art models.

Meng: That seems like it would be very useful for engineers because it gives us a metric that’s tied to what the model actually learned, not just some arbitrary statistical test.

Tom: It definitely feels like a shift in how we judge generative output quality by grounding the evaluation in the embedding space of these massive foundation models.

Conclusion: Tom: So wrapping up APEX, the main point is this assumption-free projection-based embedding examination metric for image quality assessment is a new way to measure generative output quality.

Jane: It moves away from old metrics that rely on specific assumptions about data or feature extraction methods and uses the Sliced Wasserstein Distance for a more flexible comparison.

Lu: Essentially, they’re giving us a tool to compare the actual distributions of what CLIP and DINOv2 produce in a way that is mathematically grounded.

Meng: For someone listening who only cares about how this affects their work, it means having a metric that’s less likely to break when you switch between different generative models or datasets.

Tom: It offers better consistency across domains, which is something we’ve been pushing for in image quality assessment for a long time.

Jane: The authors found that APEX is competitive with strong baselines in human alignment and provides more stable behavior when testing across different visual domains.

Lu: By using both APEX-CLIP and APEX-DINO, they give us dual perspectives on the quality of the generated image.

Tom: It’s about having a flexible evaluation framework that respects the complexity of modern AI's output space.

Jane: It’s a way to ground our quality checks in the actual embedding space that these foundation models operate in, which is where we need to be for reliable assessment moving forward.

University of Siena · AI for Good (AIGO), Istituto Italiano di Tecnologia, University of Verona

cs.CV, cs.AI

Submitted: 2026-05-08

Updated: 2026-10-08

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 83/100

The gist: The gist The APEX framework introduces an assumption-free projection-based embedding examination metric for image quality assessment by leveraging the Sliced Wasserstein Distance with CLIP and DINOv2

Key concepts

Sliced Wasserstein Distance (SWD)
SWD is a mathematical tool used to measure the distance between two probability distributions, such as those of image embeddings. It works by taking one-dimensional projections of these high-dimensional distributions and averaging the distances between them. This allows for a more stable and computationally feasible way to compare complex embedding spaces.
Foundation Model Embeddings (CLIP/DINOv2)
These are numerical representations (embeddings) created by powerful AI models like CLIP and DINOv2 when they process images. Instead of looking at pixels, these embeddings capture the high-level semantic meaning or features of an image, which APEX uses as the basis for quality assessment.
Projection-based Evaluation
This is the core technique in APEX where the distance between two complex embedding distributions is estimated by projecting them onto simpler one-dimensional lines. By measuring how these projections change, APEX can reliably estimate the true distance between images without needing to know every detail of the original high-dimensional data.

Terminology

Summary

The gist The APEX framework introduces an assumption-free projection-based embedding examination metric for image quality assessment by leveraging the Sliced Wasserstein Distance with CLIP and DINOv2 embeddings to overcome limitations in existing metrics.

How it works

  1. APEX measures the Sliced Wasserstein Distance (SWD) between image embeddings extracted by state-of-the-art foundation models, specifically CLIP and DINOv2 <ref:2605.07786#pg15> The Sliced Wasserstein Distance [Rabin et al., 2011] computes the average of the 2-Wasserstein distance on one-dimensional projections of two probability distributions µ and ν <ref:2605.07786#pg4> In practice, the SWD can be approximated by using a simple Monte Carlo scheme that uniformly draws L sample directions θl on S d−1, and replaces the integration with a finite-sample average <ref:2605.07786#pg5>.

  2. The framework is embedding-agnostic and uses two open-vocabulary foundation models, CLIP and DINOv2, as feature extractors <ref:2605.07786#pg6> APEX inherits effective scalability to high-dimensional spaces, as we prove with theoretical and empirical evidences <ref:2605.07786#pg2>. The SWD in APEX computes the distance between two distributions of image embeddings, specifically extracted either from CLIP or DINOv2 <ref:2605.07786#pg4>.

Key Contributions

** We audit failure modes and limitations of existing image quality assessment metrics, analysing how their behaviour is affected by backbone choices, distributional assumptions, kernel configurations, finite-sample effects, and cross-domain shifts <ref:2605.07786#pg2>. The work introduces APEX as a projection-based evaluation framework for comparing image distributions in foundation model embedding spaces <ref:2605.07786#pg2>. 3. We investigate the stability of evaluations by analysing the number of random projections required for a reliable SWD estimation, providing both theoretical and empirical ablations, identifying a trade-off between estimation stability and computational cost <ref:2605.07786#pg2>. 4. We benchmark APEX against established metrics via an extensive evaluation protocol spanning heterogeneous visual domains, various degradations, analysing finite-sample stability, cross-dataset consistency, and sensitivity to progressive generation and refinement <ref:2605.07786#pg2>. 5. A dedicated human perceptual study further shows that APEX is competitive with the strongest baselines in human alignment, while offering more stable behaviour across domains <ref:2605.07786#pg2>. The remainder of the paper is organized as follows <ref:2605.07786#pg2>. 3. The APEX metrics propose a feature-distribution metric that measures the SWD between image embedding extracted by state-of-the-art foundation models, specifically CLIP and DINOv2 <ref:2605.07786#pg4>. The APEX-CLIP leverages OpenAI’s CLIP ViT-L/14@336px yielding 768-dimensional [CLS] embeddings <ref:2605.07786#pg5>. Conversely, APEX-DINO captures a richer architectural hierarchy by concatenating the 1024-dimensional [CLS] tokens from layers 6, 12, and 23 of DINOv2 ViT-L/14 into a single 3,072-dimensional representation <ref:2605.07786#pg5>. These metrics can be used in synergy: APEX-CLIP provides a robust assessment of global semantic fidelity, while APEX-DINO offers diagnostic utility by isolating which level of visual abstraction is affected by specific distribution shifts <ref:2605.07786#pg5>. 3.1 Analysis on the number of projections investigates the minimal number of projections L required for a stable and robust Monte Carlo approximation of the (squared) Sliced Wasserstein Distance <ref:2605.07786#pg6>. Theorem 1 provides a bound on the estimation error, showing that if L ≥ 2D4τ 2k log 8CD 2τ − log δ squared then the estimation error SW 2(µ, ν) − SW[2] L ≤ τ with probability at least 1 − δ <ref:2605.07786#pg2>. The number of projections L needed to maintain the estimation error within a certain threshold scales linearly with k, i.e. the intrinsic dimension of the manifold where embedding distributions are supported <ref:2605.07786#pg2>. CLIP and DINOv2 have a l2-normalization layer by design, making the boundedness assumption consistent with our experimental setting <ref:2605.07786#pg2>. The number of projections L needed to maintain the estimation error within a certain threshold scales linearly with k, i.e. the intrinsic dimension of the manifold where embedding distributions are supported <ref:2605.07786#pg2>. 3.2 Evaluation Protocol defines a multi-axis evaluation protocol going beyond aggregate benchmark scores <ref:2605.07786#pg2>. All experiments were conducted on a workstation equipped with an Intel Core Ultra 9 285K CPU with 24 physical cores, an NVIDIA GeForce RTX 5080 GPU with 16 GB of VRAM and CUDA 13.0, and a 2 TB SSD <ref:2605.07786#pg2>. The protocol evaluates responsiveness to diverse pixel- and latent-level distortions (Section 4.3.1), intra-dataset stability, efficiency under finite sample sizes, invariance of the degradation signal across different domains (Section 4.3.2), capability to capture subtle details during the generative refinement phase (Section 4.3.3), and overall alignment with human perceptual judgments (Section 4.3.4) <ref:2605.07786#pg2>. The degradation signal is defined as ∆k(τ, s) = dϕ(τs(Dk)), ϕ(Dk) <ref:2605.07786#pg2>. A well-behaved metric should satisfy monotonicity relative to the degradation signal as severity increases: ∆k(τ, s1) < ∆k(τ, s2) whenever s1. 4. The experiments are conducted across five heterogeneous evaluation domains—natural images, faces, dermoscopy, radiography, and remote sensing <ref:2605.07786#pg2>. Datasets include COCO-30k (natural scenes), HAM10000 (dermoscopic images), CelebA-HQ (high-resolution facial images), NIH ChestX-ray14 (medical radiographs), and NWPU-RESISC45 (remote sensing) <ref:2605.07786#pg2>. 4.2 Image Quality Assessment Metrics compares APEX with generative evaluation standards and recent foundation-model-based baselines <ref:2605.07786#pg2>. Table 1 summarizes the technical specifications of the evaluated metrics, categorizing them into AF (Assumption-Free feature distributions), FE (Foundation model Embeddings), and OT (Optimal Transport distance, no kernel/moment matching) <ref:2605.07786#pg2>. APEX-CLIP leverages OpenAI’s CLIP ViT-L/14@336px yielding 768-dimensional [CLS] embeddings <ref:2605.07786#pg5>. Conversely, APEX-DINO captures a richer architectural hierarchy by concatenating the 1024-dimensional [CLS] tokens from layers 6, 12, and 23 of DINOv2 ViT-L/14 into a single 3,072-dimensional representation <ref:2605.07786#pg5>. APEX is formulated as a feature-distribution metric that measures the SWD between image embedding extracted by state-of-the-art foundation models, specifically CLIP and DINOv2 <ref:2605.07786#pg4>. 4.3 Evaluation Protocol assesses metrics based on five fundamental criteria: responsiveness to diverse pixel- and latent-level distortions (Section 4.3.1), intra-dataset stability, efficiency under finite sample sizes, invariance of the degradation signal across different domains (Section 4.3.2), capability to capture subtle details during the generative refinement phase (Section 4.3.3), and overall alignment with human perceptual judgments (Section 4.3.4) <ref:2605.07786#pg2>. In Section 4, we present the results of the evaluation protocol above <ref:2605.07786#pg2>. 4.3.1 Degradation Sensitivity extends the protocol of [Jayasumana et al., 2024] with pixel- and latent-space perturbations across domains <ref:2605.07786#pg2>.

Improvements for AI systems

  1. Bold header: Assumption-free similarity measurement for image quality assessment

APEX provides a mathematically grounded, assumption-free similarity measure by leveraging the Sliced Wasserstein Distance (SWD), which not relying on rigid parametric hypotheses or kernel configurations. This allows AI systems to evaluate synthetic images using modern foundation models like CLIP and DINOv2 features without being constrained by outdated feature spaces or rigid Gaussian assumptions inherent in FID.

  1. Bold header: Foundation model agnostic evaluation framework

APEX is designed to be embedding-agnostic and uses two open-vocabulary foundation models, CLIP and DINOv2, as feature extractors, meaning the system can be adapted to natively supports any modern feature space. This enables AI systems to maintain high evaluation standards across diverse generative architectures without being bottlenecked by a single pre-trained backbone like Inception-v3.

  1. Bold header: Robustness against visual degradations and domain shifts

The framework demonstrates superior robustness to visual degradations and exhibits strong intra- and cross-dataset robustness, as evidenced by the low cross-dataset consistency metric (Λ) for APEX variants compared to baselines. This capability allows AI systems to reliably assess the quality of generated content even when subjected to various pixel-level corruptions, latent distortions, or shifts between domains like medical scans and satellite imagery.

  1. Bold header: Sample efficiency and computational scalability

APEX metrics show rapid sample convergence and provide stable evaluations with significantly fewer samples compared to MMD-based baselines. Furthermore, the computation for APEX scales as O(N log N), bypassing the prohibitive quadratic overhead of MMD variants, making it highly efficient for large-scale generative model training pipelines.

  1. Bold header: Diagnostic utility via multi-scale feature analysis

The use of APEX-DINO allows for a diagnostic utility by isolating which level of visual abstraction is affected by specific distribution shifts, as shown by the per-layer analysis on APEX-DINO focusing on the response of intermediate layers to progressive degradation. This enables AI researchers to pinpoint whether failures in image quality stem from low-level texture corruption or high-level semantic inconsistencies.

Abstract

As generative models achieve unprecedented visual quality, the gold standard for image evaluation remains traditional feature-distribution metrics (e.g., FID). However, these metrics are provably hindered by the closed-vocabulary bottleneck of outdated features and the assumptive bias of rigid parametric formulations. Recent alternatives exploit modern backbones to solve the feature bottleneck, yet continue to suffer from parametric limitations. To close this gap, we introduce APEX (Assumption-free Projection-based Embedding eXamination), a novel evaluation framework leveraging the Sliced Wasserstein Distance as a mathematically grounded, assumption-free similarity measure. APEX inherits effective scalability to high-dimensional spaces, as we prove with theoretical and empirical evidences. Moreover, APEX is embedding-agnostic and uses two open-vocabulary foundation models, CLIP and DINOv2, as feature extractors. Benchmarking APEX against established baselines reveals superior robustness to visual degradations. Additionally, we show that APEX metrics exhibit intra- and cross-dataset stability, ensuring highly stable evaluations on out-of-domain datasets.

Sources

Related papers