Improved cross-validated distances for multivariate pattern analysis
Laurent Caplette, Sarah Lippé
Université de Montréal · Azrieli Research Center of the CHU Sainte-Justine
q-bio.NC, stat.ML
Submitted: 2026-08-11
Updated: 2026-08-12
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: The paper "Improved cross-validated distances for multivariate pattern analysis" by Laurent Caplette and Sarah Lippé proposes improvements to existing cross-validated distance measures used in
Terminology
Summary
The paper Improved cross-validated distances for multivariate pattern analysis
by Laurent Caplette and Sarah Lippé proposes improvements to existing cross-validated distance measures used in multivariate pattern analysis (MVPA), specifically for Euclidean and Pearson correlation distances.
The authors begin by noting that distances in MVPA are susceptible to a noise bias, which can lead to less reliable estimates. They state: distances are non-negative and will typically increase with increasing noise, even when the true distance remains small.
Cross-validation is a strategy to mitigate this bias, and the cross-validated Euclidean distance is already known to be unbiased by noise.
In Section 2, the authors show that the cross-validated Euclidean distance is equivalent to a sum of between-partition distances. They derive that:
d2Euclidean,CV (x, y) = (d2 (xA, yB) + d2 (xB, yA))/2 − (d2 (xA, xB) + d2 (yA, yB))/2
They then generalize this by including within-partition distances, assuming random partitions, to obtain:
d2Euclidean,GCV (x, y) = (d2 (xA, yB) + d2 (xB, yA) + d2 (xA, yA) + d2 (xB, yB))/4 − (d2 (xA, xB) + d2 (yA, yB))/2
They call this the Euclidean distance with generalized cross-validation (G.C.V.). They note: Because it is defined from more individual distances, this formulation has the benefit of resulting in an estimator with reduced variance.
Testing on MEG data from Cichy et al. [8], they found the G.C.V. Euclidean distance was slightly more reliable than the existing C.V. Euclidean distance
and also more accurate (i.e. closer to the ground truth distance) than its standard C.V. counterpart
in simulations.
In Section 3, the authors address the Pearson correlation distance. They critique the cross-validated version proposed by Guggenmos et al. [13], which requires regularization because the covariances in the denominator can be very small or negative, leading to extremely positive, extremely negative, or even imaginary distances.
The authors instead define a cross-validated correlation distance by leveraging the relationship between Pearson correlation and Euclidean distance on Z-transformed vectors:
dP earson,CV,Ours (x, y) = (Z(xA) − Z(yA))⊤ (Z(xB) − Z(yB))/(2n)
This can be transformed to:
dP earson,CV,Ours (x, y) = (r(xA, xB) + r(yA, yB))/2 − (r(xA, yB) + r(xB, yA))/2
And further generalized to:
dP earson,GCV (x, y) = (r(xA, xB) + r(yA, yB))/2 − (r(xA, yB) + r(xB, yA) + r(xA, yA) + r(xB, yB))/4
The authors state: Because there is no arbitrary regularization, this implementation is more accurate than the one described by Guggenmos et al. [13].
They note that although Guggenmos' distance appears slightly more reliable in its base form, this increased reliability is largely driven by the enforced regularization
(clipping distances between 0 and 2). When this clipping is removed, the G.C.V. correlation distance is more reliable most of the time. Additionally, the G.C.V. correlation distance has a more interpretable time course, with a mean distance starting around 0 and increasing after stimulus presentation.
In Section 4, the authors discuss the relationship between cross-validation and within-class correction (W.C.C.). They show that when the number of trials is 2, W.C.C. is exactly equivalent to G.C.V. For more trials, they diverge: W.C.C. computes distances between all trials pairwise and averages the distances, G.C.V. averages trials into two partitions before computing distances.
Both eliminate bias for Euclidean distances, but for correlation distances, G.C.V. is more accurate because "correlation coefficients are intrinsically biased: their value tends towards 0 with increasing noise. Hence, computing many noisy correlations and averaging them is not equivalent to averaging data before computing correlations. They observed
much greater accuracy for the G.C.V. correlation distance than for the W.C.C. correlation distance (at intermediate and high SNRs)."
The authors conclude with recommendations: "we recommend discontinuing the use of the current form of the cross-validated Euclidean distance and either use its generalized cross-validation variant or the within-class-corrected variant, because they are both more reliable and accurate. For correlation distance, they recommend
discontinuing the use of both Guggenmos et al.'s [13] implementation of the cross-validated correlation distance and the within-class-corrected correlation distance, in favor of the generalized cross-validated correlation distance, which is more accurate than both." They note their conclusions should also apply to Mahalanobis and Spearman distances. A caveat is that the generalization requires random partitions, which is usually feasible in event-related designs.
Improvements for AI systems
Improvements to AI Systems:
- Noise-Robust Similarity Metrics for Representation Learning:
-
Replace standard Euclidean or cosine distances in contrastive learning, clustering, or nearest-neighbor search with the generalized cross-validated (GCV) Euclidean distance. This reduces noise-induced bias and variance in learned embeddings, especially when training on small or noisy datasets (e.g., few-shot learning, medical imaging).
-
The improved system can produce more stable and accurate feature representations, leading to better downstream classification and retrieval performance under noisy conditions.
- Bias-Free Correlation for Time-Series and Neuroimaging Analysis:
-
Integrate the GCV Pearson correlation distance into AI models that analyze temporal or spatial patterns (e.g., EEG/MEG decoding, fMRI connectivity, or speech emotion recognition). This avoids the regularization-induced distortion of existing cross-validated correlations.
-
The improved system can estimate true neural or behavioral similarity with higher accuracy, enabling more reliable brain-computer interfaces or diagnostic tools.
- Variance-Reduced Model Evaluation and Hyperparameter Tuning:
-
Use GCV distances as a metric for comparing model outputs (e.g., generative model samples, reinforcement learning policies) against ground truth, especially when data is split into multiple partitions. This reduces variance in performance estimates.
-
The improved system can perform more reliable model selection, early stopping, or anomaly detection, with fewer false positives/negatives due to noise.
- Unbiased Distance-Based Loss Functions for Generative Models:
-
Replace standard L2 or correlation losses in training objectives (e.g., for autoencoders, GANs, or diffusion models) with the GCV Euclidean or correlation distance. This removes the positive bias that inflates loss values with noise, leading to more accurate gradient signals.
-
The improved system can generate higher-fidelity outputs (e.g., images, audio) and converge faster, especially when training on limited or noisy data.
- Improved Clustering and Outlier Detection in High-Dimensional Data:
-
Apply the GCV correlation distance in clustering algorithms (e.g., k-means, hierarchical clustering) to avoid the intrinsic bias of correlation coefficients toward zero under noise.
-
The improved system can identify more meaningful clusters and outliers in noisy, high-dimensional datasets (e.g., single-cell RNA-seq, financial transaction data), leading to more accurate biological or fraud-detection insights.
- Cross-Validation-Aware Reinforcement Learning:
-
Use GCV distances to estimate state-action similarity or reward prediction errors across different environment partitions, reducing variance in value function updates.
-
The improved system can learn more stable and sample-efficient policies in noisy or partially observable environments.
- Interpretable Model Comparison and Explainability:
-
Adopt GCV distances to compare latent representations of different models (e.g., in model distillation or interpretability analysis). This yields more interpretable and unbiased similarity scores, especially when models are trained on noisy data.
-
The improved system can provide clearer insights into which model layers or features are most aligned, aiding in debugging and transparency.
Sources
Related papers
- BrainWave: A Brain Signal Foundation Model for Clinical Applications
- Toward Robust, Reproducible, and Widely Accessible Intracranial Speech Brain-Computer Interfaces: A Comprehensive Narrative Review of Neural Mechanisms, Hardware, Algorithms, Evaluation, Clinical Pathways and Future Directions
- CytoNet: A Foundation Model for the Human Cerebral Cortex at Cellular Resolution
- Emergence of psychopathological computations in large language models
- NeuroAI and Beyond: Bridging Between Advances in Neuroscience and Artificial Intelligence
- Attraction to hierarchical feature memory explains orientation bias