Towards Truly Unsupervised Evaluation of Feature Selection -- Extended Version

arXiv:2608.12057 · cs.LG · Submitted 2026-08-12 · Read on arXiv

University of Southern Denmark

cs.LG

Submitted: 2026-08-12

Updated: 2026-09-14

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: The paper "Towards Truly Unsupervised Evaluation of Feature Selection" by Hafiz Saud Arshad, Muhammad Rajabinasab, and Arthur Zimek addresses the problem that established "unsupervised" evaluation

Terminology

Summary

The paper Towards Truly Unsupervised Evaluation of Feature Selection by Hafiz Saud Arshad, Muhammad Rajabinasab, and Arthur Zimek addresses the problem that established unsupervised evaluation techniques for feature selection are not truly unsupervised. The authors state: "Most of the methods commonly used for the unsupervised evaluation of feature selection algorithms suffer from critical design flaws which question their unsupervised nature. In this paper, we provide a critical discussion on the established allegedly unsupervised evaluation techniques, and shed light on the reasons why they are not truly unsupervised but, at best, supervised evaluation under an unsupervised downstream task."

The core issue identified is that the common unsupervised evaluation approach uses clustering (usually k-means) on the selected features and then compares the clustering results against ground-truth labels. The authors explain: "This is clearly problematic, as using any information about the labels is in contrast with the alleged unsupervised nature which we expect from the evaluation, rendering the current established so-called unsupervised way of evaluation, at best, as supervised evaluation under an unsupervised downstream task. Even not using the labels, and just using the number of distinct classes to select the number of clusters, should still be considered a violation of the unsupervised setting."

To address this gap, the authors propose a novel, truly unsupervised evaluation framework to measure the quality of the feature selection algorithms without any form of information about the labels. The proposed framework works as follows: "We apply PCA for effective dimensionality reduction and select top f principal components. To evaluate the result of some feature selection algorithm, selecting f features, we calculate an inter-dataset similarity metric IDS for the two representations of the data resulting from PCA and from some feature selection method. The framework utilizes unsupervised Principal Component Analysis, and optimal transport to measure the quality of the feature selection methods in a truly unsupervised manner."

The method evaluates feature selection algorithms by measuring the optimal transport distance between the selected feature subset and a PCA-based reference subspace of the same size. The authors state: "The objective is to identify feature selection algorithms that select features in a way that they are as close as possible to the representation obtained from PCA. The better a subset preserves the structure of the PCA representation, the higher it ranks." They use four optimal transport-based inter-dataset similarity measures: Earth Mover's Distance (OT EMD2), Entropic-Regularized Sinkhorn Distance (OT SINKHORN2), Gromov-Wasserstein Distance (OT GW2), and Sliced Wasserstein Distance (OT SLICED SW). Since these are distance measures, they convert them to similarity via the formula: OTsim = 1/OT if OT ≠ 0, else OT = +∞.

In their experiments, the authors evaluate five feature selection methods (variance-based, correlation-based, Laplacian Score, Subspace Clustering-based Feature Selection (SCFS), and Variance-Covariance Subspace Distance (VCSDFS)) plus a Random baseline, across eight high-dimensional datasets (COIL20, Isolet, ORL, lung, lung discrete, warpAR10P, warpPIE10P, and Yale). They compare nine evaluation metrics organized into four groups: supervised (ACC, AUC), pseudo-unsupervised (CLSACC, NMI), model-agnostic (AAD), and their proposed truly unsupervised metrics (four OT-based variants). They assess each method at 10 feature fractions in the low-percentage regime (0.5% to 5%).

The results show that OT GW2 produces rankings partly aligning with the supervised block by placing VCSDFS and Correlation ranked lowest and Laplacian ranked highest, while OT SINKHORN2 and OT EMD2 produce the ranking largely inverted with respect to the supervised block, rating the two weakest selectors best. The authors note that OT SLICED SW shows degraded resolution. The pairwise correlation analysis (Pearson and Spearman) reveals that "OT GW2 and OT SLICED SW exhibit positive correlation with respect to traditional metrics, especially the supervised block (Pearson coefficients: 0.29 − 0.37; Spearman coefficients: 0.26−0.33), While OT EMD2 and OT Sinkhorn2 exhibit opposite patterns. The authors interpret these findings as indicating that the proposed framework carries predictive validity: the proposed (label-free) metrics order selectors in a way that anticipates their label-based downstream performance."

The authors conclude: "We proposed a truly unsupervised evaluation framework that, instead of exploiting ground-truth labels, evaluates a feature selection algorithm by measuring the optimal transport distance between its selected feature subset and a PCA-based reference subspace of the same size. Across eight high-dimensional datasets, the proposed metrics showed consistent correlations with established evaluation metrics under both correlation measures."

The paper also acknowledges several limitations. First, optimal transport methods can be computationally expensive, particularly for large datasets. Second, "the use of PCA as the reference mechanism imposes a structural limitation on the framework. The number of principal components that can be obtained from PCA is bounded by min(n, d), where n is the number of instances and d is the number of features. Consequently, the proposed evaluation is either restricted to datasets satisfying n ≥ d, or limited to a partial view of the feature selection process when the number of selected features exceeds this bound. Third, the empirical scope of this study is limited. The experiments cover a relatively small number of datasets and feature selection methods."

For future work, the authors suggest studying the framework at finer granularity, relaxing the reliance on PCA (e.g., using autoencoders), investigating the raw values of the proposed metrics, potentially using Gromov-Wasserstein distance to compare selected features directly with the full-dimensional dataset (since it can compare distributions defined on spaces of different dimensionality), and systematically investigating the resolution loss observed for OT SLICED SW. The authors emphasize that the current results should primarily be interpreted as a proof of concept and that "the broader aim is to demonstrate the feasibility of evaluating feature selection without relying on ground-truth labels and to motivate a more rigorous, granular, and reproducible investigation of truly unsupervised evaluation in future work."

Improvements for AI systems

Improvements to AI Systems:

  1. Label-Free Model Selection for Unsupervised Learning Pipelines: AI systems that perform feature selection for downstream tasks (e.g., clustering, anomaly detection, or dimensionality reduction) can now be evaluated and tuned without any access to ground-truth labels. This enables automated model selection in fully unsupervised settings, such as exploratory data analysis on unlabeled datasets, where previously the system would have to either use pseudo-labels (risking bias) or rely on heuristics.

  2. Robust Feature Subset Ranking via Optimal Transport: The AI system can rank feature selection algorithms by computing the optimal transport distance (specifically Gromov-Wasserstein, OT GW2) between the selected feature subspace and a PCA reference subspace of the same dimensionality. This provides a principled, continuous score that reflects how well the selected features preserve the intrinsic variance structure, without requiring any label information. The system can then automatically choose the feature selector that best preserves data geometry.

  3. Predictive Validation Without Labels: The system can leverage the finding that OT GW2 and OT SLICED SW positively correlate with supervised metrics (Pearson 0.29–0.37, Spearman 0.26–0.33). This means the AI can use these label-free metrics as a proxy for downstream supervised performance, allowing it to pre-select feature subsets that are likely to yield high accuracy later, even when labels are unavailable at selection time.

  4. Avoiding Pseudo-Unsupervised Pitfalls: The AI system can be explicitly designed to reject evaluation methods that use cluster counts derived from label information (e.g., setting k-means k to the number of classes). Instead, it can adopt the proposed framework, ensuring that its evaluation pipeline is truly unsupervised, thus preventing subtle label leakage that could lead to overfitting or biased model choices.

  5. Computational Efficiency Awareness: The system can incorporate the computational cost of optimal transport (especially OT GW2) into its decision-making, automatically switching to faster variants (e.g., OT SLICED SW) for large datasets when resolution loss is acceptable, or using entropic regularization (OT SINKHORN2) for speed while being aware of its inverted ranking behavior relative to supervised metrics.

  6. Handling High-Dimensional, Low-Sample Data: The system can detect when the number of selected features exceeds min(n, d) (where n is samples, d is features) and automatically adjust the evaluation strategy—either by using Gromov-Wasserstein to compare selected features directly with the full-dimensional dataset (since it handles different dimensionalities) or by limiting the PCA reference to the available principal components, thus avoiding invalid comparisons.

  7. Automated Metric Selection for Evaluation: The AI system can dynamically choose which OT-based metric to use based on the dataset characteristics and the desired trade-off between alignment with supervised performance and computational cost. For instance, it can default to OT GW2 for high-stakes decisions (since it aligns best with supervised metrics) and use OT SLICED SW only when computational constraints dominate.

  8. Granular Feature Selection Analysis: The system can perform fine-grained evaluation at multiple feature fractions (e.g., 0.5% to 5% of features) using the proposed framework, enabling it to identify not just which algorithm is best overall, but at which feature budget each algorithm excels—useful for resource-constrained deployment scenarios.

What the Improved AI System Can Do:

  • Self-Evaluate Feature Selection in Production: Deploy an unsupervised feature selection module that continuously monitors and re-ranks its own feature choices using OT-based distances to a PCA reference, without ever needing labeled data, thus adapting to drifting data distributions in real-time.

  • Provide Explainable Unsupervised Quality Scores: Output a single, interpretable score (e.g., OT GW2 similarity) for any feature subset, allowing data scientists to compare different feature engineering choices on unlabeled data with confidence.

  • Pre-Validate Models for Supervised Tasks: When labels are scarce or expensive, the system can pre-screen feature subsets using the label-free OT metrics, predicting which subsets will likely perform well under supervised classifiers, thereby reducing the need for extensive labeled data collection.

  • Avoid Label Leakage in AutoML: In automated machine learning pipelines, the system can guarantee that feature selection evaluation is truly unsupervised, preventing the common mistake of using cluster-based metrics (like NMI) that implicitly rely on label-derived cluster counts, thus producing more honest and generalizable model rankings.

  • Scale to Large Unlabeled Datasets: By selecting computationally efficient OT variants (e.g., sliced Wasserstein) when needed, the system can evaluate feature selection on datasets with millions of samples, where traditional supervised evaluation is impossible.

  • Support Multi-View Data Comparison: Using Gromov-Wasserstein, the system can compare feature subsets of different dimensionalities directly against the original high-dimensional data, enabling evaluation of feature selection even when the selected subset is much smaller than the original feature space, without needing a PCA reference of the same size.

Abstract

Feature selection is one of the most important and fundamental tasks in data mining, tackled by a family of methods with an established set of evaluation techniques to measure the quality of a specific method. Most of the methods commonly used for the unsupervised evaluation of feature selection algorithms suffer from critical design flaws which question their unsupervised nature. In this paper, we provide a critical discussion on the established allegedly unsupervised evaluation techniques, and shed light on the reasons why they are not truly unsupervised but, at best, supervised evaluation under an unsupervised downstream task. We also propose a novel, truly unsupervised evaluation framework to measure the quality of the feature selection algorithms without any form of information about the labels. The proposed framework utilizes unsupervised Principal Component Analysis, and optimal transport to measure the quality of the feature selection methods in a truly unsupervised manner.

Sources

Related papers