Provable FDR Control for Deep Feature Selection: Deep MLPs and Beyond

arXiv:2512.04696 · stat.ML, cs.LG, math.ST, stat.TH · Submitted 2025-12-04 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Provable FDR Control for Deep Feature Selection".

Jane: A flexible feature selection framework based on deep neural networks is developed that provides a theoretical guarantee for controlling the false discovery rate (FDR),

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So Jane, we’ve been looking at this paper, "Provable FDR Control for Deep Feature Selection: Deep MLPs and Beyond," and it sounds like they've really tackled a big problem in how we choose features from complex deep learning models.

Jane: It certainly seems like the main idea is providing a theoretical guarantee for controlling the false discovery rate, which is that measure of Type-I error, in these very flexible architectures.

Lu: Exactly! I find the scope they cover is fascinating; they aren't just sticking to simple linear models or fixed structures.

Meng: That flexibility across MLPs with arbitrary width and depth, convolutional networks, and even attention mechanisms is quite something for practical application.

Lalam: From my perspective as a language model, this work suggests that we can move beyond heuristic methods like LIME or SHAP when we want rigorous error control in feature selection.

Tom: Right! So the core thesis of "Provable FDR Control for Deep Feature Selection: Deep MLPs and Beyond" is that they've developed a flexible framework based on deep neural networks that can approximately control the false discovery rate, which is a measure of Type-I error.

Jane: That means they are offering a theoretical guarantee for controlling how often you make a mistake when picking features.

Lu: The analysis underpinning this relies on building it upon a multi-index data-generating model and an asymptotic regime where the feature dimension diverges faster than the latent dimension, which is quite specific mathematically.

Meng: That level of mathematical rigor is what makes it interesting for theory, but I wonder how that translates when we actually try to deploy this in a real system with limited computational resources.

Lalam: The paper mentions their analysis is built upon a multi-index data-generating model and an asymptotic regime where the feature dimension n diverges faster than the latent dimension q star, while the sample size, the number of training iterations, network depth, and hidden layer widths are left unrestricted.

Tom: So they’re setting up a scenario where we can test this method across all those varying network parameters without worrying about them being too constrained.

Paper summary: Jane: And what they claim in this setting is that for any null feature index j in S c, the quantity sqrt n xi(t)j / P B xi(t) converges in distribution to a normal distribution with mean zero and variance one as n goes to infinity, provided q* is sublinear in n.

Meng: That asymptotic normality is a big theoretical win, but I’m always thinking about the practical implementation details—how do we actually calculate that input sensitivity xi(t) reliably in practice?

Lu: The proof establishes this by showing that the normalized quantity P B xi(t)/P B xi(t) is uniformly distributed on a unit sphere lying within the space Col(B).

Tom: That uniform distribution property is crucial because it sets up the whole mechanism for their importance statistic construction.

Jane: And they use this foundation to construct an importance statistic M j using a data-splitting aggregation scheme, which is then used with a thresholding rule to control the FDR.

Meng: So if I follow that flow, they have this statistical foundation, and then they build a specific metric M j based on the input sensitivity xi(t)j one xi(t)j two psi(xi(t)j one xi(t)j two), and then use a cutoff tau alpha derived from the distribution of M j to control the FDR at a nominal level alpha.

Tom: That entire process is what they are aiming for when they describe their feature-selection procedure.

Jane: It sounds like the key mechanism for controlling Type-I error hinges on that symmetry of M j under the null, which allows them to control FDP d(u) by counting variables in S c.

Lalam: I think what really stands out is how they tie this into existing literature, noting that FDR control represents only the minimal requirement for avoiding excessive inclusion of irrelevant variables in feature selection.

Lu: They explicitly address the complementary criterion of Type-II error, acknowledging that their analysis provides no theoretical guarantees for it, even though numerical evaluations exist.

Tom: That distinction is important; they aren't claiming to solve the entire problem of feature selection, just the false discovery rate aspect.

Jane: So when we look at the title and authors of "Provable FDR Control for Deep Feature Selection: Deep MLPs and Beyond," what they’re essentially pointing out is that this work bridges the gap between modern XAI attribution methods and true error control in feature selection.

Paper summary: Meng: It suggests that for deep models, relying solely on empirical importance scores without a theoretical safeguard against Type-I errors leaves us exposed to potential false positives.

Lu: The implications for the field are significant because they've established a method applicable across such a broad range of architectures, from simple MLPs to Transformers and CNNs.

Tom: That flexibility is what makes this work so much more useful than methods that only work for one specific type of network structure.

Jane: It means that researchers can use these theoretical guarantees to justify their selection processes in complex systems, which should lead to more trustworthy feature engineering workflows across various domains.

Meng: From an engineering standpoint, having a provable control mechanism helps us move away from just picking features that look important and instead gives us a statistical framework to ensure we aren't introducing too many noise variables into our final model pipeline.

Lalam: If we consider the broader impact, this kind of work could fundamentally improve how we interpret deep learning models by making feature selection more statistically sound, which in turn enhances the reliability of the AI systems we build.

Tom: So, to wrap up on this initial overview, "Provable FDR Control for Deep Feature Selection: Deep MLPs and Beyond" is introducing a mathematically grounded way to select features from complex deep learning models by providing a theoretical guarantee for controlling the false discovery rate.

Jane: That means we have a new tool that lets us be more confident in which features we keep.

Lu: The potential impact lies in enabling deeper analysis of feature importance within any modern, flexible neural network architecture using this framework.

Meng: For practical application, it means our selection process won't just be an educated guess anymore; it will have a statistical backing we can measure against the nominal error rate alpha.

Lalam: It really helps improve culture in AI development by pushing us toward more rigorous validation steps rather than just relying on whatever metric looks best on the surface.

Tom: We’re going to take a quick pause, but when we come back, we’re going to talk about what this means for the real-world deployment of these selection methods.

Conclusion: Tom: So, we’re wrapping up our discussion on "Provable FDR Control for Deep Feature Selection: Deep MLPs and Beyond," which is a paper that lays out a mathematical way to control the false discovery rate in deep learning feature selection. Jane, how would you summarize the big picture idea in simple terms for our listeners?

Jane: Well, Tom, essentially this paper gives us a solid mathematical proof that we can use these complex neural networks to pick features without worrying too much about accidentally picking irrelevant ones by controlling the false discovery rate. It moves feature selection from a bit of guesswork into something statistically sound.

Lu: I think the real excitement here is how they tackle the flexibility aspect; they show this works across MLPs, CNNs, and even Transformers, which means it's not just for one specific type of model anymore. That opens up so many creative avenues for feature engineering that we haven't even imagined yet.

Meng: I’m thinking about the practical side—how does this translate to what we actually deploy? If the method is theoretically sound, can we trust it more when building production systems where errors are costly? That's my main concern.

Lalam: From my viewpoint, the cultural impact here is significant because it shows us that deep learning tools aren't just black boxes; we can start applying rigorous statistical principles to them, which helps build a culture of more trustworthy and less error-prone AI systems globally.

Tom: That’s a powerful thought, Lalam. It sounds like this work isn't just about math; it’s about making the tools we use more reliable for everyone. Jane, what do you think is the main implication for how people use these models in their daily work?

Jane: The main implication is that researchers and engineers can confidently use these selection methods knowing they have a theoretical bound on how many false positives they might get, which simplifies the validation process tremendously. It gives us a much better way to justify our choices.

Lu: And this moves the focus from just achieving high accuracy in the model itself to ensuring the feature set we *use* is statistically robust, which is a necessary step for real-world deployment.

Meng: I agree with that; having that statistical backing makes it much easier for us to integrate these feature selection pipelines into our existing engineering workflows without adding layers of uncertainty.

Lalam: It really shifts the conversation toward building AI tools that are not just smart, but also statistically honest about their inputs and outputs, which is a huge step forward for the entire field.

Tom: So, to wrap up on this paper’s conclusion—"Provable FDR Control for Deep Feature Selection"—it’s a mathematical framework that lets us select features from complex deep learning models while giving us a quantifiable guarantee on controlling the false discovery rate. Jane, what's your final thought on why this specific approach is so important right now?

Jane: It’s important because it bridges the gap between high-level AI architecture and practical statistical control, offering a rigorous method that applies broadly across many different deep learning models.

Lu: The breadth of applicability is what makes the theoretical foundation so compelling; it shows the underlying principles hold up regardless of whether you're looking at a simple MLP or a massive Transformer.

Meng: I just hope the experimental results translate smoothly into stable, repeatable results in real-world, messy data environments, as that’s where our engineering work happens.

Lalam: This paper has potential to elevate the entire culture around AI development by providing this kind of statistical rigor for feature selection across diverse applications.

The University of Tokyo

stat.ML, cs.LG, math.ST, stat.TH

Submitted: 2025-12-04

Updated: 2026-10-01

Code: https://github.com/zifanzhu/DeepLINK

Importance score: 83/100

The gist: A flexible feature selection framework based on deep neural networks is developed that provides a theoretical guarantee for controlling the false discovery rate (FDR), addressing a gap between

Key concepts

Input Sensitivity ξ(t)
This measures how sensitive the model's output is to changes in a specific input feature. The theory proves that this sensitivity follows a normal distribution under certain conditions, which is crucial for the subsequent statistical selection process.
Importance Statistic Mj
This is a calculated value for each feature that quantifies its relevance. It combines the input sensitivities of two features using a specific function and sign, designed to have desirable properties that allow for FDR control during selection.
FDR Control Mechanism
A specific thresholding rule is used to select features based on their importance statistics (Mj). This mechanism ensures that the false discovery rate stays below a specified level alpha, meaning the selected features are likely true predictors.

Terminology

Summary

A flexible feature selection framework based on deep neural networks is developed that provides a theoretical guarantee for controlling the false discovery rate (FDR), addressing a gap between high-dimensional statistics and explainable AI. This method allows for feature selection in complex architectures, including MLPs, CNNs, and Transformers, while retaining modeling flexibility across various training protocols.

Theoretical Foundation and Assumptions

The analysis is built upon a multi-index data-generating model where the response vector follows a structure defined by a weight matrix B and a deterministic function g. The core theoretical result establishes the marginal asymptotic normality of the input sensitivity, denoted as input sensitivity ξ(t), under specific conditions. The key assumption is that for each null feature index j in S c, the condition y ⊥⊥ xj x−j holds (Definition 2). Under Assumptions 1–4 (including B-right orthogonal invariance of the design matrix X and stochastic gradient descent options), Theorem 1 proves that for any null feature index j ∈ S c, the quantity √nξ(t)j / P⊥B ξ(t) converges in distribution to N(0, 1) as n → ∞ while q∗ = o(n). This asymptotic normality is established by showing that P⊥B ξ(t)/P⊥B ξ(t) is uniformly distributed on the unit sphere lying in Col(B)⊥.

Feature Importance Statistic Construction

The paper constructs an importance statistic Mj for each feature using a data-splitting aggregation scheme to achieve asymptotic FDR control. The procedure involves two stages: (I) computing an importance statistic Mj for each feature that possesses desirable distributional properties, and (II) selecting features by applying an appropriate thresholding rule that controls the FDR. Specifically, the importance statistic is defined as:

Mj = sign ξ(t)j1 ξ(t)j2 ψ(ξ(t)j1, ξ(t)j2), where ψ: R≥0 × R≥0 → R≥0 is a user-specified function assumed to be non-negative, symmetric, positive homogeneous, and monotone in each input (Assumption 5). The data splitting approach relies on the symmetry under the null: if Mj is symmetric around zero under the null, then FDP(u) = FDP d(u), allowing for FDR control without needing knowledge of the asymptotic variance.

FDR Control Mechanism

The selection procedure uses a cutoff τα to control the FDR at a nominal level α. The cutoff is defined as:

τα = min u > 0: FDP d(u) ≡ n j: Mj u ∨ 1 ≤ α. The paper demonstrates that this procedure asymptotically controls the FDR at or below the nominal level α under Assumptions 1–7, leading to FDR ≤ α + op(1) and lim sup n→∞ FDR ≤ α. This control is achieved because the symmetry of Mj under the null ensures that FDP d(u) is controlled by counting variables in S c, which are assumed to have a distribution symmetric around zero.

Generalization and Robustness

The framework exhibits significant flexibility across various deep learning architectures and training protocols. The method accommodates:

  1. Flexible scope across architectures: applicable to MLPs with arbitrary width and depth, convolutional and recurrent networks, attention mechanisms, residual connections, and dropout.

  2. Agnostic to the training protocol: guarantees accommodate stochastic gradient descent with arbitrary data-independent initialization schemes and learning rates.

  3. General data-generating process: the theory is established under a multi-index model as the data-generating process with unknown nonlinearity, capturing rich latent structures beyond generalized linear models.

Furthermore, the analysis extends to designs that violate B-right orthogonal invariance by relaxing it to B-ROI (Assumption 1 (ii)), where Xd=XU for all U ∈ GB. Appendix C provides numerical evidence showing that asymptotic normality is preserved even when feature correlations are present, confirming the robustness of the results in more general settings.

Empirical Validation

Numerical experiments validate the theoretical guarantees across different network types (MLP, 1D-CNN, LSTM, Transformer) and various sample sizes (m). The results show that for a nominal FDR level α = 0.1, the proposed method achieves relatively high power under FDR control. For instance, when comparing against DeepLINK [34] and Neural Gaussian Mirror (NGM) [30], the proposed method demonstrates superior performance in controlling the false discovery rate while maintaining acceptable power levels across different data-generating processes (N(0,1), t(3), Spiked). The convergence to asymptotic normality is numerically confirmed by QQ-plots (Figure 6) and histograms of the normalized input sensitivity statistics (Figure 2).

Improvements for AI systems

Here are the specific improvements and capabilities for AI systems derived from this research paper:


)1. Robust, Theoretically Guaranteed Feature Selection in Deep Learning Models:

The core improvement is replacing heuristic feature selection methods (like simple thresholding of feature importance scores) with a procedure that provides a formal, asymptotic guarantee on controlling the False Discovery Rate (FDR).

  • A system can select features with a confidence level of at least 1% (for an FDR control level of 0.1), meaning the probability of including at least one irrelevant feature in its selected set is bounded.

  • This guarantee holds across a wide variety of modern deep learning architectures, including MLPs (arbitrary width/depth), Convolutional Networks (CNNs), Recurrent Networks (LSTMs), Attention mechanisms, Residual Connections, and Dropout layers.

)2. Architectural Agnosticism:

The method is agnostic to the specific training protocol and optimization strategy.

  • The selection procedure works regardless of whether Stochastic Gradient Descent (SGD) is used with data-independent initializations or different learning rates.

  • This allows researchers to apply the feature selection tool universally across different model training pipelines without needing to re-derive complex error bounds for each new optimizer.

)3. Stability and Reliability under High Dimensionality:

The theoretical framework establishes that the input sensitivity of deep networks admits a normal approximation when the number of features exceeds the latent dimension (i.e., as long as feature dimension grows faster than the intrinsic model complexity).

  • This allows for reliable feature selection in extremely high-dimensional datasets where traditional statistical methods fail due to sparsity or curse of dimensionality.

  • The analysis confirms that even when training is performed iteratively (at each step, using early stopping), the resulting importance metrics are asymptotically normal, ensuring the validity of the FDR control mechanism.

)4. Flexible Correlation Handling via Design Invariance:

The framework is designed around a relaxed form of design invariance known as B-right orthogonal invariance (B-ROI).

  • The system can reliably perform feature selection even when features exhibit complex correlation structures, including time dependence and heavy-tailed marginal distributions (as explored in Appendix C).

  • While the primary guarantee relies on this assumption, the paper provides a roadmap for extending the analysis to more general correlation structures, making it robust for real-world datasets where feature correlations are non-trivial.

)5. Optimized Selection via Data Splitting:

The selection procedure utilizes a data-splitting approach to construct importance statistics using the ratio of sensitivities from two independent data subsets.

  • This method is computationally simple and avoids the need to estimate complex asymptotic variance terms, making it highly practical for large-scale deployment.

  • It effectively leverages the symmetry properties of null features (those not relevant to any feature) to control Type I error without requiring knowledge of the precise null distribution's variance or convergence rate.

)6. Adaptive Thresholding and Robustness:

The procedure allows for a flexible, user-defined function ψ to construct the importance statistic, enabling adaptation to different selection criteria (e.g., selecting features with high positive scores vs. those with high magnitude regardless of sign).

  • This flexibility means the system can be tuned to prioritize different aspects of feature contribution—for instance, focusing on large contributions as defined by the user-specified function—leading to more tailored feature subsets.

)Improved AI System Capabilities:

By implementing this framework, an AI system can achieve:

  1. A high-confidence Relevance Filter: The system can automatically prune a vast set of features (e.g., thousands of input variables) down to a statistically sound subset that is highly likely to contain only truly relevant predictors, minimizing the risk of false discoveries in scientific or clinical applications.

  2. Model Interpretation with Statistical Confidence: When used with deep learning models, the system can generate feature importance scores for complex models (like Transformers or LSTMs) and provide a rigorous statistical statement about the reliability of those scores—specifically, that Type I errors are controlled at a specified level.

  3. Early Stopping Integration: Because the asymptotic normality holds at every training iteration, an adaptive system can monitor the stability of feature importance metrics across training steps to determine the optimal stopping point for model training, potentially reducing Type II error (missing relevant features) while maintaining FDR control.

  4. Generalizability Across Data Types: The framework's foundation in a multi-index model allows it to be applied to diverse data generating processes beyond standard linear or Gaussian models, such as those arising from complex biological or time-series data, provided the underlying structure is captured by the multi-index formulation.

Sources

Related papers