Just add noise: Debiasing tree-based variable importance in mixed data

arXiv:2609.14083 · stat.ML, cs.LG, stat.ME · Submitted 2026-09-12 · Read on arXiv

stat.ML, cs.LG, stat.ME

Submitted: 2026-09-12

Updated: 2026-09-12

License: http://creativecommons.org/licenses/by/4.0/

The gist: Variable importance scores from tree-based methods such as random forests favor continuous predictors over categorical ones.

Abstract

Variable importance scores from tree-based methods such as random forests favor continuous predictors over categorical ones. We present a theoretical analysis of this bias and propose a simple remedy: add a small amount of noise to each categorical predictor. The correction is demonstrated on a variety of simulated and real-world datasets and combined with integrated path stability selection to perform variable selection with false discovery control for mixed data.

Related papers