Model-Aware Data Cleaning for Tabular Foundation Models
cs.LG, cs.DB
Submitted: 2026-04-28
Updated: 2026-09-17
Comments: 14 pages, 5 figures v2: retitled (was 'Prior-Aligned Data Cleaning for Tabular Foundation Models'). Corrected framing: the reward regularizes distance to the dirty data, not the TFM prior (measured only as a diagnostic); prior alignment reframed as future work. Trained-policy (B-RL) results now evaluated leak-free under the held-out nested protocol
Code: https://github.com/LaureBerti/Learn2Clean
License: http://creativecommons.org/licenses/by/4.0/
The gist: Tabular Foundation Models (TFMs) achieve state-of-the-art zero-shot accuracy on small tabular datasets, but their in-context learning assumes approximately clean inputs: real- world missing values,
Terminology
Abstract
Tabular Foundation Models (TFMs) achieve state-of-the-art zero-shot accuracy on small tabular datasets, but their in-context learning assumes approximately clean inputs: real- world missing values, outliers, and duplicates create a prior mismatch that degrades both accuracy and calibration. We study reinforcement learning for tabular data cleaning, a learned policy that sequences cleaning operators and introduce L2C-TFM with a model-aware reward (TFMAwareReward). We are explicit about what this reward optimizes: it regularizes the Wasserstein distance between the cleaned and the original (dirty) data, a distributional-stability term, which we measure as a diagnostic. Across six experiments on ten OpenML datasets: (i) three of seven reward designs collapse to degenerate strategies, so reward engineering is non-trivial; (ii) under an 8-seed repeated-holdout protocol the model-aware reward matches a random-forest-reward baseline on accuracy (p=0.38), with a benefit confined to minority-class macro-F1 under class imbalance that is partly a reward-agnostic calibrated-threshold effect; and (iii) a policy pre-trained on one dataset transfers to held-out datasets. A diagnostic analysis shows that the two distances are distinct objectives: cleaning tends to move data away from the prior, and prior-distance, not distance-to-dirty, is what tracks downstream quality. We therefore treat prior alignment as a motivating objective and a target for future work, not a property of the reward evaluated here. Code, datasets, and the nested evaluation harness are available at https://github.com/LaureBerti/Learn2Clean/tree/master/Learn2Clean TFM.
Sources
- OpenML Benchmarking Suites
- AutoGluon-Tabular: Robust and Accurate AutoML for Structured Data
- DataComp: In search of the next generation of multimodal datasets
- Revisiting Deep Learning Models for Tabular Data
- Why do tree-based models still outperform deep learning on tabular data?
- CleanSurvival: Automated data preprocessing for time-to-event models using reinforcement learning
- CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification Tasks
- Revisiting the Calibration of Modern Neural Networks
- Can You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift
- TabICL: A Tabular Foundation Model for In-Context Learning on Large Data
- A Closer Look at Deep Learning Methods on Tabular Datasets
- Data-centric Artificial Intelligence: A Survey
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks