Balanced Adaptive Prototype Selection for Scalable TabPFN Inference on Large-Scale Tabular Data
Mahboobe Jadid, Melika Rezaye Garkani, Ali Mousavi
Islamic Azad University
cs.LG
Submitted: 2026-08-13
Updated: 2026-08-14
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: This paper introduces Balanced Adaptive Prototype Selection (BAPS), a framework for constructing compact, information-preserving contexts for scalable TabPFN inference on large-scale tabular data.
Terminology
Summary
This paper introduces Balanced Adaptive Prototype Selection (BAPS), a framework for constructing compact, information-preserving contexts for scalable TabPFN inference on large-scale tabular data. The paper identifies bounded inference context
as the primary scalability bottleneck of pretrained tabular foundation models
and formulates scalable context construction as an information-preservation problem.
BAPS operates without modifying or retraining the pretrained model
and jointly preserves representative structure, informative decision boundaries, local density, class balance, and feature-space diversity.
The proposed framework uses a Class-Aware Budget Allocation
to distribute prototypes across classes, then generates candidates from three complementary perspectives: "Representative Selection captures the dominant statistical structure of each class, Boundary Selection retains informative samples near decision boundaries, and Density Selection preserves reliable local neighborhood structures. These candidates are then integrated through
Diversity Refinement, which optimizes V(P) by removing redundant prototypes while maximizing feature-space coverage. For additional robustness, BAPS
optionally generates multiple prototype contexts using different random initializations and aggregates their prediction probabilities."
Experiments were conducted on the million-row HIGGS and SUSY datasets, showing that 512 prototypes retain strong predictive performance and reliable calibration, corresponding to an approximately 1,953-fold context compression.
Specifically, on HIGGS, BAPS Ensemble achieves a balanced accuracy of 0.697, improving over its single-context variant (0.688), Stratified Random Sampling (0.670), and remaining competitive with KMeans Medoid (0.690).
On SUSY, BAPS Ensemble obtains the highest balanced accuracy (0.782), exceeding Stratified Random Sampling (0.779), KMeans Medoid (0.771), and single-context BAPS (0.776).
The paper also reports that the paired Wilcoxon test with Holm correction confirms that the performance improvement of the complete BAPS framework over the strongest single-context baseline is statistically significant (p = 0.031).
All experiments were conducted on an Intel Core i7 CPU with 16 GB RAM and no GPU acceleration.
The paper concludes that effective context construction is not merely an optimization strategy but a fundamental requirement for scaling pretrained tabular foundation models to million-scale datasets.
The main contributions are: identifying bounded inference context as the primary scalability bottleneck, proposing BAPS as an information-preserving context-construction framework, and demonstrating that BAPS enables practical large-scale TabPFN inference on million-row datasets using only 512 prototypes and standard CPU hardware.
Improvements for AI systems
Improvements to AI Systems:
-
Add a context-construction module to pretrained tabular foundation models that selects a fixed-size, information-dense subset of training rows before inference, enabling deployment on million-row datasets without retraining or architectural changes. The improved system can process datasets that previously exceeded its context window (e.g., 1M rows) using only 512 prototypes, achieving 1,953× compression while maintaining predictive accuracy and calibration.
-
Implement multi-perspective prototype selection that jointly optimizes for class balance, decision-boundary proximity, local density, and feature-space diversity. The improved system can avoid the common failure mode of random sampling (which loses rare-class structure) and k-means medoids (which over-focus on cluster centroids), leading to more robust performance on imbalanced and high-dimensional tabular data.
-
Add an ensemble-of-contexts inference strategy that generates multiple prototype contexts via different random initializations and aggregates their prediction probabilities. The improved system can reduce variance and improve accuracy (e.g., +0.009 balanced accuracy on HIGGS, +0.006 on SUSY) without extra training, making it suitable for high-stakes classification where single-context decisions are unreliable.
-
Integrate a diversity-refinement post-processing step that removes redundant prototypes while maximizing feature-space coverage. The improved system can automatically prune redundant samples (e.g., near-duplicates in dense regions) to free up context budget for more informative regions, improving both efficiency and generalization.
-
Enable CPU-only, memory-constrained deployment of large tabular foundation models by replacing full-data inference with prototype-based inference. The improved system can run on standard hardware (e.g., Intel Core i7, 16 GB RAM, no GPU) with negligible performance loss, making it accessible to organizations without specialized compute resources.
-
Provide statistically validated performance guarantees via paired Wilcoxon tests with Holm correction. The improved system can report confidence in its context-construction gains over baselines, allowing users to trust that improvements are not due to chance, which is critical for scientific and clinical applications.
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks