Efficient Adaptive Data Acquisition via Pretrained Belief Representations
Daolang Huang, Zhuoyue Huang, Conor Hassan, Luigi Acerbi, Samuel Kaski, Tom Rainforth
cs.LG, stat.ML
Submitted: 2026-06-23
Comments: Preprint
Code: https://github.com/PriorLabs/TabPFN
License: http://creativecommons.org/licenses/by/4.0/
The gist: Learning effective policies for adaptive data acquisition remains challenging: posterior-based methods rely on surrogate models and posterior approximations that can be misspecified or biased, while
Terminology
Abstract
Learning effective policies for adaptive data acquisition remains challenging: posterior-based methods rely on surrogate models and posterior approximations that can be misspecified or biased, while direct policy-learning methods map from historical observations and fail to exploit available model representations, making learning harder. We introduce policy learning with belief representations (POLAR), based on the insight that optimal data acquisition depends on the observation history only through a sufficient belief state. Specifically, POLAR decouples representation learning from policy learning by leveraging pretrained predictive foundation models as belief-state encoders, training a policy head on top of their representations. This yields a simple, unified amortised policy learning framework for Bayesian experimental design, Bayesian optimisation, and active learning, differing only in the task-specific utility used to train the policy. Empirically, we find that POLAR outperforms state-of-the-art amortised methods across diverse tasks while requiring far fewer training samples, demonstrating a significant step in the scalability and efficiency of amortised data acquisition.
Sources
- Statistically Efficient Bayesian Sequential Experiment Design via Reinforcement Learning with Cross-Entropy Estimators
- JADAI: Jointly Amortizing Adaptive Design and Bayesian Inference
- Real-TabPFN: Improving Tabular Foundation Models via Continued Pre-training With Real-World Data
- TabPFN-2.5: Advancing the State of the Art in Tabular Foundation Models
- Constrained Bayesian Experimental Design via Online Planning
- Bayesian Active Learning for Classification and Preference Learning
- Sequential Bayesian optimal experimental design via approximate dynamic programming
- Amortized Safe Active Learning for Real-Time Data Acquisition: Pretrained Neural Policies From Simulated Nonparametric Functions
- None To Optima in Few Shots: Bayesian Optimization with MDP Priors
- Policy-Based Bayesian Experimental Design for Non-Differentiable Implicit Models
- Transformers Can Do Bayesian Inference
- TabICL: A Tabular Foundation Model for In-Context Learning on Large Data
- TabICLv2: A better, faster, scalable, and open tabular foundation model
- On Finetuning Tabular Foundation Models
- Reinforced In-Context Black-Box Optimization
- TabTune: A Unified Library for Inference and Fine-Tuning Tabular Foundation Models
- A Closer Look at TabPFN v2: Understanding Its Strengths and Extending Its Capabilities
- TabPFN: One Model to Rule Them All?
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks