Scaling Laws, Tabular Data and Actuarial Ratemaking Models
cs.LG, q-fin.RM
Submitted: 2026-09-02
Updated: 2026-09-02
Code: https://github.com/RonRichman/frmtpl-scaling-laws
Project page: https://dutangc.github.io/CASdatasets
License: http://creativecommons.org/licenses/by/4.0/
The gist: Scaling laws in modern deep learning describe how held-out loss improves as model capacity, training data, and compute increase, often following power-law trends.
Terminology
Abstract
Scaling laws in modern deep learning describe how held-out loss improves as model capacity, training data, and compute increase, often following power-law trends. We investigate whether analogous scaling regularities arise in actuarial ratemaking, where data are tabular, heterogeneous, and noisy, and where classical models such as GLMs remain strong baselines. Using a real-world motor insurance portfolio, we train models from different families across increasing fractions of the training data and multiple random seeds, evaluating out-of-sample Poisson deviance, a likelihood-based loss for Poisson count predictions in which lower values indicate better held-out fit. We find that all model families improve with additional data, but scaling exponents differ substantially: TabM exhibits markedly stronger data scaling than purely supervised tabular Transformers and standard MLP baselines. Transformer variants show weak parameter scaling unless augmented with additional inductive biases (TabM-style adaptation or self-supervision). These results provide quantitative guidance on model selection by data regime and suggest that effective scaling on actuarial tabular tasks depends on architecture and loss function objective design, with simple increases in Transformer size providing limited gains.
Sources
- TabNet: Attentive Interpretable Tabular Learning
- Broken Neural Scaling Laws
- A Hitchhiker's Guide to Scaling Law Estimation
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- TabArena: A Living Benchmark for Machine Learning on Tabular Data
- TabM: Advancing Tabular Deep Learning with Parameter-Efficient Ensembling
- Revisiting Deep Learning Models for Tabular Data
- Entity Embeddings of Categorical Variables
- A Theoretical Framework Bridging Model Validation and Loss Ratio in Insurance
- Deep Learning Scaling is Predictable, Empirically
- Training Compute-Optimal Large Language Models
- TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second
- Better by Default: Strong Pre-Tuned MLPs and Boosted Trees on Tabular Data
- TabTransformer: Tabular Data Modeling Using Contextual Embeddings
- Perceiver: General Perception with Iterative Attention
- Scaling Laws for Neural Language Models
- Adam: A Method for Stochastic Optimization
- Set Transformer: A Framework for Attention-based Permutation-Invariant Neural Networks
- Decoupled Weight Decay Regularization
- When Do Neural Nets Outperform Boosted Trees on Tabular Data?
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks