Bayesian Symbolic Regression with Entropic Reinforcement Learning
Oussama Boussif, Mohammed Mahfoud, Younesse Kaddar, Moksh Jain, Sida Li, Damiano Fornasiere, Xiaoyin Chen, Yoshua Bengio, Esmeralda S. Whitammer
Mila – Québec AI Institute · Université de Montréal · Independent · University of Oxford · University of Chicago · LawZero · CIFAR Fellow · University of Edinburgh
cs.LG
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: UAI 2026. Code available at https://github.com/jaggbow/ERRLESS
Code: https://github.com/jaggbow/ERRLESS
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: Bayesian Symbolic Regression with Entropic Reinforcement Learning Oussama Boussif, Mohammed Mahfoud, Xiaoyin Chen, Younesse Kaddar, Moksh Jain, Yoshua Bengio, Sida Li, Damiano Fornasiere, Esmeralda S.
Terminology
Summary
Bayesian Symbolic Regression with Entropic Reinforcement Learning
Oussama Boussif, Mohammed Mahfoud, Xiaoyin Chen, Younesse Kaddar, Moksh Jain, Yoshua Bengio, Sida Li, Damiano Fornasiere, Esmeralda S. Whitammer
Abstract
Symbolic regression is the problem of finding an algebraic expression describing a stochastic dependence of a target variable on a set of inputs. Unlike forms of regression that fit parameters assuming a fixed model structure, symbolic regression is a search problem over the space of expressions, represented, for example, as abstract syntax trees using a library of operators. Symbolic regression is typically used in settings with limited, noisy data in the natural sciences. However, searching for a single best-fitting expression fails to capture the epistemic uncertainty about the expression, which motivates a Bayesian perspective that enables uncertainty quantification and specification of natural priors to constrain the search space. In this work, we propose ERRLESS (Entropy-Regularized Reinforcement Learning for Expression Structure Sampling), a scalable approach for sampling from the posterior distribution over expressions given data using maximum-entropy reinforcement learning. ERRLESS learns a neural policy that constructs expressions sequentially by building up their abstract syntax trees. At convergence, the policy samples expressions from the posterior. At test time, expressions can be sampled by rollouts of this policy. We demonstrate that ERRLESS achieves competitive results on the Feynman benchmark while producing short and interpretable expressions. Additionally, we demonstrate that the mean of the posterior predictive approximated by ERRLESS achieves a higher coefficient of determination (R2) compared to an SMC baseline, which shows the value of the Bayesian perspective in symbolic regression.
Introduction
Symbolic regression (SR) is the problem of searching over a space of compositional algebraic expressions, using a certain library of primitive operators, to find a function that most closely maps the inputs observed in a dataset to their corresponding outputs. SR is a common problem in the natural sciences, where datasets are small and noisy, domain priors constrain plausible formulas, and interpretability is important. Most existing algorithms for SR have the goal of finding a Pareto front of expressions that trade off complexity with fit. With limited data, such a point estimate can be unreliable and hides the uncertainty about the expression arising from the noise and scarcity of data.
The need to model uncertainty in SR has been recognized in the literature and addressed by a Bayesian perspective on the problem. In this view, one posits a (structured) prior over expression structures and parameters and a model of observation noise. A dataset of (input, output) pairs then induces a posterior distribution over expressions, and the aim of Bayesian SR is to sample from this posterior. Previous Bayesian SR methods use reversible-jump Markov chain Monte Carlo (MCMC) or sequential Monte Carlo (SMC), but these methods rely on handcrafted proposal distributions and can be costly to scale. In this paper, we instead seek an approach that amortizes the sampling process using a neural network.
We formulate the construction of an expression as a sequential decision-making problem, turning the Bayesian SR task into the reinforcement learning (RL) problem of training a policy to sample expressions from the posterior. To achieve unbiased sampling at convergence, we train this policy using a maximum-entropy RL objective with the unnormalized posterior log-density as the reward (§3.2). The resulting system, which we call ERRLESS (Entropy-Regularized Reinforcement Learning for Expression Structure Sampling), is capable of learning a policy that samples expression structures and their real-valued parameters; once trained, the policy can be sampled to produce approximate posterior samples efficiently.
Successfully applying entropy-regularized RL to Bayesian SR requires careful design choices. ERRLESS uses a generation process that constructs expression syntax trees in a bottom-up (postorder) manner to allow effective imposition of constraints, imposes structural priors to avoid ill-formed, redundant, or unlikely subexpressions, and can incorporate additional constraints to restrict the search space to forbid dimensionally incompatible compositions of physical units (§2.1). The parametrization and training of the policy similarly require appropriate design of neural architectures and off-policy training schemes (§3.3). Finally, unlike previous Monte Carlo-based methods, ERRLESS amortizes sampling of both expression structures and parameter values into a single neural policy and avoids explicit per-candidate constant fitting.
On the Feynman Symbolic Regression Database, ERRLESS achieves a competitive prediction accuracy and produces posterior samples that improve predictions when data are few and noisy. On synthetic data, the posterior predictive mean of ERRLESS is better than PySIPS, an SMC baseline on the noisiest setting, showing the value of modeling uncertainty over expressions with a Bayesian view.
Contributions
-
We design a novel generation process (environment) for symbolic regression that enables bottom-up expression construction, incorporates dimensional analysis, and imposes structural constraints.
-
We formulate Bayesian symbolic regression, including inference of scalar parameters in expressions, as an end-to-end policy-learning problem within this environment.
-
We demonstrate that ERRLESS achieves competitive prediction accuracy compared to state-of-the-art approaches on the Feynman Symbolic Regression Database while producing short expressions.
-
We show that our approach captures the posterior distribution effectively, facilitating downstream applications in settings with scarce and noisy data.
Problem Setting
The space of expressions is defined via expression trees. A vocabulary Σ ≔ V ⊔ C ⊔ O is fixed, where V is a set of variable symbols, C is a set of constant symbols, and O is a set of operator symbols. Each operator has an arity and a semantic function. An expression tree T is a finite, rooted, ordered tree where each leaf is labeled with a symbol from V ⊔ C, and each internal node with r children is labeled with an operator of arity r. A constant assignment θ specifies numerical values for the constant symbols. A partial evaluation function f T,θ can be defined recursively.
An expression tree can be canonically represented as a sequence in reverse Polish notation by listing the labels of its nodes in postorder traversal. Unit constraints can be added by associating each variable with a physical unit vector and defining unit assignment functions for operators, which return the unit of the output if the operator can be applied to inputs with the given units, and are undefined otherwise.
For Bayesian symbolic regression, a prior probability distribution P(T) over valid expression trees and a conditional prior p(θ, σ T) over constant assignments and noise variance are fixed. Given a dataset D of input-output pairs, the posterior of (T, θ, σ) satisfies:
p(T, θ, σ D) ∝ P(T) p(θ, σ T) ∏ i=1 N N(y i; f T,θ(x i), σ2)
The objective of Bayesian symbolic regression is to sample from this posterior distribution.
Methodology: ERRLESS
Bottom-up generation. Most approaches use top-down generation, where internal nodes are sampled before leaf nodes. With top-down generation, the intermediate expression contains 'holes' to be filled, the expression can only be evaluated at the end, and enforcing physical unit constraints is inefficient. In contrast, bottom-up generation offers two main advantages: some intermediate states are valid and complete expressions, and because leaf nodes are specified from the outset, newly added operators can be constrained to be compatible with the known physical units of the operands.
Generating an expression tree is equivalent to generating its postorder sequence representation by appending one symbol at a time from left to right. A distribution π over trees is equivalent to a distribution over sequences with an autoregressive factorization:
π(w1...wn⊤) = ∏ i=1 n π(w i w<i) π(⊤ w 1...n)
Constraints on the number of nodes, and structural constraints that disallow ill-formed, redundant, or unlikely compositions, can be expressed as restrictions on the support of the next-token distributions. The constraints imposed prohibit: (i) composing functions with their inverses (e.g., √□2), (ii) nesting trigonometric functions (e.g., sin(cos(□))), (iii) nesting exponentials (e.g., e e□), and (iv) applying unary operators directly to constants.
Maximum-entropy RL training. A distribution π over tuples (T, θ, σ) is identified with a distribution over sequence representations and a conditional distribution over constant assignments and noise. The goal is to fit parameters φ so that π φ(T, θ, σ) equals the posterior p(T, θ, σ D).
Define the reward R(T, θ, σ) = log p(T, θ, σ) + ∑ i=1 N log N(y i; f T,θ(x i), σ2), so that the posterior is proportional to exp(R(T, θ, σ)). Training π φ to maximize its expected reward is equivalent to minimizing cross-entropy between π φ and the posterior, which is achieved by sampling the posterior mode. However, one can instead consider the entropy-regularized problem:
max φ E(T,θ,σ)∼π φ[R(T, θ, σ)] + H[π φ]
A key property of maximum-entropy RL is that the solution to this problem minimizes KL divergence between π φ and the distribution with density proportional to exp(R(T, θ, σ)), i.e., the posterior.
One off-policy objective is the trajectory balance (TB) objective, a core GFlowNet loss, which requires learning a scalar parameter log Z φ (the log normalizing constant). The objective associated with (T, θ, σ) is:
L TB(T, θ, σ; φ) = (log Z φ + log π φ(T, θ, σ) − R(T, θ, σ))2
Because this can be minimized to 0 for all samples simultaneously, training algorithms can optimize this objective over samples from a behavior policy that does not necessarily coincide with the current state of π φ itself.
Design choices. The policy is parameterized with a transformer-like architecture that encodes the sequence into an embedding supplied to three prediction heads: (i) a head producing logits over the next action, (ii) a head outputting the means, covariances, and mixture weights of a Gaussian mixture model for sampling constant values, and (iii) a head outputting the means, variances, and mixture weights of a mixture of log-normals for sampling the noise standard deviation σ.
A unigram prior over expression trees is chosen, based on the observation that the frequency of mathematical operators in physics equations decays exponentially with rank. Length constraints and redundancy constraints are imposed. A Gaussian prior with mean 0 and standard deviation of 10 is chosen for constant assignments, with a penalty for the number of constants used. The prior over σ is a wide half-Normal distribution with scale 2000 for the Feynman and Blackbox datasets, and a LogNormal distribution for the synthetic dataset.
To encourage exploration, off-policy training with ε-greedy exploration with annealed ε and a prioritized replay buffer is used. A modified version of TB that clips the gradients of the squared loss, analogous to the Huber loss, is used to stabilize training.
Related Works
ERRLESS differs from deep symbolic regression (DSR) and PhySO by taking a Bayesian perspective using maximum entropy RL to train the policy and using a bottom-up generative process. Unlike existing Bayesian SR methods (BSR using reversible-jump MCMC, PySIPS using SMC) that rely on handcrafted proposal distributions and computationally intensive MCMC sampling, ERRLESS learns an amortized posterior sampler with maximum entropy reinforcement learning. ERRLESS builds upon prior work on maximum entropy RL for sampling from discrete structured distributions, including the trajectory balance objective used for posterior inference over decision trees, causal models, and phylogenetic trees. ERRLESS deviates from GFN-SR by using a better generative process for constructing expressions, incorporating explicit priors, and evaluating on large-scale benchmarks complemented by improved training.
Experimental Setup
ERRLESS is compared against baselines from La Cava et al. and PhySO. All algorithms are allowed a budget of 1 million reward evaluations.
Datasets. A synthetic dataset of seven expressions is constructed to study the quality of posterior modeling, with 20 training points and 100 test points. The Feynman Symbolic Regression Database is used, comprising 100 expressions from the Feynman Lectures on Physics and 20 additional physics-inspired bonus expressions. Following Tenachi et al., 4 expressions involving arccos or arcsin are removed, leaving 116 expressions. Each expression is paired with 1 million sampled points; 10,000 are subsampled for training and 25,000 for testing. Gaussian noise is added to the training targets, with noise levels γ ∈ 0.001, 0.01, 0.1. The Blackbox benchmark is a set of 122 datasets from PMLB, using the more challenging subset of 12 datasets curated in Aldeia et al.
Metrics. The coefficient of determination R2 is used for prediction accuracy. The AUC Score, derived from the area under the curve of the success profile, is used to quantify robustness and predictive accuracy.
Results
Posterior over small expressions. On the synthetic datasets (noise level 0.1), ERRLESS is compared against the SMC-based baseline PySIPS. Drawing 1000 samples from the posterior for both methods, ERRLESS models the posterior over expression trees more accurately, as indicated by the posterior predictive mean (R2 PP). PySIPS attains lower NLL and, on most problems, higher single-expression Test R2, but its posterior-predictive mean collapses on most problems because a tail of posterior samples diverges outside the training region and dominates the mean. ERRLESS is less prone to this failure: its posterior draws yield a finite predictive mean on every problem.
Fit quality. ERRLESS achieves a competitive AUC score of 0.924 on the Feynman database, maintaining superior robustness across all noise levels compared to closely related reinforcement learning baselines such as DSR (0.873) and PhySO (0.893). ERRLESS also outperforms the Bayesian symbolic regression approach, BSR (0.702). The performance gap between ERRLESS and leading methods, such as GP-GOMEA and Operon, can be attributed to two key factors: (1) unlike most competing methods that optimize directly for reward, ERRLESS must approximate the joint posterior over expressions, constants, and noise, a comprehensive probabilistic task that typically necessitates a larger training budget than the 1M evaluations prescribed by the SRBench framework; and (2) while leading methods often achieve high fit scores, they frequently produce overly complex expressions containing hundreds of constants.
On the Blackbox benchmark, ERRLESS reaches an AUC of 0.350, which is higher than BSR (0.19) and AIFeynman (0.004) but lower than XGBoost (0.625). ERRLESS remains competitive in terms of striking a balance between the accuracy and complexity of the generated expressions. ERRLESS is an order of magnitude faster than most learning-based competitors, which can be attributed to the fact that ERRLESS does not require an inner loop to optimize the constants.
Qualitative analysis. Examining expressions discovered by the sampler, the highest-scoring candidate found by ERRLESS for the ground-truth expression ρ0/√(1 − c v) was ρ0/cos(v/c). Although these two expressions appear unrelated at first glance, their denominators have the same second-order Taylor approximation in v/c at 0. The length of the postorder representation of the ground-truth expression is greater than that of its approximate counterpart, which is unfavored by the prior. These results can be explained by the fact that ERRLESS incorporates a prior that favors concise expressions, as opposed to methods that only optimize the quality of fit. With more data samples, the likelihood would eventually dominate the prior in the reward, and the ground-truth expression would have a higher score.
Conclusion
ERRLESS is a scalable approach to Bayesian symbolic regression, using maximum-entropy reinforcement learning to amortize posterior sampling over algebraic expressions describing a stochastic dependence of a target variable on its inputs. Expression synthesis is formulated as a sequential decision-making process, and the search space is pruned by enforcing dimensional constraints during bottom-up construction. On the Feynman Symbolic Regression Database, ERRLESS achieves a competitive AUC score while approximating the full posterior distribution, rather than returning a single-point estimate.
Limitations and future work. While ERRLESS accurately models the posterior over expressions, performance can degrade on highly complex target expressions. Future work could condition the sampler on the temperatures of the log-likelihood and the priors to better trade off complexity and accuracy. Additionally, amortizing the sampler over datasets would allow posterior sampling for new datasets without the need for retraining. This naturally motivates extending ERRLESS to learn the operator library itself, allowing for reusable constructs that recur across datasets. Finally, extending the work to discovering symbolic forms of partial and ordinary differential equations would be a strong future direction given their prevalence in science.
Improvements for AI systems
Improvements to AI Systems Based on ERRLESS:
- Uncertainty-Aware Symbolic Regression for Scientific Discovery
-
Integrate ERRLESS's Bayesian posterior sampling into automated scientific discovery systems.
-
The improved system can generate multiple candidate expressions with quantified epistemic uncertainty, enabling researchers to identify robust formulas that generalize beyond noisy, small datasets.
-
It can flag regions of input space where predictions are uncertain, guiding targeted data collection.
- Amortized Posterior Inference for Structured Outputs
-
Extend ERRLESS's maximum-entropy RL framework to other structured prediction tasks (e.g., program synthesis, causal graph discovery, or neural architecture search).
-
The improved system can learn a single policy that amortizes sampling from posteriors over discrete structures (trees, graphs) and continuous parameters, avoiding expensive MCMC or SMC per new dataset.
-
It can provide fast, approximate Bayesian inference for real-time decision-making in robotics or autonomous systems.
- Physically Constrained Generation with Dimensional Analysis
-
Incorporate ERRLESS's bottom-up generation with unit constraints into AI systems for engineering design or physics-based modeling.
-
The improved system can automatically reject dimensionally inconsistent expressions during generation, reducing search space and producing physically plausible outputs.
-
It can be used in automated theorem proving or equation discovery for unknown physical laws, ensuring outputs respect fundamental units (e.g., mass, length, time).
- Prior-Driven Complexity Control
-
Use ERRLESS's unigram prior over operators and length penalties to enhance AI systems that generate interpretable models (e.g., in healthcare or finance).
-
The improved system can trade off fit vs. simplicity explicitly, producing short, human-understandable expressions that are less prone to overfitting.
-
It can be applied to auto-ML pipelines to generate parsimonious features or surrogate models for expensive simulations.
- Robust Posterior Predictive Mean for Noisy Data
-
Adopt ERRLESS's approach to posterior predictive averaging to improve AI systems in low-data regimes (e.g., drug discovery, materials science).
-
The improved system can compute a stable posterior predictive mean that avoids divergence from outlier samples, yielding more reliable predictions than single-point estimates or naive SMC.
-
It can be used for active learning, where the system selects new experiments based on predictive uncertainty to maximize information gain.
- Efficient Exploration via Off-Policy RL with Replay Buffers
-
Apply ERRLESS's training scheme (ε-greedy exploration, prioritized replay, gradient clipping) to other RL-based generative models.
-
The improved system can learn complex structured distributions faster and more stably, especially in sparse-reward environments.
-
It can be used in neural architecture search or molecule generation, where the search space is large and reward evaluation is expensive.
- Conditional Sampling for Multi-Task Generalization
-
Extend ERRLESS to condition on dataset embeddings or task metadata, enabling a single policy to sample posteriors for multiple related datasets without retraining.
-
The improved system can rapidly adapt to new scientific problems by leveraging prior knowledge from similar datasets, reducing the need for large training budgets.
-
It can serve as a foundation model for symbolic regression across domains, akin to a
language model for equations.
- Hybrid Symbolic-Numeric Optimization
-
Combine ERRLESS's structure sampling with local numerical optimization (e.g., gradient-based constant fitting) in a two-stage pipeline.
-
The improved system can first sample promising expression structures from the posterior, then refine constants for each candidate, achieving higher accuracy than pure sampling while retaining uncertainty quantification.
-
It can be used in engineering design optimization where both symbolic insight and precise numerical values are required.
- Hierarchical Expression Reuse
-
Implement ERRLESS's suggested future work of learning reusable sub-expressions (e.g., common physics motifs) across datasets.
-
The improved system can discover and cache frequently occurring substructures, accelerating search and improving interpretability by decomposing complex expressions into known components.
-
It can be applied to automated code generation or symbolic simplification tools.
- Extension to Differential Equations
-
Adapt ERRLESS's framework to sample from posteriors over ordinary/partial differential equations (ODEs/PDEs) given time-series data.
-
The improved system can discover governing equations from noisy observations, with uncertainty bounds on both the equation form and parameters.
-
It can be used in climate modeling, epidemiology, or neuroscience to infer dynamics from limited experimental data.
Abstract
Symbolic regression is the problem of finding an algebraic expression describing a stochastic dependence of a target variable on a set of inputs. Unlike forms of regression that fit parameters assuming a fixed model structure, symbolic regression is a search problem over the space of expressions, represented, for example, as abstract syntax trees using a library of operators. Symbolic regression is typically used in settings with limited, noisy data in the natural sciences. However, searching for a single best-fitting expression fails to capture the epistemic uncertainty about the expression, which motivates a Bayesian perspective that enables uncertainty quantification and specification of natural priors to constrain the search space. In this work, we propose ERRLESS (Entropy-Regularized Reinforcement Learning for Expression Structure Sampling), a scalable approach for sampling from the posterior distribution over expressions given data using maximum-entropy reinforcement learning. ERRLESS learns a neural policy that constructs expressions sequentially by building up their abstract syntax trees. At convergence, the policy samples expressions from the posterior. At test time, expressions can be sampled by rollouts of this policy. We demonstrate that ERRLESS achieves competitive results on the Feynman benchmark while producing short and interpretable expressions. Additionally, we demonstrate that the mean of the posterior predictive approximated by ERRLESS achieves a high coefficient of determination (R squared) compared to an SMC baseline, highlighting the benefits of the Bayesian perspective in symbolic regression.
Sources
- Interpretable Machine Learning for Science with PySR and SymbolicRegression.jl
- Bayesian Symbolic Regression
- GFN-SR: Symbolic Regression with Generative Flow Networks
- Learning Decision Trees as Amortized Structure Inference
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks