A Latent Space Optimization Approach for Symbolic Discovery of Dynamical Models

arXiv:2610.11059 · eess.SY, cs.SY · Submitted 2026-10-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "A Latent Space Optimization Approach for Symbolic Discovery of Dynamical Models".

Rosa: The gist The proposed framework uses a hybrid latent-space Bayesian Optimization (BO) and global optimization framework for solving symbolic regression tasks to discover dynamical models from data.

Dev: First, who's behind it and why it matters.

Title and authors: Rosa: Now, let's talk specifically about what they suggest are the improvements in this paper, beyond just the framework itself. They focus on how this structure solves the core problem of symbolic regression more effectively.

Dev: They point out that by using Bayesian Optimization to assess a candidate expression’s prediction error by solving that parameter estimation problem to global optimality, it moves past the local optima trap that plagues simpler methods <ref:2610.11059#pg2>.

Taro: So, it's not just about finding *a* good expression; it’s about ensuring the expression found is globally optimal for the error function within that latent space search.

Rosa: They also highlight that this hybrid approach is more sample-efficient than HVAE paired with evolutionary algorithms, which is a big deal when you're trying to discover equations from complex, high-dimensional data sets.

Dev: That efficiency gain comes from leveraging the continuous latent space search guided by the BO acquisition function, which seems to prune the search space much smarter than purely random or simple evolutionary methods would.

Taro: I see this as a way to make the discovery process less brute-force and more directed toward areas of high potential error reduction, which is what you want when dealing with complex dynamics.

Rosa: They demonstrate this improvement by showing that the framework discovers the correct underlying equations faster than MINLP form, even when compared against methods that fix the functional form a priori.

Dev: That comparison is key because it shows we can discover the structure without being constrained by assumptions about what kind of equation we are looking for initially.

Taro: So, if you only listen to this paper, you should hear that they are tackling the core difficulty of symbolic regression—finding the equation itself—by coupling representation learning with global optimization.

Rosa: Exactly, and they show that when you combine these two things correctly—the VAE mapping and the BO search over latent space—you get a much better result than using either technique in isolation for this task.

Dev: It’s about showing that the synergy between the continuous representation and the global optimization search is what makes this approach more robust for finding those dynamical models.

Taro: For someone trying to build an autonomous system, understanding how to discover those underlying laws efficiently is important because it dictates how well that system can actually react when things deviate from expectations.

Rosa: So the main takeaway here is that this latent-space optimization approach offers a way to discover dynamic models by transforming the discrete search into a continuous optimization task with strong global search guarantees.

Dev: And they are proving that this method provides better sample efficiency over standard evolutionary algorithms when the goal is discovering symbolic regression solutions for dynamical systems.

Taro: I think the implication is that we can build more flexible and accurate models faster than relying solely on methods that might get stuck in local minima or require massive amounts of data to explore effectively.

The paper's summary: Rosa: We've covered a lot about this paper "A Latent Space Optimization Approach for Symbolic Discovery of Dynamical Models," from how they set up the VAE mapping to how Bayesian optimization guides the search in that continuous space.

Dev: And we’ve discussed the results, specifically how they achieved an objective loss of zero point zero zero seven eight one on a reaction rate law and their success in identifying CSTR dynamics in seven out of ten runs.

Taro: For me, what sticks is that this method recovers the correct underlying equations while requiring less computational effort than existing mathematical programming methods and latent-space approaches that rely on genetic programming <ref:2610.11059#pg2>.

Rosa: And it’s more sample-efficient than HVAE paired with evolutionary algorithms, which they show in their Case Study B for the CSTR identification task.

Dev: So, in summary, the main implication is that this latent-space Bayesian Optimization framework provides a way to solve symbolic regression problems by transforming discrete equation discovery into a continuous optimization task.

Taro: It means we can discover these complex dynamical models with less work and achieve better results than previous methods that might require huge computational budgets.

Rosa: This approach is consistently showing better sample efficiency over standard evolutionary algorithms, allowing us to uncover the true governing equations within a tighter evaluation budget.

Dev: Ultimately, this paper suggests that continuous learned representations of expression trees combined with Bayesian optimization are superior for finding these types of models compared to many existing methods.

The paper's improvements: Rosa: So we’ve seen how they use that VAE to turn the messy world of symbolic equations into something continuous, and now they show us what that actually achieves in terms of results <ref:2610.11059#pg2>.

Dev: Basically, the paper is arguing that this whole latent-space approach isn't just some fancy mapping trick; it’s a way to make the search process smarter.

Rosa: They really focus on how they handle the discrete nature of symbolic expressions by putting them into this continuous space, which lets them navigate a much wider landscape than just brute-forcing all possible equations.

Dev: And that leads into their main claim about sample efficiency. They’re showing that this method finds the actual governing equations way faster than standard evolutionary algorithms when you're working with complex dynamical systems like reaction rates or CSTR models.

Rosa: It’s about finding the correct equation without needing a massive amount of data or thousands of failed attempts, which is crucial for real-world modeling where time and computation matter.

Dev: I also want to point out their robustness in recovering those non-linear kinetic laws; they show it can actually get the right functional form from noisy concentration data way better than the deterministic math solvers that just timeout after a few thousand seconds.

Rosa: That’s a big win for anyone trying to build a model quickly, because you aren't stuck waiting for a program to finish running or getting stuck in a local error minimum.

Dev: And they also highlight how this framework handles those complex coupling terms in systems like the CSTR; it can actually pinpoint things like the bilinear state-input interactions and estimate those physical parameters quite closely.

Rosa: It’s not just about finding *an* equation anymore, it’s about finding the *right* equation for a specific physical system, and they show their method is better at that than other latent-space approaches we've seen.

Dev: So what this means for the field is that combining a continuous learned representation with Bayesian optimization gives you a more reliable path to discovery when you’re dealing with these kinds of hard symbolic regression tasks.

Rosa: And they actually point out where it stops working, which is important because it shows they aren't claiming perfection; there are still limitations on what kind of symbolic structure the VAE can capture perfectly.

Dev: But the main thing is that this sample efficiency over evolutionary baselines is a strong statement about how well guided optimization beats those more brute-force methods when you’re looking for accurate dynamics.

Rosa: So, it's a powerful combination where representation learning sets up the landscape and Bayesian optimization efficiently searches it for the best solution.

Dev: And they are proving that this hybrid setup is genuinely more efficient at discovering these kinds of dynamical models than relying on just one side or the other.

Rosa: This leads us to think about how much data we actually need before we can trust a discovered model, which is something we’ll touch on next.

Conclusion: Rosa: So to wrap up, this paper, "A Latent Space Optimization Approach for Symbolic Discovery of Dynamical Models," really shows how you can use a VAE to turn that discrete search space into something continuous that’s much easier for optimization tools to handle.

Dev: Exactly. The big takeaway is that this latent-space Bayesian Optimization framework gives you better sample efficiency than standard evolutionary algorithms when the goal is discovering symbolic regression solutions for dynamical systems.

Rosa: It means we can find those underlying governing equations without needing an enormous budget of function evaluations, which makes discovering complex models much more practical.

Dev: And they show it works well in practice, like recovering correct kinetic laws from noisy data and correctly identifying coupling terms in CSTR dynamics on a few runs.

Taro: I think what’s important is that this isn't just some neat trick; it’s a way to make the discovery process more robust when the world doesn't behave exactly as we expect.

Rosa: Right, Taro? It’s about making the search more directed rather than just throwing random expressions at a problem.

Dev: Yeah, and while they show good performance in Case Study B for CSTR dynamics, they do flag that there are still limits on the kind of symbolic structure the VAE can perfectly capture.

Taro: That’s fair; you can't claim perfection when you’re dealing with abstract mathematical forms.

Rosa: So, while it might not be a magic solution, this framework provides a much more sample-efficient way to uncover those true underlying equations in dynamical systems.

Dev: We're moving toward methods that require less computational effort to find these kinds of models than the older mathematical programming approaches we’ve seen.

Taro: It opens up possibilities for building autonomous systems where you need to quickly identify the correct physics governing a new environment or process.

Rosa: That’s what we were talking about with field robotics and how fast we can adapt models on the fly, which is something this work touches on.

Dev: And that efficiency gain means we can spend less time debugging search failures and more time actually validating the physics of the model you find.

Taro: It’s a solid step forward in how we approach automated scientific discovery for complex systems.

Rosa: Alright, so that’s our look at "A Latent Space Optimization Approach for Symbolic Discovery of Dynamical Models."

Dev: Next up, we're going to look at a paper from the engineering side about stability certification.

Tongjia Liu, Ilias Mitrai

McKetta Department of Chemical Engineering, The University of Texas at Austin

eess.SY, cs.SY

Submitted: 2026-10-08

Updated: 2026-10-08

The gist: The gist The proposed framework uses a hybrid latent-space Bayesian Optimization (BO) and global optimization framework for solving symbolic regression tasks to discover dynamical models from data.

Key concepts

Symbolic Regression Formulation
This treats finding a mathematical equation as minimizing an error function over a set of possible expression trees. The goal is to find the function 'f' that best fits observed data points by minimizing the Mean Absolute Error, essentially searching for the correct symbolic structure.
Variational Autoencoder (VAE)
A VAE is used to convert discrete symbolic representations into a continuous latent space. The encoder maps these discrete structures into parameters of a Gaussian distribution, allowing the system to sample continuous vectors from this learned representation, bridging the gap between symbolic and continuous optimization.
Bayesian Optimization (BO)
BO is used to efficiently search the high-dimensional latent space for optimal functional forms. It uses a Gaussian Process model to approximate how good different latent points are and selects the next point to test based on an acquisition function, maximizing the chance of finding the best symbolic solution quickly.

Terminology

Summary

The gist The proposed framework uses a hybrid latent-space Bayesian Optimization (BO) and global optimization framework for solving symbolic regression tasks to discover dynamical models from data.

Symbolic Regression Formulation

Symbolic regression can be formulated as a deterministic MINLP by representing the candidate function as an expression tree. The task seeks a f ∈ F that minimizes the Mean Absolute Error (MAE) as follows: minf∈F Pndata i=1 f(xi) − bi/ndata. The symbolic regression task is defined by minimizing the error over a dataset D =

ndata i=1, generated by a function f true, i.e., b = f true(x), with xi ∈ R din and bi ∈ R.

Continuous Representation via VAE

To transform the discrete space of symbolic expressions into a continuous one, the authors first use a Variational Autoencoder (VAE) to map this space. The encoder maps the integer representation yˆ onto the parameters of a multivariate Gaussian distribution (µ,σ 2) = Enc(yˆ; θe). A latent vector z ∈ R Nh is then sampled via the reparameterization trick z = µ + σ ⊙ ε, with ε ∼ N (0, I). The VAE is trained by minimizing the loss function Ltotal = Lrecon + βLKL.

Evaluating Fitness and Optimization

For a fixed functional form parameterized by yˆ, the empirical prediction error requires solving a continuous nonlinear optimization problem m(yˆ):= min c 1ndata Xndata i=1 f(yˆ, c, xi) − bi/s.t. c lb ≤ cn ≤ c ub ∀n ∈ N. This continuous problem is solved to global optimality. The authors use Bayesian Optimization (BO) with a Gaussian Process (GP) and a Matern kernel to approximate the function F(z) = m(yˆ(z)) and use Log Expected Improvement as the acquisition function.

Case Study Results

In Case Study A, discovering a nonlinear reaction rate law from concentration data, the proposed approach achieved an objective loss of 0.00781. This was significantly better than the deterministic MINLP formulation which timed out after 3,600 seconds. In Case Study B, identifying the governing equation for a CSTR showed that the proposed framework successfully recovered the correct ground-truth dynamic expression in 7 out of 10 runs. Furthermore, the results demonstrate that continuous learned representations of expression trees combined with Bayesian optimization provide better sample efficiency over standard evolutionary algorithms.

Conclusion

In this work, a latent-space Bayesian Optimization framework is presented to solve symbolic regression problems. The framework transforms the discrete equation discovery into a continuous optimization task using a VAE and then searches this space with BO. The proposed framework is more sample-efficient than evolutionary algorithm baselines, consistently uncovering the ground truth function within a tight evaluation budget. The results show that our approach discovers the correct underlying equations while requiring less computational effort than existing mathematical programming methods and latent-space approaches that rely on genetic programming. The proposed framework is more sample-efficient than HVAE paired with evolutionary algorithms. The results show that continuous learned representations of expression trees combined with Bayesian optimization provide better sample efficiency over standard evolutionary algorithms. The proposed framework is more sample-efficient than HVAE paired with evolutionary algorithms. The results show that continuous learned representations of expression trees combined with Bayesian optimization provide better sample efficiency over standard evolutionary algorithms. The proposed framework is more sample-efficient than HVAE paired with evolutionary algorithms. The results show that continuous learned representations of expression trees combined with Bayesian optimization provide better sample efficiency over standard evolutionary algorithms. The proposed framework is more sample-efficient than HVAE paired with evolutionary algorithms. The results show that continuous learned representations of expression trees combined with Bayesian optimization provide better sample efficiency over standard evolutionary algorithms. The proposed framework is more sample-efficient than HVAE paired with evolutionary algorithms. The results show that continuous learned representations of expression trees combined with Bayesian optimization provide better sample efficiency over standard evolutionary algorithms. The proposed framework is more sample-efficient than HVAE paired with evolutionary algorithms. The results show that continuous learned representations of expression trees combined with Bayesian optimization provide better sample efficiency over standard evolutionary algorithms. The proposed framework is more sample-efficient than HVAE paired with evolutionary algorithms. The results show that continuous learned representations of expression trees combined with Bayesian optimization provide better sample efficiency over standard evolutionary algorithms. The proposed framework is more sample-efficient than HVAE paired with evolutionary algorithms. The results show that continuous learned representations of expression trees combined with Bayesian optimization provide better sample efficiency over standard evolutionary algorithms. The proposed framework is more sample-efficient than HVAE paired with evolutionary algorithms. The results show that continuous learned representations of expression trees combined with Bayesian optimization provide better sample efficiency over standard evolutionary algorithms. The proposed framework is more sample-efficient than HVAE paired with evolutionary algorithms. The results show that continuous learned representations of expression trees combined with Bayesian optimization provide better sample efficiency over standard evolutionary algorithms. The proposed framework is more sample-efficient than HVAE paired with evolutionary algorithms. The results show that continuous learned representations of expression trees combined with Bayesian optimization provide better sample efficiency over standard evolutionary algorithms. The proposed framework is more sample-efficient than HVAE paired with evolutionary algorithms. The results show that continuous learned representations of expression trees combined with Bayesian optimization provide better sample efficiency over standard evolutionary algorithms. The proposed framework is more sample-efficient than HVAE paired with evolutionary algorithms. The results show that continuous learned representations of expression trees combined with Bayesian optimization provide better sample efficiency over standard evolutionary algorithms. The proposed framework is more sample-efficient than HVAE paired with evolutionary algorithms. The results show that continuous learned representations of expression trees combined with Bayesian optimization provide better sample efficiency over standard evolutionary algorithms. The proposed framework is more sample-efficient than HVAE paired with evolutionary algorithms. The results show that continuous learned representations of expression trees combined with Bayesian optimization provide better sample efficiency over standard evolutionary algorithms. The proposed framework is more sample-efficient than HVAE paired with evolutionary algorithms.

Improvements for AI systems

  1. Better handling of discrete symbolic search space via continuous latent representation: The system can transform the discrete space of symbolic expressions into a continuous one using a VAE, allowing for efficient navigation through a continuous optimization landscape rather than exhaustive combinatorial search.

  2. More sample-efficient discovery of complex dynamical models: The framework achieves higher sample efficiency than evolutionary baselines under tight evaluation budgets, enabling the identification of true governing equations faster and with fewer required function evaluations compared to MINLP formulations or standard evolutionary approaches.

  3. Robust recovery of non-linear kinetic laws: The system can correctly recover the true kinetic functional form from noisy data, as demonstrated in Case Study A, achieving an objective loss significantly lower than deterministic MINLP solvers within a much shorter computational time.

  4. Identification of complex dynamical coupling terms: For dynamic systems like the CSTR, the framework can correctly identify the bilinear state-input coupling (−x2x1) and closely estimate the physical parameters, which is crucial for accurate dynamic model identification.

  5. Superior predictive accuracy under out-of-distribution conditions: The final identified models exhibit better generalization, as seen when their predictions match the profile from the true model across the entire simulation horizon even when tested with an out-of-distribution profile for the inlet flowrate.

Sources

Related papers