KKL Observer Synthesis for Nonlinear Systems via Physics-Informed Learning

arXiv:2501.11655 · eess.SY, cs.LG, cs.SY · Submitted 2026-08-11 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "KKL Observer Synthesis for Nonlinear Systems via Physics-Informed Learning".

Jane: The paper was written by M. Umar B. Niazi, John Cao, Matthieu Barreau and Karl H. Johansson from KTH Royal Institute of Technology and University of Oxford and Digital Futures.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back, everyone. I'm Tom, and joining me as always is the brilliant Jane. Today we're cracking open a fresh arXiv preprint that's got the control theory world buzzing: "KKL Observer Synthesis for Nonlinear Systems via Physics-Informed Learning."

Jane: Tom, I have to say, when I first saw that title, my eyes glazed over a little. KKL observers, physics-informed learning — it sounds like a mouthful. But the more I dig into it, the more I realize this is about something really fundamental. It's about how we can figure out what's happening inside a system when we can't measure everything directly.

Tom: Exactly. And that's the whole game, right? Think about a chemical reactor, or a self-driving car, or even the power grid. You've got sensors on the outside, but the internal state — the temperature inside the reactor, the exact position and velocity of the car, the voltage at a specific node — you can't always measure that directly. You need an observer.

Jane: Right. And the classic way to do this is with a Luenberger observer, which is this elegant trick for linear systems. You build a model of the system, you run it in parallel with the real thing, and you use the error between your model's output and the real sensor readings to correct your estimate. It works beautifully for linear systems.

Tom: But the real world isn't linear. It's nonlinear. And that's where this paper comes in. The KKL observer, named after Kazantzis, Kravaris, and Luenberger, is a way to extend that idea to nonlinear systems. The trick is to find a special coordinate transformation that makes the nonlinear system look linear, at least in a higher-dimensional space.

Jane: And finding that transformation is the hard part. It's the bottleneck. The paper describes it as solving a specific partial differential equation, and that's just not something you can do analytically for most real-world systems. It's a computational nightmare.

Tom: So what do these authors do? They say, forget solving the PDE by hand. Let's train a neural network to learn the transformation. And not just any neural network — a physics-informed one. They're embedding the actual governing equations of the system into the training process itself.

Jane: That's the part that gets me excited. It's not just a black box that's trying to fit data. It's a neural network that's being told, "Here's the physics, here's the math, and by the way, here's some data. Now go find a transformation that satisfies all of it." It's a much more constrained problem, which means it should generalize better.

Tom: And that's the promise. The authors show that this approach doesn't just work, it works better than the state-of-the-art methods, especially when you test it on states the network has never seen during training. We're talking about a huge leap in the practical ability to build observers for complex, nonlinear systems.

Jane: I love that. It's taking a beautiful piece of theory that's been around for decades and finally making it usable. This could have a massive impact on everything from robotics to process control. I can't wait to get into the details of how they actually pull this off.

Tom: Stay with us, because in the next segment, we're going to break down the core idea of the KKL observer itself and why this transformation is so crucial.

Paper summary: Jane: Welcome back. We're digging into "KKL Observer Synthesis for Nonlinear Systems via Physics-Informed Learning." Tom, let's give our listeners the big picture. What's the core problem this paper is solving?

Tom: So, imagine you have a system — let's say a pendulum. The state is its angle and its angular velocity. You can only measure the angle with a sensor. The observer's job is to estimate the angular velocity, which you can't measure. The KKL approach says, "Let's lift this problem into a higher-dimensional space where the dynamics become linear."

Jane: And that's the key insight. You create a new state, let's call it z, which is a function of the original state x. The dynamics of z are designed to be linear and stable. So, you can build a simple, linear observer for z, which is easy. Then, you just map that estimate of z back to the original coordinates to get your estimate of x.

Tom: Right. The catch is finding that map, T, that takes you from x to z. The paper calls it the transformation map. And it has to be injective, meaning it's a one-to-one mapping. If two different x's map to the same z, you can't tell them apart, and your observer is useless.

Jane: And this is where the math gets hairy. The map T has to satisfy a specific partial differential equation. For a simple pendulum, maybe you can solve it. But for a chaotic system like the Lorenz attractor, which the paper uses as a benchmark, forget it. There's no closed-form solution.

Tom: So, the authors' big idea is to learn T with a neural network. But here's the twist: they don't just train it on data. They train it to satisfy the PDE itself. They use something called a physics-informed neural network, or PINN. The loss function isn't just about matching data points; it's also about minimizing the residual of the PDE.

Jane: That's the "physics-informed" part. The network is forced to respect the underlying dynamics of the system, not just fit the training examples. That's what gives it the ability to generalize to states it hasn't seen before.

Tom: And they don't stop there. Once they've learned the forward map, they need its inverse to get the state estimate back. So, they train a second neural network to learn the inverse map. And here's a clever detail: they train them sequentially, not together.

Jane: Why is that important? Why not train them as one big autoencoder?

Tom: The paper argues that training them jointly creates conflicting gradients. The encoder wants to map x to z in a certain way, and the decoder wants to map z back to x in a certain way, and these objectives can fight each other, leading to a bad local minimum. By training the forward map first, and then freezing it, they give the inverse map a stable target to learn against.

Jane: That's a really practical insight. It's a subtle point, but it can make a huge difference in the accuracy of the final observer. So, the summary is: learn the forward map with a PINN, freeze it, then learn the inverse map with a standard network. That's the recipe.

Tom: Exactly. And the paper backs this up with theoretical guarantees, which we'll get into, and impressive simulations. But the short version is, they've found a way to make KKL observers practical for systems that were previously untouchable.

Jane: So, what are the actual guarantees? The paper promises more than just "it works in practice." We need to talk about the theory.

Tom: That's our next stop. We're going to look at the non-asymptotic learning guarantees and how they prove the observer is robust.

Improvements: Jane: Welcome back to the show. We're talking about "KKL Observer Synthesis for Nonlinear Systems via Physics-Informed Learning." Tom, we've covered the basics. Now, what does this paper actually improve upon compared to previous work?

Tom: Great question, Jane. The biggest improvement is in how they handle the training. Previous learning-based approaches, like the autoencoder method, tried to learn the forward and inverse maps simultaneously. It's like trying to learn two languages at once without a dictionary. The gradients from one task can mess up the other.

Jane: And the paper calls this out. They say the joint training leads to conflicting gradients and can get stuck in bad local minima. That's a real problem in practice.

Tom: Their fix is the sequential approach we mentioned. Learn the forward map first, using the physics-informed loss. Then, once that's solid, use it to generate data for the inverse map. This decouples the problem and makes each stage much easier to solve. It's a simple idea, but it's a big practical improvement.

Jane: And that's not the only improvement. There's also the issue of data generation. In previous methods, you had to simulate the system forward in time and hope that the transient effects of your initial conditions would die out. This paper's method is more robust because it uses the learned forward map to generate the training data for the inverse map, which is much more reliable.

Tom: Right. And the paper also provides something that a lot of learning-based methods lack: theoretical guarantees. They don't just say "trust us, it works." They prove that the approximation error of the learned inverse map is bounded, and they show how that error translates to a bound on the state estimation error.

Jane: So, they're not just improving the empirical performance; they're building a solid theoretical foundation. That's what separates a good paper from a great one. It's one thing to show it works on a few examples, but it's another to prove why it works and under what conditions.

Tom: And the conditions are pretty reasonable. They assume the system is forward complete, meaning the state stays bounded, and that it's backward distinguishable, which is a fancy way of saying you can tell different initial states apart by looking at their past outputs. These are standard assumptions in the KKL observer literature.

Jane: So, the improvements are threefold: a better training strategy, a more robust data generation process, and a rigorous theoretical framework. That's a pretty complete package.

Tom: It really is. And the simulations back it up. They test it on a bunch of benchmark systems, including chaotic ones like the Lorenz and Rössler attractors, and it consistently outperforms the state-of-the-art.

Jane: I'm eager to hear about those results. But before we get there, we should talk about the first page of the paper and the specific problem setup. Let's do that in the next segment.

Tom: Sounds good. We'll look at the formal problem statement and the assumptions they make.

First page: Jane: Welcome back. We're continuing our deep dive into "KKL Observer Synthesis for Nonlinear Systems via Physics-Informed Learning." Tom, let's look at the first page. What are the authors setting up?

Tom: The first page is all about setting the stage. They introduce the problem of state estimation for autonomous nonlinear systems. The system is described by a differential equation, and we have a sensor that gives us a partial measurement of the state. The goal is to design an observer that can reconstruct the full state from those partial measurements.

Jane: And they immediately point out the challenge. Finding the KKL transformation map is hard. It requires solving a PDE, and even if you can solve it, finding its inverse is also a nightmare. That's the core problem they're tackling.

Tom: Right. And they mention that existing techniques assume you know this map, which is a huge assumption. For most systems, you don't. So, they propose to learn it. The paper's title is "KKL Observer Synthesis for Nonlinear Systems via Physics-Informed Learning," and the abstract makes it clear they're using neural networks for this.

Jane: The abstract also highlights the two key contributions: the learning method itself and the theoretical guarantees. They're not just throwing a neural network at the problem and hoping for the best. They're providing bounds on the learning error and proving that the observer is robust to approximation errors and system uncertainties.

Tom: And they mention the applications. Observers are crucial for output feedback control, fault diagnosis, and digital twins. So, this isn't just an academic exercise. It has real-world implications for how we control and monitor complex systems.

Jane: Let's talk about that. A digital twin is a virtual replica of a physical system. To make it useful, you need to know the state of the physical system in real time. An observer like this could be the key to making that work for nonlinear systems.

Tom: Absolutely. And the paper also touches on the fact that full-state measurement is often impractical. You can't put a sensor on every component. So, you need observers to fill in the gaps. This paper makes that possible for a much wider class of systems.

Jane: The first page also sets up the notation and the structure of the paper. It's a well-organized paper, which is always a good sign. They're going to review the theoretical foundations, then the learning theory, then their method, and finally the guarantees.

Tom: And the simulations. Don't forget the simulations. They have a whole section dedicated to showing how well their method works on benchmark examples. That's where the rubber meets the road.

Jane: So, the first page is a promise. It's a promise that they have a new method that's both practical and theoretically sound. And the rest of the paper is about delivering on that promise.

Tom: And from what we've seen, they deliver. But we should bring in some other perspectives. Let's hear from Lu and Meng about the broader implications and the practical engineering challenges.

Jane: Good idea. Let's bring them in.

Conclusion: Tom: Alright, we're wrapping up our discussion on "KKL Observer Synthesis for Nonlinear Systems via Physics-Informed Learning." It's been a fantastic conversation. Jane, what's the final takeaway for our listeners?

Jane: The takeaway is that this paper has cracked a major bottleneck in nonlinear observer design. For decades, KKL observers have been a beautiful theory, but they were almost impossible to implement. This paper shows that by using physics-informed neural networks, you can learn the required transformation maps directly, making these observers practical for a wide range of systems.

Tom: And they did it with a clever sequential training strategy that avoids the pitfalls of joint training. Plus, they backed it all up with theoretical guarantees. That's a rare combination of practical insight and rigorous analysis.

Jane: The impact could be huge. Think about autonomous vehicles, robotic manipulators, chemical process control, even monitoring the power grid. Anywhere you need to estimate the internal state of a nonlinear system from limited sensor data, this method could be a game-changer.

Tom: And it's not just about the specific method. It's about the philosophy. This paper is a great example of how you can combine physics-based modeling with data-driven learning. The physics informs the learning, and the learning makes the physics usable. It's a powerful synergy.

Jane: Absolutely. We also heard from Lu and Meng about the broader implications. Lu pointed out the potential for digital twins and real-time monitoring, and Meng highlighted the practical challenges of implementation and the importance of the theoretical guarantees for safety-critical systems.

Tom: And Lalam, our in-house LLM, gave us a vision of how this could accelerate innovation across multiple fields. It's not just about one application; it's about providing a fundamental tool that can be used everywhere.

Jane: Well said. So, as we say goodbye to this paper, we're left with a sense of excitement. The tools we have for understanding and controlling nonlinear systems just got a whole lot more powerful.

Tom: And that's a great note to end on. Thanks to everyone for tuning in. We've been discussing "KKL Observer Synthesis for Nonlinear Systems via Physics-Informed Learning." Join us next time as we explore another exciting breakthrough in the world of research. Until then, keep asking questions.

Jane: And keep learning. Goodbye, everyone!

M. Umar B. Niazi, John Cao, Matthieu Barreau, Karl H. Johansson

KTH Royal Institute of Technology · University of Oxford · Digital Futures

eess.SY, cs.LG, cs.SY

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: 32 pages, 7 figures

Code: https://github.com/Mudhdhoo/PINN_KKL

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 66/100

Terminology

Summary

Summary

This paper proposes a novel learning-based approach for designing Kazantzis-Kravaris or nonlinear Luenberger (KKL) observers for autonomous nonlinear systems. The design of a KKL observer involves finding an injective map that transforms the system state into a higher-dimensional observer state, whose dynamics is linear and stable. The observer’s state is then mapped back to the original system coordinates via the inverse map to obtain the state estimate. However, finding this transformation and its inverse is quite challenging. The paper proposes learning the forward mapping using a physics-informed neural network, and then learning its inverse mapping with a conventional feedforward neural network.

The paper provides theoretical guarantees for the robustness of state estimation against approximation error and system uncertainties, including non-asymptotic learning guarantees that link approximation quality to finite sample sizes. The effectiveness of the proposed approach is demonstrated through numerical simulations on benchmark examples, showing better generalization capability outside the training domain compared to state-of-the-art methods.

The core of KKL observer design is an injective coordinate transformation that lifts the nonlinear system to a higher-dimensional space. The transformed system exhibits two key properties: it is bounded-input-bounded-state stable and linear up to output injection. The KKL observer operates by replicating the transformed system in the higher-dimensional space. To reconstruct the state estimate in the original coordinates, the left inverse of the transformation map is applied to the observer’s state. The transformation map’s injectivity property guarantees the accuracy of the estimate.

The transformation map required by a KKL observer is a solution to a specific partial differential equation (PDE). This PDE is computationally challenging to solve in practice, and finding the left inverse of the transformation map in real time is further challenging. The paper develops a framework for learning these maps using synthetic data generated from system dynamics and sensor measurements.

The paper makes several key contributions: it proposes a learning method that employs physics-informed neural networks (PINNs) to design KKL observers and establishes theoretical learning guarantees; it provides robustness guarantees for the learned KKL observer against approximation errors and system uncertainties; and through numerical simulations, it validates the effectiveness of the learning-based approach to KKL observer design and provides comparisons with state-of-the-art approaches.

The learning method is sequential. First, the transformation map is learned by a physics-informed neural network that incorporates the PDE associated with the KKL observer via automatic differentiation. Then, by generating new data and using the learned transformation map to produce corresponding features, a second neural network learns the inverse map by minimizing the empirical reconstruction risk. This sequential approach addresses two fundamental issues: the joint encoder-decoder structure is inherently susceptible to conflicting gradients between components of the objective function, which might result in the optimization getting stuck in bad local minima; and the data generation process is dependent on the truncation of the simulated trajectories to get rid of the transient effects due to the initialization of z-trajectories using the contraction property of the observer dynamics.

The paper provides non-asymptotic generalization bounds for the learned KKL observer. It proves that minimizing the empirical risk preserves the asymptotic robustness of the state estimation error to approximation errors and system uncertainties. The main theorem provides a high-probability bound on the true L2 approximation error of the learned inverse map, showing that it can be made small provided the empirical approximation errors and the estimation errors are small. The estimation errors vanish as the data sizes increase, while the approximation errors can be made small by using sufficiently rich network architectures.

The paper also provides robustness guarantees for the learned KKL observer under both approximation error and system uncertainties. It shows that the time-average steady-state estimation error is additively composed of a learning-dependent term and a design-dependent term. The learning-dependent term is the approximation error for the learned inverse map, which can be controlled via the learning process. The design-dependent term depends on the noise bounds and the observer’s design parameters, and can be minimized by choosing the observer matrix to be stable and well-conditioned, and by regularizing the learning of the inverse map to keep its Lipschitz constant small.

The effectiveness of the approach is validated through comprehensive state estimation experiments across multiple benchmark systems: the reverse Duffing oscillator, the Van der Pol oscillator, the Rössler attractor, and the Lorenz attractor. The paper demonstrates the generalization capability of the learned KKL observer by comparing it with state-of-the-art methods, including the neural ODE method, the deep model-free switching method, and the unsupervised autoencoder method. The results show that the proposed method consistently achieves the lowest symmetric mean absolute percentage error across all benchmarks, indicating better performance and generalization in both in-domain and out-of-domain validation. The better performance is because the PDE constraint serves as a physically consistent inductive bias that effectively reduces the hypothesis space for neural networks, enabling extrapolation to unseen data where purely data-driven methods fail.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems:

Improvement: Replace joint autoencoder training with a two-stage sequential learning pipeline:

  • Stage 1: Train forward map T̂θ using physics-informed loss (PDE residual + data fit)

  • Stage 2: Freeze T̂θ, generate new latent features, then train inverse map T̂η* separately

What the improved AI system can do: Avoids conflicting gradients that cause bad local minima in joint training. The inverse network sees a stationary input distribution, leading to more accurate state reconstruction. This is particularly valuable for AI systems that must learn invertible transformations (e.g., variational autoencoders, normalizing flows) where the encoder and decoder are trained jointly.

Abstract

This paper proposes a novel learning approach for designing Kazantzis-Kravaris or nonlinear Luenberger (KKL) observers for autonomous nonlinear systems. The design of a KKL observer involves finding an injective map that transforms the system state into a higher-dimensional observer state, whose dynamics is linear and stable. The observer's state is then mapped back to the original system coordinates via the inverse map to obtain the state estimate. However, finding this transformation and its inverse is quite challenging. We propose learning the forward mapping using a physics-informed neural network, and then learning its inverse mapping with a conventional feedforward neural network. Theoretical guarantees for the robustness of state estimation against approximation error and system uncertainties are provided, including non-asymptotic learning guarantees that link approximation quality to finite sample sizes. The effectiveness of the proposed approach is demonstrated through numerical simulations on benchmark examples, showing better generalization capability outside the training domain compared to state-of-the-art methods.

Sources

Related papers