Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs

arXiv:2206.14284 · stat.ML, cs.LG, cs.NA, math.NA, math.PR · Submitted 2022-06-28 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs".

Jane: The paper was written by Florian Krach, Marc Nübel and Josef Teichmann from Department of Mathematics, ETH Zurich, Switzerland.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Summary: Tom: Okay, so building on the idea of capturing these non-smooth dynamics, the paper summarizes how they actually approach this estimation process using "Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs."

Jane: If Segment one was about *what* they are modeling—those tricky paths—Segment two is getting into *how* they build the actual machinery to do that estimation, which seems quite sophisticated.

Tom: The summary really emphasizes using these specialized ODEs to estimate the underlying dynamics parameters directly from observed data, rather than needing explicit physical equations upfront.

Jane: That’s the breakthrough, isn't it? Instead of telling the model, "Gravity pulls like this," they are letting it learn the gravitational pull *from* a bunch of trajectories.

Lu: They are essentially performing an inverse problem: taking observations and deriving the governing law that must have produced those observations in the first place.

Meng: This sounds like a massive leap over traditional state-space models because you aren't constrained by assuming Gaussian noise or linearity in any way; you're learning the entire manifold of possibilities.

Lalam: Considering how much data we generate today, the ability to distill fundamental physical or behavioral laws from messy, high-dimensional sensor streams is what will fundamentally change how knowledge work operates.

Tom: And they achieve this by formulating an optimal estimation framework around these jump ODEs, which sounds mathematically intense but promises high accuracy in parameter recovery.

Jane: It’s like building a perfect digital twin of a process just by watching it run, and that's something I find incredibly reassuring for scientific modeling.

Lu: The 'optimal' part suggests they are minimizing some sort of overall error metric across the entire observed path, ensuring the model isn't just fitting one point but the whole story.

Meng: When I look at this practically, I wonder about initialization; if your starting guess for the parameters is wildly off, does the entire estimation process diverge before it can find that true optimum?

Lalam: What this means for culture is that our reliance on pre-existing, potentially flawed theoretical models might diminish. The data itself becomes the primary source of truth, guided by advanced AI inference.

Jane: So, to wrap up this segment: they are building a powerful estimation tool that uses the whole data history to figure out the underlying rules governing a system's evolution.

Tom: That sets us up perfectly for looking at how they actually tested this theory out in the next section, right? Let's see how good these estimators really are when compared against other methods.

Improvements: Tom: Moving into the improvements, the paper details

Paper discussion segment 3: Tom: So, we've seen how this Path-Dependent NJ-ODE framework is a massive generalization of the original model, moving far beyond simple, continuous processes. The real headline here is that we can now apply this tool to almost any kind of chaotic or messy system.

Jane: That’s right; it isn't just for smooth financial data anymore. The ability to handle non-Markovian or even discontinuous paths means the model can actually capture complex history, not just the last known point in time.

Lu: And that complexity is where the signature transform comes in; it allows us to encode the entire path's evolution into a fixed-size vector, which is a huge leap over simply having memory of past observations.

Meng: From an engineering standpoint, this robustness is critical; if your model doesn't assume linearity or simple Gaussian noise, you’re basically opening up the door to modeling real-world industrial or ecological systems that are far more complicated.

Lalam: I think the cultural shift here is profound; we’re moving away from models that believe in a single, clean path and toward an AI that sees the entire history of every possible paths taken by a system.

Tom: Exactly, Lu's point about complexity is key to this whole generalization. We aren't just predicting the next moment; we are finding the best possible mathematical description of how the system *should* have evolved given all past information.

Jane: It’s like building a perfect digital twin, but instead of assuming it’s always running smoothly, you let it simulate a bumpy reality that has jumps and sudden changes.

Meng: And I see this making the model far more reliable in high-frequency trading or climate modeling where assumptions about simple continuity break down all the time.

Lu: The theoretical proof shows that this optimal estimation framework converges to the true conditional expectation, which is mathematically stunning given how much we've relaxed those initial constraints.

Lalam: It suggests that our reliance on pre-existing, simplified theoretical blueprints might start to fade as data itself becomes the primary source of truth for the systemic behavior.

Tom: And finally, with this power comes a huge amount of hope for tackling real-world problems that are simply too messy or unpredictable to fit into traditional models.

Meng: We definitely need to start thinking about how our deployment pipelines will handle this level of complexity.

Conclusion: Tom: So, looking back at everything we covered today, it really struck me how much of a leap this research represents for how we model complex systems that change over time.

Jane: It’s incredible; they managed to combine these advanced concepts—the Neural ODEs and the jump processes—into one framework that's genuinely robust.

Lu: And what I keep thinking about is how powerful this makes the general theory of dynamic modeling; it moves us away from needing specific, restrictive equations and toward something truly universal.

Meng: From an engineering standpoint, that universality is great, but I’m wondering how much computational overhead this whole path-dependent jumping mechanism introduces when you scale it up to millions of data points.

Lalam: But even if the computation is heavy now, think about the kind of systems we can model—financial markets, biological processes—that were previously too messy or too abstract to predict accurately.

Tom: Exactly! It’s not just about getting a slightly better number on a test set; it's about opening up entire domains of prediction that were previously considered intractable.

Jane: I think the core message is that by treating the dynamics as continuous evolution punctuated by discrete, informative jumps, they've created an estimation method that is incredibly flexible and remarkably powerful.

Lu: This whole structure really highlights the power of combining deep learning with physical domain knowledge in a way that wasn't possible just a few years ago.

Meng: I agree with Lu; the fact that they were able to show superior performance compared to simpler, direct approximation models is huge evidence for this combined approach working optimally.

Lalam: Ultimately, what this means for culture is better decision-making across the board—allowing us to build AI systems that don't just predict correlations but truly understand dynamic processes.

Tom: It makes you excited about the future of predictive AI, Jane; it’s a genuinely exciting piece of work by the authors.

Jane: Absolutely, it gives us new tools for understanding reality itself, which is always thrilling to discuss on air.

Lu: I really hope that this opens up research into entirely new classes of stochastic processes that we haven't even thought of yet.

Meng: I’m already picturing the necessary updates to our infrastructure; we need to start building pipelines for path-dependent dynamics right away.

Lalam: If we can model these "generic dynamics," it means AI can help humanity navigate uncertainty with unprecedented grace.

Tom: Well, this has been such a fantastic deep dive into "Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs."

Jane: We feel like we could talk about this for hours, but we have to wrap up and get ready for our next topic.

Lu: I bet the next paper is going to push the boundaries even further, which keeps us excited.

Meng: Let's see what practical challenges that next paper throws at us!

Lalam: We're looking forward to seeing how another breakthrough can improve our collective understanding of the world.

Florian Krach, Marc Nübel, Josef Teichmann

Department of Mathematics, ETH Zurich, Switzerland · ETH Zurich University of Technology Zurich (ETH Zürich)

stat.ML, cs.LG, cs.NA, math.NA, math.PR

Submitted: 2022-06-28

Updated: 2026-08-25

Journal ref: Florian Krach, Marc Nübel, Josef Teichmann, Optimal estimation of generic dynamics by path-dependent neural jump ODEs, Stochastic Processes and their Applications, Volume 201, 2026, 105058, ISSN 0304-4149

DOI: 10.1016/j.spa.2026.105058

Code: https://github.com/FlorianKrach/PD-NJODE

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 95/100

The gist: The paper introduces a sophisticated framework for "Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs," addressing how complex, time-evolving systems can be modeled when both

Key concepts

Path-Dependent Neural Jump ODEs
A specialized framework used to estimate underlying system dynamics. These ODEs are 'path-dependent' and incorporate 'jumps,' allowing the model to capture complex, non-smooth, or discontinuous changes in a system's evolution.
Optimal Estimation Framework
The mathematical structure used to ensure high accuracy in parameter recovery. It involves minimizing an overall error metric across the entire observed data path, rather than just fitting single points or assuming simple noise distributions.
Inverse Problem
In this context, it means taking observational data (trajectories) and deriving the governing law or fundamental rules that must have created those observations in the first place. The model learns the physics from the data.
Non-Markovian Dynamics
A type of system behavior where the future state depends not only on the current state but also on its entire history. This ability allows models to capture complex memory effects beyond just looking at the last known point in time.

Terminology

Summary

The paper introduces a sophisticated framework for Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs, addressing how complex, time-evolving systems can be modeled when both continuous changes and discrete informational jumps occur. This methodology is critical because it allows researchers to decompose the learning problem into two intertwined sub-problems: learning the continuous evolution between observations and learning the instantaneous updates (jumps) when new information arrives.

Model Architecture and Implementation Details

The core model employed is the PD-NJ-ODE, which utilizes specific architectural choices for its components. For this architecture, the hidden size is d H = 100, and all neural networks maintain a consistent structure of 1 hidden layer with tanh activation function and 50 nodes. The signature input is utilized up to a truncation level of 2, but the model explicitly omits a recurrent jump network. The final output layer comprises a classifier network, which maps the last latent variable H tn to class probabilities (p inc, p stat, p dec) using a softmax activation, ensuring outputs are in [0, 1] cubed. In additional retraining efforts for the classifier, researchers tested various complex structures:

  • 1 hidden layer with tanh activation function and 50 nodes (matching combined training).

  • 2 hidden layers with tanh activation functions and 200 nodes.

  • 4 hidden layers with tanh activation functions and 200 nodes.

Data Preprocessing and Model Specifics

Different models require distinct preprocessing steps to ensure accurate learning. For the DeepLOB model, the dataset is standardized by z-scores as it was done in Zhang et al. (2019). Conversely, the PD-NJ-ODE model mandates that the dataset is not normalized, but instead, each sample must be shifted so that it begins at X 0 = 0 and the time starts at t 0 = 0. Furthermore, for the PD-NJ-ODE input stream, the methodology restricts inputs to only the midprice and the bid/ask prices up to level 10, excluding volume data.

Training Regimen

The training process is sequential and multi-staged. Initially, the model is first trained for 50 epochs, minimizing a combined loss function: the sum (without weighting) of the PD-NJ-ODE loss and the cross-entropy loss of the classifier, using the Adam optimizer with a batch size of 50 and a learning rate of 0.01. While training aims to forecast all inputs, in the evaluation of the MSE only the midprice is considered. Following this, the classifier is retrained for 1000 epochs with the cross-entropy loss alone, using an Adam optimizer with a batch size of 50 and a learning rate of 0.001.

Theoretical Comparison to Direct Approximation

The approach gains theoretical strength by leveraging domain knowledge, allowing it to split the problem into continuous evolution and discrete jumps. This is contrasted with directly approximating the conditional expectation F j using a neural network (referred to as the NJ-model). The authors demonstrate that this domain knowledge provides a significant practical advantage: the PD-NJ-ODE achieves a minimal evaluation metric of 5 times 10-4 on the test set, [while] the one of the NJ-model is 1 times 10-2, i.e., larger by a factor of 20. Even when testing larger architectures for the NJ-model (up to 42K trainable parameters), this performance gap persists.

Improvements for AI systems

System Improvement Focus Area 1: Hybrid Stochastic Process Modeling for High-Frequency Time Series Forecasting (Generalization of PD-NJ-ODE)

  • Improvement: Develop a generalized, adaptive architecture that seamlessly integrates Neural Ordinary Differential Equations (NODE) with discrete jump processes. Instead of assuming fixed input/output spaces or relying on pre-defined domain knowledge (like splitting learning into continuous evolution and discrete jumps), the system should dynamically determine the optimal balance between these two modes based on input entropy and local volatility metrics.

  • Implementation Details:

  1. Adaptive Switching Mechanism: Implement a meta-network (e.g., a small attention mechanism or a trained threshold function) that takes as input the current time step's raw feature vector (X t) and predicts the probability of the next change being dominated by continuous drift vs. discrete jump.

  2. Unified Latent Space: Design the NODE component to operate within a shared, high-dimensional latent space (H t). When a jump is predicted, the system should not just apply an additive shock; it must utilize a learned transformation matrix (similar to an exponential map) that projects the jump magnitude into the continuous latent space H t+, ensuring continuity and physical plausibility.

  3. Input Robustness: Generalize the input features beyond simple price levels (midprice, bid/ask up to L10). Incorporate derived microstructure features such as order book imbalance moments, liquidity decay rates, and realized volatility across different time scales (e.g., 1-second vs. 5-second windows) directly into the X t vector.

  • System Capability: The resulting system can forecast complex, non-stationary financial time series (like order book dynamics) by simultaneously modeling underlying continuous market trends (drift) and sudden, information-driven regime shifts (jumps), significantly improving predictive accuracy over models that treat these two phenomena separately.

System Improvement Focus Area 2: Interpretable Feature Extraction for Low-Dimensional Latent Dynamics

  • Improvement: Refine the architecture of the latent variable representation (H t) by integrating interpretability constraints, moving beyond simple black box latent spaces. The goal is to force the latent dimensions to correspond to known economic or physical factors (e.g., momentum, mean reversion pressure, volatility).

  • Implementation Details:

  1. Factor-Constrained Latent Space: After training the NODE component, apply a regularization penalty (L factor) on the latent variable H t. This penalty should encourage specific linear combinations of H t to correlate highly with pre-defined factor models (e.g., Fama-French factors or microstructure imbalance measures).

  2. Attention/Saliency Mapping: Integrate an attention mechanism within the NODE's structure that highlights which input features (X t k) are most responsible for the change in specific dimensions of H t. This provides a direct, quantifiable measure of feature importance at every time step.

  3. Orthogonalization: Implement an orthogonal constraint on the latent space updates to ensure that different dimensions of H t learn independent, non-redundant information, preventing the model from collapsing critical dynamics into a single dimension.

  • System Capability: The improved system provides not only a prediction (e.g., price movement or classification) but also a highly detailed causal attribution map. Users can understand why the model predicts an outcome by identifying which specific market factors (liquidity changes, imbalance shifts, etc.) are driving the change in the latent state.

System Improvement Focus Area 3: Multi-Objective, Multi-Task Training Framework for Robustness

  • Improvement: Overhaul the training regime to move away from sequential or single-target optimization. The system must be trained simultaneously on multiple, disparate tasks using a weighted, dynamically adjusted loss function.

  • Implementation Details:

  1. Joint Optimization Objective: Instead of minimizing Loss NODE + lambda times Loss Classifier, implement a generalized objective:

L Total = sum i=1 N w i(t) times L i(t, Y t)

where L i are losses for N tasks (e.g., predicting midprice movement, classifying regime change, forecasting volume spread), and the weights w i(t) are not static hyperparameters but are themselves predicted by a small auxiliary network based on the current state H t.

  1. Curriculum Learning Scheduling: Implement a dynamic curriculum scheduler during training. Start training with low-stakes, stable tasks (e.g., basic pattern recognition) and gradually increase the weight and complexity of high-stakes tasks (e.g., predicting maximum jump magnitude or extreme volatility) as the model's performance stabilizes on earlier tasks.

  2. Adversarial Training: Incorporate a Discriminator network trained to distinguish between real data samples and samples generated by the model's predicted dynamics (H t). This forces the generative part of the system to learn highly realistic, robust distributions, improving generalization beyond merely matching historical averages.

  • System Capability: The resulting AI model is significantly more robust and generalizable. By forcing it to perform multiple related tasks simultaneously and adaptively weighting their importance during training, it learns a deeper, more holistic representation of the underlying system dynamics that is less prone to overfitting on single target metrics.

Abstract

This paper studies the problem of forecasting general stochastic processes using a path-dependent extension of the Neural Jump ODE (NJ-ODE) framework. While NJ-ODE was the first framework to establish convergence guarantees for the prediction of irregularly observed time series, these results were limited to data stemming from Itô-diffusions with complete observations, in particular Markov processes, where all coordinates are observed simultaneously. In this work, we generalise these results to generic, possibly non-Markovian or discontinuous, stochastic processes with incomplete observations, by utilising the reconstruction properties of the signature transform. These theoretical results are supported by empirical studies, where it is shown that the path-dependent NJ-ODE outperforms the original NJ-ODE framework in the case of non-Markovian data. Moreover, we show that PD-NJ-ODE can be applied successfully to classical stochastic filtering problems and to limit order book (LOB) data.

Related papers