Extending SSMs with the Exponentially Weighted Signature
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Extending SSMs with the Exponentially Weighted Signature".
Jane: The Exponentially Weighted Signature (EWS) generalizes classical path signatures by replacing uniform historical weighting with bounded linear operators, offering richer temporal dynamics and cross-channel coupling.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So Jane, we're talking about this paper today titled "Extending SSMs with the Exponentially Weighted Signature." Essentially, they are taking the classical path signature and upgrading it to something much more flexible. The main claim is that this new Exponentially Weighted Signature allows for temporal weighting using bounded linear operators instead of just uniform weighting.
Jane: Exactly, Tom; what this means is that we can now model the relevance of past information in a way that isn't just a simple fade-out, but something more dynamic and structured. The authors introduce an operator pair, A = (A, B), which dictates how that weighting happens across different channels at the same time point.
Lu: From my perspective as a researcher at Tsinghua, this is really interesting because it moves us beyond what we can achieve with purely diagonal operators; it unlocks oscillatory and growth behaviors in memory dynamics that were previously inaccessible.
Meng: That sounds conceptually rich, Lu, but from an engineering standpoint, how does this move from the abstract operator stuff to something we can actually implement in a state-space model? I need to know what the practical constraints are here.
Lalam: From my processing standpoint, the structure of these operators suggests a huge potential for improving how we build cultural representations in large language models by giving them richer temporal context.
Tom: Right, so it’s not just about adding more terms; it's about fundamentally changing the mechanism for capturing history. The paper claims this EWS is the unique solution to a linear controlled differential equation on the tensor algebra, which gives it a solid mathematical foundation.
Jane: That uniqueness is important because it means we have a principled way to define this object that isn't just an arbitrary construction; it’s derived from this specific differential equation: dSA(X)s,t = −ΛASA(X)s,t dθt + SA(X)s,t ⊗ dXt.
Lu: The differential equation formulation is what really sells the theoretical power here; it connects the signature directly to dynamics on a tensor algebra. This structure naturally accommodates the complex temporal dependencies they are aiming for.
Meng: But connecting that to a practical model means we need to see how this translates into something that can handle real-world data streams efficiently, especially given that they mention the group-like structure being useful for parallel computation.
Paper summary: Lalam: That computational efficiency is significant; if the aggregation happens in parallel scans instead of sequentially, it could drastically speed up how we process these complex temporal patterns in our AI systems.
Tom: It sounds like the authors are pushing past the limitations of existing methods, specifically those that rely on fixed truncation depths or purely exponential fading, which is where this paper fits in. They show that EWS can approximate targets requiring oscillatory or growth modes much better than simpler methods do.
Jane: And they confirm this expressivity through a universal approximation theorem, suggesting that linear functionals of the signature can actually get you close to any continuous function on a compact set of paths, provided you respect the operator equivalence.
Lu: That connection to universal approximation is what makes me excited; it suggests we are building a representation that isn't just specialized for one type of pattern but is fundamentally capable of modeling a much wider range of complex temporal behaviors.
Meng: So, if we’re talking about applying this to traditional state-space models, they say the depth one term naturally captures the linear dynamics from standard SSMs, while higher-order terms handle that non-linearity. That’s a clear division of labor for implementation.
Lalam: If we can use this framework to capture more nuanced memory dynamics in AI systems, it could lead to models that retain complex contextual awareness over much longer sequences of input data than currently possible.
Tom: And that leads us nicely into the conclusion section where they discuss the broader implications of the work, moving beyond just the technical definitions we've touched on earlier. They are essentially showing how this generalized signature can extend both classical state-space models and transforms like Fourier and Laplace transforms.
Jane: It really boils down to how this EWS framework allows us to represent paths in a way that is intrinsic to the trajectory itself, rather than being dependent on the specific discretization we use when we observe those paths.
Lu: The fact that it generalizes both Laplace and Fourier transforms by relating its depth-one term to spectral filtering determined by the spectrum of A shows how deeply embedded this concept is in fundamental analysis.
Paper summary: Meng: From a practical standpoint, if this helps us build better models for coupled oscillatory SDE regression tasks, it opens up new avenues for modeling complex systems where different components interact over time in non-linear ways.
Lalam: That sounds like it could profoundly improve the robustness of our predictive AI by allowing it to anticipate system states with a richer understanding of historical context.
Tom: The authors conclude that because of this structure, the EWS is not just another signature but a powerful tool for extending SSMs, and their work shows that this extension has significant practical utility in modeling complex temporal phenomena.
Jane: They really highlight how the group-like nature of the EWS allows for efficient computation through parallel associative scans, which gives us a concrete way to think about scaling this up computationally.
Lu: The efficiency aspect is crucial because it means that even with these incredibly rich dynamics, we don't have to deal with intractable computational costs when actually running the model on large datasets.
Meng: I’m still focused on the implementation hurdle; while the theory is elegant, getting it into a production environment where we can handle high-frequency data streams efficiently is where the real challenge lies for me.
Lalam: I see this as an opportunity for our AI to develop a richer, more context-aware understanding of sequential information that goes beyond simple linear extrapolation.
Tom: So, to wrap up the summary of "Extending SSMs with the Exponentially Weighted Signature," we've seen how they generalize path signatures using bounded linear operators to capture richer memory dynamics and connect this concept to deeper mathematical structures like differential equations and transforms.
Jane: And the implications are that we can now build state-space models with a level of temporal awareness that goes beyond what simple exponential fading allows, opening doors for modeling oscillatory and growth patterns.
Lu: This work suggests a new way to represent path information intrinsically, which could lead to novel mathematical tools for analyzing complex time series data.
Meng: The practical impact is seeing how we can design more robust predictive models that account for coupled temporal dependencies in real-world systems rather than just relying on simpler linear assumptions.
Lalam: Ultimately, this framework offers a path toward developing AI that possesses a much deeper, more nuanced understanding of sequential context across long time horizons.
Conclusion: Tom: So we’ve just finished walking through how this paper expands state-space models using an exponential weighted signature, and now Jane, let's talk about what that title actually means for the people listening.
Jane: It really boils down to taking a standard way of tracking information over time and giving it a much more sophisticated memory mechanism by introducing those exponential weights. The authors are showing us how to weight the past in a way that isn't just simple decay, but something with structural complexity.
Lu: I think the real takeaway is how this new signature allows the model to capture things like oscillatory patterns or growth trends in its behavior that diagonal methods just can't touch. It opens up a whole new space for modeling dynamic systems within AI architectures.
Meng: From my side, I’m thinking about how this mathematical structure translates into actual running time and resource usage; if the computation is manageable, we could actually deploy models that capture these complex temporal interactions efficiently in production environments.
Lalam: For me, the biggest implication is how this could allow AI to develop a much richer cultural understanding of sequential information over long periods, moving beyond just short-term context to something more deep and contextual.
Tom: Exactly! So the authors are using this exponential weighted signature to create a state-space model that can handle these kinds of complex, non-linear temporal dynamics. The paper is essentially giving us a new tool to build more expressive AI systems.
Jane: And they’re pointing out that by doing this, we can move away from relying only on simple linear assumptions in our models and start incorporating much richer behaviors directly into the structure of the system itself.
Lu: I think it's fascinating because it connects these path-integral concepts to something very concrete—a controlled differential equation on a tensor algebra—which gives us a solid mathematical foundation for these powerful new dynamics.
Meng: The practical implication, I see, is that we’re moving toward state-space models that can better handle coupled systems where different parts interact in complex ways over time rather than just processing data in isolation.
Lalam: And from the perspective of language and culture, this means AI could build models that understand long-term trends and subtle shifts in context, which is something we’ve always struggled to achieve.
Tom: It sounds like the authors are paving the way for a new generation of state-space models that can be significantly more dynamic and expressive than what we have currently available.
Jane: That's right, so this paper isn't just tweaking an old method; it’s introducing a fundamentally different way to represent and model how things evolve over time in complex systems.
Lu: We should really keep an eye on how they apply the universal approximation theorem; that suggests we can actually get pretty close to almost any continuous behavior with this new signature structure.
Meng: I just hope the complexity of implementing these operators doesn't make them too difficult to tune for real-world deployment without losing all that theoretical elegance.
Lalam: It’s exciting because it suggests a future where AI can model not just what happens next, but how those interactions have shaped the entire history leading up to that point.
Mathematical Institute, University of Oxford · Department of Mathematics, Imperial College London · SKF
stat.ML, cs.LG
Submitted: 2026-03-19
Updated: 2026-09-30
Comments: 47 pages, 1 figure
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 83/100
The gist: The Exponentially Weighted Signature (EWS) generalizes classical path signatures by replacing uniform historical weighting with bounded linear operators, offering richer temporal dynamics and
Key concepts
- Exponentially Weighted Signature (EWS)
- The EWS generalizes classical signatures by introducing an operator pair A=(A, B) to govern temporal weighting. This allows for richer memory dynamics—including oscillatory and growth behaviors—that diagonal operators cannot capture. It is defined through iterated integrals incorporating this exponential weighting.
- Linear Controlled Differential Equation (CDE)
- The EWS satisfies a specific linear controlled differential equation: dSA(X)s,t = −ΛASA(X)s,t dθt + SA(X)s,t ⊗ dXt. This equation governs how the signature evolves over time and is crucial for deriving the compact recursive representation of the signature.
- Depth One Term (SA(X)(1))
- The depth one term of the EWS realizes a causal short-time Fourier transform of path increments. When A is diagonal, it reduces to the Exponentially Fading Memory Signature (EFM). This term is fundamental because it naturally encapsulates the linear dynamics found in traditional State Space Models (SSMs).
- Operator A
- The operator A defines the temporal weighting mechanism within the EWS. Learning this operator allows the model to capture complex memory dynamics, such as oscillatory or regime-dependent behavior. Its structure determines whether the resulting signature can approximate certain path shapes.
Terminology
Summary
The Exponentially Weighted Signature (EWS) generalizes classical path signatures by replacing uniform historical weighting with bounded linear operators, offering richer temporal dynamics and cross-channel coupling. This framework provides a principled, continuous-time dynamical description of paths that is amenable to gradient-based learning and extends the expressivity of state-space models beyond their classical limitations.
The Core Concept: Exponentially Weighted Signature (EWS)
The EWS generalizes the classical signature by introducing an operator pair, denoted as A = (A, B), which governs temporal weighting. The definition involves iterated integrals that incorporate this operator exponentially:
SA(X)s,t = Z t s Z tn s · · · Z t2 s e −(θt−θt1)ABdXt1 ⊗ · · · ⊗ e −(θt−θtn)ABdXtn.
This structure allows the temporal weighting to be learned through the operator A, enabling richer memory dynamics including oscillatory, growth, and regime-dependent behaviour,
which is structurally inaccessible to diagonal operators. The EWS is shown to be the unique solution to a linear controlled differential equation on the tensor algebra:
dSA(X)s,t = −ΛASA(X)s,t dθt + SA(X)s,t ⊗ dXt.
Mathematical Foundation and Dynamics
The EWS is fundamentally defined through iterated integrals within the tensor algebra T((V)), which is equipped with admissible norms. The flow operator Dh A propagates the exponential weighting across tensor levels in a multiplicative manner. The derivation operator ΛA, defined as the infinitesimal generator of the one-parameter subgroup of automorphisms G = D hA, is crucial for deriving a compact recursive representation:
ΛA(u) = Au.
This leads to the key result that the EWS satisfies a linear controlled differential equation (CDE):
dSA(X)s,t = −ΛASA(X)s,t dθt + SA(X)s,t ⊗ dXt, with Ss s = 1.
Equivalence to Classical Signatures and Transforms
The EWS admits a natural interpretation as the classical signature of a linearly transformed path. Specifically:
SA(X)s,t = S(Z[t])s,t,
where the re-weighted path Z[t] is defined by the integral:
[Z[t]r = Z r s e −(θt−θu)AdXu, r ∈ [s, t].
Furthermore, at depth one, the EWS term SA(X)(1)s,t realizes a causal short-time Fourier transform of the path increments. When A is diagonal (the EFM case), this reduces exactly to the Exponentially Fading Memory Signature (EFM). When A is non-diagonalizable, it introduces polynomial factors in the kernel, enlarging the admissible class of spectral filters beyond purely exponential modes.
Expressivity and Universal Approximation
The EWS framework demonstrates a significant expressivity gap over fixed truncation depths compared to both the classical signature and EFM. Experiments show that targets requiring oscillatory or growth modes—which are structurally inaccessible to diagonal operators—are approximated substantially better by the EWS learner, even though the optimization landscape is harder. The universal approximation theorem confirms that linear functionals of the EWS signature can approximate any continuous function on a compact set of paths, provided they respect A-equivalence.
Computational Advantages and Applications
The group-like structure of the EWS enables efficient computation via parallel associative scans, reducing complexity from O(N) sequential steps to O(log N) parallel steps for aggregating local increments. In practice, this is achieved by evaluating the EWS on sub-intervals and aggregating them using the modified Chen’s identity. Furthermore, it provides a principled extension of State Space Models (SSMs), where the depth one term naturally encapsulates the linear dynamics of traditional SSMs, while higher-order terms capture intrinsic non-linearity. The framework also allows for the representation of coupled oscillatory SDE regression tasks by learning matrices A that exhibit complex eigenvalues and substantial off-diagonal structure.
Key Properties Enumerated:
-
The EWS is defined on finite horizons, allowing temporal weighting to be learned without restricting it to purely decaying behaviour.
-
It satisfies a modified Chen’s identity: SA(X)s,t = D θt−θu A SA(X)s,u ⊗ SA(X)u,t for p=1 paths.
-
The truncated EWS is the unique solution to a linear CDE in the space of tensor algebra elements T n(V).
-
It generalizes Fourier and Laplace transforms by relating its depth-one term to spectral filtering determined by the spectrum of A.
Improvements for AI systems
Based on the provided scientific paper, here are the specific improvements to AI systems that could be achieved by adopting or integrating the Exponentially Weighted Signature (EWS) framework:
The EWS framework offers several transformative capabilities for AI systems, primarily by providing a principled, continuous-time representation of history that captures temporal context dynamically.
Here are the specific improvements and what the resulting system can do:
- -Temporal Contextual Modeling via Non-Monotone Memory Dynamics:
AI systems (especially those using sequential data like RNNs or Transformers) currently treat time implicitly through sequence indexing, leading to limitations in handling varying historical relevance (e.g., regime shifts). The EWS allows the temporal weighting mechanism (governed by the operator A) to be state-dependent and non-monotone.
AI System Improvement: Replace standard fixed-memory architectures with an EWS layer that learns a general bounded linear operator A. This allows the system to modulate how quickly past information is forgotten or emphasized based on its current state, capturing complex dynamics like oscillations, resonance, or regime transitions that are fundamentally inaccessible to Exponentially Fading Memory (EFM) models which restrict memory decay to channel-wise exponential fading.
What the improved system can do: It will outperform existing models in tasks involving financial time series (e.g., predicting market volatility during sudden regime shifts) or biological signaling, as it can dynamically tune
its sensitivity to different historical epochs.
- -Enhanced Expressivity for Coupled Oscillatory Systems:
The EWS captures cross-channel coupling at the level of temporal weighting through the matrix exponential of A, which is structurally inaccessible to diagonal operators in EFM. The numerical experiments explicitly show that when modeling coupled oscillatory SDEs (like the Duffing oscillator), EWS significantly outperforms both EFM and classical signature methods.
AI System Improvement: Design specialized neural architectures (e.g., Structured Linear Neural Controlled Differential Equations or SLiCE frameworks) where the learned operator A is a general matrix, not constrained to be diagonal.
What the improved system can do: It will accurately model complex, coupled physical systems (like those described in Section 9.2) that exhibit frequency-dependent interactions and cross-channel feedback loops, leading to superior prediction accuracy in non-linear control problems compared to models restricted by channel independence assumptions.
- -Principled Generalization of State Space Models (SSMs):
The EWS framework provides a rigorous mathematical link between the path signature and the solution of Linear Controlled Differential Equations (CDEs). It naturally encapsulates the linear dynamics of traditional SSMs within its depth-one signature term, while higher orders introduce non-linearity.
AI System Improvement: Develop a hybrid architecture where the latent state evolution is modeled not just by a standard Kalman filter/SSM equation, but by the underlying EWS CDE. This allows the system to learn temporal dynamics that are richer than simple linear state transitions (e.g., incorporating polynomial memory effects).
What the improved system can do: It will provide a more robust and mathematically rigorous foundation for modeling high-dimensional time-series data where traditional SSMs fail due to their inability to capture non-linear, long-term dependencies without excessive truncation depth.
- -Efficient and Parallel Computation via Group Structure:
The EWS possesses a group-like structure, allowing the entire iterated integral representation to be computed efficiently via parallel associative scans (modified Chen's identity), rather than sequential computations required by other signature methods.
AI System Improvement: Implement the EWS calculation within a graph neural network or parallel processing framework that leverages the algebraic properties of the tensor algebra and its flow operator.
What the improved system can do: It will enable high-throughput, near real-time inference and training for deep learning models involving high-order signature transforms, as it reduces computational complexity from potentially exponential in depth to a structure solvable in parallel with logarithmic scaling relative to the path length (Section 8).
- -Universal Approximation Guarantee:
The framework guarantees that linear functionals of the EWS can approximate any continuous function on the space of A-equivalence classes, extending the universal approximation capability of classical signatures even when adapted for temporal weighting.
AI System Improvement: Use EWS as a feature map in a learning pipeline where the readout layer is a linear functional (a neural network layer).
What the improved system can do: It ensures that if a target function respects the learned temporal equivalence structure defined by A, the model has a theoretical guarantee of finding an approximation, providing strong inductive bias for learning relevant features from complex path data.
Abstract
We introduce the exponentially weighted signature (EWS), a continuous-time model that computes iterated integrals of a path, where each increment is weighted by the matrix exponential of a learnable generator over elapsed clock time. We prove that it solves a linear controlled differential equation, keeps the group-like structure and the universality of the signature, and satisfies a modified Chen identity, enabling a parallel scan. At depth one the EWS is a state-space model (SSM), and we map linear time-invariant SSMs, Mamba channels and Mamba- 2 heads to it in closed form. The EWS extends SSMs through an arbitrary matrix generator, a clock that generalises the step size to causal functionals of the input, and higher truncation depths that are non-linear in the path within a single layer. Empirically, the EWS achieves the highest average accuracy and rank on six long time-series classification datasets, where depth generally helps. Learned clocks prove necessary for state tracking on formal language tasks, and at depth one, the EWS matches or exceeds competing SSMs on regression and forecasting with far fewer parameters.
Sources
- Exponentially Fading Memory Signature
- Universal approximation with signatures of non-geometric rough paths
- A Primer on the Signature Method in Machine Learning
- Nowcasting using regression on signatures
- Sliding-Window Signatures for Time Series: Application to Electricity Demand Forecasting
- Sparse arrays of signatures for online character recognition
- Extracting information from the signature of a financial data stream
- The Volterra signature
- Volterra equations driven by rough signals 3: Probabilistic construction of the Volterra rough path for fractional Brownian motions
- Volterra Equations Driven by Rough Signals
- Volterra equations driven by rough signals 2: higher order expansions
- Learning from the past, predicting the statistics for the future, learning an evolving system
- A Generalised Signature Method for Multivariate Time Series Feature Extraction
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey