Closing the Loop: A Control-Theoretic Framework for Provably Stable Time Series Forecasting with LLMs
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Closing the Loop: A Control-Theoretic Framework for Provably Stable Time Series Forecasting with LLMs".
Jane: The paper was written by Xingyu Zhang, Jingyao Wang, Zeen Song, Changwen Zheng and Wenwen Qiang from University of Chinese Academy of Sciences and Institute of Software, Chinese Academy of Sciences.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper with a title that just sounds like it means business: "Closing the Loop: A Control-Theoretic Framework for Provably Stable Time Series Forecasting with LLMs." Jane, when you first saw that title, what jumped out at you?
Jane: Oh, the word "control-theoretic" got me immediately, Tom. It's such a shift in mindset. Usually when we talk about large language models, we're thinking about words, grammar, maybe code. But here, the authors are saying, let's treat the whole forecasting process like an engineering control system, like a thermostat regulating a room.
Tom: A thermostat, I like that. So instead of just letting the model spit out predictions and hoping they're right, they're actively monitoring and correcting the trajectory as it goes.
Jane: Exactly. And that's what "closing the loop" means. You have an open loop, which is just a straight line from input to output. But a closed loop means you're constantly checking the output against where you want to be and feeding that error back in to adjust.
Tom: And the "provably stable" part is the real kicker, right? That's not just a claim that it works well in practice; it's a mathematical guarantee that the error won't blow up.
Jane: Right, and that's what separates this from a lot of deep learning papers. They're not just showing you a cool trick; they're giving you a theorem that says, under these conditions, the system will stay bounded. That's a big deal for trust.
Tom: So the authors are from the University of Chinese Academy of Sciences and the Institute of Software there. They're basically saying the standard way of doing autoregressive forecasting with LLMs has a fundamental flaw.
Jane: It's a structural vulnerability, they call it. And the title is their answer to that flaw. It's not just a new model; it's a new way of thinking about the inference process itself.
Tom: Well, I'm hooked. Let's get into what that flaw actually is and how they fix it.
Summary: Tom: So, Jane, we've established the title is about fixing a structural flaw. Let's get into the meat of it. What's the core problem the paper identifies with how LLMs are usually used for time series?
Jane: Okay, so imagine you're trying to predict the next week of stock prices. You give the LLM the past month. It predicts tomorrow. Then, for the day after, you feed it its own prediction from yesterday, not the real price. That's called autoregressive generation.
Tom: And that seems logical, right? You don't have the real data for the future, so you have to use your own guesses.
Jane: Sure, but here's the catch. The model was trained on real data. It was shown the correct answer at every step. So during training, it never sees its own mistakes. But during inference, it's only seeing its own mistakes. This is the "exposure bias" problem.
Tom: And in a continuous domain like time series, a tiny mistake on day one becomes a bigger mistake on day two, which becomes an even bigger mistake on day three. The paper calls it a "snowball effect."
Jane: Exactly. They formalize this. They show that if the model's sensitivity to its own errors is greater than one, which they call the spectral radius, then the error grows exponentially with the forecast horizon. It's not just a little drift; it's a guaranteed divergence.
Tom: So the model is essentially driving blind, and the road keeps getting more and more slippery with every mile.
Jane: That's a great way to put it. And the paper's solution is to stop driving blind. They propose a closed-loop system with a "residual estimator" that acts like a co-pilot, constantly looking at the recent errors and estimating what the next correction should be.
Tom: So instead of just accepting the model's prediction, you're actively adjusting it based on the pattern of past mistakes.
Jane: Precisely. And they prove that if you can make that correction strong enough, you can force the error to stay within a bounded range, no matter how long the forecast horizon is. That's the "provably stable" part.
Tom: So the summary is: naive autoregression is unstable, and they've built a control system to fix it. Now, how do they actually build this thing? What are the practical improvements?
Improvements: Tom: So we've got the problem and the high-level solution. But how do they actually make this work in practice, Jane? What are the concrete improvements they suggest?
Jane: The big one is the architecture. They call it F-LLM, which stands for Feedback-driven LLM. It's not just one big model. It's split into two main parts. First, you have the frozen LLM itself, which acts as the "plant" in control theory terms. It does the raw prediction.
Tom: And the second part is the "observer," which is the residual estimator we talked about. It's a lightweight network that watches the past errors and predicts the next correction.
Jane: Right. And the key is that this observer is trained in a clever way. They don't just train everything at once. They use a two-stage curriculum. First, they train the projection layers to align the time series with the LLM's embedding space, while also imposing a "Local Lipschitz constraint."
Tom: That's a mouthful. What does that constraint actually do?
Jane: It's a way to ensure the model's output doesn't change wildly with small changes in the input. It's like making sure the steering wheel isn't too sensitive. If it's too sensitive, a tiny bump in the road sends you flying off. This constraint makes the system "controllable," so the feedback loop can actually stabilize it.
Tom: So stage one is about making the plant predictable. Then what?
Jane: Then they freeze that and train the observer in a closed-loop mode. They let the model generate predictions, compare them to the ground truth during training, and teach the observer to predict that error. This way, the observer learns the systematic biases of the frozen LLM.
Tom: So it's not just learning to be a good predictor; it's learning to be a good *corrector* for a specific, flawed predictor.
Jane: Exactly. And the results show this works. They get consistent improvements over the base autoregressive models across datasets like electricity, weather, and traffic. And they even show it works in zero-shot settings, where you train on one dataset and test on another.
Tom: So it's not just a theoretical fix; it's a practical improvement that also preserves the LLM's ability to generalize. That's a powerful combination.
Jane: It is. And it's also efficient. The observer is just a single linear layer, so it adds almost no computational overhead. You get the stability without paying a huge price in speed.
Tom: That's the dream, right? Better results, same speed, and a proof that it won't fall apart. So where does this leave us? What's the big picture?
Conclusion: Tom: Well, we've come to the end of our time with "Closing the Loop: A Control-Theoretic Framework for Provably Stable Time Series Forecasting with LLMs." Jane, give us the final takeaway.
Jane: The takeaway is that this paper identifies a real, structural weakness in how we've been using LLMs for time series, and it offers a principled, mathematically grounded fix. It's not just a new architecture; it's a new framework for thinking about the problem.
Tom: And the beauty is that it's model-agnostic. They showed it works with GPT-two OPT, and LLaMA. So you don't need to retrain a massive model. You just add this lightweight correction loop on top of a frozen one.
Jane: Right. And that's what makes it so practical. It's a way to get more reliability and accuracy out of models that already exist, which is a huge deal for real-world applications like energy forecasting or financial planning.
Tom: It also feels like a philosophical shift. Instead of just trusting the model's output, we're treating it as a component in a larger system that we can actively control and stabilize.
Jane: Exactly. It's a move from pure prediction to active control. And the proof that the error stays bounded gives us a level of confidence that's rare in deep learning.
Tom: Well, that's a fantastic note to end on. We've said goodbye to this paper, but we're already looking forward to the next one. Thanks for joining us, and we'll see you next time.
Xingyu Zhang, Jingyao Wang, Zeen Song, Changwen Zheng, Wenwen Qiang
University of Chinese Academy of Sciences · Institute of Software, Chinese Academy of Sciences
cs.LG
Submitted: 2026-08-17
Updated: 2026-08-18
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 85/100
The gist: The paper identifies a critical theoretical flaw in applying Large Language Models (LLMs) to time series forecasting.
Key concepts
- Control-Theoretic Framework
- This is a shift in mindset where the forecasting process is treated like an engineering control system. Instead of just generating predictions, the model constantly monitors its output against the desired target and feeds any resulting error back into the system to actively make necessary adjustments.
- Exposure Bias
- This is a structural flaw in LLM inference. The model is trained on correct data but sees only its own previous predictions during forecasting. This means it never learns from its own mistakes, leading to compounding errors as the forecast progresses.
- F-LLM Architecture
- The practical solution involves splitting the system into two parts: a frozen LLM (the 'plant' that makes raw predictions) and a lightweight 'observer' (the residual estimator). This observer learns to predict and correct the systematic biases of the frozen LLM.
Terminology
Summary
The paper identifies a critical theoretical flaw in applying Large Language Models (LLMs) to time series forecasting. While LLMs have shown potential in this domain by leveraging sequential reasoning capabilities, existing approaches employ a naive autoregressive generation strategy
that operates in an open-loop manner during inference, consuming its own generated outputs recursively. This leads to inevitable error accumulation (exposure bias), where minor early deviations cascade into significant trajectory drift over long horizons.
The authors formalize this issue: "Standard LLMs are trained under a 'Teacher Forcing' regime, where the model is conditioned on ground-truth history. Yet, during inference, the model operates in an open-loop autoregressive mode, consuming its own generated outputs recursively to predict subsequent steps. This discrepancy creates a distribution shift known as Exposure Bias. In the continuous domain, even a microscopic error ϵ at step t introduces a covariate shift for step t + 1. Without ground-truth correction, these errors do not cancel out; instead, they propagate and amplify through the feedback loop, exhibiting compound growth."
The authors propose F-LLM (Feedback-driven LLM), a novel closed-loop framework that reformulates autoregressive forecasting through the lens of control theory. The core insight is to view the forecasting error not merely as a performance metric, but as a state disturbance that must be actively rejected.
Since the true error is unobservable during inference, they introduce a learnable Residual Estimator that functions as a system Observer, inferring the latent error state from the current context. This estimate is fed into a Feedback Controller to calibrate the trajectory on-the-fly.
The paper provides rigorous mathematical analysis of error propagation:
-
Proposition 4.1 (Exponential Error Growth): "Consider the linearized error dynamics ∆xt+1 ≈ Jg∆xt + ϵt. If the spectral radius ρ(Jg) > 1, the upper bound of the expected error norm grows exponentially with the horizon H." The expected error norm scales as O(ρ(Jg) H).
-
Theorem 4.1 (Bounded Stability via Feedback): "Assume the single-step modeling error is bounded, i.e., ∥ϵt∥ ≤ γ for some finite γ. If there exists a feedback gain L and a constant q ∈ [0, 1) such that the closed-loop operator satisfies the contraction condition ∥Jg(x) − L∥2 ≤ q for all x, then the cumulative error sequence is uniformly bounded: lim sup ∥∆xt∥ ≤ γ/(1−q)."
The authors emphasize that Theorem 4.1 serves as more than a theoretical proof; it acts as a blueprint for our architectural design.
The theorem requires two conditions: (1) access to previous error, necessitating the Residual Estimator as an observer; (2) bounded and well-conditioned Jacobian, mandating Local Lipschitz Regularization on the base predictor.
F-LLM consists of three modules:
-
Patch-wise Autoregressive Predictor (Plant): Segments input into non-overlapping patches, uses a learnable embedding layer, frozen LLM backbone, and learnable inverse embedding layer.
-
Residual Estimator (Observer): A lightweight network that "takes the sequence of past estimated residuals ∆P<t as input and predicts the correction term for the current step,
parameterizing
the optimal control law derived in Eq.(6)." -
Closed-Loop Interaction (Controller): The corrected patch p̃t = p̂t + ∆p̂t is appended to the context for the next autoregressive step, preventing the
snowball effect.
The authors employ a Two-Stage Curriculum Learning approach:
-
Stage 1 (Open-Loop Alignment): Freeze the Residual Estimator, train only projection layers with Teacher Forcing, and impose a Local Lipschitz constraint on the actuator fhead via the loss: Llip = E[∥fhead(h+δ) − fhead(h)∥2/∥δ∥2 − κ]+.
-
Stage 2 (Closed-Loop Feedback Learning): Freeze projection layers, activate the Residual Estimator, and supervise it to predict deviations: L2 = (1/N)Σ∥(pt − p̂t) − rψ(∆p̂<t)∥2.
Long-term Forecasting: Evaluated on seven real-world datasets (ETTh1, ETTh2, ETTm1, ETTm2, ECL, Weather, Traffic) with prediction horizons 96, 192, 336, 720. F-LLM achieves the best performance in most horizons,
reducing MSE by 5.6% compared to the strongest baseline on challenging long-horizon cases.
Zero-Shot Forecasting: Tested on M3 and M4 competition datasets. F-LLM attains the lowest SMAPE on 7 out of 8 transfer settings and is second best on the remaining one,
demonstrating preserved zero-shot generalization.
Ablation Studies: Removing either the feedback module or the LLM backbone consistently degrades performance. The gap between F-LLM w/o Feedback
and F-LLM (Full)
quantifies the systematic bias correction benefit.
Generality: F-LLM works with GPT-2, OPT, and LLaMA backbones, showing consistent performance gains across all backbones.
Efficiency: F-LLM introduces less than 5% additional inference time relative to the frozen LLM baseline
and is even faster than AutoTimes despite slightly more tunable parameters.
Comparison with LoRA: F-LLM outperforms parameter-efficient fine-tuning (LoRA), indicating correcting structured prediction bias at inference time is more effective than lightweight parameter adaptation for mitigating error accumulation.
The paper concludes: "We identified the structural vulnerability of Error Accumulation inherent in applying autoregressive Large Language Models to continuous time series forecasting. To mitigate the inevitable trajectory drift caused by open-loop inference, we proposed F-LLM, a novel framework that reformulates forecasting as a closed-loop control problem. By integrating a Residual Estimator as a system observer and enforcing Local Lipschitz constraints for controllability, F-LLM effectively rejects disturbances and actively calibrates the generative path. Our theoretical analysis provides a rigorous guarantee of uniformly bounded errors."
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in AI systems, along with what the improved system can do:
-
Implementation: Add a lightweight residual estimator (a single linear layer) that runs alongside any frozen autoregressive LLM. This estimator takes the sequence of past prediction errors as input and outputs a correction term.
-
Integration: During inference, after each autoregressive step, add the estimated correction to the raw prediction before feeding it back into the model's context for the next step.
-
Implementation: During training of the encoder/decoder projection layers (not the frozen LLM), add a penalty term that constrains the Lipschitz constant of the output projection head. Specifically, perturb the latent representation with small noise and penalize output amplification beyond a threshold.
-
Stage 1: Train only projection layers using teacher forcing with the Lipschitz constraint (open-loop alignment).
-
Stage 2: Freeze projections, then train the residual estimator using closed-loop student forcing, where the model generates predictions and the estimator learns to predict the deviation from ground truth.
-
Implementation: Add a feedback gain matrix (learnable) that scales the estimated residual before injection. This ensures the closed-loop operator satisfies the contraction condition (spectral radius < 1), guaranteeing bounded error.
-
Stable long-horizon predictions: Maintain trajectory alignment over 720+ time steps without exponential error drift, even when the base LLM has spectral radius > 1.
-
Zero-shot generalization: Transfer across domains (e.g., M4 → M3) with 5-10% SMAPE improvement over standard autoregressive LLMs, without any fine-tuning on target data.
-
Model-agnostic correction: Work with any decoder-only LLM (GPT-2, OPT, LLaMA) without modifying the backbone, adding less than 5% inference overhead.
-
Exposure bias mitigation: Reduce the train-inference distribution mismatch in any autoregressive generative model (text, audio, video) by actively correcting drift during generation, not just during training.
-
Controllable generation: Provide a mathematical guarantee that generation error remains bounded (≤ γ/(1−q)) under mild Lipschitz conditions, making outputs more reliable for safety-critical applications.
-
Robust prediction in dynamic environments: Reject disturbances (e.g., sudden regime shifts, noise) by treating them as system disturbances that the feedback controller actively cancels.
-
Adaptive correction: Learn systematic biases of the base model on-the-fly and correct them without retraining, enabling rapid deployment on new data distributions.
Abstract
Large Language Models (LLMs) have recently shown exceptional potential in time series forecasting, leveraging their inherent sequential reasoning capabilities to model complex temporal dynamics. However, existing approaches typically employ a naive autoregressive generation strategy. We identify a critical theoretical flaw in this paradigm: during inference, the model operates in an open-loop manner, consuming its own generated outputs recursively. This leads to inevitable error accumulation (exposure bias), where minor early deviations cascade into significant trajectory drift over long horizons. In this paper, we reformulate autoregressive forecasting through the lens of control theory, proposing F-LLM (Feedback-driven LLM), a novel closed-loop framework. Unlike standard methods that passively propagate errors, F-LLM actively stabilizes the trajectory via a learnable residual estimator (Observer) and a feedback controller. Furthermore, we provide a theoretical guarantee that our closed-loop mechanism ensures uniformly bounded error, provided the base model satisfies a local Lipschitz constraint. Extensive experiments demonstrate that F-LLM significantly mitigates error propagation, achieving good performance on time series benchmarks.
Sources
- Sequence Level Training with Recurrent Neural Networks
- Deep multi-scale video prediction beyond mean square error
- Extracting Seasonal Gradual Patterns from Temporal Sequence Data Using Periodic Patterns Mining
- Enhancing Representation Learning for Periodic Time Series with Floss: A Frequency Domain Regularization Approach
- iTransformer: Inverted Transformers Are Effective for Time Series Forecasting
- Fourier Neural Operator for Parametric Partial Differential Equations
- Representation Learning via Invariant Causal Mechanisms
- Not All Frequencies Are Created Equal:Towards a Dynamic Fusion of Frequencies in Time-Series Forecasting
- TimeXer: Empowering Transformers for Time Series Forecasting with Exogenous Variables
- CITRAS: Covariate-Informed Transformer for Time Series Forecasting
- Technology Trends for Massive MIMO towards 6G
- PromptCast: A New Prompt-based Learning Paradigm for Time Series Forecasting
- Large Language Models Are Zero-Shot Time Series Forecasters
- One Fits All:Power General Time Series Analysis by Pretrained LM
- UniTime: A Language-Empowered Unified Model for Cross-Domain Time Series Forecasting
- Time-LLM: Time Series Forecasting by Reprogramming Large Language Models
- TEMPO: Prompt-based Generative Pre-trained Transformer for Time Series Forecasting
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Language Models are Few-Shot Learners
- OPT: Open Pre-trained Transformer Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks