CastFSR: A Fast–Slow–Reflect Agentic Reasoning Framework for Context-Aware Time Series Forecasting

summary

Video file (mp4)

The gist

CastFSR is an agentic framework that formulates context-aware time series forecasting as a Fast–Slow–Reflect workflow.

In short

The episode explores CastFSR, a framework for time series forecasting that moves beyond simple pattern matching. It uses an agentic approach to predict future values by incorporating external factors like weather or holidays. The process involves three stages: a fast numerical guess, slow contextual reasoning, and a final reflection check against reality to produce accurate and plausible forecasts.

Key concepts

Time Series Forecasting
This is the task of predicting future values in a sequence of data. Unlike simple pattern extension, real-world data requires understanding external factors like weather or market conditions to ensure the prediction is accurate.
CastFSR Framework
A three-stage agentic framework for forecasting. It combines a quick numerical baseline with slow, contextual reasoning (using external data) and a final reflection stage to check results against physical constraints.
Agentic Reasoning
The system operates like a project manager rather than just a calculator. It decides which specialized tools to use (e.g., statistical or deep learning models) and when to incorporate external data, treating forecasting as a sequence of decisions.
Reflection Stage
The final step where the system validates its prediction against constraints like temporal consistency or domain rules. This prevents physically impossible results, such as a negative value for wind power, which a pure numerical model would not catch.

Terminology used across episodes

This episode discusses

The paper

CastFSR: A Fast--Slow--Reflect Agentic Reasoning Framework for Context-Aware Time Series Forecasting · Read on arXiv

Xiaoyu Tao, Mingyue Cheng, Bokai Pan, Chuang Jiang, Huanjian Zhang, Tian Gao, Yaguo Liu, Qi Liu, Enhong Chen

University of Science and Technology of China

Time series forecasting is fundamental to decision-making in complex systems, where future dynamics are influenced not only by historical observations but also by evolving contextual features. Recent advances in large language models (LLMs) have extended forecasting beyond numerical extrapolation toward context-aware reasoning. However, existing approaches often lack explicit mechanisms to identify relevant contexts, reason about their impacts, and validate forecasts against temporal and domain constraints. In this work, we propose CastFSR, an agentic framework that formulates context-aware forecasting as a Fast--Slow--Reflect workflow. In fast thinking, CastFSR profiles observations and selects lightweight forecasters to construct a data-driven forecast prior. In slow deliberation, it retrieves contextual evidence, adaptively determines informative look-back windows, and reasons about how contexts reshape future dynamics. In reflection, it iteratively refines forecasts to ensure temporal, contextual, and domain consistency. CastFSR supports both training-free inference with off-the-shelf LLMs and efficient deployment through a two-stage SFT and reinforcement learning strategy that transfers its orchestration capability to compact LLMs. Extensive experiments on public datasets demonstrate that CastFSR consistently outperforms representative baselines. Our code is available at https://github.com/Xiaoyu-Tao/CastFSR.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "CastFSR: A Fast–Slow–Reflect Agentic Reasoning Framework for Context-Aware Time Series Forecasting".

Jane: The paper was written by Xiaoyu Tao, Mingyue Cheng, Bokai Pan, Chuang Jiang, Huanjian Zhang et al. from University of Science and Technology of China.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everybody. Today we're looking at a paper that's got a mouthful of a title: "CastFSR: A Fast–Slow–Reflect Agentic Reasoning Framework for Context-Aware Time Series Forecasting." Jane, I'm going to need you to unpack that title for me because I'm already lost.

Jane: Happy to, Tom. So think of time series forecasting as trying to predict the next few numbers in a sequence. The old way is you look at the numbers you have and you just extend the pattern. But this paper says that's not enough, because real-world data like electricity prices or wind power doesn't just follow its own history. It reacts to things like weather, holidays, and market conditions.

Tom: Right, so it's not just about the numbers themselves, but the world around those numbers.

Jane: Exactly. And the title breaks down into three parts. Fast is the quick numerical guess, Slow is the careful reasoning about context, and Reflect is checking your work against reality. The paper calls this an agentic framework, which just means the system makes a series of decisions instead of doing one single calculation.

Lu: If I can jump in here, Jane, what excites me is that this isn't a new neural network architecture. It's a way of organizing how different tools work together. The large language model isn't doing the math. It's acting like a project manager, deciding which forecasting tool to use, when to bring in weather data, and when to question the result.

Tom: So the AI is less like a calculator and more like a coordinator?

Lu: Precisely. And that's a big shift. Most forecasting models are a single pipeline. You put history in, you get a prediction out. This paper treats forecasting as a sequence of decisions, which is much closer to how a human expert would actually approach it.

Meng: From my side, the practical question is whether this actually runs. And the paper addresses that directly. You can run it with a big commercial model like DeepSeek or GPT, but they also trained a smaller model, a four-billion parameter one, to do the same job. That matters for companies that can't send their data to an external API.

Jane: And that's the part I love, Meng. The framework itself is separate from the brain that runs it. You can swap in different language models and the workflow stays the same. It's like having a great recipe that works no matter who's doing the cooking.

Tom: So the title is really promising a system that thinks before it predicts.

Jane: That's the idea. And the results in the paper suggest it works, but we'll get into those numbers in a bit. For now, just remember the three steps: fast guess, slow reasoning, and then double-checking.

Tom: Alright, I'm hooked. Let's keep going and see what the abstract actually promises.

Summary: Tom: So we've got the title unpacked. Now let's talk about what the paper actually claims to achieve. Jane, what's the big picture here?

Jane: The big picture is that forecasting is hard because the future doesn't just repeat the past. The paper's summary makes this clear with a concrete example. Imagine predicting wind power. If you only look at historical wind power readings, you might miss that tomorrow's forecast calls for a big drop in wind speed. The numbers alone won't tell you that. You need context.

Tom: And that's where the three stages come in.

Jane: Right. Stage one is fast thinking. The system looks at the historical data and picks a lightweight forecasting model from a pool. It might choose a statistical model like ARIMA, or a deep learning model like PatchTST, depending on what the data looks like. This gives a quick numerical baseline.

Lu: What's clever here is that the system doesn't just pick one model and stick with it. It profiles the data first. It checks for trends, seasonality, and data quality. Then it routes the task to the model that fits those characteristics. The paper shows that this adaptive selection beats using any single fixed model.

Tom: So it's like choosing the right tool for the job instead of using a hammer for everything.

Lu: Exactly. And then stage two is the slow deliberation. The system looks at contextual features like weather forecasts or calendar events and decides whether they should change the numerical baseline. If the context doesn't matter, it leaves the forecast alone. If it does matter, it adjusts specific parts of the forecast.

Meng: The part I appreciate is that it doesn't just blindly adjust everything. The paper talks about adaptive look-back windows. Different types of context operate on different time scales. Weather might matter for the next few hours, but a holiday effect might matter for a whole day. The system figures out the right window for each factor.

Jane: And then stage three is reflection. The system checks its candidate forecast against three things: temporal consistency, meaning the trend and seasonality look right; contextual consistency, meaning the adjustments are supported by evidence; and domain constraints, like making sure wind power isn't negative.

Tom: Wait, negative wind power? That sounds like a silly mistake.

Jane: It happens more than you'd think. The paper shows a case study where the initial forecast produced negative values for wind power, which is physically impossible. The reflection stage caught that and corrected it. That's the kind of check that a pure numerical model wouldn't do.

Lu: And that's the philosophical point. The paper treats forecasting as a reasoning problem, not just a pattern-matching problem. The numbers come from specialized tools, but the judgment about whether those numbers make sense comes from the reasoning engine.

Meng: The results back this up. Across ten datasets, the full framework beats most baselines on most metrics. And the version with the smaller trained model, CastFSR-R1, does even better than the version using the big commercial model on several benchmarks.

Tom: So the summary is that this framework adds judgment to forecasting.

Jane: That's exactly it. And next we should talk about what specific improvements they made to get there.

Improvements: Tom: Alright, so we know the framework works. But what did the authors actually do differently? What are the concrete improvements over existing methods?

Jane: The biggest improvement is the separation of concerns. Most LLM-based forecasting methods try to get the language model to directly output numbers. That's a problem because language models aren't great at precise arithmetic. CastFSR instead uses the LLM as an orchestrator. It delegates the number crunching to specialized forecasting models.

Lu: And that's a fundamental design choice. The paper calls it a model-agnostic instantiation. You can take the same workflow and run it with GPT, DeepSeek, or a smaller fine-tuned model. The experiments show that different LLMs all perform reasonably well, which proves the framework isn't dependent on one particular brain.

Meng: The training strategy is also a real contribution. They took the trajectories generated by the big model and used them to train a compact four-billion parameter model. First they do supervised fine-tuning to teach the basic workflow, then they apply reinforcement learning to optimize the decisions. The ablation study shows both stages matter.

Tom: So removing either one hurts performance?

Meng: Yes. Without the supervised fine-tuning, the model doesn't know how to follow the three-stage process. Without the reinforcement learning, it doesn't learn to make good decisions about when to adjust forecasts and when to leave them alone. You need both.

Jane: Another improvement is the adaptive look-back window. Traditional methods use a fixed window of history. CastFSR searches through long-range contextual histories and picks different windows for different types of context. The case study shows that when the selected window has relevant evidence, the forecast improves. When the window is misleading, the forecast suffers.

Lu: Which tells us something important. The framework isn't magic. It's only as good as the evidence it retrieves. But the fact that it can identify when the evidence is weak and choose to make conservative adjustments is a real step forward.

Tom: So it knows when to be confident and when to be cautious?

Lu: That's the aspiration. The reflection stage is what enables that. It checks whether the adjustments are supported by evidence and whether they violate any constraints. If something looks wrong, it can go back and reconsider.

Meng: And the results show this pays off. On the long-term forecasting benchmarks, CastFSR-R1 gets the best MSE on ETTh1 and ETTh2. On the short-term benchmarks, it's best on NP and PJM. The improvements over the best fixed model are substantial in several cases.

Jane: The paper also shows that removing any single stage hurts. Removing the fast-thinking stage causes the biggest drop, which makes sense because the numerical prior is the foundation. But removing the slow reasoning or the reflection also degrades performance.

Tom: So the improvements are about structure and judgment, not just bigger models.

Jane: Exactly. And that's a much more sustainable path forward. Let's wrap up with what this means for the future.

Conclusion: Tom: We've spent some time with "CastFSR: A Fast–Slow–Reflect Agentic Reasoning Framework for Context-Aware Time Series Forecasting," and I think we should pull it all together. Jane, what's the one thing listeners should remember?

Jane: The one thing is that forecasting doesn't have to be a single blind calculation. CastFSR shows that you can break it into three steps: a fast numerical guess, a slow contextual reasoning pass, and a final reflection check. Each step has a clear job, and together they produce forecasts that are both accurate and physically plausible.

Lu: I'd add that this reframes what we ask AI to do. Instead of asking a language model to be a calculator, we ask it to be a decision-maker. That's a much better use of its strengths. The reasoning, the tool coordination, the judgment about when to trust evidence, those are things LLMs are genuinely good at.

Meng: And from a deployment standpoint, the fact that they trained a four-billion parameter model to do this is significant. You don't need a massive commercial API to get these benefits. A compact model with the right training can internalize the workflow and run efficiently.

Tom: So this isn't just a research curiosity. It's something that could actually be deployed.

Meng: Right. Energy companies forecasting load, grid operators predicting wind output, traders estimating electricity prices, these are all domains where context matters and where a system like this could be used in practice.

Jane: And the paper is honest about limitations. The case study on look-back windows shows that when the retrieved context is misleading, the forecast suffers. So the framework is only as good as the evidence it finds. That's a real constraint.

Lu: But it also points to the next step. If you can build better context retrieval, the framework gets better. The architecture is solid. The improvements will come from feeding it better information.

Tom: Before we say goodbye, let's give the paper its full name one more time. "CastFSR: A Fast–Slow–Reflect Agentic Reasoning Framework for Context-Aware Time Series Forecasting." It's a framework that treats forecasting as a thinking process, not just a calculation.

Jane: And that's a shift worth paying attention to. We're moving from models that predict to models that reason about what they're predicting.

Tom: Thanks for joining us, everyone. We'll be back with the next paper soon. Until then, keep questioning the numbers.

Jane: And remember, context matters. See you next time.

More episodes

← Home