Macroeconomic Forecasting with Large Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Macroeconomic Forecasting with Large Language Models".
Jane: The paper was written by Andrea Carriero, Davide Pettenuzzo and Shubhranshu Shekhar from Queen Mary University of London and Brandeis University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back, everyone. We're looking at a fascinating new paper today, and it's called "Macroeconomic Forecasting with Large Language Models." Jane, this one feels like it's right at the intersection of two worlds that don't usually talk to each other.
Jane: Absolutely, Tom. On one side you have the traditional economists with their models that have been refined over decades, and on the other side you have these massive language models that have been trained on basically everything. The paper asks a pretty direct question: can these new models actually predict things like unemployment, inflation, and industrial production better than the old methods?
Tom: And the authors here are Andrea Carriero, Davide Pettenuzzo, and Shubhranshu Shekhar. They're not just dipping their toes in either. They took the FRED-MD database, which is this huge collection of over a hundred monthly macroeconomic variables, and they ran a serious comparison.
Jane: Right, and that's what I love about this. They didn't just pick one fancy model and show it working on one variable. They tested five different time series language models against the standard econometric toolkit. Bayesian VARs, factor models, neural networks, even regression trees. It's a proper bake-off.
Tom: A bake-off with a lot at stake, because if these language models can genuinely forecast the economy, that changes who gets to do this kind of work. You wouldn't need a PhD in econometrics to build a forecasting model. You'd just need access to a pretrained model and some data.
Jane: Exactly. But here's the catch that I think is really important for our listeners to understand. These language models are trained on data that includes the very series they're being asked to forecast. So the paper has to deal with this tricky issue of whether the models have already "seen" the answers, which could make them look better than they really are.
Tom: That's the elephant in the room for sure. But the authors are upfront about it, and they even run an experiment where they give the same unfair advantage to the traditional models to level the playing field. We'll get into that in a bit.
Jane: So stick around, because the results are genuinely surprising. Some of these models are competitive, some are not, and the paper ends up proposing a hybrid approach that might be the real winner here.
Tom: Let's get into the details.
Paper discussion segment 2: Tom: So Jane, we set the stage. Now let's talk about what actually happened when they ran these models. The paper's summary is pretty clear: most of the time series language models just aren't good enough on their own.
Jane: Right. Out of the five models they tested, only two consistently beat a simple autoregressive benchmark. That's Salesforce's Moirai and Google's TimesFM. The other three, LagLlama, TTM, and Time-GPT, they struggled. Time-GPT especially had most of its results above the benchmark line, meaning it was often worse than just using a basic AR model.
Tom: And that's a big deal because the AR model is the simplest thing you can do. It just says, "tomorrow will look a lot like today." For a model with billions of parameters and months of training to lose to that, it's a humbling result.
Jane: It is, but it's also not the whole story. When they compared Moirai and TimesFM against the serious econometric models, the Bayesian VARs and factor models, the performance was actually comparable. Not better, but comparable. And there's a really interesting pattern in the errors.
Tom: What pattern is that?
Jane: The econometric models had a tight distribution of errors. They were consistently decent across all one hundred twenty variables. The language models had a wider spread. They could be spectacular on some series, like housing starts, where TimesFM predicted the two thousand eight crash remarkably well. But then they'd also produce these wild forecasts on other series, like interest rate spreads, that were just way off.
Tom: So it's a reliability problem. The language models are like a brilliant but erratic student, while the econometric models are the steady, dependable one.
Jane: Exactly. And there's another layer to this. The paper breaks down the results by how persistent the series is. For highly persistent variables, like interest rates, the language models really struggled. Their forecasts had much more variance. But for less persistent series, they held their own.
Tom: And then there's the pandemic period. When they looked at two thousand twenty to two thousand twenty-two the language models actually did relatively better, especially Moirai. But the authors are careful to point out that these models were trained on data that includes the pandemic, so it's not a fair real-time test.
Tom: Right, and that's the contamination issue we flagged earlier. The models have seen the future, at least partially.
Jane: Precisely. So the summary is: these models are promising, they can capture nonlinear patterns that linear models miss, but they're not ready to replace the econometric toolkit on their own.
Paper discussion segment 3: Jane: So Tom, we've established that the language models are good but unreliable. The paper's real contribution, though, is what they do about that. They propose a hybrid approach they call the TSLM-BVAR.
Tom: And that's the part I found most exciting. Instead of asking which model is better, they ask how to make them work together. The idea is pretty elegant. You run the Bayesian VAR first, get its forecast, and then you look at the residuals, the mistakes it made.
Jane: Right. And then you feed those residuals to the language model. The language model tries to find patterns in the errors that the VAR missed. Then you add that correction back to the VAR's forecast.
Tom: It's like having a careful, experienced colleague do the main work, and then bringing in a creative, unconventional thinker to catch the blind spots. And the results are genuinely impressive.
Jane: They are. The hybrid approach shifts the whole error distribution down. The median relative RMSFE drops below the BVAR alone, and crucially, the worst-case errors shrink dramatically. Remember how TimesFM had those wild forecasts with relative errors of one point four or even one point six? In the hybrid, the maximum comes down to around one point one, right in line with the BVAR.
Tom: So the hybrid keeps the reliability of the econometric model while adding the flexibility of the language model. You get the best of both worlds. The paper shows that the hybrid beats both components when used alone, across most horizons.
Jane: And I think that's the real message here. The future isn't about choosing between traditional econometrics and machine learning. It's about building bridges between them. The domain knowledge embedded in the BVAR, things like how monetary policy affects inflation, that's not something a language model can learn from a generic corpus of time series data.
Tom: But the language model can pick up on nonlinear dynamics and unusual patterns that the BVAR's linear structure would miss. The housing starts example comes to mind. The hybrid approach lets each model do what it's best at.
Jane: And practically speaking, this is huge. It means you don't need to retrain a massive language model from scratch. You can use the pretrained model as-is, just feed it residuals. That's computationally feasible for a lot of organizations that couldn't afford to train their own foundation model.
Tom: So the path forward is clear. We'll see more of these hybrid systems, I think, not just in macro forecasting but across the board.
Conclusion: Tom: Alright, Jane, let's wrap this up. We've spent this whole episode on "Macroeconomic Forecasting with Large Language Models," and I think we've got a clear picture now.
Jane: We do. The paper is a rigorous, honest evaluation. Five language models tested against the standard econometric toolkit on one hundred twenty macroeconomic variables. Most of the language models couldn't beat a simple benchmark, and the two that could, Moirai and TimesFM, were still less reliable than the Bayesian VARs and factor models.
Tom: But the key takeaway isn't that language models are useless for macro forecasting. It's that they're a tool, and like any tool, they work best when used correctly. The hybrid approach they propose, where the language model corrects the residuals of the econometric model, is the real innovation here.
Jane: And it works. The hybrid consistently beats both components in isolation, and it tames those wild forecasts that made the language models risky to use on their own. That's a practical, actionable result.
Tom: So what does this mean for the world? I think it means we're going to see more collaboration between econometricians and machine learning researchers. The days of treating these as competing approaches are over. The best forecasts are going to come from combining the strengths of both.
Jane: And for anyone listening who works with economic data, this is a signal that you don't need to throw away your existing models. You can augment them. The barrier to entry is lower than you might think, because you can use the pretrained language models as-is.
Tom: Alright, we've covered the title, the summary, the improvements, and the implications. Time to say goodbye to this paper and get ready for the next one.
Jane: Thanks for joining us, everyone. We'll be back soon with another paper, another discussion, and hopefully another insight worth taking with you.
Tom: Take care, and keep forecasting.
Andrea Carriero, Davide Pettenuzzo, Shubhranshu Shekhar
Queen Mary University of London · Brandeis University
econ.EM, cs.CL, cs.LG
Submitted: 2026-08-17
Updated: 2026-08-18
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 79/100
The gist: This paper presents a comparative analysis evaluating the accuracy of Time Series Language Models (TSLMs) against traditional macro time series forecasting approaches.
Key concepts
- Macroeconomic Forecasting with Large Language Models
- This paper investigates whether large language models can predict economic indicators like unemployment and inflation better than traditional economic models. The authors tested five different language models against standard econometric tools using a database of over a hundred monthly macroeconomic variables.
- Hybrid Approach (TSLM-BVAR)
- This proposed method combines two approaches: first, running a Bayesian VAR to get the main forecast, and second, feeding the model's errors (residuals) into a language model. The language model then tries to find patterns in those errors to correct the initial forecast.
- Reliability vs. Flexibility
- The traditional econometric models provided steady, reliable forecasts with tight error distributions across many variables. In contrast, the language models showed high flexibility and could be better on some series but produced wild, erratic forecasts on others.
Terminology
Summary
This paper presents a comparative analysis evaluating the accuracy of Time Series Language Models (TSLMs) against traditional macro time series forecasting approaches. The authors conduct a rigorous evaluation of TSLMs against traditional macro-prediction methods, using as common ground the FRED-MD database,
finding that their results provide valuable insights into the strengths and limitations of TSLMs in the prediction of macroeconomic time series, shedding light on their applicability in real world scenarios.
Building on these insights, they "propose a novel hybrid approach that integrates the structure and interpretability of econometric models with the flexibility and nonlinearity of TSLMs, yielding the most robust and accurate forecasts across a range of horizons and variables."
The paper has three main contributions. First, it provides a thorough investigation on how TSLMs perform in predicting macroeconomic time series.
Second, it offers "a detailed comparison of their performance versus state-of-the-art time series methods such as Bayesian Vector Autoregressions (BVARs) and Factor Models, as well as Neural Networks and Bayesian Additive Regression Trees (BART), which have both become very popular nonlinear models for macroeconomic forecasting. Third, it develops
a novel hybrid approach that combines the structural rigor of traditional econometric models with the adaptability of TSLMs, leading to consistent improvements in forecast accuracy across variables and horizons."
The empirical application focuses on forecasting all variables contained in the FRED monthly database using data ranging from 1960 to 2022.
The authors evaluate five TSLMs: LagLlama, Moirai, TTM (Tiny Time Mixers), Time-GPT, and TimesFM. The picture emerging from the analysis is that only two of the five TSLMs we consider are competitive against a simple AR benchmark (Salesforce's Moirai and Google's TimesFM).
Moreover, when these more competitive TSLMs are stacked against Bayesian VARs and Factor Models, their forecasting performance is broadly comparable, if not slightly inferior.
The authors find that "the forecasting gains achieved by the econometric models tend to be more stable while TSLMs can perform very well for a handful of series but also show less reliability at times, as they seem to be more prone to generating the occasional unreasonable forecast. On the other hand,
TSLMs seem to work relatively better in the post-Covid-19 era (with the important caveat that the training set of these models does include information from the pandemic and post-pandemic period). The results are based on zero-shot forecasting, and the authors find that
fine-tuning the TSLM models does not yield to significant improvements in their forecast accuracy."
Regarding the econometric models, both the BVAR and the factor model are consistently showing a good performance relative to the AR model.
The performance of BART is consistently worse than that of the benchmark,
while the NNAR model produces better forecast than BART, and is overall slightly better than the AR(1) benchmark.
At longer horizons, NNAR is outperformed by both TSLMs and the linear econometric models,
suggesting that while non-linearities do play an important role, TSLMs seem to have an edge at capturing them, over the simpler non-linear models we considered.
The authors investigate the impact of different information sets on forecast accuracy, noting that some of the TSLM models do include in the training set the FRED-MD database, which means they potentially have access to the perfect forecast.
When re-estimating the econometric models using the entire sample up to December 2022, the forecasting performance of BVAR and Factor models is dramatically enhanced by the use of the extra information, with their respective box plots lying entirely below one,
and they "perform significantly better than the TSLMs, which appeared to perform comparably to the econometric models only when the latter were estimated without the benefit of hindsight."
The paper also examines how TSLMs handle persistent series. For series with low to moderate persistence, they seem to perform relatively better, still on par with BVARs and factor models.
However, when considering highly persistent variables, one can observe a definite increase in the variation of the performance of TSLM,
with the same pattern applying to the non-linear econometric models NNAR and BART.
The authors propose a TSLM-Augmented BVAR (TSLM-BVAR) hybrid approach with three steps: (1) estimate a BVAR and generate forecasts, computing in-sample residuals; (2) train a foundational time series model on the residuals and generate a zero-shot forecast; (3) construct the hybrid forecast by combining the econometric prediction and the TSLM forecast. The results show that the hybrid approach shifts the TSLM error distribution downward, placing most of its mass below one—the threshold for outperforming the AR benchmark.
The maximum relative RMSFEs from the TSLMs are greatly reduced: at one step ahead, the error of Moirai decreases from 1.436 to 1.139, and that of TimesFM from 1.482 to 1.150,
figures very close to the maximum error of the BVAR at 1.158. Overall, the hybrid approach improves upon both the econometric model and the TSLMs used in isolation,
as using the econometric model in the first step prevents TSLMs from producing subpar forecasts, while using the TSLMs on the residuals enhances the performance of the econometric model.
The authors conclude that while TSLMs can offer valuable insights, they are not yet a clear replacement for state-of-the-art econometric models.
A key challenge is the lack of control over their training data,
as "many of these models are pretrained on the FRED-MD dataset, hence already contain the macroeconomic series that are the focus of our forecasting exercise, introducing potential biases and making it difficult to conduct a clean pseudo out-of-sample forecasting exercise. Despite these limitations,
TSLMs exhibit some promising features, particularly in capturing nonlinearities and adapting to evolving economic conditions, and
the TSLM-BVAR approach demonstrates that integrating structured econometric information with flexible machine learning models can substantially improve forecasting performance."
Improvements for AI systems
Based on the paper, here are specific improvements I can implement in AI systems, along with what the improved system can do:
Improvement: Implement a two-stage forecasting pipeline where:
-
Stage 1: A Bayesian VAR (BVAR) generates baseline forecasts and in-sample residuals
-
Stage 2: A TSLM (e.g., Moirai-large or TimesFM) models the residuals to capture nonlinear patterns
-
Stage 3: Final forecast = BVAR forecast + TSLM residual prediction
What the improved system can do:
-
Reduce maximum forecast errors by 20-30% compared to using TSLMs alone (e.g., from RMSFE 1.48 down to 1.15)
-
Eliminate
hallucination
episodes where TSLMs produce unreasonable forecasts -
Maintain interpretability of econometric models while capturing nonlinear dynamics
-
Improve left-tail performance, achieving up to 12% better forecasts than BVAR alone for series where TSLMs excel
Sources
- GPT-4 Technical Report
- Chronos: Learning the Language of Time Series
- Surveying Generative AI's Economic Expectations
- ChatGPT and Deepseek: Can They Predict the Stock Market and Macroeconomy?
- Tiny Time Mixers (TTMs): Fast Pre-trained Models for Enhanced Zero/Few-Shot Forecasting of Multivariate Time Series
- On $\tau$-tilting finite Borel-Schur algebras
- Monash Time Series Forecasting Archive
- Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey
- LibCity: A Unified Library Towards Efficient and Comprehensive Urban Spatial-Temporal Prediction
- Fine-Tuning can Distort Pretrained Features and Underperform Out-of-Distribution
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Transformers in Time Series: A Survey
- Unified Training of Universal Time Series Forecasting Transformers
Related papers
- SLIM: Stochastic Learning and Inference in Overidentified Models
- High-dimensional censored MIDAS logistic regression for corporate survival forecasting
- Cross-Fitting-Free Debiased Machine Learning with Multiway Dependence
- Can large language models assist choice modelling? Insights into prompting strategies and current models' capabilities
- Causal Inference in Possibly Nonlinear Factor Models
- Mining Causality: AI-Assisted Search for Instrumental Variables