Macroeconomic Forecasting with Large Language Models

summary

Video file (mp4)

The gist

This paper presents a comparative analysis evaluating the accuracy of Time Series Language Models (TSLMs) against traditional macro time series forecasting approaches.

In short

The episode discusses a paper comparing large language models to traditional econometric models for macroeconomic forecasting. The authors tested five language models against standard methods using 120 variables and found that most were unreliable alone. The key contribution is a hybrid approach, TSLM-BVAR, which combines the econometric model's reliability with the language model's ability to capture nonlinear patterns in the errors.

Key concepts

Macroeconomic Forecasting with Large Language Models
This paper investigates whether large language models can predict economic indicators like unemployment and inflation better than traditional economic models. The authors tested five different language models against standard econometric tools using a database of over a hundred monthly macroeconomic variables.
Hybrid Approach (TSLM-BVAR)
This proposed method combines two approaches: first, running a Bayesian VAR to get the main forecast, and second, feeding the model's errors (residuals) into a language model. The language model then tries to find patterns in those errors to correct the initial forecast.
Reliability vs. Flexibility
The traditional econometric models provided steady, reliable forecasts with tight error distributions across many variables. In contrast, the language models showed high flexibility and could be better on some series but produced wild, erratic forecasts on others.

Terminology used across episodes

This episode discusses

The paper

Macroeconomic Forecasting with Large Language Models · Read on arXiv

Andrea Carriero, Davide Pettenuzzo, Shubhranshu Shekhar

Queen Mary University of London · Brandeis University

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Macroeconomic Forecasting with Large Language Models".

Jane: The paper was written by Andrea Carriero, Davide Pettenuzzo and Shubhranshu Shekhar from Queen Mary University of London and Brandeis University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back, everyone. We're looking at a fascinating new paper today, and it's called "Macroeconomic Forecasting with Large Language Models." Jane, this one feels like it's right at the intersection of two worlds that don't usually talk to each other.

Jane: Absolutely, Tom. On one side you have the traditional economists with their models that have been refined over decades, and on the other side you have these massive language models that have been trained on basically everything. The paper asks a pretty direct question: can these new models actually predict things like unemployment, inflation, and industrial production better than the old methods?

Tom: And the authors here are Andrea Carriero, Davide Pettenuzzo, and Shubhranshu Shekhar. They're not just dipping their toes in either. They took the FRED-MD database, which is this huge collection of over a hundred monthly macroeconomic variables, and they ran a serious comparison.

Jane: Right, and that's what I love about this. They didn't just pick one fancy model and show it working on one variable. They tested five different time series language models against the standard econometric toolkit. Bayesian VARs, factor models, neural networks, even regression trees. It's a proper bake-off.

Tom: A bake-off with a lot at stake, because if these language models can genuinely forecast the economy, that changes who gets to do this kind of work. You wouldn't need a PhD in econometrics to build a forecasting model. You'd just need access to a pretrained model and some data.

Jane: Exactly. But here's the catch that I think is really important for our listeners to understand. These language models are trained on data that includes the very series they're being asked to forecast. So the paper has to deal with this tricky issue of whether the models have already "seen" the answers, which could make them look better than they really are.

Tom: That's the elephant in the room for sure. But the authors are upfront about it, and they even run an experiment where they give the same unfair advantage to the traditional models to level the playing field. We'll get into that in a bit.

Jane: So stick around, because the results are genuinely surprising. Some of these models are competitive, some are not, and the paper ends up proposing a hybrid approach that might be the real winner here.

Tom: Let's get into the details.

Paper discussion segment 2: Tom: So Jane, we set the stage. Now let's talk about what actually happened when they ran these models. The paper's summary is pretty clear: most of the time series language models just aren't good enough on their own.

Jane: Right. Out of the five models they tested, only two consistently beat a simple autoregressive benchmark. That's Salesforce's Moirai and Google's TimesFM. The other three, LagLlama, TTM, and Time-GPT, they struggled. Time-GPT especially had most of its results above the benchmark line, meaning it was often worse than just using a basic AR model.

Tom: And that's a big deal because the AR model is the simplest thing you can do. It just says, "tomorrow will look a lot like today." For a model with billions of parameters and months of training to lose to that, it's a humbling result.

Jane: It is, but it's also not the whole story. When they compared Moirai and TimesFM against the serious econometric models, the Bayesian VARs and factor models, the performance was actually comparable. Not better, but comparable. And there's a really interesting pattern in the errors.

Tom: What pattern is that?

Jane: The econometric models had a tight distribution of errors. They were consistently decent across all one hundred twenty variables. The language models had a wider spread. They could be spectacular on some series, like housing starts, where TimesFM predicted the two thousand eight crash remarkably well. But then they'd also produce these wild forecasts on other series, like interest rate spreads, that were just way off.

Tom: So it's a reliability problem. The language models are like a brilliant but erratic student, while the econometric models are the steady, dependable one.

Jane: Exactly. And there's another layer to this. The paper breaks down the results by how persistent the series is. For highly persistent variables, like interest rates, the language models really struggled. Their forecasts had much more variance. But for less persistent series, they held their own.

Tom: And then there's the pandemic period. When they looked at two thousand twenty to two thousand twenty-two the language models actually did relatively better, especially Moirai. But the authors are careful to point out that these models were trained on data that includes the pandemic, so it's not a fair real-time test.

Tom: Right, and that's the contamination issue we flagged earlier. The models have seen the future, at least partially.

Jane: Precisely. So the summary is: these models are promising, they can capture nonlinear patterns that linear models miss, but they're not ready to replace the econometric toolkit on their own.

Paper discussion segment 3: Jane: So Tom, we've established that the language models are good but unreliable. The paper's real contribution, though, is what they do about that. They propose a hybrid approach they call the TSLM-BVAR.

Tom: And that's the part I found most exciting. Instead of asking which model is better, they ask how to make them work together. The idea is pretty elegant. You run the Bayesian VAR first, get its forecast, and then you look at the residuals, the mistakes it made.

Jane: Right. And then you feed those residuals to the language model. The language model tries to find patterns in the errors that the VAR missed. Then you add that correction back to the VAR's forecast.

Tom: It's like having a careful, experienced colleague do the main work, and then bringing in a creative, unconventional thinker to catch the blind spots. And the results are genuinely impressive.

Jane: They are. The hybrid approach shifts the whole error distribution down. The median relative RMSFE drops below the BVAR alone, and crucially, the worst-case errors shrink dramatically. Remember how TimesFM had those wild forecasts with relative errors of one point four or even one point six? In the hybrid, the maximum comes down to around one point one, right in line with the BVAR.

Tom: So the hybrid keeps the reliability of the econometric model while adding the flexibility of the language model. You get the best of both worlds. The paper shows that the hybrid beats both components when used alone, across most horizons.

Jane: And I think that's the real message here. The future isn't about choosing between traditional econometrics and machine learning. It's about building bridges between them. The domain knowledge embedded in the BVAR, things like how monetary policy affects inflation, that's not something a language model can learn from a generic corpus of time series data.

Tom: But the language model can pick up on nonlinear dynamics and unusual patterns that the BVAR's linear structure would miss. The housing starts example comes to mind. The hybrid approach lets each model do what it's best at.

Jane: And practically speaking, this is huge. It means you don't need to retrain a massive language model from scratch. You can use the pretrained model as-is, just feed it residuals. That's computationally feasible for a lot of organizations that couldn't afford to train their own foundation model.

Tom: So the path forward is clear. We'll see more of these hybrid systems, I think, not just in macro forecasting but across the board.

Conclusion: Tom: Alright, Jane, let's wrap this up. We've spent this whole episode on "Macroeconomic Forecasting with Large Language Models," and I think we've got a clear picture now.

Jane: We do. The paper is a rigorous, honest evaluation. Five language models tested against the standard econometric toolkit on one hundred twenty macroeconomic variables. Most of the language models couldn't beat a simple benchmark, and the two that could, Moirai and TimesFM, were still less reliable than the Bayesian VARs and factor models.

Tom: But the key takeaway isn't that language models are useless for macro forecasting. It's that they're a tool, and like any tool, they work best when used correctly. The hybrid approach they propose, where the language model corrects the residuals of the econometric model, is the real innovation here.

Jane: And it works. The hybrid consistently beats both components in isolation, and it tames those wild forecasts that made the language models risky to use on their own. That's a practical, actionable result.

Tom: So what does this mean for the world? I think it means we're going to see more collaboration between econometricians and machine learning researchers. The days of treating these as competing approaches are over. The best forecasts are going to come from combining the strengths of both.

Jane: And for anyone listening who works with economic data, this is a signal that you don't need to throw away your existing models. You can augment them. The barrier to entry is lower than you might think, because you can use the pretrained language models as-is.

Tom: Alright, we've covered the title, the summary, the improvements, and the implications. Time to say goodbye to this paper and get ready for the next one.

Jane: Thanks for joining us, everyone. We'll be back soon with another paper, another discussion, and hopefully another insight worth taking with you.

Tom: Take care, and keep forecasting.

More episodes

← Home