Projected random forests and conformal prediction of circular data

summary

Video file (mp4)

The gist

The paper applies conformal prediction techniques to regression problems with circular responses, producing prediction sets with adaptive arc length and finite-sample coverage guarantees for any

In short

The episode discusses a paper titled "Projected random forests and conformal prediction of circular data." The hosts explain how to apply machine learning to cyclical data, such as wind direction. They detail a method using projected random forests and out-of-bag predictions to provide precise, reliable prediction intervals for real-world applications.

Key concepts

Circular Data
Data that exists on a circle, like wind direction or time of day. Standard linear models struggle with this because they cannot account for the cyclical nature, such as treating 359 degrees and 1 degree as being very close instead of far apart.
Conformal Prediction
A method that takes any predictive model and wrapping its output in a guarantee. It provides a range of possible answers (an arc) where the true answer is expected to lie, ensuring the prediction set contains the truth at a specified percentage of time.
Projected Random Forests
A technique used for circular data. It converts a circular problem into two linear components: predicting the cosine and sine of the angle. This allows standard random forests to handle the data, then combining them back into an angle using arctangent.
Out-of-Bag Mechanism
A way to measure model error without sacrificing training data. The authors use predictions from trees that were not included in a specific data point's training set, providing a calibration method for the prediction intervals.

Terminology used across episodes

This episode discusses

The paper

Projected random forests and conformal prediction of circular data · Read on arXiv

Paulo C. Marques F., Rinaldo Artes, Helton Graziadei

Insper Institute of Education and Research · Federal University of São Carlos

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Projected random forests and conformal prediction of circular data".

Jane: The paper was written by Paulo C. Marques F., Rinaldo Artes and Helton Graziadei from Insper Institute of Education and Research and Federal University of São Carlos.

Tom: Stay tuned as we take you through the paper and discuss its implications.

First Impressions and the Core Problem: Tom: Welcome back to the show, everyone. Today we're diving into a paper that's got a mouthful of a title: "Projected random forests and conformal prediction of circular data." Jane, I'll be honest, when I first saw "circular data," I thought it was about recycling statistics.

Jane: Ha, that's a good guess, Tom, but it's actually way more interesting. Circular data is data that lives on a circle — think wind directions, the time of day something happens, or even the orientation of a bird in flight. The key issue is that three hundred fifty-nine degrees and one degree are basically the same direction, but a regular linear model would think they're three hundred fifty-eight degrees apart.

Tom: Oh, that's a nasty problem. So if you're predicting wind direction and the true answer is north, a model might say "south" because it's thinking in straight lines, not around a circle.

Jane: Exactly. And that's where this paper comes in. The authors — Paulo Marques F., Rinaldo Artes, and Helton Graziadei — they're tackling this by taking a really powerful tool from machine learning, random forests, and adapting it to handle this circular nature. But they don't just stop at making predictions; they also want to give you a measure of how confident those predictions are.

Tom: And that's where conformal prediction comes in, right? That's the part that gives you a range of possible answers instead of just one guess.

Jane: You got it. Conformal prediction is like a safety net. It takes any predictive model and wraps it in a guarantee: "If you use this prediction set, the true answer will be inside it, say, ninety percent of the time." That's a huge deal for real-world applications where you need to know when the model might be wrong.

Tom: So we're not just getting a point estimate, we're getting a whole arc of possible directions, and we know exactly how often that arc will actually contain the truth. That sounds like the kind of thing meteorologists would love for wind forecasting.

Jane: Or anyone dealing with animal movement, geological fault lines, even psychological traits measured on a circumplex. The applications are broad. But the real magic here is how they make it work without needing a separate calibration dataset, which is usually a big requirement for conformal prediction. They found a clever workaround using the random forest's built-in out-of-bag mechanism.

Tom: Oh, that's slick. So they're getting the safety guarantee without sacrificing precious data. I can't wait to hear how they actually pulled that off. Let's get into the details.

The Method and the Machinery: Jane: So, Tom, we left off with this idea that they're using out-of-bag predictions to avoid needing a separate calibration set. Let's unpack how that actually works, because it's pretty clever.

Tom: Please, walk me through it. I'm picturing a forest of trees, but I need the details.

Jane: So a random forest is a bunch of decision trees, each trained on a random sample of the data, with replacement. That means for any given data point, there's a good chance it was left out of some of those trees' training sets. Those trees are called "out-of-bag" for that point.

Tom: Right, and normally you'd use those out-of-bag predictions to test the model internally. But these authors are using them for something else.

Jane: Exactly. Instead of setting aside a whole chunk of data just to calibrate the prediction intervals, they use the out-of-bag predictions for every point in the training set. This gives them a way to measure how wrong the model is on data it hasn't seen, without sacrificing any of that data for training.

Tom: So they're getting the calibration for free, essentially. But how do they handle the circular part? A regular random forest can't just predict an angle.

Jane: That's the "projected" part of the title. They take the circular response — say, wind direction — and split it into two linear components: the cosine and the sine of the angle. Then they train two separate random forests: one to predict the cosine, and one to predict the sine.

Tom: Oh, I see. So they're turning a circular problem into two linear problems, which standard random forests can handle.

Jane: Precisely. Then, when they want a prediction, they take the two outputs, cosine and sine, and combine them back into an angle using an arctangent function. It's a neat trick that lets them use all the powerful, well-tested machinery of standard random forests on a problem they weren't designed for.

Tom: And it works? I mean, the paper says they tested it against a specialized circular forest and a projected normal linear model.

Jane: It does. On both synthetic data and a real wind direction dataset, their projected random forest produced prediction intervals with a shorter median arc length. That means the prediction sets were tighter, more precise, while still hitting the target coverage rate of about ninety percent.

Tom: So they're getting the same guarantee of catching the true answer ninety percent of the time, but with a smaller, more useful range of directions. That's a win-win. But I'm curious, does this out-of-bag trick actually have a solid theoretical foundation, or is it just an empirical observation?

Jane: That's a great question, and it's actually the subject of the appendix. They provide a theoretical result showing that under certain conditions, the out-of-bag conformal prediction sets will have coverage close to the nominal level. It's not an exact guarantee like the split conformal method, but it's a strong justification for why it works so well in practice.

Tom: So we have a method that's both practical and theoretically grounded. I'm starting to see why this paper is generating buzz. But what does this mean for people actually building systems? Let's bring in Meng to get an engineer's take.

Practical Impact and Engineering Challenges: Meng: Hey Tom, Jane. I've been listening in, and I have to say, the part that excites me most is the practical side. We're always looking for ways to get uncertainty estimates without paying a huge data penalty.

Jane: That's a great point, Meng. The fact that you don't need to hold out a calibration set is huge for small datasets, where every observation counts.

Meng: Exactly. And the projection trick is something we can implement right now. We have solid libraries for random forests, and the cosine/sine transformation is just a few lines of code. It's not some exotic new algorithm we have to build from scratch.

Tom: So it's practical to deploy. But what about the computational cost? You're training two random forests for the mean prediction, and then two more for the variability model. That's four forests total.

Meng: Right, it's not free. But random forests are embarrassingly parallel, so you can train those four forests on separate cores or machines. The wall-clock time might not be that bad. And the payoff is a model that's specifically tuned for circular data, which is something we didn't have easy access to before.

Lu: If I can jump in here, Meng. The computational cost is real, but the bigger win is the flexibility. This projection approach isn't just for random forests. The paper explicitly says it's a general procedure that can turn any linear-response model into a circular one. You could use gradient boosting, neural networks, anything.

Meng: That's a good point, Lu. So the core idea is a template, not a one-off solution. That makes it much more valuable for a production environment where you might want to experiment with different base models.

Jane: And the wind direction dataset in the paper shows it's not just a toy problem. They used real hourly weather data from Brazil, with features like temperature, humidity, and wind speed from the previous hour. That's a real-world scenario.

Lu: And it makes me think about the broader implications. This isn't just about wind. Think about any system that deals with cyclical patterns — traffic flow at different times of day, energy demand cycles, even biological rhythms. The ability to give a calibrated prediction interval for these kinds of problems could be transformative.

Meng: Sure, but I want to make sure the coverage guarantee holds up in the messy real world. The paper shows it works on this dataset, but what about when the data isn't perfectly exchangeable? Say, if there's a seasonal drift?

Jane: That's a fair concern, Meng. The theoretical guarantee relies on exchangeability, which is a bit stronger than just independence. But the paper does show a calibration plot for the wind data, and the empirical coverage tracks the nominal level quite well across different confidence levels. It's a good sign, but more testing in different conditions would be needed.

Tom: So we have a practical method, a theoretical backing, and a real-world demonstration. What's the big picture here? Lu, you mentioned broader implications. Let's dig into that.

The Big Picture and Cultural Impact: Lu: I love this paper because it's a perfect example of how good engineering can unlock new scientific possibilities. We've had circular statistics for decades, but it was often seen as a niche, specialized field. This paper makes it accessible to anyone who knows how to use a random forest.

Tom: So it democratizes the toolset. You don't need to be a directional statistics expert to get good results on circular data.

Lu: Exactly. And that's where I see the cultural impact. Consider how we understand health. Sleep cycles are circular. Circadian rhythms are circular. If we can build better models that give us confidence intervals for these cycles, we can build better health apps that tell you not just "you slept at eleven PM," but "your sleep onset is likely between ten:thirty and eleven:thirty PM, with ninety percent confidence."

Jane: That's a powerful vision. It's about moving from point predictions to a range of possibilities, which is how we actually experience the world. Nothing is ever a single exact moment.

Lalam: If I may add to that, Lu. The cultural shift here is about embracing uncertainty in a constructive way. In many fields, from weather forecasting to public health, we're often forced to give a single number because that's what the tools produce. This approach gives us a vocabulary for saying "we're not sure, but here's the range where the truth likely lies."

Meng: And that's not just a scientific benefit. It's a communication benefit. When a forecast says "there's a ninety percent chance the wind will be from this direction," people can make better decisions about, say, flying a drone or planning a construction job.

Tom: So it's about building trust through honesty. Instead of pretending we know the exact answer, we're giving people a reliable measure of our uncertainty.

Jane: And that's the real takeaway from "Projected random forests and conformal prediction of circular data." It's not just a new algorithm; it's a philosophy of how to make predictions more useful and more honest.

Lu: And it opens the door for future work. The paper mentions extending this to directional data in three dimensions, which would be huge for things like tracking the orientation of objects in space or modeling animal movement in three dee.

Meng: I'd also love to see this integrated into more automated machine learning pipelines, where the choice of base model is optimized automatically. The projection trick is so clean, it could easily become a standard module.

Tom: Well, it's been a fantastic discussion. We've covered the problem of circular data, the clever projection solution, the out-of-bag trick for calibration, and the broader implications for how we handle uncertainty.

Jane: And we should say goodbye to this paper. It's given us a lot to think about. For our listeners, if you work with any kind of cyclical data, this is a must-read.

Tom: Absolutely. Thanks to Lu, Meng, and Lalam for joining us. And thanks to our listeners for tuning in. We'll be back soon with another paper. Until then, keep your predictions calibrated and your intervals honest.

More episodes

← Home