neuralGAM: An R Package for Fitting Generalized Additive Neural Networks
summary
The gist
The neuralGAM package implements a Neural Network topology based on Generalized Additive Models, allowing users to fit an independent Neural Network to estimate the contribution of each feature to
In short
The episode discusses 'neuralGAM,' an R package by Ines Ortega-Fernandez and Marta Sestelo. It addresses the trade-off between neural network power and interpretability by implementing a Generalized Additive Neural Network. The hosts conclude that this tool allows users to achieve both high predictive performance and transparency, making it valuable for real-world decision-making.
Key concepts
- Generalized Additive Models (GAM)
- A classic statistical tool where the model is built by adding up the contributions of individual variables. This structure allows users to see exactly how each feature affects the final outcome, providing high interpretability.
- Backfitting Algorithm
- An iterative approach used in neuralGAM. Instead of training one large network, it updates each small, independent neural network sequentially using the part of the response that hasn't been explained by previous networks.
Terminology used across episodes
This episode discusses
- neuralGAM: An R Package for Fitting Generalized Additive Neural Networks · Paper Radio
- Adam: A Method for Stochastic Optimization
- Searching for Activation Functions
- Intriguing properties of neural networks
The paper
neuralGAM: An R Package for Fitting Generalized Additive Neural Networks · Read on arXiv
Ines Ortega-Fernandez, Marta Sestelo
Galician Research and Development Center in Advanced Telecommunications · Galician Centre for Mathematical Research and Technology · University of Vigo
Nowadays, Neural Networks are considered one of the most effective methods for various tasks such as anomaly detection, computer-aided disease detection, or natural language processing. However, these networks suffer from the ``black-box'' problem which makes it difficult to understand how they make decisions. In order to solve this issue, an R package called neuralGAM is introduced. This package implements a Neural Network topology based on Generalized Additive Models, allowing to fit an independent Neural Network to estimate the contribution of each feature to the output variable, yielding a highly accurate and interpretable Deep Learning model. The neuralGAM package provides a flexible framework for training Generalized Additive Neural Networks, which does not impose any restrictions on the Neural Network architecture. We illustrate the use of the neuralGAM package in both synthetic and real data examples.
DOI: 10.32614/RJ-2026-016
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "neuralGAM: An R Package for Fitting Generalized Additive Neural Networks".
Jane: The paper was written by Ines Ortega-Fernandez and Marta Sestelo from Galician Research and Development Center in Advanced Telecommunications and Galician Centre for Mathematical Research and Technology and University of Vigo.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title and Authors: Tom: Alright, welcome back to the show, everyone. Today we're looking at a paper that's been making the rounds on arXiv, and it's called "neuralGAM: An R Package for Fitting Generalized Additive Neural Networks." Jane, I gotta say, just the title alone tells you we're in for something interesting.
Jane: Absolutely, Tom. And the authors here are Ines Ortega-Fernandez from the Gradiant research center in Spain, and Marta Sestelo from the University of Vigo. They've put together something that's trying to bridge two worlds that don't usually talk to each other.
Tom: Two worlds that don't talk? Which ones are those?
Jane: Well, on one side you've got the statisticians who love their interpretable models — the kind where you can see exactly what each variable is doing. And on the other side, you've got the deep learning folks who love neural networks because they can learn almost anything, but they're basically black boxes.
Tom: And this package is trying to make them shake hands?
Jane: Exactly. The name is a clue. GAM stands for Generalized Additive Models, which is a classic statistical tool. And neural — well, that's the neural network part. So they're taking the structure of a GAM and replacing the traditional smooth functions with little neural networks.
Tom: So instead of one big neural network that looks at everything at once, you get a bunch of small ones, each responsible for one feature?
Jane: You got it. And that's what makes it interpretable. You can look at each small network and see exactly how that feature affects the outcome, because the model forces them to work independently and then just adds up their contributions.
Tom: That sounds like a really clever way to get the best of both worlds. And the fact that it's an R package means it's accessible to a huge community of data scientists who might not want to switch to Python just to use neural networks.
Jane: Right, and that's a big deal. R is still the language of choice for a lot of statistical modeling, especially in academia and fields like biostatistics. So having a tool like this available there could really change how people approach their problems.
Tom: I'm curious, Jane, is this the first time anyone's tried this? Combining GAMs with neural networks?
Jane: No, there have been attempts before. The paper actually mentions a few, like Neural Additive Models from Google, and something called GAMI-Net. But those are in Python. And there are some R packages that do similar things, like deepregression. But the authors argue that their approach is different because it uses completely independent networks trained through a classic algorithm called backfitting.
Tom: So they're not just porting someone else's idea to R, they've got their own twist on it.
Jane: That seems to be the claim. And we'll get into the details of how it actually works as we go through the paper. But for now, I think the big picture is that this is about making powerful models more transparent, which is something that matters more and more as these tools get used in real-world decisions.
Tom: Transparency in AI is definitely a hot topic. And this package seems like a step in that direction. Alright, let's dig into the actual summary and see what the authors are promising.
Summary of the Paper: Jane: So we've got the title and the authors down. Now let's talk about what this paper is actually claiming to do. The summary tells us that neuralGAM is a package that lets you fit a Generalized Additive Neural Network, and the core idea is that you train an independent neural network for each feature in your model.
Tom: Right, and I remember you said that's what makes it interpretable. But how does that actually work in practice? Like, if I have three features, I just train three separate networks?
Jane: Almost. The trick is that you can't just train them separately and then add up their outputs, because each network might try to explain the same variation in the response. So the authors use a clever iterative approach. They start by making a guess, then they update each network one at a time, using the part of the response that hasn't been explained yet.
Tom: So it's like each network gets a turn to explain what it can, and then the next one gets the leftovers?
Jane: Exactly. That's the backfitting algorithm, and it's been around in statistics for decades. The authors are essentially taking that old algorithm and using neural networks as the smoother inside it, instead of the traditional splines.
Tom: And what kinds of problems can this handle? Is it just for continuous outcomes, or can it do classification too?
Jane: The paper says it supports three families: Gaussian for continuous data, binomial for binary classification, and Poisson for count data. So it covers a lot of the common use cases. And the link functions — the way you connect the linear predictor to the response — those are standard too, like logit for binomial and log for Poisson.
Tom: That's pretty comprehensive. And the summary also mentions something about uncertainty. That's usually a weak spot for neural networks, right?
Jane: Big time. Neural networks are great at predictions, but they're not great at telling you how confident they are. The authors address this using something called Monte Carlo Dropout. Basically, during prediction, you randomly turn off some neurons and run the network multiple times. The variation in the outputs gives you an estimate of the uncertainty.
Tom: So you get confidence intervals for each feature's effect. That's huge for interpretability.
Jane: It really is. You can plot each feature's contribution with a shaded band around it, and you can see where the model is confident and where it's not. That's something you just don't get from a typical black-box neural network.
Tom: And the summary also mentions that the package is built on top of Keras and TensorFlow, so it's using serious deep learning infrastructure.
Jane: Right, which means you get all the flexibility of modern neural networks — different architectures, activation functions, regularizers — but within a framework that keeps the model interpretable.
Tom: I'm starting to see why this could be a big deal. But I want to know more about the actual methodology. How does this backfitting thing really work under the hood?
Improvements Suggested by the Paper: Jane: So we've covered the basics, but the paper also talks about improvements over existing methods. And I think this is where it gets really interesting, Tom.
Tom: I'm all ears. What are they improving on?
Jane: Well, the paper compares itself to a few other approaches. There's deepregression, which is another R package that combines GAMs with neural networks. And there's also the black-box neural network approach, which is just a regular feedforward network. The authors argue that their method offers a better balance between interpretability and flexibility.
Tom: So they're saying they beat deepregression?
Jane: Not exactly beat, but they're offering something different. The paper shows results on both simulated and real data. On the simulated data, neuralGAM actually recovers the true underlying functions more accurately than deepregression. The shapes are smoother and closer to what was actually generating the data.
Tom: And on the real data? They used that flight delay dataset, right?
Jane: Yeah, the NYC flights data. They're trying to predict whether a flight will be delayed based on things like departure delay, air time, temperature, and humidity. And here's the interesting part: deepregression actually got a slightly higher AUC — that's a measure of classification performance — than neuralGAM.
Tom: So deepregression won on the real data?
Jane: On raw predictive performance, yes. But the gap was small. And the authors make a really good point about what you're giving up. The black-box neural network also performed well, but you can't see what it's doing. With neuralGAM, you get these beautiful plots showing exactly how each feature affects the probability of a delay.
Tom: So it's a trade-off between a tiny bit of performance and a whole lot of understanding?
Jane: Exactly. And for many real-world applications, that understanding is critical. If you're in healthcare or finance or any regulated industry, you need to be able to explain why your model made a particular prediction. A black-box model, no matter how accurate, is hard to deploy.
Tom: That makes sense. And the paper also mentions some flexibility improvements, right? Like you can customize the architecture for each feature individually?
Jane: Yes, that's a nice touch. You can say, "I want a two-layer network for this feature, but a single layer for that one." You can even use different activation functions. So you have a lot of control over how each feature is modeled.
Tom: And they also added a diagnostic function, which is something you don't always see in neural network packages.
Jane: Right, the diagnose function gives you residual plots and QQ plots, so you can check whether your model assumptions are holding. That's very much in the spirit of traditional statistical modeling, and it's a nice bridge between the two worlds.
Tom: So the improvements are really about making neural networks more like proper statistical models — with diagnostics, uncertainty, and interpretability.
Jane: That's the core message. And I think that's a really valuable contribution, because it gives practitioners a tool that's both powerful and trustworthy.
First Page Discussion: Tom: Alright, we've talked about the summary and the improvements. Let's go back to the very beginning of the paper and look at what the authors set up as their motivation. Jane, what's the big problem they're trying to solve?
Jane: The first page really sets the stage by talking about the black-box problem in neural networks. They mention that while neural networks are incredibly effective, it's often hard to understand how they make decisions. And that's a real issue when you're using these models in high-stakes situations.
Tom: And they mention two approaches to fix that, right? Post-hoc and ante-hoc?
Jane: Exactly. Post-hoc methods try to explain a black-box model after it's been trained. You might use something like SHAP values or LIME to figure out what the model is paying attention to. But those are approximations, and they don't always capture what's really going on.
Tom: And ante-hoc is the opposite?
Jane: Ante-hoc means you build the interpretability in from the start. You design the model so that it's inherently transparent. And that's what neuralGAM does. Instead of training a big black box and then trying to explain it, they train a model that's additive by construction.
Tom: So the interpretability isn't an afterthought, it's baked into the architecture.
Jane: Right. And the paper also gives a nice history of this idea. They mention that people have been trying to combine GAMs and neural networks since the late nineties. There was a model by Potts in one thousand nine hundred ninety-nine and then more recent ones like Neural Additive Models from Google in two thousand twenty-one.
Tom: So this isn't a new idea, but the authors think their implementation is better?
Jane: They argue that their approach is different because of how they train the networks. They use independent networks for each feature and train them through the local scoring and backfitting algorithms. This is a more classical statistical approach, and they claim it gives smoother, more stable estimates.
Tom: And they also mention that this is, as far as they know, the only R package that does this — using fully independent deep neural networks per feature with backfitting.
Jane: That's a strong claim, but it seems to hold up. The other R packages they compare against, like deepregression, take a different approach. They embed spline bases directly into a neural network architecture, which is clever but doesn't give you the same kind of per-feature network independence.
Tom: So neuralGAM is filling a specific niche in the R ecosystem.
Jane: Exactly. And I think that's important, because R is still the language of choice for a lot of statistical work. Having this tool available there could really help bridge the gap between traditional statistics and modern deep learning.
Tom: I also noticed they mention that the package relies on Keras and TensorFlow. So users get the full power of those frameworks, but within a more interpretable structure.
Jane: Right. And they've even provided a helper function to install all the Python dependencies, which is a nice touch for R users who might not be familiar with setting up Python environments.
Tom: So the first page really sets up the problem, gives the history, and positions their contribution. It's a solid foundation for the rest of the paper.
Conclusion: Tom: Well, we've covered a lot of ground today on "neuralGAM: An R Package for Fitting Generalized Additive Neural Networks." Jane, what's the big takeaway for our listeners?
Jane: I think the big takeaway is that you don't have to choose between interpretability and performance anymore. This package shows that you can have a neural network that's both powerful and transparent, and that's a really valuable combination.
Tom: And it's available in R, which means a whole community of statisticians can start using it without having to learn a whole new language.
Jane: Exactly. The authors have done a great job of packaging this up with all the tools you need — visualization, uncertainty estimation, diagnostics. It's not just a research prototype, it's a usable tool.
Tom: We also saw that while it might not always beat other methods on raw predictive performance, it offers something they don't: the ability to see exactly what each feature is doing.
Jane: And in many real-world applications, that understanding is worth more than a small boost in accuracy. If you're making decisions that affect people's lives, you need to be able to explain those decisions.
Tom: The authors also laid out some interesting future directions — things like multi-class classification, interaction terms, and even conformal prediction for better uncertainty estimates.
Jane: Those would be great additions. But even as it stands now, this is a solid contribution to the field of interpretable machine learning.
Tom: Alright, let's say goodbye to this paper and get ready for the next one. Thanks for joining us, and we'll see you next time.
Jane: Take care, everyone.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language