neuralGAM: An R Package for Fitting Generalized Additive Neural Networks

arXiv:2505.08610 · stat.ML, cs.LG, stat.CO, stat.ME · Submitted 2026-08-09 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "neuralGAM: An R Package for Fitting Generalized Additive Neural Networks".

Jane: The paper was written by Ines Ortega-Fernandez and Marta Sestelo from Galician Research and Development Center in Advanced Telecommunications and Galician Centre for Mathematical Research and Technology and University of Vigo.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title and Authors: Tom: Alright, welcome back to the show, everyone. Today we're looking at a paper that's been making the rounds on arXiv, and it's called "neuralGAM: An R Package for Fitting Generalized Additive Neural Networks." Jane, I gotta say, just the title alone tells you we're in for something interesting.

Jane: Absolutely, Tom. And the authors here are Ines Ortega-Fernandez from the Gradiant research center in Spain, and Marta Sestelo from the University of Vigo. They've put together something that's trying to bridge two worlds that don't usually talk to each other.

Tom: Two worlds that don't talk? Which ones are those?

Jane: Well, on one side you've got the statisticians who love their interpretable models — the kind where you can see exactly what each variable is doing. And on the other side, you've got the deep learning folks who love neural networks because they can learn almost anything, but they're basically black boxes.

Tom: And this package is trying to make them shake hands?

Jane: Exactly. The name is a clue. GAM stands for Generalized Additive Models, which is a classic statistical tool. And neural — well, that's the neural network part. So they're taking the structure of a GAM and replacing the traditional smooth functions with little neural networks.

Tom: So instead of one big neural network that looks at everything at once, you get a bunch of small ones, each responsible for one feature?

Jane: You got it. And that's what makes it interpretable. You can look at each small network and see exactly how that feature affects the outcome, because the model forces them to work independently and then just adds up their contributions.

Tom: That sounds like a really clever way to get the best of both worlds. And the fact that it's an R package means it's accessible to a huge community of data scientists who might not want to switch to Python just to use neural networks.

Jane: Right, and that's a big deal. R is still the language of choice for a lot of statistical modeling, especially in academia and fields like biostatistics. So having a tool like this available there could really change how people approach their problems.

Tom: I'm curious, Jane, is this the first time anyone's tried this? Combining GAMs with neural networks?

Jane: No, there have been attempts before. The paper actually mentions a few, like Neural Additive Models from Google, and something called GAMI-Net. But those are in Python. And there are some R packages that do similar things, like deepregression. But the authors argue that their approach is different because it uses completely independent networks trained through a classic algorithm called backfitting.

Tom: So they're not just porting someone else's idea to R, they've got their own twist on it.

Jane: That seems to be the claim. And we'll get into the details of how it actually works as we go through the paper. But for now, I think the big picture is that this is about making powerful models more transparent, which is something that matters more and more as these tools get used in real-world decisions.

Tom: Transparency in AI is definitely a hot topic. And this package seems like a step in that direction. Alright, let's dig into the actual summary and see what the authors are promising.

Summary of the Paper: Jane: So we've got the title and the authors down. Now let's talk about what this paper is actually claiming to do. The summary tells us that neuralGAM is a package that lets you fit a Generalized Additive Neural Network, and the core idea is that you train an independent neural network for each feature in your model.

Tom: Right, and I remember you said that's what makes it interpretable. But how does that actually work in practice? Like, if I have three features, I just train three separate networks?

Jane: Almost. The trick is that you can't just train them separately and then add up their outputs, because each network might try to explain the same variation in the response. So the authors use a clever iterative approach. They start by making a guess, then they update each network one at a time, using the part of the response that hasn't been explained yet.

Tom: So it's like each network gets a turn to explain what it can, and then the next one gets the leftovers?

Jane: Exactly. That's the backfitting algorithm, and it's been around in statistics for decades. The authors are essentially taking that old algorithm and using neural networks as the smoother inside it, instead of the traditional splines.

Tom: And what kinds of problems can this handle? Is it just for continuous outcomes, or can it do classification too?

Jane: The paper says it supports three families: Gaussian for continuous data, binomial for binary classification, and Poisson for count data. So it covers a lot of the common use cases. And the link functions — the way you connect the linear predictor to the response — those are standard too, like logit for binomial and log for Poisson.

Tom: That's pretty comprehensive. And the summary also mentions something about uncertainty. That's usually a weak spot for neural networks, right?

Jane: Big time. Neural networks are great at predictions, but they're not great at telling you how confident they are. The authors address this using something called Monte Carlo Dropout. Basically, during prediction, you randomly turn off some neurons and run the network multiple times. The variation in the outputs gives you an estimate of the uncertainty.

Tom: So you get confidence intervals for each feature's effect. That's huge for interpretability.

Jane: It really is. You can plot each feature's contribution with a shaded band around it, and you can see where the model is confident and where it's not. That's something you just don't get from a typical black-box neural network.

Tom: And the summary also mentions that the package is built on top of Keras and TensorFlow, so it's using serious deep learning infrastructure.

Jane: Right, which means you get all the flexibility of modern neural networks — different architectures, activation functions, regularizers — but within a framework that keeps the model interpretable.

Tom: I'm starting to see why this could be a big deal. But I want to know more about the actual methodology. How does this backfitting thing really work under the hood?

Improvements Suggested by the Paper: Jane: So we've covered the basics, but the paper also talks about improvements over existing methods. And I think this is where it gets really interesting, Tom.

Tom: I'm all ears. What are they improving on?

Jane: Well, the paper compares itself to a few other approaches. There's deepregression, which is another R package that combines GAMs with neural networks. And there's also the black-box neural network approach, which is just a regular feedforward network. The authors argue that their method offers a better balance between interpretability and flexibility.

Tom: So they're saying they beat deepregression?

Jane: Not exactly beat, but they're offering something different. The paper shows results on both simulated and real data. On the simulated data, neuralGAM actually recovers the true underlying functions more accurately than deepregression. The shapes are smoother and closer to what was actually generating the data.

Tom: And on the real data? They used that flight delay dataset, right?

Jane: Yeah, the NYC flights data. They're trying to predict whether a flight will be delayed based on things like departure delay, air time, temperature, and humidity. And here's the interesting part: deepregression actually got a slightly higher AUC — that's a measure of classification performance — than neuralGAM.

Tom: So deepregression won on the real data?

Jane: On raw predictive performance, yes. But the gap was small. And the authors make a really good point about what you're giving up. The black-box neural network also performed well, but you can't see what it's doing. With neuralGAM, you get these beautiful plots showing exactly how each feature affects the probability of a delay.

Tom: So it's a trade-off between a tiny bit of performance and a whole lot of understanding?

Jane: Exactly. And for many real-world applications, that understanding is critical. If you're in healthcare or finance or any regulated industry, you need to be able to explain why your model made a particular prediction. A black-box model, no matter how accurate, is hard to deploy.

Tom: That makes sense. And the paper also mentions some flexibility improvements, right? Like you can customize the architecture for each feature individually?

Jane: Yes, that's a nice touch. You can say, "I want a two-layer network for this feature, but a single layer for that one." You can even use different activation functions. So you have a lot of control over how each feature is modeled.

Tom: And they also added a diagnostic function, which is something you don't always see in neural network packages.

Jane: Right, the diagnose function gives you residual plots and QQ plots, so you can check whether your model assumptions are holding. That's very much in the spirit of traditional statistical modeling, and it's a nice bridge between the two worlds.

Tom: So the improvements are really about making neural networks more like proper statistical models — with diagnostics, uncertainty, and interpretability.

Jane: That's the core message. And I think that's a really valuable contribution, because it gives practitioners a tool that's both powerful and trustworthy.

First Page Discussion: Tom: Alright, we've talked about the summary and the improvements. Let's go back to the very beginning of the paper and look at what the authors set up as their motivation. Jane, what's the big problem they're trying to solve?

Jane: The first page really sets the stage by talking about the black-box problem in neural networks. They mention that while neural networks are incredibly effective, it's often hard to understand how they make decisions. And that's a real issue when you're using these models in high-stakes situations.

Tom: And they mention two approaches to fix that, right? Post-hoc and ante-hoc?

Jane: Exactly. Post-hoc methods try to explain a black-box model after it's been trained. You might use something like SHAP values or LIME to figure out what the model is paying attention to. But those are approximations, and they don't always capture what's really going on.

Tom: And ante-hoc is the opposite?

Jane: Ante-hoc means you build the interpretability in from the start. You design the model so that it's inherently transparent. And that's what neuralGAM does. Instead of training a big black box and then trying to explain it, they train a model that's additive by construction.

Tom: So the interpretability isn't an afterthought, it's baked into the architecture.

Jane: Right. And the paper also gives a nice history of this idea. They mention that people have been trying to combine GAMs and neural networks since the late nineties. There was a model by Potts in one thousand nine hundred ninety-nine and then more recent ones like Neural Additive Models from Google in two thousand twenty-one.

Tom: So this isn't a new idea, but the authors think their implementation is better?

Jane: They argue that their approach is different because of how they train the networks. They use independent networks for each feature and train them through the local scoring and backfitting algorithms. This is a more classical statistical approach, and they claim it gives smoother, more stable estimates.

Tom: And they also mention that this is, as far as they know, the only R package that does this — using fully independent deep neural networks per feature with backfitting.

Jane: That's a strong claim, but it seems to hold up. The other R packages they compare against, like deepregression, take a different approach. They embed spline bases directly into a neural network architecture, which is clever but doesn't give you the same kind of per-feature network independence.

Tom: So neuralGAM is filling a specific niche in the R ecosystem.

Jane: Exactly. And I think that's important, because R is still the language of choice for a lot of statistical work. Having this tool available there could really help bridge the gap between traditional statistics and modern deep learning.

Tom: I also noticed they mention that the package relies on Keras and TensorFlow. So users get the full power of those frameworks, but within a more interpretable structure.

Jane: Right. And they've even provided a helper function to install all the Python dependencies, which is a nice touch for R users who might not be familiar with setting up Python environments.

Tom: So the first page really sets up the problem, gives the history, and positions their contribution. It's a solid foundation for the rest of the paper.

Conclusion: Tom: Well, we've covered a lot of ground today on "neuralGAM: An R Package for Fitting Generalized Additive Neural Networks." Jane, what's the big takeaway for our listeners?

Jane: I think the big takeaway is that you don't have to choose between interpretability and performance anymore. This package shows that you can have a neural network that's both powerful and transparent, and that's a really valuable combination.

Tom: And it's available in R, which means a whole community of statisticians can start using it without having to learn a whole new language.

Jane: Exactly. The authors have done a great job of packaging this up with all the tools you need — visualization, uncertainty estimation, diagnostics. It's not just a research prototype, it's a usable tool.

Tom: We also saw that while it might not always beat other methods on raw predictive performance, it offers something they don't: the ability to see exactly what each feature is doing.

Jane: And in many real-world applications, that understanding is worth more than a small boost in accuracy. If you're making decisions that affect people's lives, you need to be able to explain those decisions.

Tom: The authors also laid out some interesting future directions — things like multi-class classification, interaction terms, and even conformal prediction for better uncertainty estimates.

Jane: Those would be great additions. But even as it stands now, this is a solid contribution to the field of interpretable machine learning.

Tom: Alright, let's say goodbye to this paper and get ready for the next one. Thanks for joining us, and we'll see you next time.

Jane: Take care, everyone.

Ines Ortega-Fernandez, Marta Sestelo

Galician Research and Development Center in Advanced Telecommunications · Galician Centre for Mathematical Research and Technology · University of Vigo

stat.ML, cs.LG, stat.CO, stat.ME

Submitted: 2026-08-09

Updated: 2026-08-11

Comments: 19 pages, 4 figures. Accepted version; published in The R Journal 18(1):253-276

Journal ref: The R Journal 18(1) (2026) 253-276

DOI: 10.32614/RJ-2026-016

Code: https://github.com/inesortega/neuralGAM

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 68/100

The gist: The neuralGAM package implements a Neural Network topology based on Generalized Additive Models, allowing users to fit an independent Neural Network to estimate the contribution of each feature to

Key concepts

Generalized Additive Models (GAM)
A classic statistical tool where the model is built by adding up the contributions of individual variables. This structure allows users to see exactly how each feature affects the final outcome, providing high interpretability.
Backfitting Algorithm
An iterative approach used in neuralGAM. Instead of training one large network, it updates each small, independent neural network sequentially using the part of the response that hasn't been explained by previous networks.

Terminology

Summary

The neuralGAM package implements a Neural Network topology based on Generalized Additive Models, allowing users to fit an independent Neural Network to estimate the contribution of each feature to the output variable, yielding a highly accurate and interpretable deep learning model. The neuralGAM package provides a flexible framework for training Generalized Additive Neural Networks, which does not impose any restrictions on the per-term neural network architecture, while the overall model remains additive.

The package trains an ensemble of independent neural networks (one per feature) to learn each feature’s contribution to the response. By enforcing an additive decomposition via backfitting and local scoring, it prioritizes interpretability over the flexibility of unconstrained deep networks: the aim is not to match their full expressive power but to strike a practical balance, delivering competitive predictive performance while maintaining the GAM-style additivity that supports clear feature-level interpretation.

The methodology extends the classical Generalized Additive Model (GAM) framework by using independent neural networks to estimate the contribution of each covariate to the response, ensuring additivity through the local scoring and backfitting algorithms. This white-box modeling approach enables the visualization and interpretation of each feature’s partial effect while maintaining the ability to learn complex, non-linear patterns from data.

The package supports Gaussian, binomial, and Poisson response distributions, making it suitable for multiple applications. It relies on keras and TensorFlow for the implementation of neural networks, and also relies on other R packages for visualization (ggplot2) and numerical utilities.

The main function of the package is neuralGAM, which fits a Generalized Additive Neural Network to estimate the contribution of each smooth function to the response. The package includes methods for the generic summary, print, plot, autoplot, predict, and diagnose functions, as well as a plot history function for visualizing training and validation loss. A helper function install neuralGAM assists the user in installing the required Python dependencies in a custom conda environment.

The package implements uncertainty estimation via the uncertainty method argument, allowing the computation of epistemic uncertainty. To quantify epistemic uncertainty, the package leverages Monte Carlo (MC) Dropout, a widely used technique for estimating model uncertainty in neural networks by interpreting dropout as a Bayesian approximation. During training, dropout layers randomly drop a fraction of the units in each layer of the network, acting as a regularizer to prevent overfitting. At inference time, dropout remains active to generate stochasticity, and the network is evaluated repeatedly through B stochastic forward passes, producing a distribution of predictions that can be used to quantify epistemic uncertainty for each individual term and for the additive predictor in the link and response scales.

The package also includes cross-validation support during neural network training via the validation split argument, and allows users to visualize how the training and validation loss evolves after each backfitting iteration using the plot history function. Beyond global hyperparameters, neuralGAM allows the user to specify architectural choices individually for each smooth term, and allows the user to fit the model with any custom-built loss function.

The paper illustrates the use of the neuralGAM package in both synthetic and real data examples. In the simulated scenario with a Gaussian response, both neuralGAM and deepregression were applied to a dataset of size n = 30625, and neuralGAM achieved slightly better results across all metrics, with the deviance explained above 90% in both cases. The neuralGAM approach provides the added benefit of direct confidence interval estimation via epistemic uncertainty quantification using MC Dropout, and the shapes obtained with neuralGAM are smoother and more closely aligned with the true functions.

In the real-life application related to flight delay prediction based on weather and flight conditions' data from the NYC Flights 13 data set, the paper compares three approaches: neuralGAM (binomial neural additive model with per-term subnetworks and MC–Dropout for epistemic CIs), a distributional deep regression model fitted with deepregression, and a fully-connected, black-box neural network (Multi-Layer Perceptron, MLP) implemented in keras. The deviance explained by deepregression (34.09%) was similar to that obtained by neuralGAM (30.89%). Regarding performance on the test set, the AUC-ROC achieved by deepregression was 0.838, slightly higher than the value of 0.817 achieved by neuralGAM and comparable to the black-box neural network (0.830). The neuralGAM results exhibit smooth, stable, and interpretable effects due to its neural network based regularization, while deepregression tends to produce more irregular and wiggly smooths, especially in sparse data regions.

Overall, neuralGAM offers a strong compromise between interpretability and the flexibility provided by neural network-based estimation, producing stable, smooth, and easily interpretable effect estimates with uncertainty quantification. As future work, several research directions have been identified to further extend the proposed algorithm, including the generalization of the model to allow multinomial logistic regression, the incorporation of interaction terms, support for automatic feature selection procedures, support for cyclic and spatial features, and the inclusion of additional uncertainty estimation methods, including conformal prediction, bootstrap-based methods, or the use of empirical quantiles to obtain confidence intervals from MC Dropout.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems and what the improved system can do:

  • Improvement: Implement a Generalized Additive Neural Network (GANN) framework where each feature is modeled by an independent neural network, combined through local scoring and backfitting algorithms.

  • What it can do: The AI system can learn complex, non-linear relationships between each input feature and the output while maintaining full transparency. Users can visualize exactly how each feature contributes to predictions, eliminating the black-box problem.

  • Improvement: Integrate MC-Dropout at inference time with multiple stochastic forward passes (e.g., 100–500 passes) to estimate epistemic uncertainty for each feature contribution and the overall prediction.

  • What it can do: The system can provide confidence intervals for every prediction and for each feature's partial effect. This is critical for high-stakes applications (e.g., medical diagnosis, financial risk assessment) where knowing the model's certainty is as important as the prediction itself.

  • Improvement: Allow each feature's neural network to have its own architecture (number of layers, units, activation functions, regularization) rather than a one-size-fits-all model.

  • What it can do: The system can allocate more capacity to complex features (e.g., using deep networks with 256+ units and Swish activation) while using lightweight networks for simpler features. This improves both accuracy and computational efficiency.

  • Improvement: Enforce additivity through backfitting with centering constraints (ensuring each feature's effect sums to zero across the dataset), combined with dropout regularization in each subnetwork.

  • What it can do: The system produces smooth, stable, and interpretable effect estimates that are less prone to overfitting, especially in sparse data regions, compared to unconstrained neural networks or spline-based methods.

  • Improvement: Support multiple response distributions (Gaussian, binomial, Poisson) and allow custom loss functions (e.g., Huber loss for robustness) passed directly from Keras/TensorFlow.

  • What it can do: The system can handle regression, binary classification, and count data problems, and can be tailored to specific objectives (e.g., robust regression) without modifying the core architecture.

  • Improvement: Include built-in validation split monitoring (tracking training vs. validation loss per backfitting iteration) and diagnostic plots (Q-Q plots, residual vs. fitted, observed vs. predicted).

  • What it can do: The system can automatically detect overfitting during training and provide visual tools to verify model assumptions (e.g., normality of residuals, homoscedasticity), enabling users to make informed adjustments.

  • Improvement: Propagate uncertainty from the link scale to the response scale using the delta method (e.g., for logistic or Poisson models).

  • What it can do: The system provides confidence intervals directly on the probability or count scale, which are immediately interpretable to end-users without requiring them to understand link functions.

  • Improvement: Design the system to work seamlessly with standard R resampling packages (e.g., rsample, caret) for k-fold cross-validation.

  • What it can do: The system can be rigorously evaluated for generalization performance, producing reliable AUC-ROC, RMSE, or deviance metrics across multiple data splits, which is essential for model selection and deployment decisions.

  1. Explainable Predictions: For a given input, the system outputs not just a prediction but also a breakdown of how each feature contributed (e.g., departure delay contributed +0.3 to the log-odds of arrival delay) with confidence intervals.

  2. High Accuracy with Transparency: Achieve predictive performance comparable to black-box neural networks (e.g., AUC-ROC of 0.82–0.84 on flight delay prediction) while remaining fully interpretable.

  3. Reliable Uncertainty in Sparse Regions: Automatically widen confidence intervals where data is scarce (e.g., for flights with >600 minutes air time), alerting users to low-confidence predictions.

  4. Customizable Complexity: Users can specify different network depths and activation functions per feature, optimizing the trade-off between accuracy and interpretability for their specific dataset.

  5. Model Validation in One Call: Automatically generate diagnostic plots and training/validation loss curves, allowing users to verify model assumptions and detect overfitting without writing additional code.

  6. Deployment-Ready Predictions: Provide predictions on the original response scale (e.g., probability of delay, expected count) with standard errors, directly usable in production systems where uncertainty matters.

  7. Scalable to Real-World Problems: Handle datasets with tens of thousands of samples and multiple features (e.g., NYC flight data with weather conditions) in reasonable training times, while maintaining per-feature interpretability.

  8. Robust to Non-Linearity: Capture sharp non-linearities (e.g., temperature effects on flight delays) that linear models miss, while avoiding the wiggly overfitting seen in spline-based approaches.

Abstract

Nowadays, Neural Networks are considered one of the most effective methods for various tasks such as anomaly detection, computer-aided disease detection, or natural language processing. However, these networks suffer from the ``black-box'' problem which makes it difficult to understand how they make decisions. In order to solve this issue, an R package called neuralGAM is introduced. This package implements a Neural Network topology based on Generalized Additive Models, allowing to fit an independent Neural Network to estimate the contribution of each feature to the output variable, yielding a highly accurate and interpretable Deep Learning model. The neuralGAM package provides a flexible framework for training Generalized Additive Neural Networks, which does not impose any restrictions on the Neural Network architecture. We illustrate the use of the neuralGAM package in both synthetic and real data examples.

Sources

Related papers