Bio-Inspired Artificial Neural Networks based on Predictive Coding

arXiv:2508.08762 · stat.ML, cs.LG · Submitted 2025-08-12 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Bio-Inspired Artificial Neural Networks based on Predictive Coding".

Jane: The paper was written by Davide Casnici, Charlotte Frenkel and Justin Dauwels from Delft University of Technology.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Alright, welcome back to the show, everyone! Today we're diving into a fascinating new paper that just hit arXiv, and it's called "Bio-Inspired Artificial Neural Networks based on Predictive Coding." Jane, I have to say, the title alone got me excited.

Jane: Oh, absolutely, Tom. And I think for our listeners who might be new to this, let's just break that title down. We all know artificial neural networks — that's the technology behind everything from your phone's face unlock to large language models. The "bio-inspired" part is the key twist here.

Tom: Right, because for decades, the way we train these networks has been through something called backpropagation. It's incredibly powerful, but there's a catch that's been bugging neuroscientists for years.

Jane: And that catch is that backpropagation isn't really how brains work. The paper makes this point really clearly. When you train a deep network with backpropagation, the error signal has to travel all the way from the output layer back to the early layers. It's a global signal.

Tom: But in the brain, neurons only talk to their immediate neighbors. That's the locality principle. So the question is, how can you train a network using only local information? And that's where predictive coding comes in.

Jane: Exactly. The paper is essentially a tutorial — a really well-written one — that walks you through predictive coding from the ground up. It's by Davide Casnici, Charlotte Frenkel, and Justin Dauwels at TU Delft, and they're positioning this as a biologically plausible alternative to backpropagation.

Tom: And I love that they trace the history too. This isn't a brand-new idea. It goes back to the 1950s with P. Elias and signal compression, then Rao and Ballard in the 90s for the visual cortex, and then Karl Friston formalized it under the free energy principle.

Jane: Right, so the core idea is that the brain isn't passively receiving signals. It's actively predicting what it expects to see, and then it only pays attention to the difference between its prediction and reality. That difference is the prediction error.

Tom: And that error signal is local! It's computed right there at the synapse, between the prediction and the actual input. So you get this beautiful framework where learning happens with information that's available on the spot.

Jane: The paper does a fantastic job of laying out the math, starting from information theory, moving through variational inference, and then showing how predictive coding networks emerge from those principles. It's heavy math, but they make it accessible.

Tom: And they connect it to things we already know. They show how predictive coding can approximate backpropagation under certain conditions, and they even link it to the Kalman Filter, which is huge for signal processing folks.

Jane: So this isn't just a neuroscience curiosity. It's a bridge between how brains work and the algorithms we already use in engineering. I'm really curious to see how they actually implement this and what the performance looks like.

Tom: Me too, because that's where the rubber meets the road. We'll get into the experiments and the results in a bit, but first, let's bring in our resident AI researcher, Lu, to give us a bigger picture on why this matters.

Lu: Thanks, Tom. I think the most exciting implication here is that we might be able to build AI systems that learn the way brains do — with local rules, with energy efficiency, and with the ability to handle uncertainty. Backpropagation has given us incredible results, but it requires massive compute and massive data. Predictive coding offers a different path.

Jane: And that path might be more robust, too. The paper mentions that predictive coding can automatically scale gradients based on uncertainty, which is something backpropagation doesn't do.

Lu: Exactly. And when you think about edge devices — phones, sensors, robots — you can't run backpropagation there. It's too expensive. But local learning rules? That could be a game-changer for on-device learning.

Tom: Alright, I'm hooked. Let's dig into the actual mechanics of how predictive coding works in the next segment.

Summary: Jane: So we've established that "Bio-Inspired Artificial Neural Networks based on Predictive Coding" is about a biologically plausible alternative to backpropagation. Tom, let's get into the actual summary of what this paper teaches us.

Tom: Absolutely. The paper builds everything from scratch, and I think the most elegant part is how they derive the predictive coding network from variational inference. They start with a generative model — essentially, the brain's model of the world — and then they want to infer the causes of sensory input.

Jane: And the problem is that computing the true posterior distribution is intractable. You'd have to integrate over all possible causes, which is impossible for anything real. So they use variational inference to approximate it.

Tom: Right, and they make a clever choice. They use a Dirac delta distribution for the variational posterior. That sounds scary, but it just means they're approximating the posterior with a single point — the most likely value. So instead of tracking a whole distribution, they just track the mode.

Jane: And that simplifies the math enormously. The objective becomes what they call the negative free energy, and it reduces to a sum of precision-weighted prediction errors across layers.

Tom: Let me put that in plain English. Imagine you're watching a friend throw a ball. Your brain predicts where the ball will land. If the prediction is wrong, you get an error signal. That error gets weighted by how confident you are in your prediction. If you're very confident, the error matters a lot. If you're unsure, you don't overreact.

Jane: That's a great analogy. And the network has two types of neurons: value neurons that encode the predictions, and error neurons that encode the mismatch. The error neurons are shown as red triangles in their figures, and they only connect to adjacent layers.

Tom: So the update rule for the value neurons is beautifully local. Each neuron looks at the error from the layer above and the error from the layer below, and it adjusts its activity to reduce both. That's it. No global signal needed.

Jane: And then there's a second phase where the generative parameters — the weights — are updated. And that update is also local. It's essentially a Hebbian rule: neurons that fire together, wire together, but modulated by the prediction error.

Lu: I want to jump in here, because this is where the paper really shines. They show that if you let the inference phase converge — if you let the neural activity settle to equilibrium — then the error terms in predictive coding become mathematically equivalent to the error terms in backpropagation.

Tom: That's the big result. Under certain conditions, predictive coding computes exactly the same gradients as backpropagation, but using only local information.

Lu: And they cite the work by Whittington and Bogacz, and later by Song et al., who showed that predictive coding can exactly implement backpropagation. That's a profound theoretical result.

Jane: But the paper also points out that predictive coding is more general than backpropagation. Because of those precision weights — the covariance matrices — the network can automatically modulate how much it trusts each error signal. Backpropagation treats all errors equally.

Meng: I'm hearing a lot of theory, which is great, but I'm an engineer. How does this actually run? What's the computational cost?

Jane: Great question, Meng. The paper has a whole section on that. Let's get into the experiments in the next segment.

Improvements: Tom: Welcome back. We're still talking about "Bio-Inspired Artificial Neural Networks based on Predictive Coding," and now we're getting to the good stuff — the actual experiments and the improvements they demonstrate.

Jane: Right, and Meng just asked the key question: does this actually work in practice? The paper runs experiments on MNIST, FashionMNIST, and CIFAR10, using both multilayer perceptrons and convolutional networks.

Meng: And what did they find? I'm guessing there's a trade-off.

Tom: There is, and it's a fascinating one. In the compression task — think autoencoders — predictive coding achieves nearly identical performance to backpropagation, but with half the generative parameters. Half!

Meng: How is that possible?

Jane: Because predictive coding uses a single network that folds onto itself. In backpropagation, you need a separate encoder and decoder. In predictive coding, the same network does both jobs. The top layer encodes the compressed representation, and the forward pass decodes it.

Meng: That's elegant. So you get the same reconstruction quality with a smaller model?

Tom: Exactly. On MNIST, backpropagation gets a test MSE of about nine point three nine times ten to the minus three, and predictive coding gets one point one zero times ten to the minus two. Very close. On FashionMNIST, they're essentially tied. And on CIFAR10, predictive coding actually beats backpropagation — one point two nine versus one point zero three, wait, let me check that.

Jane: Actually, Tom, let me correct that. On CIFAR10, backpropagation gets one point zero three and predictive coding gets one point two nine, so backpropagation is slightly better there. But the point stands — they're very close.

Tom: Right, thanks for keeping me honest. The key takeaway is that predictive coding matches backpropagation's performance while using half the parameters. That's a real improvement in model efficiency.

Meng: But what about training time? I saw in the table that predictive coding takes longer per epoch.

Jane: That's the trade-off. Because of the inference phase — the iterative updates to the neural activity — predictive coding needs more computation. On MNIST, an epoch takes about seventeen seconds versus seven seconds for backpropagation. On CIFAR10 with the convolutional network, it's thirty-nine seconds versus eight seconds.

Meng: So it's about two to five times slower. And the FLOPs are much higher too — one hundred thirty-eight million versus four million on MNIST.

Tom: Right, but here's the thing. The inference phase is where the biological plausibility comes from. The brain doesn't do a single forward pass — it settles into a state over time. So the extra computation is the price you pay for local learning.

Lu: And I'd argue that this is the right trade-off to study. The paper is honest about the current limitations, but it also points to a future where we might have specialized hardware that exploits these local rules. On a neuromorphic chip, that inference phase could be extremely efficient.

Meng: That's a fair point. If you can map these local updates onto hardware that does parallel, event-driven computation, the wall-clock time could be very different from what you see on a GPU.

Jane: And the paper also shows that predictive coding can approximate backpropagation gradients. They explain how you can scale the covariance matrix in the final layer to suppress the error signal, which keeps the neural activity close to the forward pass, and then the weight updates match backpropagation.

Tom: So it's not just a biological curiosity — it's a flexible framework that can interpolate between local learning and exact backpropagation, depending on how you set the precision.

Lu: That flexibility is what excites me most. You can imagine training a network with backpropagation-like gradients for the first few epochs, then gradually transitioning to more local, uncertainty-aware updates as the model refines its understanding.

Meng: So the improvements here are really about efficiency in parameters and flexibility in learning rules, at the cost of compute.

Tom: Exactly. And the paper is honest about the scalability challenges. For very deep networks and complex datasets, backpropagation still wins. But this is a young field, and the paper provides a clear roadmap for improvement.

Jane: Let's wrap up with the bigger picture in our final segment.

Conclusion: Tom: We've spent this whole episode on "Bio-Inspired Artificial Neural Networks based on Predictive Coding," and I think it's fair to say this paper is a gift to the community.

Jane: It really is. It's a tutorial that takes you from first principles — information theory, variational inference — all the way to working code in PyTorch. That's rare. Most papers assume you already know the background.

Tom: And they make the connections to existing algorithms explicit. We talked about backpropagation and the Kalman Filter, but they also frame predictive coding within the free energy principle, which is a big deal in computational neuroscience.

Lu: I think the most impactful contribution is showing that predictive coding is not just a toy model. It matches backpropagation on standard benchmarks, it uses half the parameters for compression tasks, and it does all of this with local learning rules. That's a strong proof of concept.

Meng: As an engineer, I appreciate that they published the code. That means we can actually test these ideas, build on them, and see where the bottlenecks really are.

Jane: And the bottlenecks are real — the inference phase is computationally expensive. But that's also where the opportunity is. If we can build hardware that matches the brain's architecture, those costs could drop dramatically.

Tom: The paper also raises an interesting cultural point, Lalam. How do you see this changing the way we think about AI?

Lalam: I think it shifts the narrative from "AI as a black box trained by brute force" to "AI as a system that actively models its world." Predictive coding gives us a framework where machines can express uncertainty, update their beliefs, and learn continuously from streaming data. That's closer to how humans learn — through prediction and correction, not through massive offline training runs.

Tom: That's a beautiful way to put it. And it has implications beyond just performance. It's about building systems that are more transparent, more robust, and more aligned with how biological intelligence works.

Jane: And the authors — Casnici, Frenkel, and Dauwels — they've done a service to the field by making this accessible. The lecture notes format, the step-by-step derivations, the GitHub repository — it lowers the barrier to entry for anyone who wants to explore predictive coding.

Lu: I'd add that this paper will likely become a standard reference for researchers entering the field. It's the kind of paper you hand to a new PhD student and say, "Read this first."

Meng: And then tell them to benchmark it against their favorite architecture. The paper gives you all the tools to do that.

Tom: So as we say goodbye to "Bio-Inspired Artificial Neural Networks based on Predictive Coding," I want to leave our listeners with this: the brain has been solving the credit assignment problem for millions of years, and we're finally starting to understand how. This paper is a beautiful step in that direction.

Jane: And it's a step that connects neuroscience, machine learning, and signal processing in a way that few papers do. We'll be watching this space closely.

Tom: Absolutely. Thanks for joining us, everyone. Next up, we've got a paper on neuromorphic hardware that pairs perfectly with what we discussed today. See you then.

Jane: Take care, everyone.

Davide Casnici, Charlotte Frenkel, Justin Dauwels

Delft University of Technology

stat.ML, cs.LG

Submitted: 2025-08-12

Updated: 2026-08-18

Code: https://github.com/cogsys-tudelft/PredictiveCodingTutorial

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 58/100

Key concepts

Predictive Coding
A framework suggesting the brain actively predicts incoming sensory data. Learning occurs by analyzing the 'prediction error'—the difference between what was expected and what was actually received. This process uses local information.
Backpropagation
A traditional method for training neural networks where the error signal must travel globally from the output layer back to the early layers. The paper notes this is not how biological neurons communicate.
Locality Principle
The idea that in biological brains, neurons only communicate with their immediate neighbors. Predictive coding utilizes this principle by computing error signals locally at synapses.
Variational Inference
A mathematical technique used in the paper to approximate complex probability distributions (like the true posterior). It simplifies the problem by approximating the distribution with a single, most likely value.

Terminology

Summary

Summary

This paper provides a tutorial-style introduction to Predictive Coding (PC) as a biologically plausible alternative to backpropagation (BP) for training artificial neural networks (ANNs). The authors note that while BP is the current backbone training algorithm for ANNs, it relies on a global error signal generated at the extreme of the network hierarchy, which contradicts the Hebbian model of synaptic plasticity in the brain, where weight updates should be local and determined only by the activity of pre- and post-synaptic neurons. PC, originating from P. Elias’s 1950s work on signal compression and later proposed by Rao and Ballard as a model of the visual cortex, updates network weights using only local information. K. Friston subsequently formalized PC under the free energy principle (FEP), grounding it within Bayesian inference and dynamical systems frameworks.

The paper covers the mathematical foundations of PC, including information theory concepts such as Shannon’s entropy, cross entropy, and the Kullback-Leibler divergence (DKL). It then introduces variational inference (VI), explaining how intractable posterior distributions can be approximated with tractable variational distributions. The authors derive the variational free energy (FE) and the negative free energy (NFE), also known as the evidence lower bound (ELBO), showing how optimizing the NFE simultaneously bounds two intractable quantities: the DKL and the model log-evidence, thereby recasting an intractable inference problem into a tractable optimization problem.

The paper then specifies the PC model, assuming a Gaussian generative model with a hierarchical structure, a Gaussian variational posterior that reduces to a Dirac delta distribution, and a mean field approximation. Under these assumptions, the NFE simplifies to a sum of precision-weighted error terms between variational posterior values and generative model predictions. The authors derive update rules for both variational parameters (neural activity) and generative parameters (synaptic weights). The variational parameters are updated through gradient ascent, resulting in dynamics that depend on inhibitory feedforward signals from error nodes above and excitatory feedback signals from error nodes below, enabling local computation. The generative parameters are updated proportionally to the derivative of the NFE, using the optimal variational parameters as a proxy for the true posterior, following an Expectation-Maximization (EM) algorithm approach. The paper also derives a possible learning rule for the covariance matrices, though it notes challenges related to computational locality and numerical stability.

The paper compares PC with BP, showing that under certain conditions, PC can approximate BP gradients. Specifically, by setting the PC gradient to zero at equilibrium and scaling the covariance matrix in the final layer, PC error terms approximate those of BP, and the weight updates can be matched by rescaling PC gradients. The authors reference studies demonstrating that PC can compute exact BP gradients with additional modifications and can be extended to any graph topology.

The paper also connects PC with Kalman Filtering (KF), showing that by augmenting a PC network with recurrent connections, it can process temporal data while maintaining local parameter updates in both space and time. In this setting, PC becomes analogous to Bayesian filtering, and for linear Markov processes with Gaussian noise, it reduces to the KF algorithm. The authors derive the NFE for this recurrent PC configuration, showing that the gradient with respect to the variational parameters has the same form as in the standard PC case, and that the state transition and emission matrices can be learned with update rules analogous to those for PC generative parameters.

Finally, the paper presents a computational example comparing PC and BP on classification and compression tasks using MNIST, FashionMNIST, and CIFAR10 datasets. For compression tasks, PC requires only half the generative parameters compared to BP while maintaining comparable performance, but requires greater computational resources due to its inner optimization process. For classification tasks, PC needs the same number of generative parameters as BP, plus a small percentage of additional variational parameters, and introduces non-negligible computational overhead. The authors note that while PC is competitive with BP for small networks and datasets, it tends to underperform on larger datasets and more complex architectures, highlighting scalability limitations that remain an open research challenge.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems:

Improvement: Implement PC-based training instead of BP for neural networks, enabling weight updates using only pre- and post-synaptic information (local Hebbian plasticity).

What the improved system can do:

  • Train deep networks without requiring a global error signal propagated backward through all layers

  • Achieve comparable accuracy to BP on standard benchmarks (e.g., 98.12% vs 98.31% on MNIST classification)

  • Perform compression tasks with only half the generative parameters (326k vs 652k) while maintaining equivalent reconstruction quality

  • Scale gradients automatically based on uncertainty via precision-weighted errors (Σ−1 terms), potentially improving robustness to noisy data

Sources

Related papers