
The Brain’s Learning Algorithm Isn’t Backpropagation
What this covers
To try everything Brilliant has to offer—free—for a full 30 days, visit https://brilliant.org/ArtemKirsanov . You’ll also get 20% off an annual premium subscription
===== My name is Artem, I'm a neuroscience PhD student at Harvard University. 🌎 Website and Social links: https://kirsanov.ai/ 📥 "Receptive Field" neuro-newsletter: https://artemkirsanov.substack.com/ ✨ Support me on Patreon to get access to Discord community: https://patreon.com/artemkirsanov =====
In this video we explore Predictive Coding – a biologically plausible alternative to the backpropagation algorithm, deriving it from first principles.
Backpropagation video: https://youtu.be/SmZmBKc7Lrs?si=qduaq9hRYIymFOiL
🕒 OUTLINE: 00:00 Introduction 01:15 Credit Assignment Problem 02:49 Problems with Backprop 06:05 Foundations of Predictive Coding 08:07 Energy Formalism 11:08 Activity Update Rule 15:12 Neural Connectivity 17:42 Weight Update Rule 20:58 Putting all together 25:15 Brilliant 26:27 Outro
📚 FURTHER READING & REFERENCES: Bogacz, R., 2017. A tutorial on the free-energy framework for modelling perception and learning. Journal of Mathematical Psychology 76, 198–211. https://doi.org/10.1016/j.jmp.2015.11.003 Friston, K., 2018. Does predictive coding have a future? Nat Neurosci 21, 1019–1021. https://doi.org/10.1038/s41593-018-0200-7 Huang, Y., Rao, R.P.N., 2011. Predictive coding. WIRES Cognitive Science 2, 580–593. https://doi.org/10.1002/wcs.142 Keller, G.B., Mrsic-Flogel, T.D., 2018. Predictive Processing: A Canonical Cortical Computation. Neuron 100, 424–435. https://doi.org/10.1016/j.neuron.2018.10.003 Lillicrap, T.P., Santoro, A., Marris, L., Akerman, C.J., Hinton, G., 2020. Backpropagation and the brain. Nat Rev Neurosci 21, 335–346. https://doi.org/10.1038/s41583-020-0277-3 Marino, J., 2021. Predictive Coding, Variational Autoencoders, and Biological Connections. https://doi.org/10.48550/arXiv.2011.07464 Millidge, B., Salvatori, T., Song, Y., Bogacz, R., Lukasiewicz, T., 2022a. Predictive Coding: Towards a Future of Deep Learning beyond Backpropagation? Millidge, B., Seth, A., Buckley, C.L., 2022b. Predictive Coding: a Theoretical and Experimental Review. https://doi.org/10.48550/arXiv.2107.12979 Millidge, B., Song, Y., Salvatori, T., Lukasiewicz, T., Bogacz, R., 2023. A THEORETICAL FRAMEWORK FOR INFERENCE AND LEARNING IN PREDICTIVE CODING NETWORKS. Millidge, B., Tschantz, A., Buckley, C.L., 2022c. Predictive Coding Approximates Backprop Along Arbitrary Computation Graphs. Neural Computation 34, 1329–1368. https://doi.org/10.1162/neco_a_01497 Millidge, B., Tschantz, A., Seth, A., Buckley, C.L., 2020. Relaxing the Constraints on Predictive Coding Models. https://doi.org/10.48550/arXiv.2010.01047 Rao, R.P.N., Ballard, D.H., 1999. Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effects. Nat Neurosci 2, 79–87. https://doi.org/10.1038/4580 Rosenbaum, R., 2022. On the relationship between predictive coding and backpropagation. PLoS ONE 17, e0266102. https://doi.org/10.1371/journal.pone.0266102 Salvatori, T., Mali, A., Buckley, C.L., Lukasiewicz, T., Rao, R.P.N., Friston, K., Ororbia, A., 2025. A Survey on Brain-Inspired Deep Learning via Predictive Coding. https://doi.org/10.48550/arXiv.2308.07870 Salvatori, T., Song, Y., Lukasiewicz, T., Bogacz, R., Xu, Z., 2023. Reverse Differentiation via Predictive Coding. https://doi.org/10.48550/arXiv.2103.04689 Salvatori, T., Song, Y., Yordanov, Y., Millidge, B., Xu, Z., Sha, L., Emde, C., Bogacz, R., Lukasiewicz, T., 2024. A Stable, Fast, and Fully Automatic Learning Algorithm for Predictive Coding Networks. https://doi.org/10.48550/arXiv.2212.00720 Song, Y., Lukasiewicz, T., Xu, Z., Bogacz, R., n.d. Can the Brain Do Backpropagation? — Exact Implementation of Backpropagation in Predictive Coding Networks. Song, Y., Millidge, B., Salvatori, T., Lukasiewicz, T., Xu, Z., Bogacz, R., 2024. Inferring neural activity before plasticity as a foundation for learning beyond backpropagation. Nat Neurosci 27, 348–358. https://doi.org/10.1038/s41593-023-01514-1 Whittington, J.C.R., Bogacz, R., 2019. Theories of Error Back-Propagation in the Brain. Trends in Cognitive Sciences 23, 235–250. https://doi.org/10.1016/j.tics.2018.12.005 Whittington, J.C.R., Bogacz, R., 2017. An Approximation of the Error Backpropagation Algorithm in a Predictive Coding Network with Local Hebbian Synaptic Plasticity. Neural Computation 29, 1229–1262. https://doi.org/10.1162/NECO_a_00949
===== This video was sponsored by Brilliant =====
*Disclaimer:* This channel is my personal project. The views and content expressed here are my own and are separate from my research role at Harvard University.
#PredictiveCoding #Neuroscience #AI
_Description remastered: February 2026. Links & Bio updated; original context preserved._
Source description (no synthesized summary yet).
Predictive coding is a biologically plausible alternative to backpropagation that solves critical constraints (discontinuous processing and lack of local autonomy) incompatible with neural physiology, while potentially offering computational advantages for artificial intelligence.
- Backpropagation requires separated forward and backward phases with global coordination that contradicts continuous parallel processing in biological brains
- Predictive coding uses locally autonomous neurons that minimize prediction errors at their own layer and the layer they predict, without centralized control
- Predictive coding's local update rules may reduce catastrophic forgetting and better preserve learned knowledge compared to backpropagation's global output-focused optimization
This asset isn't compiled yet
You're seeing its claims, ranked. Compile it to build the argument threads, weight them, and check each claim against your library — the full view.
The backpropagation algorithm, while remarkably successful in machine learning, contradicts essential biological principles of brain function, making its exact implementation in neural tissue virtually impossible.
“However, there is a fundamental problem. The back propagation algorithm contradicts essential biological principles of brain function, making its exact implementation in neural tissue virtually impossible.”
Backpropagation requires discontinuous processing where neurons must freeze their feedforward activity values while error signals flow backward, requiring the brain to pause information processing for hundreds of milliseconds before updating connections.
“For this process to work, neurons must essentially freeze their feed forward activity values like taking snapshots of activity and holding on to them while error signals flow backward.”
As neuroscience and artificial intelligence continue to inform each other, predictive coding stands as a compelling bridge between the remarkable learning capabilities of biological brains and the next generation of neural network architectures.
“As neuroscience and artificial intelligence continue to inform each other, predictive coding stands as a compelling bridge between the remarkable learning capabilities of biological brains and the next generation of neural network architectures.”
Individual neurons and synapses in the brain function as autonomous agents, modifying their states based solely on information physically available at their specific locations, operating as a massively parallel locally autonomous system where computation and learning occurs simultaneously throughout the network in a distributed manner without centralized control.
“Instead, individual neurons and synapses mostly function as autonomous agents, modifying their states based solely on information physically available at their specific locations. The brain operates in a massively parallel locally autonomous system where computation and learning occurs simultaneously throughout the network in a distributed manner without centralized control.”
Predictive coding originated from mid-twentieth century research proposing that the brain's fundamental objective is to predict incoming sensory information, which enhances survival by allowing organisms to anticipate threats and interpret noisy observations.
“This framework originated from midentth century research, proposing that the brain's fundamental objective is to predict incoming sensory information. From an evolutionary perspective, prediction enhances survival by allowing organisms to anticipate threats and interpret noisy observations.”
The credit assignment problem is the fundamental challenge that computational systems must solve: determining which parameters to adjust and by how much to achieve a desired output
“The fundamental challenge that computational systems must solve is called credit assignment. When you have a system with numerous parameters like connection weights between neurons that can be adjusted to achieve a desired output such as recognizing objects in an image or executing appropriate actions. How do you determine which parameters to adjust and by how much?”
Information flows bidirectionally through the predictive coding hierarchy with top-down connections carrying predictions from higher levels to lower levels, while bottom-up connections carry prediction errors (differences between predictions and actual activity).
“Information flows birectionally through this hierarchy. Top-down connections carry predictions from higher levels to lower levels, while bottom up connections carry prediction errors, differences between predictions and the actual activity.”
Predictive coding's local autonomy makes the algorithm extremely parallelizable and in certain settings more efficient than back propagation, and theoretical considerations suggest that resulting updates may lead to better solutions than back propagation.
“The local autonomy makes the algorithm extremely parallelizable and in certain settings more efficient than back propagation. Theoretical considerations suggest that resulting updates may actually lead to better solutions than back propagation.”
The predictive coding framework can be approached as an energy-based model by associating each possible network state with a single number representing abstract energy, then deriving rules for how the system should evolve to reduce this energy, paralleling physical systems that naturally progress towards minimum energy states.
“We'll approach our network as a so-called energy- based model. Essentially, this means associating each possible network state with a single number representing some form of abstract energy. We can then derive rules for how the system should evolve to reduce this energy. This framework parallels physical systems that naturally progress towards minimum energy states like a ball rolling downhill to minimize gravitational potential energy or proteins folding to minimize atomic interaction energy.”
While some coordinating mechanisms like theta and gamma rhythms, attentional systems and neuromodulators such as dopamine influence broad neural populations, they operate at much coarser temporal and spatial scales than would be required for backpropagation which relies on cell-by-cell precision
“While there are some coordinating mechanisms, oscillations like theta and gamma rhythms, attentional systems and neurom modulators like dopamine that influence broad populations. These mechanisms operate at much coarser temporal and spatial scales than would be required for back propagation which relies on cellby cell precision.”
From the update rules, a representational neuron must be inhibited by its corresponding error neuron and excited by error neurons sending feedback signals from the layer below, elegantly mapping the mathematical formulation onto biological circuitry.
“With this structure in mind, we can directly read off the required neural connectivity from our update rule. A representational neuron X subi must be inhibited by its corresponding error neuron and excited by error neurons sending feedback signals from the layer below. This elegantly maps our mathematical formulation onto biological circuitry.”
During the system's evolution, it effectively rolls downhill on the energy surface defined in high-dimensional space, and mathematically this corresponds to moving in the direction of steepest descent opposite to the gradient, where the gradient vector points in the direction of steepest ascent and is composed of derivatives with respect to each parameter
“During the systems evolution, it effectively rolls downhill on the energy surface defined in a highdimensional space where each coordinate represents a parameter such as neural activity or synaptic weight. Mathematically, this downhill roll corresponds to moving in the direction of steepest descent opposite to what's called the gradient of the function where the gradient vector points in the direction of steepest asend and is composed of derivatives with respect to each parameter.”
In real models with nonlinear activation functions, the update rules for opposing synapses are not mathematically identical, but research suggests perfect symmetry may not be essential; approximate symmetry emerging from independent learning is sufficient for effective network function.
“I should note that in real models though there is a nonlinear activation function which we have been sweeping under the rug. When these nonlinearities are included, the updates for the two sapses are not mathematically identical. Fortunately, research suggests that perfect symmetry may not be essential. Even when feed forward and feedback signapses learn independently with slightly different update rules, the approximate symmetry that emerges is sufficient for the network to function effectively.”
If predictive coding networks were allowed to freely adjust every parameter (both neural activities and weights), they would naturally settle to a zero energy state which would be trivial and not perform any meaningful computation, so in practical implementations and likely in the brain, certain neurons are clamped to specific values
“Let's now put everything together and see how this framework operates as a complete system. If we allow the network to freely adjust every parameter, both neural activities and the weights, it would naturally settle to a zero energy state. However, this solution would be trivial and not perform any meaningful computation. In practical implementations of predictive coding and likely in the brain itself, certain neurons are kind of clamped to specific values.”
A neuron's predicted activity is given by the weighted sum of activities of upstream neurons, determined by synaptic weights which are typically followed by a nonlinear activation function (like sigmoid or ReLU), though the speaker simplifies this for pedagogical clarity
“We can visualize it as rods connecting neuron nodes on the layer above to the platforms at a current level positioned at variable angles corresponded to synaptic weights which determine how other neurons activities influence the prediction. The sum of activities from all neurons in the layer above multiplied by synaptic weights connecting them. Note that typically activities pass through a nonlinear activation function like sigmoid or relu, but I'm omitting it here for simplicity.”
Repeating the iterative relaxation process across diverse examples gradually refines the network's internal model of the world, developing compressed representations of data
“Repeating this process across diverse examples gradually refineses the network's internal model of the world. Through this process, the network develops compressed representations of data.”
Backpropagation with gradient descent is the workhorse algorithm that powers virtually the entire field of machine learning today
“Their efforts led to back propagation with gradient descent, the workhorse algorithm that powers virtually the entire field of machine learning today.”