REVIEW 2 major objections 5 minor 27 references
Extended predictive coding framework as variational free-energy minimisation under exponential-family assumption
T0 review · 2 major / 5 minor · reviewed 2026-07-12 · grok-4.5
Pith's one-line read Exponential-family posteriors and priors keep free-energy minimisation in predictive-coding form while allowing nonlinear, heterogeneous, non-negative neural activations.
desk verdict Solid EFD extension of FEP–PC that legitimately allows heterogeneous monotone activations and local plasticity; the third-cumulant drop is the real (and already flagged) soft spot, not a hidden collapse. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The EFD–FEP model: the approximate gradient of hierarchical free energy under exponential-family assumptions, which takes the predictive-coding shape τ η̇_q = G(η_q){−η_q + η_p + W_pred^⊤ ε̄} (or its natural-gradient version without G) once the third cumulant is neglected.
What would settle it
Construct a network whose posterior is strongly skewed (e.g., high-rate Poisson) and check whether free energy still decreases under the proposed predictive-coding dynamics; if free energy rises, the neglected third-cumulant term is decisive and the claim fails.
Extended reading notes
Core claim
Under hierarchical factorisation, exponential-family posterior and prior (shared base measure), fixed-variance Gaussian likelihood and linear prediction map, the negative gradient of layer-wise variational free energy with respect to natural parameters is, after dropping the third posterior cumulant, exactly of predictive-coding form: the update of each natural parameter is driven by a prior-attracting term plus bottom-up prediction error, optionally re-weighted by the Fisher information matrix. The same free-energy objective produces three local plasticity rules for the prediction, recurrent and top-down weights.
Load-bearing premise
The third-order cumulant of the posterior is small enough that ignoring it still leaves an update that both looks like predictive coding and continues to reduce free energy.
Editorial extensions
If this is right
- Cortical circuits can implement free-energy minimisation with neurons that possess diverse, nonlinear, non-negative F–I curves without leaving the free-energy principle.
- Heterogeneity of response properties becomes a computational resource that enlarges the class of representable posteriors rather than an obstacle.
- Prediction and prior pathways can be learned by local rules that map onto basal Hebbian plasticity and apical calcium-mediated plasticity of pyramidal cells.
- The ordinary-versus-natural gradient choice supplies a concrete efficiency–stability trade-off that can be tested by varying network noise.
Reading between the lines
- If the third-cumulant approximation remains accurate for the distributions actually used by cortex, free-energy theory no longer forces linear Gaussian units and can therefore be confronted with measured F–I diversity.
- Geometric orthogonality between prior-regulating and error-feedback subspaces offers a testable signature: high-level beliefs should occupy directions invisible to bottom-up error alone.
- The same free-energy construction may admit non-Gaussian likelihoods once surrogate gradients that still guarantee free-energy decrease are identified, potentially linking free energy to motifs beyond classical predictive coding.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper derives an extended predictive-coding (PC) implementation of variational free-energy (VFE) minimisation under the free-energy principle (FEP). Assuming hierarchical factorisation, exponential-family (EFD) posterior and prior (shared base measure), fixed-variance Gaussian likelihood, and a linear prediction map, the negative gradient of layer-wise VFE with respect to natural parameters η_q is shown to take PC form after neglecting the third posterior cumulant: τ η̇_q = G_q(η_q){-η_q + η_p + W_pred^ op ε̄} (ordinary gradient) or the natural-gradient version without G_q (Eqs. 14–15). The same objective yields local plasticity rules for prediction, recurrent and top-down weights (Eqs. 21–23). Type-A (factorised) models map onto heterogeneous, nonlinear F–I curves (activation functions abla A) that remain non-negative, while Type-B models are discussed more cautiously. Stochastic variants and biological correspondences (pyramidal apical/basal compartments, BAC firing) are proposed.
Significance. If the approximation is controlled, the result meaningfully widens the FEP–PC correspondence beyond the Gaussian/Laplace regime that has dominated the literature. It supplies a normative account of nonlinear and heterogeneous neuronal response properties, eliminates negative firing rates without ad-hoc rectification, and derives local Hebbian-like and dendritic-plateau-compatible plasticity from a single objective. These features address long-standing biological criticisms of Gaussian PC models and therefore strengthen FEP as an explanatory theory of cortical perceptual inference. The geometric and information-thermodynamic remarks in the Discussion are suggestive but secondary. The derivation itself is analytic and transparent; no machine-checked proofs or numerical validation are supplied.
major comments (2)
- [Section 3.2.1 / Appendix A] Section 3.2.1 (Eqs. 12–14) and Appendix A: the claimed PC form of free-energy descent is obtained only after discarding the third-cumulant remainder Δ_k = (1/2) M^{ij} T_ijk. Appendix A itself notes that for Poisson (and other non-sub-Gaussian) posteriors every cumulant scales as e^η, so |Δ| need not be small, and that the residual is not guaranteed to keep the approximate flow a descent direction for F. The appeal to uncertainty-weighted decay of W_pred keeping M small is heuristic, not proved. Because the central claim is precisely that the dynamics remain VFE-reducing while taking PC form, a concrete regime (bounds on |T_ijk| relative to G_q and G_φ, restriction to distributions with vanishing third cumulants, or a Lyapunov argument for the approximate vector field) is required; otherwise the FEP–PC correspondence is only formal, not variational.
- [Section 3.2 / Appendix B] The likelihood is kept fixed-variance Gaussian throughout the main derivation (Eq. 9). Appendix B correctly shows that a non-Gaussian EFD likelihood replaces the simple additive error with an expectation E_q[(z-μ_q)A_φ(ξ)] that lacks a local neural implementation. Consequently the advertised “extension to the exponential family” is only partial; the title and abstract should more clearly delimit that the PC correspondence still rests on a Gaussian observation model, or the main text should supply a controlled approximation for non-Gaussian likelihoods that preserves descent.
minor comments (5)
- [Figure 1] Figure 1 caption and surrounding text mix “firing activities” ˇx with both spike counts and rates; a single consistent interpretation (or an explicit statement that both are admissible) would help.
- [Section 3.2.1] The distinction between ordinary and natural gradient descent (OGD vs NGD) is introduced cleanly, yet the biological preference for one or the other is left as an efficiency–stability trade-off without even a schematic simulation; a short numerical illustration of convergence speed versus noise robustness would strengthen the claim.
- [Section 3.1.2] Notation for the Legendre dual occasionally switches between μ = abla A(η) and E_q[ˇx]; a single convention after Eq. 6 would reduce cognitive load.
- [Section 4.3] Section 4.3 lists four biologically awkward properties of error-coding neurons; the feedback-alignment suggestion is plausible but remains an empirical observation from machine learning. A brief citation to any cortical evidence for approximate weight symmetry would be useful.
- Typos: “i.e.the” (missing space) in the Abstract; “V oss” in Helmholtz reference; occasional missing spaces after periods in the arXiv header.
Circularity Check
No significant circularity: PC-form dynamics and local plasticity are obtained by direct differentiation of the stated hierarchical VFE under explicit EFD + linear-prediction + Gaussian-likelihood assumptions, with the third-cumulant drop flagged as an approximation rather than hidden by definition.
full rationale
The central derivation (Sections 3.1–3.2) begins from a hierarchical factorisation of the generative model and approximate posterior (Eq. 4), places both posterior and prior in the exponential family with shared base measure (Eq. 5), keeps the likelihood fixed-variance Gaussian (Eq. 9), and imposes a linear prediction map ξ_ϕ = W_pred x̌. The layer-wise VFE is then written in closed form via the Legendre dual (Eq. 10). Its exact gradient with respect to the natural parameters η_q splits into a KLD piece that is exactly G_q(η_q)(−η_q + η_p) (Eq. 11) and a log-likelihood piece whose only non-PC term is the third-cumulant contribution (1/2)∇_η Tr[G_ϕ Σ_ξ] (Eq. 12). Neglecting that term yields the ordinary- and natural-gradient PC flows (Eqs. 14–15) and, by the same objective, the three local plasticity rules (Eqs. 21–23). None of these steps equates the claim to its own definition, fits free parameters to data that are later “predicted,” or rests on a self-citation uniqueness theorem. The biological correspondences (F–I curves as ∇A, apical-tuft plasticity, etc.) are post-hoc interpretations of already-derived equations and do not enter the formal argument. Appendix A itself records that the neglected remainder need not be small for Poisson-like posteriors, so the approximation is transparent rather than circular. The paper is therefore self-contained against its own stated assumptions; the only residual risk is correctness of the approximation, not circularity of the derivation chain.
Assumptions & free parameters
free parameters (4)
- τ_inference, τ_plasticity
- Choice of log-partition A_q (per neuron or population)
- Likelihood covariance G_ϕ (fixed)
- Architecture of W_pred, W_rec, W_TD (dimensionality and sparsity)
assumptions (8)
- domain assumption Hierarchical Markov factorisation of generative model and mean-field factorised approximate posterior (Eq. 4).
- domain assumption Posterior and prior belong to the exponential family with shared base measure h; natural parameters fully parameterise them.
- domain assumption Likelihood is fixed-variance Gaussian so A_ϕ is quadratic (Eq. 9).
- domain assumption Linear prediction map ξ_ϕl(x̌_l) = W_pred_l x̌_l.
- ad hoc to paper Third posterior cumulant ∇³A_q is negligible in the log-likelihood gradient.
- domain assumption Prior natural parameters are linear in same-layer and higher-layer mean (or sample) activities (Eq. 18).
- domain assumption Type-A: A_q additively decomposable so G_q is diagonal and each neuron has an independent monotone activation.
- standard math Standard EFD identities: μ = ∇A, G = ∇²A = Cov, D_KL[f∥g] = A*(μ_A) + B(η_B) − η_B^⊤ μ_A.
invented entities (3)
-
EFD–FEP model (and OGD/NGD/SOGD/SNGD variants)
-
Type-A vs Type-B EFD–FEP subtypes
-
Internal state η_q of a representational neuron (distinct from raw input current)
Cite this review
Pith. "Pith review of Extended predictive coding framework as variational free-energy minimisation under exponential-family assumption." pith.science (2026). https://pith.science/paper/CDO4CPYS
@misc{pith2026260530882,
author = {Pith},
title = {Pith review of: Extended predictive coding framework as variational free-energy minimisation under exponential-family assumption},
year = {2026},
howpublished = {\url{https://pith.science/paper/CDO4CPYS}},
note = {Machine review of arXiv:2605.30882}
}
read the original abstract
The sensory cortices of the brain perform perceptual inference efficiently through their complex networks of neurons. One of the theoretical accounts of this process is the free-energy principle (FEP), which postulates that the brain performs variational Bayesian inference. Pioneering studies have shown that FEP can correspond to the predictive coding (PC) hypothesis under the Gaussian assumption and Laplace approximation. However, PC-based implementations of FEP within such a limited Gaussian regime have failed to capture several properties of biological neural networks, such as nonlinearity and heterogeneity of input--output properties within a network, and the biological implausibility of negative firing rates. This study shows that, when a broader class of probability distributions, namely the exponential family of distributions (EFD), is assumed for the variational posterior and prior, these missing characteristics are exhibited within the network, maintaining the FEP--PC correspondence up to the second cumulant of the posterior. We also show that the proposed model can be trained by biologically plausible local plasticity rules. Our results enrich the explanatory power of FEP regarding neural dynamics involved in perception as variational inference.
Figures
Reference graph
Works this paper leans on
-
[1]
The free-energy principle: a rough guide to the brain?Trends Cogn
Karl Friston. The free-energy principle: a rough guide to the brain?Trends Cogn. Sci., 13(7): 293–301, July 2009
2009
-
[2]
Predictive coding under the free-energy principle.Philos
Karl Friston and Stefan Kiebel. Predictive coding under the free-energy principle.Philos. Trans. R. Soc. Lond. B Biol. Sci., 364(1521):1211–1221, May 2009
2009
-
[3]
The free-energy principle: a unified brain theory?Nat
Karl Friston. The free-energy principle: a unified brain theory?Nat. Rev. Neurosci., 11(2): 127–138, February 2010. 16
2010
-
[4]
A tutorial on the free-energy framework for modelling perception and learning
Rafal Bogacz. A tutorial on the free-energy framework for modelling perception and learning. J. Math. Psychol., 76(Pt B):198–211, February 2017
2017
-
[5]
9 ofAllgemeine Encyklopädie der Physik
Hermann von Helmholtz.Handbuch der physiologischen Optik, volume Bd. 9 ofAllgemeine Encyklopädie der Physik. Leopold V oss, Leipzig, 1867
-
[6]
Cambridge University Press, September 1996
David C Knill, Whitman Richards, Whitman Richard, D C Knill, D Kersten, A Yuille, D Mum- ford, A Jepson, W Richards, D C Knill, J Feldman, A L Yuille, H H Bülthoff, B M Bennett, D D Hoffman, C Prakash, S N Richman, P Mamassian, A Blake, D Sheinberg, P N Belhumeur, W T Freeman, K Nakayama, S Shimojo, E H Adelson, A P Pentland, and H Barlow.Perception as Ba...
1996
-
[7]
MIT Press, 2007
Kenji Doya, Shin Ishii, Alexandre Pouget, and Rajesh P N Rao.Bayesian Brain: Probabilistic Approaches to Neural Coding. MIT Press, 2007
2007
-
[8]
Bayesian brain theory: Computational neuroscience of belief.Neuroscience, 566:198–204, February 2025
Hugo Bottemanne. Bayesian brain theory: Computational neuroscience of belief.Neuroscience, 566:198–204, February 2025
2025
Show all 27 references
-
[9]
Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effects.Nat
R P Rao and D H Ballard. Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effects.Nat. Neurosci., 2(1):79–87, January 1999
1999
-
[10]
A new cellular mechanism for coupling inputs arriving at different cortical layers.Nature, 398(6725):338–341, March 1999
M E Larkum, J J Zhu, and B Sakmann. A new cellular mechanism for coupling inputs arriving at different cortical layers.Nature, 398(6725):338–341, March 1999
1999
-
[11]
A cellular mechanism for cortical associations: an organizing principle for the cerebral cortex.Trends Neurosci., 36(3):141–151, March 2013
Matthew Larkum. A cellular mechanism for cortical associations: an organizing principle for the cerebral cortex.Trends Neurosci., 36(3):141–151, March 2013
2013
-
[12]
Conjunctive input processing drives feature selectivity in hippocampal CA1 neurons.Nat
Katie C Bittner, Christine Grienberger, Sachin P Vaidya, Aaron D Milstein, John J Macklin, Junghyup Suh, Susumu Tonegawa, and Jeffrey C Magee. Conjunctive input processing drives feature selectivity in hippocampal CA1 neurons.Nat. Neurosci., 18(8):1133–1142, August 2015
2015
-
[13]
Implications of neuronal diversity on population coding
Maoz Shamir and Haim Sompolinsky. Implications of neuronal diversity on population coding. Neural Comput., 18(8):1951–1986, August 2006
1951
-
[14]
Intrinsic biophysical diversity decorrelates neuronal firing while increasing information content.Nat
Krishnan Padmanabhan and Nathaniel N Urban. Intrinsic biophysical diversity decorrelates neuronal firing while increasing information content.Nat. Neurosci., 13(10):1276–1282, October 2010
2010
-
[15]
Population diversity and function of hyperpolarization- activated current in olfactory bulb mitral cells.Sci
Kamilla Angelo and Troy W Margrie. Population diversity and function of hyperpolarization- activated current in olfactory bulb mitral cells.Sci. Rep., 1(1):50, July 2011
2011
-
[16]
Multivariate analysis of electrophysiological diversity of xenopus visual neurons during development and plasticity.Elife, 4, November 2015
Christopher M Ciarleglio, Arseny S Khakhalin, Angelia F Wang, Alexander C Constantino, Sarah P Yip, and Carlos D Aizenman. Multivariate analysis of electrophysiological diversity of xenopus visual neurons during development and plasticity.Elife, 4, November 2015
2015
-
[17]
Diversity amongst human cortical pyramidal neurons revealed via their sag currents and frequency preferences.Nat
Homeira Moradi Chameh, Scott Rich, Lihua Wang, Fu-Der Chen, Liang Zhang, Peter L Carlen, Shreejoy J Tripathy, and Taufik A Valiante. Diversity amongst human cortical pyramidal neurons revealed via their sag currents and frequency preferences.Nat. Commun., 12(1):2497, May 2021
2021
-
[18]
Learning probability distributions of sensory inputs with Monte Carlo predictive coding.PLOS Computational Biology, 20(10): e1012532, October 2024
Gaspard Oliviers, Rafal Bogacz, and Alexander Meulemans. Learning probability distributions of sensory inputs with Monte Carlo predictive coding.PLOS Computational Biology, 20(10): e1012532, October 2024. ISSN 1553-7358. doi: 10.1371/journal.pcbi.1012532. URL https:// journals...
2024 doi
-
[19]
Active inference and agency.Cogn
Karl Friston. Active inference and agency.Cogn. Neurosci., 5(2):119–121, April 2014
2014
-
[20]
The markov blankets of life: autonomy, active inference and the free energy principle.J
Michael Kirchhoff, Thomas Parr, Ensor Palacios, Karl Friston, and Julian Kiverstein. The markov blankets of life: autonomy, active inference and the free energy principle.J. R. Soc. Interface, 15(138):20170792, January 2018
2018
-
[21]
Active inference
Thomas Parr, Giovanni Pezzulo, and Karl J Friston. Active inference. https://mitpress. mit.edu/9780262362283/active-inference/, December 2021. Accessed: 2026-3-7. 17
2021
-
[22]
Life as we know it.J
Karl Friston. Life as we know it.J. R. Soc. Interface, 10(86):20130475, September 2013
2013
-
[23]
Applied Mathematical Sciences
Shun-Ichi Amari.Information Geometry and Its Applications. Applied Mathematical Sciences. Springer, Tokyo, Japan, 1 edition, February 2016
2016
-
[24]
Laws of thermodynamics for exponential families.arXiv [cond-mat.stat- mech], January 2025
Akshay Balsubramani. Laws of thermodynamics for exponential families.arXiv [cond-mat.stat- mech], January 2025
2025
-
[25]
Thermodynamics of prediction.Phys
Susanne Still, David A Sivak, Anthony J Bell, and Gavin E Crooks. Thermodynamics of prediction.Phys. Rev. Lett., 109(12):120604, September 2012
2012
-
[26]
On the thermodynamics of prediction under dissipative adaptation.arXiv [q-bio.NC], September 2020
Kai Ueltzhöffer. On the thermodynamics of prediction under dissipative adaptation.arXiv [q-bio.NC], September 2020
2020
-
[27]
C. Beck. Superstatistics: theory and applications.Continuum Mechanics and Thermodynamics, 16(3):293–304, March 2004. ISSN 1432-0959. doi: 10.1007/s00161-003-0145-1. URL http://dx.doi.org/10.1007/s00161-003-0145-1. A On third cumulant-neglecting approximation In Section 3.2.1, ...
2004 doi
Reviewed July 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.