Pith. sign in

REVIEW 2 major objections 5 minor 27 references

Extended predictive coding framework as variational free-energy minimisation under exponential-family assumption

T0 review · 2 major / 5 minor · reviewed 2026-07-12 · grok-4.5

Pith's one-line read Exponential-family posteriors and priors keep free-energy minimisation in predictive-coding form while allowing nonlinear, heterogeneous, non-negative neural activations.

desk verdict Solid EFD extension of FEP–PC that legitimately allows heterogeneous monotone activations and local plasticity; the third-cumulant drop is the real (and already flagged) soft spot, not a hidden collapse. read the letter →

arxiv 2605.30882 v2 pith:CDO4CPYS submitted 2026-05-29 q-bio.NC

classification q-bio.NC
keywords free-energyprinciplepredictivecodingexponentialfamilyvariationalinferenceneuraldynamicssynapticplasticityheterogeneousactivations
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The free-energy principle is widely taken to explain perception as variational Bayesian inference, but earlier derivations that recover predictive coding required Gaussian assumptions and produced linear, homogeneous, and often negative-valued firing rates. This paper shows that the same free-energy objective still yields predictive-coding message passing when the approximate posterior and prior are taken from the much larger exponential family, provided the third posterior cumulant is neglected. The resulting networks can host neurons with arbitrary monotonic activation functions (sigmoidal, exponential, hyperbolic, etc.), can mix those functions inside one layer, and automatically stay non-negative. The identical free-energy gradient also supplies local synaptic plasticity rules that map onto basal and apical dendritic mechanisms. A sympathetic reader therefore obtains a single normative account that both preserves the free-energy–predictive-coding link and matches several electrophysiological facts that Gaussian models could not accommodate.

What carries the argument

The EFD–FEP model: the approximate gradient of hierarchical free energy under exponential-family assumptions, which takes the predictive-coding shape τ η̇_q = G(η_q){−η_q + η_p + W_pred^⊤ ε̄} (or its natural-gradient version without G) once the third cumulant is neglected.

What would settle it

Construct a network whose posterior is strongly skewed (e.g., high-rate Poisson) and check whether free energy still decreases under the proposed predictive-coding dynamics; if free energy rises, the neglected third-cumulant term is decisive and the claim fails.

Watch

Extended reading notes

Core claim

Under hierarchical factorisation, exponential-family posterior and prior (shared base measure), fixed-variance Gaussian likelihood and linear prediction map, the negative gradient of layer-wise variational free energy with respect to natural parameters is, after dropping the third posterior cumulant, exactly of predictive-coding form: the update of each natural parameter is driven by a prior-attracting term plus bottom-up prediction error, optionally re-weighted by the Fisher information matrix. The same free-energy objective produces three local plasticity rules for the prediction, recurrent and top-down weights.

Load-bearing premise

The third-order cumulant of the posterior is small enough that ignoring it still leaves an update that both looks like predictive coding and continues to reduce free energy.

Editorial extensions

If this is right

  • Cortical circuits can implement free-energy minimisation with neurons that possess diverse, nonlinear, non-negative F–I curves without leaving the free-energy principle.
  • Heterogeneity of response properties becomes a computational resource that enlarges the class of representable posteriors rather than an obstacle.
  • Prediction and prior pathways can be learned by local rules that map onto basal Hebbian plasticity and apical calcium-mediated plasticity of pyramidal cells.
  • The ordinary-versus-natural gradient choice supplies a concrete efficiency–stability trade-off that can be tested by varying network noise.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the third-cumulant approximation remains accurate for the distributions actually used by cortex, free-energy theory no longer forces linear Gaussian units and can therefore be confronted with measured F–I diversity.
  • Geometric orthogonality between prior-regulating and error-feedback subspaces offers a testable signature: high-level beliefs should occupy directions invisible to bottom-up error alone.
  • The same free-energy construction may admit non-Gaussian likelihoods once surrogate gradients that still guarantee free-energy decrease are identified, potentially linking free energy to motifs beyond classical predictive coding.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper derives an extended predictive-coding (PC) implementation of variational free-energy (VFE) minimisation under the free-energy principle (FEP). Assuming hierarchical factorisation, exponential-family (EFD) posterior and prior (shared base measure), fixed-variance Gaussian likelihood, and a linear prediction map, the negative gradient of layer-wise VFE with respect to natural parameters η_q is shown to take PC form after neglecting the third posterior cumulant: τ η̇_q = G_q(η_q){-η_q + η_p + W_pred^ op ε̄} (ordinary gradient) or the natural-gradient version without G_q (Eqs. 14–15). The same objective yields local plasticity rules for prediction, recurrent and top-down weights (Eqs. 21–23). Type-A (factorised) models map onto heterogeneous, nonlinear F–I curves (activation functions abla A) that remain non-negative, while Type-B models are discussed more cautiously. Stochastic variants and biological correspondences (pyramidal apical/basal compartments, BAC firing) are proposed.

Significance. If the approximation is controlled, the result meaningfully widens the FEP–PC correspondence beyond the Gaussian/Laplace regime that has dominated the literature. It supplies a normative account of nonlinear and heterogeneous neuronal response properties, eliminates negative firing rates without ad-hoc rectification, and derives local Hebbian-like and dendritic-plateau-compatible plasticity from a single objective. These features address long-standing biological criticisms of Gaussian PC models and therefore strengthen FEP as an explanatory theory of cortical perceptual inference. The geometric and information-thermodynamic remarks in the Discussion are suggestive but secondary. The derivation itself is analytic and transparent; no machine-checked proofs or numerical validation are supplied.

major comments (2)
  1. [Section 3.2.1 / Appendix A] Section 3.2.1 (Eqs. 12–14) and Appendix A: the claimed PC form of free-energy descent is obtained only after discarding the third-cumulant remainder Δ_k = (1/2) M^{ij} T_ijk. Appendix A itself notes that for Poisson (and other non-sub-Gaussian) posteriors every cumulant scales as e^η, so |Δ| need not be small, and that the residual is not guaranteed to keep the approximate flow a descent direction for F. The appeal to uncertainty-weighted decay of W_pred keeping M small is heuristic, not proved. Because the central claim is precisely that the dynamics remain VFE-reducing while taking PC form, a concrete regime (bounds on |T_ijk| relative to G_q and G_φ, restriction to distributions with vanishing third cumulants, or a Lyapunov argument for the approximate vector field) is required; otherwise the FEP–PC correspondence is only formal, not variational.
  2. [Section 3.2 / Appendix B] The likelihood is kept fixed-variance Gaussian throughout the main derivation (Eq. 9). Appendix B correctly shows that a non-Gaussian EFD likelihood replaces the simple additive error with an expectation E_q[(z-μ_q)A_φ(ξ)] that lacks a local neural implementation. Consequently the advertised “extension to the exponential family” is only partial; the title and abstract should more clearly delimit that the PC correspondence still rests on a Gaussian observation model, or the main text should supply a controlled approximation for non-Gaussian likelihoods that preserves descent.
minor comments (5)
  1. [Figure 1] Figure 1 caption and surrounding text mix “firing activities” ˇx with both spike counts and rates; a single consistent interpretation (or an explicit statement that both are admissible) would help.
  2. [Section 3.2.1] The distinction between ordinary and natural gradient descent (OGD vs NGD) is introduced cleanly, yet the biological preference for one or the other is left as an efficiency–stability trade-off without even a schematic simulation; a short numerical illustration of convergence speed versus noise robustness would strengthen the claim.
  3. [Section 3.1.2] Notation for the Legendre dual occasionally switches between μ = abla A(η) and E_q[ˇx]; a single convention after Eq. 6 would reduce cognitive load.
  4. [Section 4.3] Section 4.3 lists four biologically awkward properties of error-coding neurons; the feedback-alignment suggestion is plausible but remains an empirical observation from machine learning. A brief citation to any cortical evidence for approximate weight symmetry would be useful.
  5. Typos: “i.e.the” (missing space) in the Abstract; “V oss” in Helmholtz reference; occasional missing spaces after periods in the arXiv header.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: PC-form dynamics and local plasticity are obtained by direct differentiation of the stated hierarchical VFE under explicit EFD + linear-prediction + Gaussian-likelihood assumptions, with the third-cumulant drop flagged as an approximation rather than hidden by definition.

full rationale

The central derivation (Sections 3.1–3.2) begins from a hierarchical factorisation of the generative model and approximate posterior (Eq. 4), places both posterior and prior in the exponential family with shared base measure (Eq. 5), keeps the likelihood fixed-variance Gaussian (Eq. 9), and imposes a linear prediction map ξ_ϕ = W_pred x̌. The layer-wise VFE is then written in closed form via the Legendre dual (Eq. 10). Its exact gradient with respect to the natural parameters η_q splits into a KLD piece that is exactly G_q(η_q)(−η_q + η_p) (Eq. 11) and a log-likelihood piece whose only non-PC term is the third-cumulant contribution (1/2)∇_η Tr[G_ϕ Σ_ξ] (Eq. 12). Neglecting that term yields the ordinary- and natural-gradient PC flows (Eqs. 14–15) and, by the same objective, the three local plasticity rules (Eqs. 21–23). None of these steps equates the claim to its own definition, fits free parameters to data that are later “predicted,” or rests on a self-citation uniqueness theorem. The biological correspondences (F–I curves as ∇A, apical-tuft plasticity, etc.) are post-hoc interpretations of already-derived equations and do not enter the formal argument. Appendix A itself records that the neglected remainder need not be small for Poisson-like posteriors, so the approximation is transparent rather than circular. The paper is therefore self-contained against its own stated assumptions; the only residual risk is correctness of the approximation, not circularity of the derivation chain.

Assumptions & free parameters 4 free parameters · 8 assumptions · 3 invented entities

The central PC-form claim rests on hierarchical VFE factorisation, EFD posteriors/priors with shared base measure, fixed-variance Gaussian likelihood, linear prediction map, and explicit neglect of the third posterior cumulant. Biological conclusions further rest on Type-A factorisability and interpretive maps from dual parameters to membrane/firing and from plasticity equations to apical/basal compartments. No data-fitted constants drive the main theorem; free parameters are structural (timescales, choice of A, weight architecture).

free parameters (4)
  • τ_inference, τ_plasticity
    Timescale constants in the continuous-time dynamics and learning rules; chosen by modeller, not derived.
  • Choice of log-partition A_q (per neuron or population)
    Selects the posterior family and activation function; free modelling choice that determines claimed heterogeneity.
  • Likelihood covariance G_ϕ (fixed)
    Assumed fixed Gaussian precision; scales error feedback and the decay term in W_pred learning.
  • Architecture of W_pred, W_rec, W_TD (dimensionality and sparsity)
    Linear maps are postulated; their ranks and sparsity shape which drives are orthogonal (geometric discussion in §4.1).
assumptions (8)
  • domain assumption Hierarchical Markov factorisation of generative model and mean-field factorised approximate posterior (Eq. 4).
    Standard in hierarchical PC/FEP; required to write layer-wise VFE and local message passing.
  • domain assumption Posterior and prior belong to the exponential family with shared base measure h; natural parameters fully parameterise them.
    Core modelling assumption of the paper (Section 3.1.2–3.2).
  • domain assumption Likelihood is fixed-variance Gaussian so A_ϕ is quadratic (Eq. 9).
    Retained from classical FEP–PC; Appendix B shows relaxing it breaks clean PC form.
  • domain assumption Linear prediction map ξ_ϕl(x̌_l) = W_pred_l x̌_l.
    Needed to map predictions onto synaptic weights and obtain additive error feedback.
  • ad hoc to paper Third posterior cumulant ∇³A_q is negligible in the log-likelihood gradient.
    Explicit approximation that produces the PC form (Section 3.2.1); Appendix A quantifies the residual.
  • domain assumption Prior natural parameters are linear in same-layer and higher-layer mean (or sample) activities (Eq. 18).
    Introduces W_rec and W_TD so learning rules can be written as local products.
  • domain assumption Type-A: A_q additively decomposable so G_q is diagonal and each neuron has an independent monotone activation.
    Required for the clean one-neuron-per-dimension biological mapping in Section 3.4.
  • standard math Standard EFD identities: μ = ∇A, G = ∇²A = Cov, D_KL[f∥g] = A*(μ_A) + B(η_B) − η_B^⊤ μ_A.
    Textbook exponential-family geometry used throughout Section 3.
invented entities (3)
  • EFD–FEP model (and OGD/NGD/SOGD/SNGD variants)
    purpose: Name the family of dynamics obtained by VFE descent under EFD posterior/prior.
    Organising label for the derived equations; not an extra physical object, but the paper’s central construct.
  • Type-A vs Type-B EFD–FEP subtypes
    purpose: Separate factorised (diagonal G) models with easy neural maps from non-factorised ones.
    Classification invented for implementability discussion; Type-B biological support is case-by-case speculation.
  • Internal state η_q of a representational neuron (distinct from raw input current)
    purpose: Map natural parameters onto a membrane-related dynamical variable dual to firing rate μ_q.
    Interpretive substrate; paper suggests e.g. after-hyperpolarisation recovery rate, without independent measurement protocol.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Extended predictive coding framework as variational free-energy minimisation under exponential-family assumption." pith.science (2026). https://pith.science/paper/CDO4CPYS

@misc{pith2026260530882,
  author       = {Pith},
  title        = {Pith review of: Extended predictive coding framework as variational free-energy minimisation under exponential-family assumption},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CDO4CPYS}},
  note         = {Machine review of arXiv:2605.30882}
}
read the original abstract

The sensory cortices of the brain perform perceptual inference efficiently through their complex networks of neurons. One of the theoretical accounts of this process is the free-energy principle (FEP), which postulates that the brain performs variational Bayesian inference. Pioneering studies have shown that FEP can correspond to the predictive coding (PC) hypothesis under the Gaussian assumption and Laplace approximation. However, PC-based implementations of FEP within such a limited Gaussian regime have failed to capture several properties of biological neural networks, such as nonlinearity and heterogeneity of input--output properties within a network, and the biological implausibility of negative firing rates. This study shows that, when a broader class of probability distributions, namely the exponential family of distributions (EFD), is assumed for the variational posterior and prior, these missing characteristics are exhibited within the network, maintaining the FEP--PC correspondence up to the second cumulant of the posterior. We also show that the proposed model can be trained by biologically plausible local plasticity rules. Our results enrich the explanatory power of FEP regarding neural dynamics involved in perception as variational inference.

Figures

Figures reproduced from arXiv: 2605.30882 by the authors.

Figure 1
Figure 1. Schematic illustration of EFD–FEP model derived in Section 3. Representational neurons [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

27 extracted references

  1. [1]

    The free-energy principle: a rough guide to the brain?Trends Cogn

    Karl Friston. The free-energy principle: a rough guide to the brain?Trends Cogn. Sci., 13(7): 293–301, July 2009

  2. [2]

    Predictive coding under the free-energy principle.Philos

    Karl Friston and Stefan Kiebel. Predictive coding under the free-energy principle.Philos. Trans. R. Soc. Lond. B Biol. Sci., 364(1521):1211–1221, May 2009

  3. [3]

    The free-energy principle: a unified brain theory?Nat

    Karl Friston. The free-energy principle: a unified brain theory?Nat. Rev. Neurosci., 11(2): 127–138, February 2010. 16

  4. [4]

    A tutorial on the free-energy framework for modelling perception and learning

    Rafal Bogacz. A tutorial on the free-energy framework for modelling perception and learning. J. Math. Psychol., 76(Pt B):198–211, February 2017

  5. [5]

    9 ofAllgemeine Encyklopädie der Physik

    Hermann von Helmholtz.Handbuch der physiologischen Optik, volume Bd. 9 ofAllgemeine Encyklopädie der Physik. Leopold V oss, Leipzig, 1867

  6. [6]

    Cambridge University Press, September 1996

    David C Knill, Whitman Richards, Whitman Richard, D C Knill, D Kersten, A Yuille, D Mum- ford, A Jepson, W Richards, D C Knill, J Feldman, A L Yuille, H H Bülthoff, B M Bennett, D D Hoffman, C Prakash, S N Richman, P Mamassian, A Blake, D Sheinberg, P N Belhumeur, W T Freeman, K Nakayama, S Shimojo, E H Adelson, A P Pentland, and H Barlow.Perception as Ba...

  7. [7]

    MIT Press, 2007

    Kenji Doya, Shin Ishii, Alexandre Pouget, and Rajesh P N Rao.Bayesian Brain: Probabilistic Approaches to Neural Coding. MIT Press, 2007

  8. [8]

    Bayesian brain theory: Computational neuroscience of belief.Neuroscience, 566:198–204, February 2025

    Hugo Bottemanne. Bayesian brain theory: Computational neuroscience of belief.Neuroscience, 566:198–204, February 2025

Show all 27 references
  1. [9]

    Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effects.Nat

    R P Rao and D H Ballard. Predictive coding in the visual cortex: a functional interpretation of some extra-classical receptive-field effects.Nat. Neurosci., 2(1):79–87, January 1999

  2. [10]

    A new cellular mechanism for coupling inputs arriving at different cortical layers.Nature, 398(6725):338–341, March 1999

    M E Larkum, J J Zhu, and B Sakmann. A new cellular mechanism for coupling inputs arriving at different cortical layers.Nature, 398(6725):338–341, March 1999

  3. [11]

    A cellular mechanism for cortical associations: an organizing principle for the cerebral cortex.Trends Neurosci., 36(3):141–151, March 2013

    Matthew Larkum. A cellular mechanism for cortical associations: an organizing principle for the cerebral cortex.Trends Neurosci., 36(3):141–151, March 2013

  4. [12]

    Conjunctive input processing drives feature selectivity in hippocampal CA1 neurons.Nat

    Katie C Bittner, Christine Grienberger, Sachin P Vaidya, Aaron D Milstein, John J Macklin, Junghyup Suh, Susumu Tonegawa, and Jeffrey C Magee. Conjunctive input processing drives feature selectivity in hippocampal CA1 neurons.Nat. Neurosci., 18(8):1133–1142, August 2015

  5. [13]

    Implications of neuronal diversity on population coding

    Maoz Shamir and Haim Sompolinsky. Implications of neuronal diversity on population coding. Neural Comput., 18(8):1951–1986, August 2006

  6. [14]

    Intrinsic biophysical diversity decorrelates neuronal firing while increasing information content.Nat

    Krishnan Padmanabhan and Nathaniel N Urban. Intrinsic biophysical diversity decorrelates neuronal firing while increasing information content.Nat. Neurosci., 13(10):1276–1282, October 2010

  7. [15]

    Population diversity and function of hyperpolarization- activated current in olfactory bulb mitral cells.Sci

    Kamilla Angelo and Troy W Margrie. Population diversity and function of hyperpolarization- activated current in olfactory bulb mitral cells.Sci. Rep., 1(1):50, July 2011

  8. [16]

    Multivariate analysis of electrophysiological diversity of xenopus visual neurons during development and plasticity.Elife, 4, November 2015

    Christopher M Ciarleglio, Arseny S Khakhalin, Angelia F Wang, Alexander C Constantino, Sarah P Yip, and Carlos D Aizenman. Multivariate analysis of electrophysiological diversity of xenopus visual neurons during development and plasticity.Elife, 4, November 2015

  9. [17]

    Diversity amongst human cortical pyramidal neurons revealed via their sag currents and frequency preferences.Nat

    Homeira Moradi Chameh, Scott Rich, Lihua Wang, Fu-Der Chen, Liang Zhang, Peter L Carlen, Shreejoy J Tripathy, and Taufik A Valiante. Diversity amongst human cortical pyramidal neurons revealed via their sag currents and frequency preferences.Nat. Commun., 12(1):2497, May 2021

  10. [18]

    Learning probability distributions of sensory inputs with Monte Carlo predictive coding.PLOS Computational Biology, 20(10): e1012532, October 2024

    Gaspard Oliviers, Rafal Bogacz, and Alexander Meulemans. Learning probability distributions of sensory inputs with Monte Carlo predictive coding.PLOS Computational Biology, 20(10): e1012532, October 2024. ISSN 1553-7358. doi: 10.1371/journal.pcbi.1012532. URL https:// journals...

  11. [19]

    Active inference and agency.Cogn

    Karl Friston. Active inference and agency.Cogn. Neurosci., 5(2):119–121, April 2014

  12. [20]

    The markov blankets of life: autonomy, active inference and the free energy principle.J

    Michael Kirchhoff, Thomas Parr, Ensor Palacios, Karl Friston, and Julian Kiverstein. The markov blankets of life: autonomy, active inference and the free energy principle.J. R. Soc. Interface, 15(138):20170792, January 2018

  13. [21]

    Active inference

    Thomas Parr, Giovanni Pezzulo, and Karl J Friston. Active inference. https://mitpress. mit.edu/9780262362283/active-inference/, December 2021. Accessed: 2026-3-7. 17

  14. [22]

    Life as we know it.J

    Karl Friston. Life as we know it.J. R. Soc. Interface, 10(86):20130475, September 2013

  15. [23]

    Applied Mathematical Sciences

    Shun-Ichi Amari.Information Geometry and Its Applications. Applied Mathematical Sciences. Springer, Tokyo, Japan, 1 edition, February 2016

  16. [24]

    Laws of thermodynamics for exponential families.arXiv [cond-mat.stat- mech], January 2025

    Akshay Balsubramani. Laws of thermodynamics for exponential families.arXiv [cond-mat.stat- mech], January 2025

  17. [25]

    Thermodynamics of prediction.Phys

    Susanne Still, David A Sivak, Anthony J Bell, and Gavin E Crooks. Thermodynamics of prediction.Phys. Rev. Lett., 109(12):120604, September 2012

  18. [26]

    On the thermodynamics of prediction under dissipative adaptation.arXiv [q-bio.NC], September 2020

    Kai Ueltzhöffer. On the thermodynamics of prediction under dissipative adaptation.arXiv [q-bio.NC], September 2020

  19. [27]

    C. Beck. Superstatistics: theory and applications.Continuum Mechanics and Thermodynamics, 16(3):293–304, March 2004. ISSN 1432-0959. doi: 10.1007/s00161-003-0145-1. URL http://dx.doi.org/10.1007/s00161-003-0145-1. A On third cumulant-neglecting approximation In Section 3.2.1, ...

Pith tools

Reviewed July 12, 2026 · model on record in the stance chip above.