REVIEW 3 major objections 4 minor 38 references
State-Dependent Observation Noise Reintroduces Epistemic Value in Linear-Gaussian Active Inference
T0 review · 3 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read Letting observation noise depend on the state restores the epistemic drive in linear-Gaussian active inference.
desk verdict Clean pinning lemma and honest H3' caveat; a genuinely new observation-side boundary for the EFE collapse, though the general epistemic-value claim is conditional and the exact-inference question is left open. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the state-dependent observation noise covariance R(x) evaluated at the predicted belief mean μ⁻, the first-order Gaussian filter's evaluation rule. The action-dependence enters through μ⁻: a policy moves the mean, the mean selects the noise level, and the noise level sets the innovation covariance S = CΣ⁻Cᵀ + R(μ⁻), the posterior covariance Σ⁺ = Σ⁻ − Σ⁻CᵀS⁻¹CΣ⁻, and the gain. The per-step epistemic value has the closed form ε = ½ ln det(CΣ⁻Cᵀ + R(μ⁻)) − ½ ln det R(μ⁻), which depends on policy precisely through μ⁻; this identity is what the constancy collapse exploits when R is fixed and what the paper exploits when R varies. The supporting mechanism is a pinning/in
What would settle it
Run the paper's scalar example (A = B = C = Q = 1, prior variance 1, R(x) = 1 + x²) and compare one-step actions u = 0 and u = 2: the claim predicts posterior variances 2/3 and 2/7 and one-step epistemic values of ½ ln 3 ≈ 0.55 and ½ ln(7/5) ≈ 0.17 nats. Alternatively, exhibit any fixed linear-Gaussian filter—a policy-independent schedule of noise covariances—that reproduces the agent's posterior means and covariances on every observation sequence; the theorem asserts none exists, and the paper's own witness treats such an exhibit as a refutation.
Extended reading notes
Core claim
The paper's central claim is that state-dependent observation noise is a genuine, minimal escape from the linear-Gaussian flattening result. In the standard model—linear dynamics, linear observation map, Gaussian noises, additive control—replace the fixed observation covariance R with a continuous state-dependent R(x), and let the agent run the usual first-order Gaussian filter in which R is evaluated at the predicted mean μ⁻. Because actions shift μ⁻, they shift the innovation covariance S = CΣ⁻Cᵀ + R(μ⁻), and therefore the posterior covariance and the gain. A pinning lemma shows that any candidate linear-Gaussian filter that reproduces the agent's posterior covariance at one step must be u
Load-bearing premise
The load-bearing premise is that the agent's beliefs are correctly described by the standard first-order Gaussian filter—R evaluated at the predicted mean and a Gaussian posterior covariance maintained throughout—rather than by the exact non-Gaussian posterior under state-dependent noise; if exact Bayesian inference is demanded, the flattening notion, the pinning lemma, and the theorem would all need to be re-derived, a step the paper explicitly leaves open.
Editorial extensions
If this is right
- A fixed observation noise R is effectively a commitment to a certainty-equivalent agent: the epistemic term is constant, so any active-inference model that keeps R constant has an information drive no policy can actually exercise.
- No fixed linear-Gaussian filter, even one permitted a policy-independent schedule of noise covariances chosen in advance, can reproduce the beliefs of an agent with reachable state-dependent observation noise: the model class cannot be flattened.
- For scalar observations, reachable non-constancy of R alone is sufficient for non-constant epistemic value, because the epistemic map is strictly monotone in the noise variance; no extra visibility condition is needed.
- The existence of a planning reduction, the stepwise constancy of epistemic value, and the absence of a second-order dual effect are equivalent in this model class, so reintroducing epistemic value and reintroducing the dual effect are the same act.
- The boundary of the linear-Gaussian collapse sits on the observation model: only the fact that R depends on a controllable state is needed to make information gain action-sensitive—not nonlinear dynamics, not multiplicative control, not a modified objective.
Reading between the lines
- Inference: if this theorem generalizes to the hierarchical, nonlinear settings where state-dependent sensory precision is already used as a model of attention, then attention and information-seeking would be the same mechanism whenever precision sits downstream of controllable states—something the paper raises as a conjecture.
- Inference: a natural companion result would treat nonlinear observation maps h(x) with fixed noise—the other classic source of the dual effect; an analogous pinning argument should show that it too blocks flattening, completing the observation-side anatomy.
- Inference: if an agent has to learn R(x) from experience rather than receive it, curiosity could emerge as a by-product of fitting the noise model, since estimating where the sensor is reliable automatically makes some actions more informative than others—a testable prediction about learning dynamics.
- Inference: the paper's horizon-1 numbers (an information pull of about 1.7 nats toward the informative cue versus a pragmatic gradient of about 4.5 nats) make a concrete quantitative prediction that a two-step planner should eventually take the detour, turning the claimed local epistemic force into a global one.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies a linear-Gaussian state-space model in which the observation noise covariance is a state-dependent function R(x) and the agent uses a first-order Gaussian filter that evaluates R at the predicted mean (Section 3.2). Its main result (Theorem 1, Section 3.3) states that under positive definiteness, full observation rank, and reachable non-constancy of R, the posterior covariance and Kalman gain depend on the policy; no fixed linear-Gaussian reduction (Definition 1) exists; and under an additional informational non-constancy assumption H3' the per-step epistemic value of the expected free energy is non-constant across policies. Corollary 1 removes H3' for scalar observations. Corollary 2 relates planning reductions to the absence of a second-order dual effect. The paper also presents cpomdp, an archived software witness that raises an IncompatibleLinearizationError for such models, and two demonstrations.
Significance. If established, the paper provides a clean, observation-side counterexample to the Koudahl et al. collapse: a model that remains linear-Gaussian in dynamics, observation map, and noise family, yet whose agent has action-dependent information gain. The proof is concise and transparent; Lemma 1's pinning argument is the right tool, and the executable witness with check gates is a genuine reproducibility asset. The paper is also careful to flag its own limitations (Remarks 2–3, Discussion), which is commendable. The main caveats are the scope of the formal result to the first-order Gaussian filter and the conditional nature of part (iii) for n_o>1.
major comments (3)
- [Theorem 1(iii), Eq. (10), H3' in §3.3] Theorem 1(iii) is, for n_o>1, a restatement of the assumption H3' rather than a derivation. H3' is defined as the existence of policies whose per-step epistemic values differ at the first separating time, and the proof says only that 'H3' says precisely that these two quantities differ'. The scalar Corollary 1 gives a genuine sufficient condition, and Remark 6 gives a Loewner-comparability condition, but neither is the theorem's general hypothesis. Since Example 1 shows H3 does not imply H3', the abstract's phrase 'under a mild non-degeneracy condition' is accurate only if H3' is itself accepted as that condition — but H3' is the claimed phenomenon. I recommend stating part (iii) as an explicit conditional, or proving a nontrivial sufficient condition on R and C that implies H3'.
- [§3.2 and Remark 2; Discussion limitations] The central claim is proved for the first-order Gaussian filter that evaluates R at the predicted mean μ^-, not for the exact non-Gaussian posterior of the generative model (6). Remark 2 explicitly leaves open whether the exact conditional covariance carries a strict Bar-Shalom–Tse dual effect, and the Discussion's fourth limitation says the distance to the exact filter is unquantified. The title and abstract nevertheless state unqualifiedly that state-dependent observation noise itself reintroduces epistemic value in linear-Gaussian active inference. This matters because a skeptic could attribute the non-constancy to the evaluation rule. My own reading is that the mechanism is probably robust — for R(x)=1+x^2 the exact predictive entropy includes E log R(x), which depends on the controllable mean — but the paper should either prove that or restrict the title/abstract claims to the first
- [Definition 1 vs abstract, §1 and §3.2] Definition 1's 'fixed linear-Gaussian reduction' is specifically a Kalman filter with a policy-independent noise schedule (A,B,C,Q,R̄_k). The abstract and Section 1 say 'no fixed linear-Gaussian filter reproduces the agent'. A general linear-Gaussian filter with arbitrary policy-independent gains is not ruled out by the theorem as stated; the theorem rules out the Kalman-filter-with-named-covariance-schedule object. The proof is correct for the stated definition, but the unqualified phrase in the abstract is stronger. Please either broaden the theorem to arbitrary policy-independent linear filters or add the Definition 1 qualifier to the abstract and introduction.
minor comments (4)
- [§2.1, §2.2] Minor wording: 'readers from active inference will find Section 2.1 familiar, whilst control theorists Section 2.2' is missing a verb after 'control theorists'.
- [Remark 3] The claim that Gaussian smoothing is 'injective, so the smoothed landscape is non-constant exactly when R is' is asserted without proof or reference. It is plausible, but not immediate for arbitrary continuous R; please add a short argument or citation.
- [§5, T-maze] The T-maze demonstration is not formally covered by Theorem 1, as the paper acknowledges. Consider labeling it explicitly as a numerical illustration rather than an executable version of the theorem, to avoid confusion with the 'witness' language used for the single-chain model.
- [Notation, Table 1] The symbol ε_k(π) is used in Theorem 1 before its definition in Table 1; a forward pointer in §3.3 would help.
Circularity Check
Theorem 1(iii) assumes H3', which is defined as the very non-constancy it 'proves'; the general epistemic-value claim is circular-by-definition, while the scalar corollary and no-flattening result remain genuine.
-
self definitional
[Section 3.3, H3' and Theorem 1(iii); proof of Theorem 1(iii) in Section 4]
"H3' Informational non-constancy. Grant H3 and let k∗ be the first time at which R differs across the predicted means of two policies. There exist policies π, π′ whose per-step epistemic values differ there: ε_k∗(π) ≠ ε_k∗(π′). ... H3' says precisely that these two quantities differ between π and π′ at k = k∗, so ε_k∗(π) ≠ ε_k∗(π′)."
H3' is not derived; it is defined as the existence of two policies with differing per-step epistemic values. Theorem 1(iii) then assumes H3' and concludes exactly that existence, with the proof stating 'H3' says precisely that...'. For n_o > 1, the epistemic-value claim is thus a restatement of the assumption, not a derivation from the state-dependent noise mechanism. The paper's own Remark 4 acknowledges the H3/H3' gap, but acknowledgment does not remove the definitional character. The non-circular content is Corollary 1 (scalar monotonicity supplies H3' from H3) and parts (i)-(ii); the abstract's general claim of restored epistemic value under a 'non-degeneracy condition' is the target repackaged as an assumption.
full rationale
The paper's main derivation chain is otherwise self-contained. Lemma 1's pinning argument is proved from the covariance update (8)-(9), and Theorem 1(i)-(ii) genuinely follow from H1-H3 without importing the conclusion. Corollary 1 is a real result: for scalar observations, strict monotonicity makes H3 imply H3', so epistemic non-constancy is established rather than assumed. The cpomdp witness is software built for this paper, but it is not used to prove the theorem, so the self-citation is not load-bearing. The paper also transparently limits the result to the agent's maintained Gaussian recursion rather than the exact non-Gaussian posterior (Remark 2; Section 6 limitations), and I do not count that scope limitation as circularity. The one clear reduction-by-construction is H3': the central general epistemic-value claim of Theorem 1(iii) is true by definition because H3' already asserts the existence of policies with differing per-step epistemic values. This warrants a score of 6 rather than higher because the non-flattening theorem and the scalar case are independent, substantive results.
Assumptions & free parameters
assumptions (6)
- domain assumption H1: R(x) is continuous and positive definite for every x.
- domain assumption H2: the observation map C has full row rank.
- domain assumption H3: R is non-constant on the reachable set of predicted means.
- domain assumption H3': there exist policies with differing per-step epistemic values at the first separating time.
- ad hoc to paper The agent's inference is the first-order Gaussian filter: R_k = R(mu_k^-), and beliefs are the Gaussian posterior mean/covariance from the standard Kalman update.
- standard math Standard closed forms for Gaussian mutual information / EFE and the information-form Kalman update.
Cite this review
Pith. "Pith review of State-Dependent Observation Noise Reintroduces Epistemic Value in Linear-Gaussian Active Inference." pith.science (2026). https://pith.science/paper/SO2MXHKQ
@misc{pith2026260720306,
author = {Pith},
title = {Pith review of: State-Dependent Observation Noise Reintroduces Epistemic Value in Linear-Gaussian Active Inference},
year = {2026},
howpublished = {\url{https://pith.science/paper/SO2MXHKQ}},
note = {Machine review of arXiv:2607.20306}
}
read the original abstract
Recent work established that under active inference, linear-Gaussian state-space models lose their epistemic drive (any incentive to act so as to gain information) "under any circumstances". The epistemic term of the Expected Free Energy becomes constant: the agent flattens to a Kalman filter whose gain sequence is fixed in advance, regardless of action. The minimal departure that restores the drive is unknown; the only established route is control entering the dynamics multiplicatively; the observation side of this boundary is unexplored. We show that state-dependent observation noise is such a departure: a covariance R(x) that varies with the state x, representing a sensor's accuracy degrading with range. The agent runs the standard first-order Gaussian filter of this literature, R evaluated at the predicted mean. Coupling R(x) to a controllable latent mean makes the posterior covariance, and hence the effective Kalman gain, depend on the action. Consequently, no fixed linear-Gaussian filter reproduces the agent and, under a mild rank condition on the observation map and a non-degeneracy condition on R(x), epistemic value is no longer constant; for scalar observations, reachable non-constancy alone is needed. This is a minimal constructive instance of the Bar-Shalom-Tse dual effect in the agent's maintained covariance: actions now influence the quality of future estimates, not merely the state. Our library cpomdp detects the incompatibility from model specification alone and raises a typed IncompatibleLinearizationError. The theorem ships with an executable witness: exhibiting any fixed filter that reproduced the agent's beliefs would refute both theorem and witness at once. Together this offers a precise, observation-side characterisation of curiosity in a Gaussian agent, bridging dual control and active inference.
Reference graph
Works this paper leans on
-
[1]
On epistemics in expected free energy for linear Gaussian state space models
Koudahl, M.T.; Kouw, W.M.; de Vries, B. On epistemics in expected free energy for linear Gaussian state space models. Entropy 2021, 23, 1565. https://doi.org/10.3390/e23121565 41
-
[2]
Koudahl, M.T.; van de Laar, T.; de Vries, B. Realising synthetic active inference agents, Part I: Epistemic objectives and graphical specification language. arXiv 2023, arXiv:2306.08014
arXiv 2023
-
[3]
Realizing synthetic active inference agents, Part II: Variational message updates
van de Laar, T.; Koudahl, M.; de Vries, B. Realizing synthetic active inference agents, Part II: Variational message updates. Neural Comput. 2024, 37, 38–75. https://doi.org/10.1162/neco_a_01713
-
[4]
Dual effect, certainty equivalence, and separation in stochastic control
Bar-Shalom, Y.; Tse, E. Dual effect, certainty equivalence, and separation in stochastic control. IEEE Trans. Autom. Control 1974, 19, 494–500. https://doi.org/10.1109/TAC.1974.1100635
arXiv 1974
-
[5]
Nonlinear estimation with state-dependent Gaussian observation noise
Spinello, D.; Stilwell, D.J. Nonlinear estimation with state-dependent Gaussian observation noise. IEEE Trans. Autom. Control 2010, 55, 1358–1366. https://doi.org/10.1109/TAC.2010.2042006
arXiv 2010
-
[6]
Dual control theory
Feldbaum, A.A. Dual control theory. I. Avtom. Telemekh. 1960, 21, 1240–1249
1960
-
[7]
Feldbaum, A.A. Dual control theory problems. IFAC Proc. Vol. 1963, 1, 541–550. https://doi.org/10.1016/S1474-6670(17)69687-3
-
[8]
Active inference on discrete state- spaces: A synthesis
Da Costa, L.; Parr, T.; Sajid, N.; Veselic, S.; Neacsu, V.; Friston, K. Active inference on discrete state- spaces: A synthesis. J. Math. Psychol. 2020, 99, 102447. https://doi.org/10.1016/j.jmp.2020.102447
arXiv 2020
Show all 38 references
-
[9]
Whence the expected free energy? Neural Comput
Millidge, B.; Tschantz, A.; Buckley, C.L. Whence the expected free energy? Neural Comput. 2021, 33, 447–482. https://doi.org/10.1162/neco_a_01354
2021 doi
-
[10]
Active inference and epistemic value in graphical models
van de Laar, T.; Koudahl, M.; van Erp, B.; de Vries, B. Active inference and epistemic value in graphical models. Front. Robot. AI 2022, 9, 794464. https://doi.org/10.3389/frobt.2022.794464
2022
-
[11]
Dual effect, certainty equivalence, and separation revisited: A counterexample and a relaxed characterization for optimality
Derpich, M.S.; Yüksel, S. Dual effect, certainty equivalence, and separation revisited: A counterexample and a relaxed characterization for optimality. IEEE Trans. Autom. Control 2023, 68, 1259–1266. https://doi.org/10.1109/TAC.2022.3151189
2023
-
[12]
A counterexample in stochastic optimum control
Witsenhausen, H.S. A counterexample in stochastic optimum control. SIAM J. Control 1968, 6, 131–
1968
-
[13]
A generalization of the Kalman filter for models with state-dependent observation variance
Zehnwirth, B. A generalization of the Kalman filter for models with state-dependent observation variance. J. Am. Stat. Assoc. 1988, 83, 164–167. https://doi.org/10.1080/01621459.1988.10478582
1988
-
[14]
State-dependent Kalman filters for robust engine control
Dutka, A.S.; Javaherian, H.; Grimble, M.J. State-dependent Kalman filters for robust engine control. In Proceedings of the 2006 American Control Conference, Minneapolis, MN, USA, 14–16 June 2006; pp. 1185–1190. https://doi.org/10.1109/ACC.2006.1656378
2006 arXiv
-
[15]
Revisiting active perception
Bajcsy, R.; Aloimonos, Y.; Tsotsos, J.K. Revisiting active perception. Auton. Robots 2018, 42, 177–196. https://doi.org/10.1007/s10514-017-9615-3
2018 doi
-
[16]
The belief roadmap: Efficient planning in belief space by factoring the covariance
Prentice, S.; Roy, N. The belief roadmap: Efficient planning in belief space by factoring the covariance. Int. J. Robot. Res. 2009, 28, 1448–1465. https://doi.org/10.1177/0278364909341659
2009 doi
-
[17]
Belief space planning assuming maximum likelihood observations
Platt, R.; Tedrake, R.; Kaelbling, L.P.; Lozano-Pérez, T. Belief space planning assuming maximum likelihood observations. In Proceedings of Robotics: Science and Systems VI, Zaragoza, Spain, 27–30 June 2010. https://doi.org/10.15607/RSS.2010.VI.037
2010 doi
-
[18]
Motion planning under uncertainty using iterative local optimization in belief space
van den Berg, J.; Patil, S.; Alterovitz, R. Motion planning under uncertainty using iterative local optimization in belief space. Int. J. Robot. Res. 2012, 31, 1263–1278. https://doi.org/10.1177/0278364912456319
2012 doi
-
[19]
Attention, uncertainty, and free-energy
Feldman, H.; Friston, K.J. Attention, uncertainty, and free-energy. Front. Hum. Neurosci. 2010, 4, 215. https://doi.org/10.3389/fnhum.2010.00215
2010 arXiv
-
[20]
Perceptions as hypotheses: Saccades as experiments
Friston, K.; Adams, R.A.; Perrinet, L.U.; Breakspear, M. Perceptions as hypotheses: Saccades as experiments. Front. Psychol. 2012, 3, 151. https://doi.org/10.3389/fpsyg.2012.00151
2012 arXiv
-
[21]
Working memory, attention, and salience in active inference
Parr, T.; Friston, K.J. Working memory, attention, and salience in active inference. Sci. Rep. 2017, 7, 14678. https://doi.org/10.1038/s41598-017-15249-0 43
2017 doi
-
[22]
Introducing a Bayesian model of selective attention based on active inference
Mirza, M.B.; Adams, R.A.; Friston, K.; Parr, T. Introducing a Bayesian model of selective attention based on active inference. Sci. Rep. 2019, 9, 13915. https://doi.org/10.1038/s41598-019-50138- 8
2019 doi
-
[23]
Codes on graphs: Normal realizations
Forney, G.D. Codes on graphs: Normal realizations. IEEE Trans. Inf. Theory 2001, 47, 520–548. https://doi.org/10.1109/18.910573
2001 doi
-
[24]
Factor graphs and the sum-product algorithm
Kschischang, F.R.; Frey, B.J.; Loeliger, H.-A. Factor graphs and the sum-product algorithm. IEEE Trans. Inf. Theory 2001, 47, 498–519. https://doi.org/10.1109/18.910572
2001 doi
-
[25]
An introduction to factor graphs
Loeliger, H.-A. An introduction to factor graphs. IEEE Signal Process. Mag. 2004, 21, 28–41. https://doi.org/10.1109/MSP.2004.1267047
2004 arXiv
-
[26]
The factor graph approach to model-based signal processing
Loeliger, H.-A.; Dauwels, J.; Hu, J.; Korl, S.; Ping, L.; Kschischang, F.R. The factor graph approach to model-based signal processing. Proc. IEEE 2007, 95, 1295–1322. https://doi.org/10.1109/JPROC.2007.896497
2007
-
[27]
Simulating active inference processes by message passing
van de Laar, T.W.; de Vries, B. Simulating active inference processes by message passing. Front. Robot. AI 2019, 6, 20. https://doi.org/10.3389/frobt.2019.00020
2019
-
[28]
Variational message passing and local constraint manipulation in factor graphs
Şenöz, İ.; van de Laar, T.; Bagaev, D.; de Vries, B. Variational message passing and local constraint manipulation in factor graphs. Entropy 2021, 23, 807. https://doi.org/10.3390/e23070807
2021 doi
-
[29]
RxInfer: A Julia package for reactive real-time Bayesian inference
Bagaev, D.; Podusenko, A.; de Vries, B. RxInfer: A Julia package for reactive real-time Bayesian inference. J. Open Source Softw. 2023, 8, 5161. https://doi.org/10.21105/joss.05161
2023 doi
-
[30]
Generalised filtering
Friston, K.; Stephan, K.; Li, B.; Daunizeau, J. Generalised filtering. Math. Probl. Eng. 2010, 2010, 621670. https://doi.org/10.1155/2010/621670
2010 doi
-
[31]
Kalman filters as the steady-state solution of gradient descent on variational free energy
Baltieri, M.; Isomura, T. Kalman filters as the steady-state solution of gradient descent on variational free energy. arXiv 2021, arXiv:2111.10530
2021 arXiv
-
[32]
Optimal Filtering; Prentice-Hall: Englewood Cliffs, NJ, USA, 1979
Anderson, B.D.O.; Moore, J.B. Optimal Filtering; Prentice-Hall: Englewood Cliffs, NJ, USA, 1979. 44
1979
-
[33]
Linear Estimation; Prentice-Hall: Upper Saddle River, NJ, USA, 2000
Kailath, T.; Sayed, A.H.; Hassibi, B. Linear Estimation; Prentice-Hall: Upper Saddle River, NJ, USA, 2000
2000
-
[34]
Bayesian Filtering and Smoothing, 2nd ed.; Cambridge University Press: Cambridge, UK, 2023
Särkkä, S.; Svensson, L. Bayesian Filtering and Smoothing, 2nd ed.; Cambridge University Press: Cambridge, UK, 2023. https://doi.org/10.1017/9781108917407
2023 doi
-
[35]
cpomdp, version 0.4.2; Zenodo, 2026
Corva, D. cpomdp, version 0.4.2; Zenodo, 2026. https://doi.org/10.5281/zenodo.21429863
2026 doi
-
[36]
pymdp: A Python library for active inference in discrete state spaces
Heins, C.; Millidge, B.; Demekas, D.; Klein, B.; Friston, K.; Couzin, I.D.; Tschantz, A. pymdp: A Python library for active inference in discrete state spaces. J. Open Source Softw. 2022, 7, 4098. https://doi.org/10.21105/joss.04098
2022 doi
-
[37]
Active inference and epistemic value
Friston, K.; Rigoli, F.; Ognibene, D.; Mathys, C.; Fitzgerald, T.; Pezzulo, G. Active inference and epistemic value. Cogn. Neurosci. 2015, 6, 187–214. https://doi.org/10.1080/17588928.2015.1020053
2015
-
[147]
https://doi.org/10.1137/0306011 42
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.