Pith. sign in

REVIEW 4 major objections 5 minor 40 references

Mutual Information Rate -- Linear Noise Approximation and Exact Computation

T0 review · 4 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read For discrete biochemical networks the Gaussian approximation never reaches the exact information rate, and in nonlinear continuous systems its bias direction depends on the linearization route.

desk verdict Useful, clearly written, and honest, but the missing convergence diagnostics for the PWS ground truth make the novel quantitative claims—especially the low-copy-number optimum—provisional. read the letter →

arxiv 2508.21220 v1 pith:Z7SVVCOF submitted 2025-08-28 q-bio.MN cs.ITmath.ITphysics.bio-ph

classification q-bio.MNcs.ITmath.ITphysics.bio-ph
keywords mutualinformationrateGaussianapproximationlinearnoisePathWeightSamplingstochasticreactionnetworksLangevindynamicstheorybiochemicalsignaling
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tests when the standard shortcut for computing the mutual information rate—the amount of information shared per unit time between input and output trajectories—can be trusted for systems that are not actually Gaussian. Using Path Weight Sampling, an exact Monte Carlo method, the authors show that for a simple discrete biochemical reaction network the Gaussian approximation is only a lower bound and stays a lower bound even when copy numbers are large and statistics look nearly Gaussian. For a continuous nonlinear Langevin system, the Gaussian approximation is biased in opposite directions depending on how it is applied: linearizing the dynamics analytically overestimates the true rate, while estimating spectra from simulations underestimates it. The practical payoff is a set of concrete usage conditions: rely on the Gaussian approximation for weak nonlinearity or slow output responses, and switch to exact methods for low input copy numbers or fast, strongly nonlinear systems.

What carries the argument

The load-bearing object is the mutual information rate R(S,X), the long-time limit of trajectory mutual information. Path Weight Sampling (PWS) estimates it exactly by Monte Carlo: sample input trajectories s_T from P(s_T), output trajectories x_T from P(x_T|s_T), and average log P(x_T|s_T) − log P(x_T); for Langevin dynamics the path weight is the Onsager–Machlup action. The Gaussian approximation replaces trajectory statistics by a jointly Gaussian model, reducing the rate to a frequency integral over the coherence φ_sx(ω)=|S_sx(ω)|²/(S_ss(ω)S_xx(ω)). The two variants differ in how the spectra are obtained: analytically from the linear noise approximation (LNA), or numerically from simulat

What would settle it

Recompute the four-reaction network's information rate with an independent exact method—e.g., numerical integration of the stochastic filtering equation—at large copy numbers (s̄,x̄ ≥ 100) and compare with the LNA closed form; agreement would refute the claim that the Gaussian bound never tightens. For the nonlinear Langevin system, scan parameters with fast response and high gain and check whether the empirical Gaussian estimate ever exceeds the PWS rate; the paper predicts it never does.

Watch

Extended reading notes

Core claim

The paper's central claim is that the Gaussian approximation of the mutual information rate—exact only for linear systems with additive Gaussian noise—systematically misestimates the true rate in non-Gaussian systems. In a four-reaction linear network, the LNA-based Gaussian rate R=(λ/2)(√(1+ρ/λ)−1) is a strict lower bound that does not tighten as copy numbers grow; a reaction-based discrete approximation R=(λ/2)(√(1+2ρ/λ)−1) matches exact PWS results over nearly all parameters, except low input copy numbers where the true rate peaks above both. In a continuous Langevin system with a Hill-type nonlinearity, linearizing dynamics via the LNA overestimates the exact rate, while estimating spect

Load-bearing premise

The conclusions treat the Path Weight Sampling results as ground truth, but PWS is exact only with infinitely many Monte Carlo samples and, for the Langevin extension, an infinitesimal time step; the paper does not report convergence checks, so any systematic bias in the PWS estimates would change the measured deviations of the Gaussian approximation.

Editorial extensions

If this is right

  • For linear discrete reaction networks, the Gaussian approximation is a strict lower bound on the true information rate and never becomes tight at high copy numbers; estimates from it should be interpreted as a floor, not a prediction.
  • The discrete approximation of [4] is accurate over a wide parameter range (output copy numbers down to ≪1 and input copy numbers ≳10), making it a cheap replacement for exact computation in those regimes.
  • For continuous nonlinear systems, the empirical Gaussian estimate is a lower bound for Gaussian inputs, while the LNA-based estimate is not bounded; across the tested parameters the empirical route is closer to the exact rate.
  • Practical usage rule: Gaussian approximations are reliable for small gain or slow output response; exact methods like PWS are needed for low input copy numbers (≲10) in discrete networks and for fast, strongly nonlinear continuous systems.
  • The true information rate of the discrete network has an optimum at low input copy number (s̄ ≈ 1), above both analytic approximations, so reducing input copy number can increase information transmission in this regime.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The distinction between reaction-based and state-based readouts implies the Gaussian bound is the information accessible to a decoder that cannot resolve individual reaction events; a testable prediction is that a downstream decoder with finer time resolution approaches the higher PWS rate.
  • The opposite bias directions likely extend beyond Hill kinetics to any saturating input–output relation, so repeating the comparison with Michaelis–Menten or exponential saturation would test whether the asymmetry is generic.
  • The low-input-copy-number peak suggests intrinsic noise can carry information rather than only degrade it; if real signaling networks tune input copy numbers near this optimum, that would be evidence of noise exploitation.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The manuscript evaluates the Gaussian approximation (GA) to the mutual information rate in two model problems, using Path Weight Sampling (PWS) as an exact benchmark. For a linear discrete reaction network, the LNA-based GA is shown to be a lower bound that is not tight even at large copy numbers, whereas a recently proposed 'discrete approximation' agrees with PWS except at low input copy numbers, where PWS indicates a nonmonotonic maximum. For a nonlinear Langevin system with Hill-type activation, the LNA-based GA overestimates and an empirical Gaussian estimate underestimates the exact rate, with errors increasing with gain and decreasing with response time. The paper derives closed-form expressions, discusses the mechanism behind the discreteness effect, and gives practical guidance on when approximate methods suffice.

Significance. If the numerical results are reliable, the paper provides a useful quantitative caution against routine use of the Gaussian approximation and gives concrete guidance. Its analytic derivations are self-consistent, and benchmarking against an independent exact Monte Carlo method is appropriate. The extension of PWS to Langevin dynamics is a useful methodological contribution. However, the central quantitative claims—especially the low-copy-number optimum and the signed deviations in the nonlinear system—rest entirely on PWS estimates for which no convergence evidence or uncertainty is reported. As it stands, the paper's empirical conclusions are plausible but not fully substantiated.

major comments (4)
  1. [§II.C, Eqs. (9), (11); §III.A–B] The PWS benchmark is defined by three limits: N→∞ in Eq. (9), T→∞ in Eq. (4), and Δt→0 in Eq. (11). None of these limits is documented for any figure: no sample sizes, trajectory lengths, time steps, convergence checks, or error bars are reported. This is load-bearing because the central quantitative claims—the non-tight lower bound in Fig. 1a, the nonmonotonic optimum in Fig. 1b, and the signed deviations in Figs. 3–4—are deviations from PWS. Without evidence that the PWS estimates have converged, the reported deviations could be numerical artifacts. Please add convergence diagnostics and error bars/confidence intervals for all PWS points.
  2. [§III.A, Fig. 1b] The low-copy-number maximum of the information rate is a central new observation, yet the paper itself states that 'a precise characterization of this finding' is left for future work. Because both analytic approximations are independent of κ, the entire phenomenon rests on the PWS points for small κ. The figure shows no error bars and no dependence on trajectory length or sample size. The authors should show that the nonmonotonicity is stable under increasing T and N, and provide uncertainty estimates, before concluding that low input copy numbers maximize the rate.
  3. [§III.B, Eq. (24), Fig. 4] The empirical Gaussian rate is computed by Welch's method from simulated trajectories. Finite-sample spectral and coherence estimates are biased; the manuscript does not report window length, overlap, number of independent realizations, or statistical uncertainty. The claimed lower-bound behavior relies on the asymptotic Mitra–Stark argument, so the numerical verification in Fig. 4 must be shown not to be an artifact of spectral estimation bias. Please provide the estimation parameters and error bars, and ideally a check that the empirical Gaussian estimate converges to the true spectral coherence from below as data length increases.
  4. [§II.C, Eq. (11)] For diffusive systems, the path weight Eq. (11) is valid only in the limit Δt→0 and, for multiplicative noise, up to a normalization constant that depends on x_i and therefore does not cancel in Eq. (9). The paper does not report Δt or a convergence study for the Euler–Maruyama discretization. While the case studies appear to use additive noise (constant σ), the general claim that PWS has been extended to diffusive dynamics needs a statement of the discretization error and a convergence check for the specific systems studied.
minor comments (5)
  1. [§II.C] Typo: 'Euler-Mayurama' should be 'Euler–Maruyama'.
  2. [§II.C, Eq. (9)] The displayed estimator appears to be missing an explicit 1/N prefactor; as written it reads like a sum rather than an average.
  3. [Figs. 3–4] The colorbar labels are hard to parse ('10 1 100 101'); they should be formatted as 10^{-1}, 10^0, 10^1, and the units of τx should be stated.
  4. [General] No code or data availability statement is provided. For reproducibility, the PWS implementation, spectral estimation parameters, and simulation details should be included in an appendix or public repository.
  5. [Abstract/Introduction] The phrase 'even when the system's statistics are nearly Gaussian' is not quantified. A brief measure of non-Gaussianity (e.g., skewness of the stationary distribution or a path-space diagnostic) would strengthen the claim.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: approximations are independent of the PWS benchmark and not fitted to it.

full rationale

The paper's central comparisons use Path Weight Sampling (PWS) from Ref. [5] as an exact Monte Carlo benchmark. PWS is published independently, is not fitted to any of the approximate formulas, and does not incorporate the target results of this paper. The Gaussian approximation formulas (Eqs. 18, 23) and the discrete approximation (Eq. 19) are derived analytically from the underlying stochastic models (LNA and filtering approximations, respectively) before comparison with PWS. No parameter is calibrated to PWS data, and the reported deviations are observed rather than enforced by construction. The claimed lower-bound property of the empirical Gaussian estimate is supported by an independent theorem (Mitra and Stark), not by a self-citation. Self-citations to [4] and [24] provide context and mechanistic explanation, but the load-bearing quantitative evidence comes from PWS, an external exact method. The absence of explicit PWS convergence statistics (sample size, trajectory length, time step) is a reproducibility/correctness concern, not circularity. Therefore the derivation chain is self-contained with respect to the evaluated approximations.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central claims rest on the exactness of PWS and standard modeling assumptions. No free parameters are fitted to data; the model parameters are inputs used to explore regimes. The only potentially fragile assumptions are the numerical convergence of PWS and the discretized path weight.

assumptions (4)
  • domain assumption Path Weight Sampling (PWS) computes the exact mutual information rate for the considered systems
    The paper uses PWS as ground truth; exactness is established in Ref. [5], not re-derived here.
  • standard math The Onsager-Machlup discretized path weight (Eq. 11) accurately represents the conditional trajectory likelihood for the Langevin systems
    The path weight formula is standard stochastic calculus, but its numerical accuracy with the chosen time step is not quantified.
  • domain assumption The processes are stationary and ergodic, so the mutual information rate in Eq. (4) exists and can be estimated from finite trajectory Monte Carlo
    Stationarity is assumed for the rate definition; no mixing or ergodicity test is provided.
  • domain assumption The chemical master equation is the exact model for the discrete reaction system
    Standard representation of reaction kinetics used to define the true discrete dynamics.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Mutual Information Rate -- Linear Noise Approximation and Exact Computation." pith.science (2026). https://pith.science/paper/Z7SVVCOF

@misc{pith2026250821220,
  author       = {Pith},
  title        = {Pith review of: Mutual Information Rate -- Linear Noise Approximation and Exact Computation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Z7SVVCOF}},
  note         = {Machine review of arXiv:2508.21220}
}
read the original abstract

Efficient information processing is crucial for both living organisms and engineered systems. The mutual information rate, a core concept of information theory, quantifies the amount of information shared between the trajectories of input and output signals, and enables the quantification of information flow in dynamic systems. A common approach for estimating the mutual information rate is the Gaussian approximation which assumes that the input and output trajectories follow Gaussian statistics. However, this method is limited to linear systems, and its accuracy in nonlinear or discrete systems remains unclear. In this work, we assess the accuracy of the Gaussian approximation for non-Gaussian systems by leveraging Path Weight Sampling (PWS), a recent technique for exactly computing the mutual information rate. In two case studies, we examine the limitations of the Gaussian approximation. First, we focus on discrete linear systems and demonstrate that, even when the system's statistics are nearly Gaussian, the Gaussian approximation fails to accurately estimate the mutual information rate. Second, we explore a continuous diffusive system with a nonlinear transfer function, revealing significant deviations between the Gaussian approximation and the exact mutual information rate as nonlinearity increases. Our results provide a quantitative evaluation of the Gaussian approximation's performance across different stochastic models and highlight when more computationally intensive methods, such as PWS, are necessary.

Figures

Figures reproduced from arXiv: 2508.21220 by the authors.

Figure 1
Figure 1. FIG. 1. The mutual information rate of a simple linear re [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. FIG. 2. The dynamical input output relationship of the non [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. FIG. 3. The information rate of a non-linear system as a function of its gain over a range of response timescales. We vary [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (2 more)
Figure 2
Figure 2. Figure 2: We observe that the linearized input-output relation closely matches the slope of the true nonlinear dynam￾ical input-output relation at s = ¯s = 100, but overall it does not correspond to a (least-squares) linear fit of the nonlinear dynamical input-output relation. F…
Figure 4
Figure 4. Figure 4: FIG. 4. Deviation from the exact information rate for the [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

40 extracted references · 40 canonical work pages

  1. [1]

    The independent Gaussian white noise process ηs(t) sum- marizes all reactions that contribute to fluctuations in S

    Signal The input signal is generated by a birth-death process, ∅ κ − −⇀↽− −λ S, (A1) Its dynamics in Langevin form are, ˙s = κ − λs(t) + ηs(t), (A2) yielding the steady state signal concentration ¯ s = κ/λ. The independent Gaussian white noise process ηs(t) sum- marizes all reactions that contribute to fluctuations in S. The strength of the noise term in ...

  2. [2]

    (A5) We define the activation level a(s) to be a Hill function, a(s) = s(t)n K n + s(t)n

    Linear approximation We now consider the readout X, which is produced via a nonlinear activation function a(s): S ρ a (s) − − − →S + X, X µ − − → ∅. (A5) We define the activation level a(s) to be a Hill function, a(s) = s(t)n K n + s(t)n . (A6) Such a dependency, in which K sets the concentration of S at which the activation is half-maximal and n sets the...

  3. [3]

    If the intrinsic noise of the network is not cor- related to the process that drives the signal, the power spectrum of the network output obeys the spectral ad- dition rule [36]

    Information rate Following Tostevin & Ten Wolde [1, 2], we can express the Gaussian information rate as follows, R(S; X ) = 1 4π Z ∞ −∞ dω log 1 + |K(ω)|2 |N(ω)|2 Sss(ω) , (A13) where |K(ω)|2 is the frequency dependent gain and |N(ω)|2 is the frequency dependent noise of the output process X . If the intrinsic noise of the network is not cor- related to t...

  4. [4]

    Any difference between the exact information rate and the Gaussian information rate must then be a result of the Gaussian approximation

    Linear network To disambiguate differences in the information rate caused by the linear approximation of our nonlinear reac- tion network on the one hand and the Gaussian approx- imation of the underlying jump process on the other, we consider the information rate of a linear network. Any difference between the exact information rate and the Gaussian info...

  5. [5]

    Mutual information between in- and output trajectories of biochemical networks

    F. Tostevin and P. R. ten Wolde, Mutual Information between Input and Output Trajectories of Biochemical Networks, Physical Review Letters 102, 218101 (2009), 0901.0280

  6. [6]

    Tostevin and P

    F. Tostevin and P. R. ten Wolde, Mutual information in time-varying biochemical systems, Phys. Rev. E 81, 061917 (2010)

  7. [7]

    H. H. Mattingly, K. Kamino, B. B. Machta, and T. Emonet, Escherichia coli chemotaxis is information limited, Nature Physics 17, 1426 (2021)

  8. [8]

    Moor and C

    A.-L. Moor and C. Zechner, Dynamic information trans- fer in stochastic biochemical networks, Physical Review Research 5, 013032 (2023)

Show all 40 references
  1. [9]

    Reinhardt, G

    M. Reinhardt, G. Tkaˇ cik, and P. R. ten Wolde, Path Weight Sampling: Exact Monte Carlo Computation of the Mutual Information between Stochastic Trajectories, Physical Review X 13, 041017 (2023)

  2. [10]

    L. Hahn, A. M. Walczak, and T. Mora, Dynamical In- formation Synergy in Biochemical Signaling Networks, Physical Review Letters 131, 128401 (2023)

  3. [11]

    J. E. Segall, S. M. Block, and H. C. Berg, Temporal com- parisons in bacterial chemotaxis., Proceedings of the Na- tional Academy of Sciences 83, 8987 (1986)

  4. [12]

    M. W. Covert, T. H. Leung, J. E. Gaston, and D. Bal- timore, Achieving stability of lipopolysaccharide-induced nf-κb activation, Science 309, 1854 (2005)

  5. [13]

    Strong, R

    S. Strong, R. D. R. Van Steveninck, W. Bialek, R. Koberle, et al., On the application of information the- ory to neural spike trains, in Pac Symp Biocomput, Vol. 1998 (1998) pp. 621–632

  6. [14]

    C. E. Shannon, A Mathematical Theory of Communica- tion, Bell System Technical Journal 27, 379 (1948)

  7. [15]

    E. Ziv, I. Nemenman, and C. H. Wiggins, Optimal signal processing in small stochastic biochemical networks, PloS one 2, e1077 (2007)

  8. [16]

    Cheong, A

    R. Cheong, A. Rhee, C. J. Wang, I. Nemenman, and A. Levchenko, Information transduction capacity of noisy biochemical signaling networks, science 334, 354 (2011)

  9. [17]

    J. O. Dubuis, G. Tkaˇ cik, E. F. Wieschaus, T. Gregor, and W. Bialek, Positional information, in bits, Proceedings of the National Academy of Sciences 110, 16301 (2013)

  10. [18]

    S. E. Palmer, O. Marre, M. J. Berry, and W. Bialek, Pre- dictive information in a sensory population, Proceedings of the National Academy of Sciences 112, 6908 (2015)

  11. [19]

    Chalk, O

    M. Chalk, O. Marre, and G. Tkaˇ cik, Toward a unified theory of efficient, predictive, and sparse coding, Pro- ceedings of the National Academy of Sciences 115, 186 (2018)

  12. [20]

    Bauer, M

    M. Bauer, M. D. Petkova, T. Gregor, E. F. Wieschaus, and W. Bialek, Trading bits in the readout from a ge- netic network, Proceedings of the National Academy of Sciences 118, e2109011118 (2021)

  13. [21]

    Sachdeva, T

    V. Sachdeva, T. Mora, A. M. Walczak, and S. E. Palmer, Optimal prediction with resource constraints using the information bottleneck, PLoS computational biology 17, e1008743 (2021)

  14. [22]

    A. J. Tjalma, V. Galstyan, J. Goedhart, L. Slim, N. B. Becker, and P. R. ten Wolde, Trade-offs between cost and information in cellular prediction, Proceedings of the National Academy of Sciences 120, e2303078120 (2023)

  15. [23]

    A. J. Tjalma and P. R. ten Wolde, Predicting concen- tration changes via discrete receptor sampling, Physical Review Research 6, 033049 (2024)

  16. [24]

    T. M. Cover and J. A. Thomas, Elements of Information Theory, 2nd ed. (John Wiley & Sons, 2006)

  17. [25]

    Tkaˇ cik, C

    G. Tkaˇ cik, C. G. Callan, Jr., and W. Bialek, Information flow and optimization in transcriptional regulation, Pro- ceedings of the National Academy of Sciences 105, 12265 (2008), 0705.0313

  18. [26]

    W. H. de Ronde, F. Tostevin, and P. R. ten Wolde, Effect of feedback on the fidelity of information transmission of time-varying signals, Physical Review E—Statistical, Nonlinear, and Soft Matter Physics 82, 031914 (2010)

  19. [27]

    P. P. Mitra and J. B. Stark, Nonlinear limits to the infor- mation capacity of optical fibre communications, Nature 411, 1027 (2001)

  20. [28]

    A.-L. Moor, A. Tjalma, M. Reinhardt, P. R. ten Wolde, and C. Zechner, State- versus Reaction-Based Information Processing in Biochemical Networks, arXiv 10.48550/arxiv.2505.13373 (2025), 2505.13373

  21. [29]

    Meijers, S

    M. Meijers, S. Ito, and P. R. ten Wolde, Behavior of information flow near criticality, Physical Review E 103, L010102 (2021), 1906.00787

  22. [30]

    Fan and A

    R. Fan and A. Hilfinger, Characterizing the nonmono- tonic behavior of mutual information along biochemical reaction cascades, Physical Review E110, 034309 (2024), 2309.10162

  23. [31]

    Van Kampen, Stochastic processes in physics and chemistry (North Holland, Amsterdam, 1992)

    N. Van Kampen, Stochastic processes in physics and chemistry (North Holland, Amsterdam, 1992)

  24. [32]

    Onsager and S

    L. Onsager and S. Machlup, Fluctuations and Irreversible Processes, Physical Review 91, 1505 (1953)

  25. [33]

    E. A. Nadaraya, On Estimating Regression, Theory of Probability & Its Applications 9, 141 (1964)

  26. [34]

    G. S. Watson, Smooth Regression Analysis, Sankhy¯ a: The Indian Journal of Statistics, Series A (1961-2002) 26, 359 (1964)

  27. [35]

    Malaguti and P

    G. Malaguti and P. R. ten Wolde, Theory for the opti- mal detection of time-varying signals in cellular sensing systems, eLife 10, 1 (2021)

  28. [36]

    A. V. Oppenheim and R. W. Schafer, Digital Signal Pro- cessing (Prentice-Hall, 1975)

  29. [37]

    P. B. Warren, S. T˘ anase-Nicola, and P. R. ten Wolde, Ex- act results for noise power spectra in linear biochemical reaction networks, The Journal of Chemical Physics125, 144904 (2006), https://pubs.aip.org/aip/jcp/article- pdf/doi/10.1063/1.2356472/15393314/144904 1 online.pdf

  30. [38]

    S. A. Cepeda-Humerez, J. Ruess, and G. Tkaˇ cik, Esti- mating information in time-varying signals., PLoS com- putational biology 15, e1007290 (2019)

  31. [39]

    The bound specifically re- quires the Gaussian model to use the covariance of the full, original system

    Note that this argument does not apply to the Linear Noise Approximation (LNA). The bound specifically re- quires the Gaussian model to use the covariance of the full, original system. When the system is first linearized using the LNA, the resulting linear model does not retai...

  32. [40]

    T˘ anase-Nicola, P

    S. T˘ anase-Nicola, P. B. Warren, and P. R. ten Wolde, Sig- nal detection, modularity, and the correlation between extrinsic and intrinsic noise in biochemical networks, Physical review letters 97, 068102 (2006)

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.