Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Sensing-Limited Control of Noiseless Linear Systems Under Nonlinear Observations

T0 review · 3 major / 5 minor · reviewed 2026-08-03 · deepseek-v4-flash

Pith's one-line read This paper establishes that a noiseless linear plant can be estimated and stabilized only if the causal information rate from its unstable state to a nonlinear sensor reaches the expansion rate log2|det Au|, and that a strict excess, under

desk verdict Necessary directed-information conditions are clean and correct; the sufficiency proof breaks at the entropy-to-MSE step, so 'closing the gap' overreaches. read the letter →

arxiv 2601.12782 v2 pith:UNZGQXVE submitted 2026-01-19 eess.SY cs.ITcs.SYmath.IT

classification eess.SYcs.ITcs.SYmath.IT
keywords directedinformationsensing-limitedcontrolmean-squareobservabilitystabilizabilitynonlinearobservationsdata-ratetheoremlog-concavityentropyrecursion
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks what a nonlinear sensor must tell a controller, in bits per step, for a noiseless linear system with unstable modes to be observable and stabilizable. Its necessary result: if estimation error stays bounded in mean square, the average directed information rate from the unstable state to the measurements must be at least the expansion rate log2|det Au|. Its sufficient result, under regularity assumptions including persistent sensing curvature and uniformly well-conditioned error covariance, is that a strictly larger rate forces the posterior entropy to -infinity, which forces the error covariance to zero. Read together, these show that the bottleneck is the sensing geometry, not the communication channel, and that the classical data-rate threshold survives in the nonlinear-observation regime as an information-flow constraint.

What carries the argument

Directed information rate, defined as the average of the sum of I(z_t^u; y_t | y^{t-1}), and the entropy recursion it induces for unstable dynamics. The recursion converts stability of estimation error into a lower bound on information flow, while the sufficiency direction relies on strong log-concavity of the posterior obtained from persistent block-wise curvature (Assumption 1) and on the log-concave entropy-covariance equivalence to turn entropy divergence into variance decay.

What would settle it

Find an explicit observation channel that satisfies Assumptions 1-3 and a linear unstable plant where the directed information rate is strictly above Rexp but the minimum mean-square error does not go to zero; that would refute Theorem 3. Since the paper gives no concrete model satisfying Assumption 1, the first step is to exhibit such a channel; if none can be constructed, the sufficiency result has no instantiation.

Watch

Extended reading notes

Core claim

Define the expansion rate Rexp as the sum of log2|λi| over eigenvalues of A of magnitude at least 1. The paper proves, using the directed-information recursion h(z_{t+1}^u|y^t) = h(z_t^u|y^{t-1}) - I(z_t^u; y_t | y^{t-1}) + Rexp, that mean-square observability forces the average directed information rate from the unstable state to the observations to be at least Rexp; the same necessity holds for mean-square stabilizability. Under Assumptions 1-3 the reverse holds in the stronger asymptotic sense: a rate strictly above Rexp drives the posterior entropy to -infinity, and since the posterior is eventually strongly log-concave, the error covariance converges to zero, giving asymptotic observabi

Load-bearing premise

The sufficiency proof collapses if the sensor can go through short stretches with low curvature—for example, saturation or quantization plateaus—because Assumption 1's uniform block-wise negative curvature is what forces the posterior to become strongly log-concave, and without that, entropy divergence does not imply variance decay.

Editorial extensions

If this is right

  • In any noiseless linear plant with unstable modes, a sensor that cannot sustain a directed information rate of at least log2|det Au| cannot support bounded mean-square estimation, regardless of controller design.
  • If the rate is strictly above the threshold and the sensing channel has persistent curvature and well-conditioned error covariance, a certainty-equivalence controller renders the closed loop asymptotically mean-square stable.
  • The result extends the classical data-rate theorem from communication-constrained to sensing-constrained control, replacing encoder capacity with the geometry of the physical sensor.
  • The threshold depends only on the unstable eigenvalues of A, not on the nonlinearity, so sensor nonlinearities affect achievability but not the fundamental lower bound.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial: the necessary bound does not require any regularity assumptions, so it holds even for pathological nonlinear sensors; the sufficiency theorem, by contrast, inherits the strength of Assumption 1, and no explicit nonlinear observation model satisfying Assumption 1 is exhibited in the paper.
  • Editorial: because the control input enters the directed-information expression only through the distribution of the observations, the recursion suggests a sensing-aware control design principle: choose inputs that increase the mutual information each step gives about the unstable state.
  • Editorial: the same entropy-recursion argument should carry over to systems with process noise by adding the noise-injection entropy term to the recursion; the paper lists this as future work, and the necessary direction would then include the entropy of the driving noise.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper studies noiseless linear systems with nonlinear observations and derives information-theoretic conditions for mean-square observability and stabilizability. Theorems 1–2 establish that mean-square observability/stabilizability implies a lower bound on the average directed information rate from the unstable state to the observations, namely at least the expansion rate R_exp. Theorems 3–4 claim the converse: under three regularity assumptions (persistent sensing curvature, bounded prior Hessian, and bounded condition number of the conditional covariance), a directed information rate strictly larger than R_exp ensures asymptotic mean-square observability and, with stabilizability, asymptotic mean-square stabilizability. The necessity proofs are simple entropy accounting; the sufficiency proofs rely on log-concavity bounds and a sketched lemma on asymptotic strong log-concavity of the posterior.

Significance. If the sufficiency theorems were correct, the paper would constitute a meaningful extension of classical data-rate theorems to nonlinear sensing channels, using directed information as the relevant measure of causal information flow. The necessary conditions are elegant and appear sound, and the use of external inequalities (Bobkov–Madiman) is appropriate. However, the central sufficiency claim is not established because of a load-bearing gap in the entropy-to-variance step (Section IV.B), a severely incomplete proof of Lemma 2, and an unverified strong Assumption 3. These issues affect the paper's main contribution—'closing the gap' between information rates and estimation performance—and require substantial revision.

major comments (3)
  1. [Theorem 3, Section IV.B, after Eq. (11)] The proof claims that h(e_{T+1}|y^{T+1})→−∞ implies log det(Σ^u_{T+1|T+1})→−∞ and hence E||e_{T+1}||^2→0. This step is invalid. Inequality (9) applies to a fixed log-concave density and its covariance. Applied conditionally for each realization y, it gives h(e|Y=y) ≥ (1/2)log((2πe)^n det Σ_y)−nC, so taking expectations yields only E[log det Σ_Y]→−∞. Since log det is concave, Jensen's inequality gives log det(EΣ) ≥ E[log det Σ], which does not force EΣ→0. A scalar counterexample: let Σ take values 1/n and 1 each with probability 1/2; then E[log Σ]→−∞ while EΣ→1/2. Thus the displayed implication is false. The sufficiency proofs of Theorems 3 and 4 therefore do not establish mean-square convergence. An additional argument (e.g., uniform almost-sure control of large-covariance events, or a direct proof that det Σ→0 a.s.) is needed.
  2. [Lemma 2, Appendix A] The proof of Lemma 2 is only sketched. The remainder term R_t is introduced but never bounded; the partition argument and the asymptotic dominance for Jordan blocks are asserted, not proved. In particular, the statement that for marginally unstable modes the cumulative negative term has order O(t^{2K−1}) and dominates the positive term of order O(t^{2K−2}) is not substantiated. Since Lemma 2 is essential for applying the log-concavity bound (9), this is a load-bearing gap. Additionally, Assumption 1 is an abstract curvature condition; the paper gives no concrete nonlinear observation model (e.g., saturated, quantized, or logistic) that satisfies it for all realizations and all blocks of length L. This weakens the claim that the results cover a 'broad class' of nonlinear observations.
  3. [Theorem 3, stability part] The proof concludes 'Stability follows from convergence' but does not establish the uniform δ property required in Definition 4(1). Convergence of E||e_t||^2→0 for the given initial condition shows attractivity for that initial condition, but not that for every ε>0 there exists δ>0 such that E||e_0||^2≤δ implies E||e_t||^2≤ε for all t. The estimator is nonlinear and the error dynamics are not shown to be Lyapunov stable. A separate argument is needed.
minor comments (5)
  1. [Abstract and Theorems 1–2] The abstract says the information flow 'must exceed' the expansion rate, but Theorems 1–2 state liminf ≥ R_exp. The strict inequality appears only in the sufficiency theorems. Please clarify the distinction.
  2. [Notation throughout] The notation for the unstable state is inconsistent: z_u^T vs z^u_T, and the directed information I(z_u^T → y^T) uses different superscript/subscript conventions. Please unify.
  3. [Section II.B] Typo: 'decomposotion' should be 'decomposition'.
  4. [Assumption 1] The Hessian notation ∇^2_{z^u_t} log p(y_k|z^u_k) is ambiguous because the log-likelihood is a function of z^u_k, not z^u_t. The parenthetical 'where past states depend on z^u_t via inverse dynamics' should be formalized with an explicit chain-rule expression, including the role of control inputs.
  5. [References] References [12] and [13] are both by Bobkov and Madiman; consider citing them in a consistent style and clarifying which result (Corollary 4.2 of [13]) establishes the specific bound (9).

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the derivation is a self-contained entropy-accounting argument with independent external inequalities; the flagged concern is a proof gap, not circularity.

full rationale

The paper's derivation chain is self-contained and non-circular. Theorem 1 is direct entropy accounting: using h(z^u_{t+1}|y^t) = h(z^u_t|y^t) + R_exp, mean-square observability is shown to imply a lower bound on the directed information rate; no fitted parameter or self-definition is involved. Theorem 3 uses the same recursion to convert the assumed rate condition into divergence of prediction entropy, then invokes the external Bobkov–Madiman log-concavity bound (9) and Assumptions 1–3 to relate entropy divergence to covariance convergence. The rate condition is not defined in terms of observability or stabilizability; it is an independent information-theoretic quantity. There are no self-citations, no fitted inputs presented as predictions, no ansatz smuggled in via the authors' own prior work, and no renaming of a known empirical pattern. The cited inequality (9) comes from independent published work ([12],[13]) and is not a load-bearing self-citation. The main weakness in the paper is a mathematical gap in Theorem 3's final step: h(e_{T+1}|y^{T+1}) is an expected conditional entropy, while the covariance Σ^u_{T+1|T+1} is a random matrix, so h(e|y)→−∞ does not by itself force det Σ→0 or E||e||^2→0. This is a correctness/rigor concern, not a circularity, because it does not make the conclusion equivalent to the assumptions by construction. Under the stated circularity criteria, the appropriate finding is no significant circularity.

Assumptions & free parameters 4 free parameters · 8 assumptions · 0 invented entities

All theorems rest on standard entropy identities plus three strong domain assumptions (persistent log-likelihood curvature, smooth prior, bounded condition number). No constants are fitted to data; the assumption constants L, α, β, κ are existential and uninstantiated, which limits the practical force of the sufficiency results.

free parameters (4)
  • L = not specified
    Assumption 1 postulates existence of block length L≥1 such that cumulative Hessian over L consecutive observations is ⪯ -αI; no construction or bound is given.
  • α = not specified (>0)
    Curvature margin in Assumption 1; the sufficiency proof needs α>0 but no value is provided or derived from a sensing model.
  • β = not specified (>0)
    Upper bound on the Hessian of log-initial-density in Assumption 2; needed for Lemma 2.
  • κ = not specified
    Uniform condition-number bound on posterior error covariance in Assumption 3; needed to convert determinant decay into mean-square error decay.
assumptions (8)
  • standard math Differential entropy transformation h(Az|y)=h(z|y)+log|det A| for invertible A (translation and scaling)
    Used in Eq. (7)–(8) and Eq. (10) for the entropy recursion.
  • standard math Conditioning reduces differential entropy: h(X|Y) ≤ h(X)
    Used to bound h(e|y) by h(e) in Theorem 1 and to upper-bound posterior entropy in Theorem 3.
  • standard math Gaussian distributions maximize differential entropy under a covariance constraint
    Used in Theorem 1 to bound h_{T+1} by Hmax.
  • standard math Bobkov–Madiman log-concave entropy bounds (Eq. (9))
    External result from [12],[13] bridging entropy of log-concave densities to Gaussian entropy; central to Theorem 3.
  • domain assumption Assumption 1: cumulative Hessian of observation log-likelihood over L consecutive steps is ⪯ -αI
    Guarantees persistent sensing curvature; used in Lemma 2 to prove posterior strong log-concavity. Not verified for any concrete sensing model.
  • domain assumption Assumption 2: log-initial-density has Hessian bounded above by βI
    Smoothness of the prior; used in Lemma 2.
  • domain assumption Assumption 3: conditional error covariance condition number is uniformly bounded by κ
    Needed to go from det(Σ)→0 to λ_max(Σ)→0; may be violated in practice.
  • domain assumption The unstable subsystem matrix A_u is invertible
    Needed for entropy scaling by log|det A_u|; holds because unstable eigenvalues have |λ|≥1, but not stated explicitly.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Sensing-Limited Control of Noiseless Linear Systems Under Nonlinear Observations." pith.science (2026). https://pith.science/paper/UNZGQXVE

@misc{pith2026260112782,
  author       = {Pith},
  title        = {Pith review of: Sensing-Limited Control of Noiseless Linear Systems Under Nonlinear Observations},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UNZGQXVE}},
  note         = {Machine review of arXiv:2601.12782}
}
read the original abstract

This paper investigates the fundamental information-theoretic limits for the control and sensing of noiseless linear dynamical systems subject to a broad class of nonlinear observations. We analyze the interactions between the control and sensing components by characterizing the minimum information flow required for stability. Specifically, we derive necessary conditions for mean-square observability and stabilizability, demonstrating that the average directed information rate from the state to the observations must exceed the intrinsic expansion rate of the unstable dynamics. Furthermore, to address the challenges posed by non-Gaussian distributions inherent to nonlinear observation channels, we establish sufficient conditions by imposing regularity assumptions, specifically log-concavity, on the system's probabilistic components. We show that under these conditions, the divergence of differential entropy implies the convergence of the estimation error, thereby closing the gap between information-theoretic bounds and estimation performance. By establishing these results, we unveil the fundamental performance limits imposed by the sensing layer, extending classical data-rate constraints to the more challenging regime of nonlinear observation models.

Figures

Figures reproduced from arXiv: 2601.12782 by the authors.

Figure 1
Figure 1. Closed-loop control system with nonlinear sensing. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Sensing-Limited Control Under Non-Designable Observation Mechanisms

    eess.SY 2026-06 unverdicted novelty 7.0 of 10

    Derives that the directed information rate from the unstable state process to the observation process must exceed the open-loop expansion rate of unstable modes for mean-square stabilizability under non-designable sensing.

Reference graph

Works this paper leans on

13 extracted references · 1 linked inside Pith · cited by 1 Pith paper

  1. [1]

    On stabilization of linear systems with limited informa- tion,

    D. Liberzon, “On stabilization of linear systems with limited informa- tion,”IEEE Trans. Autom. Control, vol. 48, no. 2, pp. 304–307, Feb 2003

  2. [2]

    Control under communication constraints,

    S. Tatikonda and S. Mitter, “Control under communication constraints,” IEEE Trans. Autom. Control, vol. 49, no. 7, pp. 1056–1068, Jul 2004

  3. [3]

    Feedback control under data rate constraints: An overview,

    G. N. Nair, F. Fagnani, S. Zampieri, and R. J. Evans, “Feedback control under data rate constraints: An overview,”Proc. IEEE, vol. 95, no. 1, pp. 108–137, Jan 2007

  4. [4]

    Control over noisy channels,

    S. Tatikonda and S. Mitter, “Control over noisy channels,”IEEE Trans. Autom. Control, vol. 49, no. 7, pp. 1196–1201, Jul 2004

  5. [5]

    Anytime information theory,

    A. Sahai, “Anytime information theory,” Ph.D. dissertation, Mas- sachusetts Institute of Technology, Cambridge, MA, 2001

  6. [6]

    Stabilizability of stochastic linear systems with finite feedback data rates,

    G. Nair and R. Evans, “Stabilizability of stochastic linear systems with finite feedback data rates,”SIAM J. Control Optim., vol. 43, pp. 413–436, 2004

  7. [7]

    Evaluating channels for control: capacity reconsidered,

    A. Sahai, “Evaluating channels for control: capacity reconsidered,” in Proc. Amer. Control Conf., vol. 4, 2000, pp. 2358–2362

  8. [8]

    Learning perceptive humanoid locomotion over challenging terrain,

    W. Sun, B. Cao, L. Chen, Y . Su, Y . Liu, Z. Xie, and H. Liu, “Learning perceptive humanoid locomotion over challenging terrain,”arXiv preprint arXiv:2503.00692, Mar 2025

Show all 13 references
  1. [9]

    Amp: Adversarial motion priors for stylized physics-based character control,

    X. B. Peng, Z. Ma, P. Abbeel, S. Levine, and A. Kanazawa, “Amp: Adversarial motion priors for stylized physics-based character control,” ACM Trans. Graph., vol. 40, no. 4, pp. 1–20, Jul 2021

  2. [10]

    Autonomous vehicles: theoretical and practical challenges,

    M. Mart ´ınez-D´ıaz and F. Soriguera, “Autonomous vehicles: theoretical and practical challenges,”Transp. Res. Procedia, vol. 33, pp. 275–282, 2018

  3. [11]

    Beyond transmitting bits: Context, semantics, and task-oriented communications,

    D. G ¨und¨uz, Z. Qin, I. E. Aguerri, H. S. Dhillon, Z. Yang, A. Yener, K. K. Wong, and C.-B. Chae, “Beyond transmitting bits: Context, semantics, and task-oriented communications,”IEEE J. Sel. Areas Commun., vol. 41, no. 1, pp. 5–41, Jan 2022

  4. [12]

    The entropy per coordinate of a random vector is highly constrained under convexity conditions,

    S. Bobkov and M. Madiman, “The entropy per coordinate of a random vector is highly constrained under convexity conditions,”IEEE Trans. Inf. Theory, vol. 57, no. 8, pp. 4940–4954, Aug 2011

  5. [13]

    Reverse Brunn–Minkowski and reverse entropy power inequalities for convex measures,

    ——, “Reverse Brunn–Minkowski and reverse entropy power inequalities for convex measures,”J. Funct. Anal., vol. 262, no. 7, pp. 3309–3339, Apr 2012

Pith tools

Reviewed August 3, 2026 · model on record in the stance chip above.