{"id":"38b005c2-c10b-4d69-a8bb-d3ff3827a109","arxiv_id":"2412.19705","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Both the certainty-equivalence and the fixed-regularization semidefinite programs for direct data-driven LQR produce the zero gain under process noise, so they are inconsistent estimators.","lead":"This paper proves that a popular direct data-driven LQR method returns the zero control gain, with probability one, whenever the training data contain any nonzero noise. It also proves that a regularized variant collapses to zero gain in the long-data limit, so neither method is statistically consistent.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified — the central inconsistency results are sound under the stated Gaussian noise model; only minor proof formality in Lemma 8 and unproven remarks remain.","rationale":"Reader's verdict CONDITIONAL is sound. I examined the four load-bearing steps: (i) Lemma 8's persistency-of-excitation proof; (ii) Lemma 3's full-rank transfer via the augmented system; (iii) Lemma 4's equivalence between the SDP and the linear system (17); (iv) Theorem 3's high-probability bounds. Step (iii) is correct: for any feasible (X,Y), trace(QX0Y) ≥ trace(Q), and the Schur complement makes the X-part tight, so the optimal value is trace(Q) exactly when Y solves (17). Uniqueness of the gain follows because X0Y=I. Step (iv) is also correct; the only delicate inequality, bounding trace((√R U0Y)^T √R U0Y), requires the largest singular value of X0Y, and the stated bound on that largest singular value follows from trace(QX0Y) ≤ trace(Q)+O(1/T). The notation in the plain text is ambiguous but the intended inequality is valid. The weakest point, as the reader notes, is the full-rank condition in Lemma 3. It is, however, an explicit assumption of the theorem (σw>0, iid Gaussian, independent input), and under that model the condition holds with probability one. Degenerate or input-correlated noise would indeed break the proof, but the paper does not claim those cases. The unproven remarks and the informal density argument in Lemma 8 are present in the manuscript and should be tightened in revision, which supports the CONDITIONAL verdict; they do not affect the central claims.","tokens_in":12,"tokens_out":43958,"duration_ms":736001,"concrete_test":"Run a Monte Carlo check for a small random LTI system (n=m=1) at T=(m+n)(n+1)+n with iid Gaussian noise: record whether rank(D_T)=2n+m in every sample and whether the CE SDP (7) returns Kce=0; a single failure would contradict Lemma 3 or Theorem 2.","verdict_should_be":"UNCHANGED","load_bearing_attack":"After re-deriving the algebra in Sections IV and V, I find the central claims correct. Theorem 2 follows from Lemma 4 once Lemma 3 supplies full row rank; Lemma 4's Schur-complement reduction to (17) is valid, and the lower bound objective trace(Q) forces X0Y = I, U0Y = 0, and X1Y = 0 for every optimal Y. Lemma 3 correctly applies the fundamental lemma to the augmented system x_{t+1}=Ax_t+[B I][u_t;w_t], whose input is persistently exciting with probability one by Lemma 8. Theorem 3's fixed-η collapse is supported by the minimum-norm feasible solution (31) and the high-probability lower bound on σ(D_T). The key bound (36) is correct when the multiplying factor is read as the largest singular value of X0Y (the bar over sigma is lost in plain-text transcription); with that reading, (37) follows from (33)-(35). The non-load-bearing items are: Lemma 8's 'continuous density' sentence is informal (a non-constant polynomial of a Gaussian vector can have a density that is not continuous, but the zero set still has Gaussian measure zero), and Remarks 1 and the Section VI-B statement about RP solutions converging to (17) are asserted without proof. These do not undermine the main theorems.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies the noise sensitivity of two semidefinite programs (SDPs) for direct data-driven LQR: the certainty-equivalent DDD LQR (7) and the robustness-promoting regularized DDD LQR (9). Under the assumption that the process noise w_t is iid Gaussian with positive definite covariance, independent of the iid Gaussian input and the initial state, the authors prove two negative results. Theorem 2 shows that for T ≥ (m+n)(n+1)+n, every optimal solution of (7) yields the zero gain K_ce = 0_{m×n} with probability one, no matter how small σ_w > 0 is. Theorem 3 shows that for a fixed regularization parameter η, the gain K_rp(T) obtained from (9) converges to zero in probability as T → ∞. Consequently, neither SDP is a statistically consistent estimator of the true LQR gain. The proofs use a change of variables that treats noise as an additional input, the fundamental lemma, and concentration inequalities for the smallest singular value of the data matrix.","tokens_in":14291,"tokens_out":24819,"duration_ms":214504,"significance":"If the results hold, they provide a rigorous and somewhat surprising negative result for a widely cited SDP-based direct data-driven control method: the CE DDD LQR is discontinuous in the noise level, collapsing to a zero gain for any nonzero Gaussian noise, and the natural regularized variant does not fix the inconsistency as T grows. This is an important caution for the control community and clarifies the need for robust or otherwise regularized formulations. The main theorems are supported by elementary but careful algebraic arguments, and the paper includes numerical experiments that corroborate the theoretical predictions. The contributions are clearly bounded: the results cover Gaussian iid process noise and a particular regularized SDP, and the authors note that alternative formulations are not affected.","major_comments":[{"comment":"The proof of Theorem 3 contains a gap in the derivation of the key bound (37). The paper's notation defines σ(M) as the least singular value. Inequality (34) is obtained from trace(QX0Y) ≥ trace(Q)σ(X0Y) ≥ σ(Q)σ(X0Y), so it bounds the least singular value of X0Y*_rp(T). However, inequality (36) and the subsequent bound on ||U0Y*_rp(T)|| require an upper bound on the largest singular value of X0Y*_rp(T): indeed, ||R^{1/2}U0Y||_F^2 ≤ σ_max(X0Y) · trace(R^{1/2}U0Y (X0Y)^{-1} (R^{1/2}U0Y)^T). The manuscript does not provide such an upper bound on σ_max(X0Y); a bound on the smallest singular value does not help. This is load-bearing for Theorem 3. The gap is repairable: from (33) and Q ≽ σ(Q)I, obtain trace(X0Y) ≤ [trace(Q)+η(2n+m)/σ(D_T D_T^T)]/σ(Q), and hence σ_max(X0Y) ≤ trace(X0Y). Combining this with (36) yields (37). Please correct the singular value notation and supply this step.","section":"Section V-B, Step 3 (Eqs. (34)-(37))"}],"minor_comments":[{"comment":"In (36) and (37), the symbol σ is used where the largest singular value is intended. Given the convention defined in the notation section, the authors should use \\bar{\\sigma} for the largest singular value of X0Y and U0Y; this will also make the repaired Step 3 clearer.","section":"Section V-B, Eqs. (36)-(37)"},{"comment":"The proof of Lemma 8 states that the density of the determinant g is continuous and concludes P(g=0)=0 from P(g≤0)−P(g<0). This is not generally true: a non-constant polynomial of jointly Gaussian random variables need not have a continuous density (e.g., g=x^2 with x∼N(0,1)). The intended conclusion P(g=0)=0 is correct, but it should be justified by the standard fact that the zero set of a non-zero polynomial has Lebesgue measure zero and the Gaussian vector has an absolutely continuous distribution.","section":"Appendix, Lemma 8"},{"comment":"The assertion that any optimal solution of RP DDD LQR converges in probability to a solution of equation (17) is stated without proof; please provide a proof or explicitly label it as a conjecture.","section":"Section VI-B, last paragraph"},{"comment":"The measurement-noise extension in Remark 1 is stated without proof; please add a proof or qualify it as a conjecture.","section":"Remark 1"},{"comment":"The generalization to arbitrary continuous input and noise distributions is asserted without proof; if retained, a short proof should be supplied.","section":"Section IV-B, after Lemma 4"},{"comment":"Minor typographical issues: 'Y ALMIP' should be 'YALMIP'; reference [13] is a submission and should be updated if a preprint is available; check the phrase 'controled' and other minor grammar errors.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid negative result, and the central claims are likely correct, but the proof of Theorem 3 needs a small but real repair in Step 3. The self-citation to [17] for Lemma 6 is appropriate and not circular. The manuscript fits the journal's scope. I recommend major revision to address the singular-value gap and the unproven remarks."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead the paper. The main result is real: Theorem 2 proves that with any nonzero Gaussian process noise, the CE SDP for direct data-driven LQR returns the zero gain with probability one, once the trajectory is long enough. Theorem 3 shows the regularized version also collapses to zero gain in probability as T grows, for fixed regularization. This is a strong negative result, and it's new. De Persis and Tesi observed low-gain behavior and gave a sufficient stabilization condition, but they did not prove inconsistency. The paper's reduction of the SDP to the linear system (17) is clean, and the proofs check out. I re-derived the Schur complement steps and the bounds in Section V; the stress-test confirms the algebra. The paper is honest that scaling η with T may prevent collapse, and the experiments back the theory.\n\nSoft spots are minor. Two remarks are asserted without proof: the measurement-noise variant and the claim that RP solutions converge to a solution of (17). Neither is load-bearing for the main theorems. There's also a typo in Section VI-B that says the noiseless ‖Krp‖ approaches zero, which contradicts the figure; it should say 'with noise'. Lemma 8's 'continuous density' phrasing is informal, but the conclusion stands: a non-constant polynomial of a Gaussian vector has measure-zero zero set. These are polish issues.\n\nThe citation pattern is fine. The only self-citation is a concentration bound from Oymak and Ozay, which is standalone and legitimate.\n\nThis paper is for the data-driven control community, especially anyone using the De Persis–Tesi SDPs. It's a cautionary result that should stop people from assuming certainty equivalence works here. I'd send it to peer review. With minor revisions, it will be a useful reference.\n\nRecommendation: accept with minor revisions.","headline":"Proof that the standard data-driven LQR SDP collapses to zero gain under any noise—clean negative result, deserves review.","tokens_in":15022,"tokens_out":2137,"would_cite":true,"duration_ms":26290,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["93C55","90C22","93E20","93B52"],"pacs":[],"model":"deepseek-v4-flash","headline":"Under any nonzero Gaussian process noise, the standard SDP for direct data-driven LQR returns the zero gain with probability one, and the regularized SDP collapses in probability as data grows.","keywords":["direct data-driven control","LQR","semidefinite program","noise sensitivity","statistical consistency","certainty equivalence","persistency of excitation","zero-gain collapse"],"falsifier":"Run the SDP (7) on data drawn exactly from the paper's model for small $n,m$ and $T\\ge(m+n)(n+1)+n$, using a high-precision solver; any optimal solution with $K_{\\mathrm{ce}}\\neq0_{m\\times n}$ would contradict the claim that $K_{\\mathrm{ce}}=0$ with probability one. A symbolic check for the smallest nontrivial $(n,m,T)$ can confirm whether the optimality conditions truly force $U_0Y^*=0$.","tokens_in":13826,"feed_emoji":"📉","tokens_out":13454,"duration_ms":115555,"temperature":0.7,"pith_summary":"This paper establishes a negative statistical result for two semidefinite-programming (SDP) formulations of direct data-driven linear quadratic regulation (LQR). It proves that if state-input data are generated by a linear system with any nonzero independent Gaussian process noise, then, for a long enough trajectory, every optimal solution of the certainty-equivalence SDP yields the trivial gain $K_{\\mathrm{ce}}=0_{m\\times n}$ with probability one. It also proves that the robustness-promoting regularized SDP, with a fixed regularization parameter, produces gains $K_{\\mathrm{rp}}(T)$ that converge to $0_{m\\times n}$ in probability as the horizon $T$ grows. The consequence is that neither SDP is a statistically consistent estimator of the true LQR gain, even though the same SDPs recover the true gain exactly in the noise-free case. This matters because these methods are otherwise attractive: they synthesize controllers directly from raw data without explicit system identification, and the paper shows their noise-free success is a knife-edge phenomenon.","feed_headline":"Noise forces data-driven LQR SDP to zero gain","feed_subtitle":"Even the regularized semidefinite program collapses to the trivial controller as the data horizon grows.","key_machinery":"The central object is the combined data matrix $D_T=[X_0^\\top\\ U_0^\\top\\ X_1^\\top]^\\top$ together with the linear system $D_TY=[I_n;0_{m\\times n};0_{n\\times n}]$. The decisive fact (Lemma 3) is that with Gaussian noise and $T\\ge(m+n)(n+1)+n$, $D_T$ has full row rank $2n+m$ with probability one; this follows by treating the noise as an extra input in the lifted system $x_{t+1}=Ax_t+[B\\ I_n][u_t^\\top\\ w_t^\\top]^\\top$ and invoking the persistency-of-excitation lemma. Full row rank makes the linear system solvable, solvability makes the SDP constraints tight at a zero-gain solution, and Lemma 4 shows every optimal solution must satisfy that linear system. For the regularized program, the minimum-norm solution of the same linear system is feasible enough that the smallest singular value of $D_T$ growing like $\\sqrt{T}$ drives the gain to zero in probability.","core_discovery":"Under the data-generating model $x_{t+1}=Ax_t+Bu_t+w_t$ with $w_t \\stackrel{\\mathrm{i.i.d.}}{\\sim} \\mathcal{N}(0,\\sigma_w^2 I_n)$, independent of the Gaussian input $u_t$, and with $T\\ge (m+n)(n+1)+n$, the paper proves (Theorem 2) that $\\mathbb{P}_T(K_{\\mathrm{ce}}=0_{m\\times n})=1$ for every optimal solution of the SDP (7). For the regularized SDP (9) with fixed $\\eta>0$, it proves (Theorem 3) that $K_{\\mathrm{rp}}(T)\\xrightarrow{p}0_{m\\times n}$ as $T\\to\\infty$. The algebraic mechanism is that the combined data matrix $D_T=[X_0^\\top\\ U_0^\\top\\ X_1^\\top]^\\top$ is full row rank with probability one, so the underdetermined linear system $D_TY=[I_n;0_{m\\times n};0_{n\\times n}]$ has a solution; such a solution makes $X_0Y=I_n$, $U_0Y=0$, and $X_1Y=0$, which attains the SDP's unconstrained lower bound $\\mathrm{trace}(Q)$ and therefore is optimal. The regularized result extends the same forcing through quantitative persistency-of-excitation bounds that make the smallest singular value of $D_T$ grow like $\\sqrt{T}$.","pith_inferences":["The mechanism is broader than the Gaussian assumption: the proof only needs the combined data matrix to be full row rank almost surely, so any continuous, input-independent noise distribution that makes the lifted data persistently exciting should produce the same zero-gain collapse.","This suggests the pathology belongs to the SDP encoding rather than to certainty equivalence itself, which sharpens the contrast with model-based certainty-equivalence LQR and points toward alternative direct formulations that provably mimic model-based solutions.","The paper's upper-bound argument hints at a concrete remedy: if the regularization parameter grows with the horizon, say $\\eta=cT$, the objective no longer collapses to $\\mathrm{trace}(Q)$, and the numerical experiments with $\\eta=10T$ show nonzero gains; a rigorous consistency analysis of that scaling is a natural next test.","The same full-rank forcing should appear under measurement noise (as Remark 1 notes), suggesting that output-feedback or reduced-measurement variants may only escape the zero-gain trap if they break the exact solvability of $D_TY=[I_n;0;0]$."],"forward_implications":["With any nonzero independent Gaussian process noise, the certainty-equivalence SDP (7) returns the zero gain $K_{\\mathrm{ce}}=0_{m\\times n}$ with probability one once $T\\ge(m+n)(n+1)+n$, so it never recovers the true LQR gain.","The regularized SDP (9) with a fixed $\\eta>0$ has gains $K_{\\mathrm{rp}}(T)$ whose spectral norm converges to zero in probability as $T\\to\\infty$; adding more data does not restore consistency.","The data-to-controller map is discontinuous: the same SDP gives the exact $K_{\\mathrm{lqr}}$ in the noise-free case, while arbitrarily small noise flips it to the trivial zero gain.","Under the theorem's hypotheses, the sufficient stabilizability condition of [8] is almost surely $\\Psi=AA^\\top$, so for open-loop unstable systems that condition fails with probability one and cannot certify stabilization in the noisy setting."],"supporting_citations":[{"why":"Introduces the certainty-equivalence DDD LQR SDP (7) and the noise-free recovery result that Theorem 2 overturns in the noisy setting.","marker":"[1]"},{"why":"Introduces the regularized SDP (9) and the stabilizability sufficient condition analyzed in Corollary 1.","marker":"[8]"},{"why":"Provides the original fundamental lemma on persistency of excitation that underpins Lemma 2.","marker":"[14]"},{"why":"Gives the state-space form of the fundamental lemma used directly as Lemma 2.","marker":"[15]"},{"why":"Supplies the quantitative persistency-of-excitation bound used in Lemma 5.","marker":"[16]"},{"why":"Provides the high-probability lower bound on the smallest singular value of a Hankel matrix of Gaussian inputs used in Lemma 6.","marker":"[17]"},{"why":"Establishes the statistical consistency of model-based certainty-equivalence LQR, the contrast motivating the paper's question.","marker":"[3]"}],"fun_headline_variants":["Even arbitrarily small noise dooms data-driven LQR SDP","Data-driven LQR SDP collapses to zero gain with noise","Regularization can't fix noise fragility of LQR SDP","Noise makes LQR SDPs output trivial zero-gain controller","Data-driven LQR SDP noise-sensitive: zero gain inevitable"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the data are random in a way that makes the combined matrix $[X_0^\\top\\ U_0^\\top\\ X_1^\\top]^\\top$ full row rank with probability one: the process noise must be continuous, independent of the input, and nondegenerate, and the trajectory must be long enough ($T\\ge(m+n)(n+1)+n$); if that rank condition fails, the SDP is no longer forced to the zero-gain solution.","fun_headline_variants_meta":{"raw":{"variants":["Even arbitrarily small noise dooms data-driven LQR SDP","Data-driven LQR SDP collapses to zero gain with noise","Regularization can't fix noise fragility of LQR SDP","Noise makes LQR SDPs output trivial zero-gain controller","Data-driven LQR SDP noise-sensitive: zero gain inevitable"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000502,"raw_usage":{"total_tokens":2466,"prompt_tokens":972,"completion_tokens":1494,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":588,"completion_tokens_details":{"reasoning_tokens":1406}},"tokens_in":588,"tokens_out":1494,"duration_ms":12932,"temperature":1.0,"reasoning_tokens":1406,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T23:59:35.344680+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the SDP (7) on data drawn exactly from the paper's model for small $n,m$ and $T\\ge(m+n)(n+1)+n$, using a high-precision solver; any optimal solution with $K_{\\mathrm{ce}}\\neq0_{m\\times n}$ would contradict the claim that $K_{\\mathrm{ce}}=0$ with probability one. A symbolic check for the smallest nontrivial $(n,m,T)$ can confirm whether the optimality conditions truly force $U_0Y^*=0$.","supporting_citations":[{"cited_title":"Formulas for data-driven contr ol: Stabilization, optimality, and robustness,","cited_arxiv_id":null,"evidence_quote":"Introduces the certainty-equivalence DDD LQR SDP (7) and the noise-free recovery result that Theorem 2 overturns in the noisy setting."},{"cited_title":"Low-complexity learning of lin ear quadratic regulators from noisy data,","cited_arxiv_id":null,"evidence_quote":"Introduces the regularized SDP (9) and the stabilizability sufficient condition analyzed in Corollary 1."},{"cited_title":"Willems’ fundamental lemma for state-space systems and its extensio n to multiple datasets,","cited_arxiv_id":null,"evidence_quote":"Gives the state-space form of the fundamental lemma used directly as Lemma 2."},{"cited_title":"A quantitative notion of persistency of excitation and the robust fundamen tal lemma,","cited_arxiv_id":null,"evidence_quote":"Supplies the quantitative persistency-of-excitation bound used in Lemma 5."},{"cited_title":"Non-asymptotic identiﬁcation of LTI systems from a single trajectory,","cited_arxiv_id":null,"evidence_quote":"Provides the high-probability lower bound on the smallest singular value of a Hankel matrix of Gaussian inputs used in Lemma 6."},{"cited_title":"Certainty equivalence is e fﬁcient for linear quadratic control,","cited_arxiv_id":null,"evidence_quote":"Establishes the statistical consistency of model-based certainty-equivalence LQR, the contrast motivating the paper's question."}],"review_version":1}