{"id":"1b89e938-d38c-47b9-9ca5-07dc09f3f9ca","arxiv_id":"2607.03393","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Two data-driven natural-gradient controllers, certified λ-contractive by LMIs, synthesize robust linear policies from input-state data without identifying A and B.","lead":"The paper gives a way to design robot controllers straight from input-state data by forcing the closed loop to follow a natural-gradient flow, with stability certificates from SDPs. It matters because it replaces hard-to-tune LQR weights with a single step-size that encodes uncertainty via the Fisher information matrix.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The LMI certificates guarantee only mean-dynamics contraction under the linear data-based map; they do not certify the physical nonlinear plant used for all hardware claims.","rationale":"The reader already flags Assumption 3 together with the linearization (50)–(51) as the weakest assumption. That diagnosis is correct and load-bearing: the entire stability argument is internal to the linear data-based maps, while every practical claim is made on the nonlinear robot. No additional formal gap of comparable severity appears; the convex relaxations are only sufficient (as the reader notes) and the single-parameter α story is empirically illustrated, but those are secondary. The concrete residual/contraction check on a large-heading trajectory would settle whether the linear certificate still has predictive value on the plant that was actually controlled. Because the paper already presents the result as conditional on the linear regime, the verdict remains CONDITIONAL; the stress test simply sharpens the same concern rather than introducing a new one.","tokens_in":20929,"tokens_out":593,"duration_ms":5229,"concrete_test":"Collect a short closed-loop trajectory on the physical ROSbot under the Theorem-1 gain (α=10^{-5}, λ=0.9) that deliberately drives |ϕ| beyond ≈0.3 rad; compute the empirical one-step residual ||x_{k+1}−(X_1 G)x_k|| and the realized contraction factor of ||μ_k||_P. If either residual exceeds the process-noise level used in the LMI or the observed contraction factor exceeds the designed λ, the hardware claim is unsupported by the certificate.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Theorems 1–2 establish that a feasible LMI solution yields a gain K such that the data-based closed-loop mean obeys the NGD recursion μ_{k+1}=(I−2αΣP)μ_k and is therefore λ-contractive in expectation. That certificate is derived entirely under the linear data-based representations (11) and (19) that rest on the constant-A=I linearization (50)–(51) and on the rank condition of Assumption 3. All hardware and Gazebo experiments, however, are performed on the true nonlinear Mecanum kinematics (48)–(49). When heading ϕ leaves the small-angle regime used both for data collection and for the linear model, the map realized by the recovered K is no longer the map whose contraction was certified. Consequently the strongest claim—that the LMI-derived K produces a stable NGD closed-loop on the physical platform—rests on an unquantified linearization residual that is never bounded or validated against the nonlinear plant. The Monte-Carlo and SNR studies remain inside the same linear (or lightly-perturbed linear) model, so they do not close the gap.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper develops a direct data-driven natural-gradient-descent (NGD) control framework for unknown stochastic LTI systems. Using two data-based closed-loop parameterizations (raw input-state and sample-covariance), it embeds the Fisher information matrix (inverse closed-loop covariance) as a preconditioner so that the mean dynamics exactly reproduce the NGD recursion μ_{k+1}=(I-2αΣP)μ_k. Theorems 1 and 2 supply LMI/SDP conditions that certify λ-contractiveness of this mean dynamics (hence stability) and recover the linear gain K; supporting sample-complexity and iteration-complexity lemmas are given. The approach is demonstrated in Monte-Carlo SNR studies, (N,α) sweeps, Gazebo simulation, and hardware experiments on a ROSbot XL Mecanum platform, with comparisons to model-based LQR and existing data-driven LQR baselines that emphasize single-parameter (α) interpretability.","tokens_in":21258,"tokens_out":1340,"duration_ms":22712,"significance":"If the claims hold, the work supplies a geometrically motivated, single-scalar-tuned alternative to classical Q/R shaping for data-driven LQR-like design, together with explicit uncertainty-aware covariance recursions (eqs. 13, 21) and SDP certificates. The hardware demonstration on a real mobile robot and the systematic (N,α) trade-off tables are concrete strengths that go beyond purely theoretical data-driven LMI papers. The contribution is therefore of genuine interest to the data-driven and learning-based control communities, provided the linearization gap and the conservatism of the convex relaxations are clarified.","major_comments":[{"comment":"Theorems 1–2 certify λ-contractiveness only for the linear data-based maps (11) and (19) that rest on the constant-A=I linearization (50)–(51) and Assumption 3. All hardware, Gazebo, and nonlinear SNR results (Table II, Section V, Appendix A) are obtained on the true Mecanum kinematics (48)–(49). No residual bound, maximum heading excursion, or empirical validation that the realized closed-loop remains inside the certified linear regime is supplied; consequently the abstract claim of “stability-guaranteed policy synthesis … on a ROSbot XL platform” is not rigorously supported by the theory. Either restrict the claims, quantify the linearization error, or add a supporting nonlinear argument.","section":"Theorems 1–2, Section V, Appendix A, eqs. (48)–(51)"},{"comment":"The convex relaxations M ≻ GΣGᵀ and Z ≻ YΣ^{-1}Y (eqs. 28c–e and the analogous set in Theorem 2) are only sufficient. The manuscript never checks tightness, reports the duality gap, or verifies a posteriori that the recovered K satisfies the original stationary-covariance equality rather than merely the relaxed upper bound. Without such evidence the certificates may be arbitrarily conservative, especially for the small data sets (N=24) used on hardware.","section":"Theorems 1–2, eqs. (28c)–(28e), (36)–(37)"},{"comment":"Lemma 4 supplies a high-probability sample-size bound under Gaussian noise and the designed closed-loop, yet the hardware experiments use only N=24 samples for a 7-dimensional regressor and never report the realized condition number of Φ or D_0, nor verify the BMSB constants. Given that Theorem 2 is already shown to be highly sensitive to small N and tiny α (Tables III–IV), the practical reliability of the rank and positive-definiteness assumptions under the collected excitation remains unquantified.","section":"Lemma 4, Section V.F, Tables III–IV, Appendix B"}],"minor_comments":[{"comment":"Notation for the two parameterizations (G versus H, X_1 versus X-bar_1) is introduced cleanly but then occasionally mixed in the algorithm box and the recovery formulas; a short consistency pass would help.","section":"Algorithm 1, Theorems 1–2"},{"comment":"Figures 1 and 5–14 would benefit from explicit legends that identify which curve belongs to which controller/α; several captions simply say “various α” without listing the values.","section":"Section V, Appendix"},{"comment":"The free parameters α, λ and W are acknowledged, yet the text never states how W is chosen for the hardware runs (estimated or hand-tuned). A one-sentence clarification would remove ambiguity.","section":"Assumption 2, Section V.D"},{"comment":"A few typographical inconsistencies appear (e.g., “Linköping” vs. “Link ¨oping”, missing spaces around “λ-contractive”). None affect readability but should be cleaned.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The theoretical core is a natural and useful data-driven extension of the authors’ own model-based NGD papers [9],[10]; the novelty therefore resides mainly in the uncertainty-aware data parameterizations and the hardware results. The linearization gap is the single most important issue for a control-systems journal; once it is addressed (even by a careful discussion and residual plots) the paper should be publishable. Scope fit for TCST is good."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The new piece is the pair of uncertainty-aware data parameterizations (the extra Tr(GΣGᵀ)W terms in the covariance recursions) plus the two SDPs that recover an NGD gain K directly from input-state data. That is a clean, usable extension of their earlier model-based NGD papers, not a re-packaging.\n\nWhat works: Theorems 1–2 are standard Schur-complement LMIs with explicit convex relaxations; the mean dynamics are forced to match μ_{k+1}=(I−2αΣP)μ_k and the λ-contractiveness claim follows under the data-based linear maps. The sample- and iteration-complexity lemmas are useful. The (N,α) tables and SNR Monte-Carlos line up with the predicted trade-offs, and the single scalar α is genuinely easier to interpret than Q/R tuning. Hardware on the ROSbot XL and Gazebo comparisons against DDLQR and model LQR are more than most theory papers deliver.\n\nSoft spots, in proportion. The certificates live entirely inside the linear data-based representations that rest on the constant-A=I, small-heading linearization. The physical plant is nonlinear Mecanum kinematics; when ϕ leaves the linear regime the certified map is no longer the realized map. That residual is never bounded. Theorem 2 is also noticeably more fragile under limited data. No code is released and the relaxations are only sufficient. These are real limits on how far the hardware claims can be pushed, but they do not break the linear-data theory or the practical tuning story.\n\nThis is for people already working in direct data-driven LQR / LMI control who want a geometry-aware alternative with one-parameter aggressiveness. It is not a foundational result, but it is carefully done and the math is inside standard territory. I would send it to referees; the linearization caveat should be stated more clearly, but the paper deserves a full review rather than a desk reject. Worth reading if you care about interpretable data-driven linear policies.","headline":"Solid data-driven extension of the authors' NGD idea with clean LMIs and real robot runs; the linearization gap is real but does not erase the contribution.","tokens_in":21895,"tokens_out":520,"would_cite":true,"duration_ms":5083,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Direct data-driven natural gradient control forces closed-loop states to follow an uncertainty-aware descent path without identifying a model.","keywords":["data-driven control","natural gradient descent","Fisher Information Matrix","linear matrix inequalities","closed-loop covariance","direct parameterization","robotics"],"falsifier":"Collect a data set that deliberately violates full row rank or drive the robot with large heading angles that leave the linear regime; if the LMI-synthesized gain still produces the predicted natural-gradient contraction and matches hardware trajectories, the claim is false.","tokens_in":21809,"feed_emoji":"🤖","tokens_out":650,"duration_ms":5334,"temperature":0.7,"pith_summary":"The paper claims that a linear feedback gain can be synthesized directly from input-state data so that the closed-loop mean state evolves exactly as a natural-gradient step. The Fisher Information Matrix of the Gaussian state is the inverse covariance; using it as a preconditioner makes the update large in uncertain directions and small in well-known ones. Two data-based representations of the closed-loop map (raw snapshots and sample covariances) are turned into linear matrix inequalities that enforce this geometry and certify contraction. A single scalar step-size then trades speed against smoothness, replacing the usual multi-matrix LQR tuning. Hardware trials on a Mecanum robot show that the resulting policies are stable, intuitive, and competitive with classical and data-driven LQR baselines under limited data.","feed_headline":"Data alone steers robots along natural-gradient paths","feed_subtitle":"One scalar step-size replaces LQR weight tuning and certifies contraction from input-state samples","key_machinery":"The Fisher Information Matrix of a Gaussian state (G=Σ^{-1}) together with the two data-parameterizations X_1 G = I−2αΣP (or the covariance analogue); these identities force the closed-loop mean to follow natural-gradient flow while the accompanying LMIs certify contraction and recover the gain K from data alone.","core_discovery":"Given sufficiently rich input-state data, a feasible solution of the stated LMIs yields a linear gain K such that the data-based closed-loop mean dynamics equal the natural-gradient recursion μ_{k+1}=(I−2αΣP)μ_k and are therefore λ-contractive in expectation; the same LMIs also upper-bound the stationary covariance so that the Fisher Information Matrix used for preconditioning remains consistent with the closed-loop uncertainty.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Data alone sets natural-gradient robot gains via LMI contraction","Input-state samples drive NGD policies that stay λ-contractive","Fisher-preconditioned data control skips models yet certifies stability","Closed-loop NGD from raw trajectories matches mean recursion and covariance","LMI solutions yield data-based K that equals natural-gradient updates"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The collected data matrix must have full row rank and the plant must stay inside the linear small-heading model used to gather that data; if either fails, the LMI certificates no longer apply to the physical system.","fun_headline_variants_meta":{"raw":{"variants":["Data alone sets natural-gradient robot gains via LMI contraction","Input-state samples drive NGD policies that stay λ-contractive","Fisher-preconditioned data control skips models yet certifies stability","Closed-loop NGD from raw trajectories matches mean recursion and covariance","LMI solutions yield data-based K that equals natural-gradient updates"]},"model":"grok-4.5","effort":"low","cost_usd":0.003122,"raw_usage":{"total_tokens":1071,"prompt_tokens":730,"num_sources_used":0,"completion_tokens":74,"cost_in_usd_ticks":31220000,"prompt_tokens_details":{"text_tokens":730,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":267,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":730,"tokens_out":74,"duration_ms":3206,"temperature":1.0,"reasoning_tokens":267,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T02:48:22.291194+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Collect a data set that deliberately violates full row rank or drive the robot with large heading angles that leave the linear regime; if the LMI-synthesized gain still produces the predicted natural-gradient contraction and matches hardware trajectories, the claim is false.","supporting_citations":[],"review_version":1}