{"id":"dd34dbc3-b8b9-4bbd-b6ae-2eafc35a7119","arxiv_id":"2607.03251","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"One-step policy-gradient LQR updates with normalized sliding-window least-squares stabilize unknown slowly varying and piecewise-constant linear systems and track frozen-time optima on average.","lead":"The paper gives a lightweight adaptive controller that updates a linear feedback gain by one policy-gradient step per time, using a local model fit from a short sliding window of closed-loop data. It proves practical stability for slowly drifting and for switched unknown linear systems, with average tracking of the frozen-time LQR optimum.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the PE/smallness conditions already flagged by the reader.","rationale":"The reader’s strongest claim accurately restates Theorems 16 and 18. The weakest-assumption paragraph correctly isolates continuous normalized PE under feedback and the non-constructive smallness conditions; those are the genuine load-bearing hypotheses, not hidden gaps in the sequential-stability or dwell-time arguments. The appendix proofs are complete and follow the standard LTI PGAC template with the necessary LTV perturbations (normalized LS bias, cost jumps at switches, two-layer contraction). Simulations are consistent with the theory. No independent algebraic contradiction or missing step that would overturn the claims was found. Therefore the ACCEPT verdict stands; confidence remains moderate because the constants are non-constructive and PE is assumed rather than proved under feedback.","tokens_in":23639,"tokens_out":574,"duration_ms":6958,"concrete_test":"Independently re-derive the key descent inequality (C.3) / (D.3) from Lemmas 21–23 and the exact-gradient inequality (C.2) without invoking the uniform cost bound C of Lemma 25; if the polynomial bounds p_i remain finite only after assuming C_t(K_t)≤C, confirm that the induction in Lemma 25 closes for the stated thresholds on δ,ε,η. A second quick check: verify that the sequential-stability similarity factors H_t=Σ_t^{1/2} indeed satisfy ∥H_{t+1}^{-1}H_t∥≤1+α/2 once ∥Σ_{t+1}-Σ_t∥ is controlled by (C.8).","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claims of Theorems 16 and 18 rest on sequential stability of the adaptive policy sequence under one-step certainty-equivalence policy gradient with normalized sliding-window LS. The proofs (Appendix C–D) carefully close the loop among identification error (Lemmas 14/17), cost boundedness (Lemma 25), sequential stability (Lemma 26 / Lemma 27), and the resulting PES / interval-wise PES bounds, using standard Lyapunov-perturbation and gradient-dominance arguments from the LTI PGAC literature. The only places where the argument can fail are exactly those already identified: collapse of the uniform normalized PE lower bound γ under closed-loop adaptation (Assumption 11), or violation of the non-constructive smallness conditions on δ, ε (or ε_sw), and η relative to constants that depend on (a_m,b_m,Q,R,K_t0). These are standard and disclosed; they do not appear to hide an internal inconsistency in the algebra as written.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes policy gradient adaptive control (PGAC) for unknown LTV systems: at each time the state-feedback gain is updated by one gradient step of a frozen-time LQR cost built from a normalized sliding-window least-squares model estimate. For slowly varying systems (Assumptions 2–4, 11), Theorem 16 shows sequential stability of the policy sequence and practical exponential stability of the closed loop without a dwell-time condition, plus an average frozen-time optimality-gap bound. For piecewise-constant systems (Assumptions 2, 3, 7, 11), Theorem 18 gives within-mode decay and interval-wise PES under a dwell-time condition, with a corresponding optimality-gap bound that depends on switching frequency. Identification-error lemmas (14, 17), cost boundedness by induction, sequential-stability arguments, and telescoping gap estimates are developed in the appendix; three numerical examples (continuous LTV, switched LTV, nonlinear planar quadrotor) illustrate the method.","tokens_in":23851,"tokens_out":951,"duration_ms":7849,"significance":"The contribution is a useful middle ground between one-shot certainty-equivalence LQR/SDP updates and purely direct data-driven schemes: the first-order update is computationally light and the stepsize directly limits policy chatter under noisy estimates. Extending sequential-stability certificates to continuous slow variation without dwell time, and obtaining uniform constants for infinitely many switches in the piecewise-constant case, are genuine advances over the authors’ preliminary switched result and over several recent LTV data-driven methods. The normalized sliding-window estimator cleanly separates variation, noise, and excitation in the error bounds. The analysis is fully spelled out in the appendix and the numerical examples (including a nonlinear plant) support the claims. The PE and smallness conditions are standard for this literature and are disclosed.","major_comments":[{"comment":"Assumption 11 requires a uniform lower bound γ on the normalized data Gramian for every window while the gain is adapting. Lemmas 14 and 17 and all subsequent thresholds in Theorems 16 and 18 collapse if closed-loop excitation fails. The manuscript should either (i) give a concrete, checkable construction of e_t that preserves γ under the sequential-stability bounds already proved, or (ii) state explicitly that PE is an external hypothesis and discuss how γ can be monitored online. Without one of these, the certificates remain conditional on an assumption that is not automatically inherited from the closed-loop dynamics.","section":null},{"comment":"The admissible ranges for δ, ε (or ε_sw) and η in Theorems 16 and 18 are expressed via existential constants ν_i that depend on (a_m, b_m, Q, R, K_t0) but are never made explicit or estimated. While this is common in sequential-stability arguments, a short remark (or a numerical illustration of how large η may be taken for the examples of Section 5) would make the results more usable and would clarify that the conditions are not vacuous for the simulated systems.","section":null}],"minor_comments":[{"comment":"In Algorithm 1 the data window used for identification is written (X_{t+1}, U_{t+1}, X_{t+2}) while the surrounding text uses the indexing of (12); a one-line clarification of the time shift would avoid confusion.","section":null},{"comment":"Figures 2–5 would benefit from a short caption note on the magnitude of the residual set relative to e_m and w_m, so that the practical-stability claim is visually linked to the theory.","section":null},{"comment":"A few typos remain (e.g., “D¨ orfler”, “Z¨ urich” in the author block; “the pol-icy update” line break in Section 4.2). A light copy-edit pass is enough.","section":null},{"comment":"The discussion of DeePO and discounted LQR in Section 4.4 is interesting but could be shortened; the main theorems already stand without it.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The paper is a solid incremental contribution that fits Automatica’s scope. The PE issue is the only load-bearing soft spot; once the authors address it with a short constructive remark or an explicit caveat, the manuscript is ready for acceptance. No novelty or citation concerns."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a careful, usable extension of the authors’ LTI PGAC line to unknown LTV plants. The punchline is that one gradient step of the frozen-time LQR cost, driven by normalized sliding-window least squares, is enough to get practical exponential stability without dwell time under slow drift (Thm 16), and interval-wise PES under dwell time for piecewise-constant jumps (Thm 18), plus average frozen-time optimality-gap bounds in both cases.\n\nWhat is actually new is the combination: normalized LS that removes the a-priori state-bound circularity of ordinary LS, sequential-stability certificates that do not need dwell time for continuous slow variation, and uniform constants that allow infinitely many switches in the piecewise case (improving their ECC switched note). The algorithm is light—one gradient step, recursive rank-one update—and the stepsize directly limits policy chatter from noisy estimates, which is a genuine practical advantage over online Riccati/SDP certainty-equivalence. Appendix proofs close the loop (ID error → cost boundedness by induction → sequential stability → ISS-style state bounds → telescoping gaps) with standard Lyapunov-perturbation and gradient-dominance tools. Simulations are honest: slow LTV, switched LTV, and a nonlinear quadrotor where the LTV model is only local.\n\nSoft spots are the usual ones for this literature and are disclosed: continuous PE of the normalized Gramian under adapting feedback (Assumption 11), an initial stabilizing gain, and existential smallness of δ, ε (or ε_sw), and η relative to non-constructive ν_i depending on (a_m,b_m,Q,R,K_t0). No shipped code. The contribution is incremental on their prior PGAC work rather than a foundational rewrite. None of that breaks the logic as written; the stress-test correctly finds no hidden algebraic inconsistency.\n\nThis is for people who already care about data-driven / adaptive LQR and want a lightweight online alternative with explicit certificates. It deserves a serious referee. I would engage with it and expect it to clear peer review with the usual tightening of constants and PE discussion.","headline":"Solid Automatica-style LTV extension of PGAC: one-step gradient updates plus normalized window LS, with real PES and average optimality-gap theorems for slow drift and switched modes.","tokens_in":24476,"tokens_out":541,"would_cite":true,"duration_ms":4978,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["93C40","93B52","49N10","93C05"],"pacs":[],"model":"grok-4.5","headline":"One gradient step per time update stabilizes unknown linear time-varying systems while tracking frozen-time LQR policies.","keywords":["adaptive control","linear quadratic regulator","linear time-varying systems","policy gradient","data-driven control","sequential stability","normalized least-squares"],"falsifier":"Run the algorithm on a slowly varying LTV plant whose continuous drift exceeds the paper's (non-constructive) variation threshold while keeping excitation, noise and step-size fixed; if the state still decays practically exponentially, the slow-variation claim is false. Alternatively, drop the probing signal so that the normalized Gramian falls below the assumed floor and check whether the identification error and subsequent stability guarantees remain valid.","tokens_in":24503,"feed_emoji":"⚙️","tokens_out":697,"duration_ms":6288,"temperature":0.7,"pith_summary":"Unknown linear plants whose parameters drift or jump need a feedback gain that keeps changing from live closed-loop data. Most existing data-driven schemes recompute a full optimizer (Riccati equation or SDP) at every step; that is expensive and can swing wildly when the model estimate is noisy. This paper instead inserts a single policy-gradient step of the linear-quadratic cost into the feedback loop, using a local model estimated by normalized sliding-window least-squares. The resulting algorithm is cheap, its step-size directly limits how much the gain can jump, and the paper proves that the closed loop remains practically exponentially stable for two standard classes of time-varying systems. For slow continuous drift the proof needs no dwell-time condition; for abrupt switches a dwell-time contraction across modes is enough. Average frozen-time optimality gaps are also bounded, so the adapted gain tracks the ideal LQR of the current frozen model.","feed_headline":"One gradient step stabilizes unknown drifting linear plants","feed_subtitle":"Lightweight policy-gradient updates give practical exponential stability without full re-optimization","key_machinery":"Policy-gradient adaptive control (PGAC): at each instant form a normalized sliding-window least-squares model of the recent closed-loop trajectory, evaluate the LQR policy gradient of that frozen surrogate, and take exactly one descent step; sequential stability of the resulting gain sequence then supplies the closed-loop certificates.","core_discovery":"Under small enough variation (or long enough dwell), small enough normalized identification error, and a sufficiently small gradient step-size, one-step policy-gradient adaptive control produces a sequentially stable gain sequence for slowly varying unknown LTV systems and therefore practical exponential stability without dwell time; for piecewise-constant LTV systems the same controller yields within-mode decay plus interval-wise practical exponential stability under a dwell-time condition, together with explicit average frozen-time LQR optimality-gap bounds in both cases.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["One gradient step gives practical stability for unknown LTV systems","PGAC: incremental policy gradients stabilize drifting unknown plants","Normalized window LS + one-step gradient for adaptive LQR control","No full re-opt: sliding-window gradient adapts feedback for LTV plants","Policy-gradient adaptive control: certificates for slow & switched LTV"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The closed-loop data must stay persistently exciting under the adapting gain itself; if the normalized excitation level collapses, the identification-error bounds fail and the stability thresholds become empty.","fun_headline_variants_meta":{"raw":{"variants":["One gradient step gives practical stability for unknown LTV systems","PGAC: incremental policy gradients stabilize drifting unknown plants","Normalized window LS + one-step gradient for adaptive LQR control","No full re-opt: sliding-window gradient adapts feedback for LTV plants","Policy-gradient adaptive control: certificates for slow & switched LTV"]},"model":"grok-4.5","effort":"low","cost_usd":0.001658,"raw_usage":{"total_tokens":875,"prompt_tokens":802,"num_sources_used":0,"completion_tokens":73,"cost_in_usd_ticks":16580000,"prompt_tokens_details":{"text_tokens":802,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":0,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":802,"tokens_out":73,"duration_ms":1382,"temperature":1.0,"reasoning_tokens":0,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T03:48:06.200514+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the algorithm on a slowly varying LTV plant whose continuous drift exceeds the paper's (non-constructive) variation threshold while keeping excitation, noise and step-size fixed; if the state still decays practically exponentially, the slow-variation claim is false. Alternatively, drop the probing signal so that the normalized Gramian falls below the assumed floor and check whether the identification error and subsequent stability guarantees remain valid.","supporting_citations":[],"review_version":1}