{"id":"b63e8081-1961-4114-843e-5e08d24a9e3c","arxiv_id":"2604.06158","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The multistage DRRO-LQR problem over linear disturbance-feedback policies admits an exact SDP reformulation whose solution is the nominal certainty-equivalent LQR law plus a strictly causal empirical-mean correction.","lead":"The paper introduces the first tractable multistage distributionally robust regret optimization formulation for finite-horizon LQR control under shared ambiguity in disturbance statistics. This yields a less conservative controller than standard DRO while preserving regret guarantees, via an SDP-reformulable problem.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"The reader's weakest assumption treats the linear-policy restriction as jointly guaranteeing tractability and optimality. The paper's claim, however, only asserts tractability and an explicit form inside that class; it does not claim global optimality over all (possibly nonlinear) policies. Because the scoped claim contains no detectable flaw in its stated domain, the reader's identified assumption is not load-bearing for the central result.","tokens_in":1708,"tokens_out":320,"duration_ms":45166,"concrete_test":"For horizon T=2 and state dimension n=2, implement both the claimed SDP and a direct numerical min-max formulation (optimize over linear policy parameters and over distributions in the Gelbrich ball via moment constraints); if the optimal values and recovered policies agree to within 1e-6 relative error, the exact SDP reformulation holds for this instance.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that, when the decision maker is restricted to linear disturbance-feedback policies, the multistage DRRO-LQR problem under common stage-law Gelbrich ambiguity admits an exact SDP reformulation whose solution is the nominal CE-LQR law plus a strictly causal correction term. The abstract states this result directly and notes that worst-case distributions are non-unique. No internal inconsistency, hidden duality gap, or unstated assumption that would invalidate exactness of the SDP is visible from the claim as formulated. The restriction to linear policies is explicitly scoped to the tractability result rather than asserted as globally optimal.","agreement_with_reader":"disagree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript presents a distributionally robust regret optimization (DRRO) formulation for finite-horizon LQR under common stage-law ambiguity, where disturbances are independent but share an unknown distribution whose mean and covariance lie in a Gelbrich ball around nominal values. The central result is that, when the decision maker is restricted to linear disturbance-feedback policies, the multistage DRRO-LQR problem admits an exact semidefinite programming reformulation. The optimal controller takes the form of the nominal certainty-equivalent LQR law plus a strictly causal correction term driven by the empirical mean. The paper further characterizes the worst-case distributions (showing non-uniqueness for the DRRO-optimal policy) and reports numerical experiments indicating that DRRO is substantially less conservative than the corresponding DRO controller under the same ambiguity set while preserving the regret guarantee.","tokens_in":1828,"tokens_out":755,"duration_ms":43425,"significance":"If the derivations hold, the result is significant because it delivers the first tractable multistage ex-ante DRRO formulation for stochastic control, converting a generally NP-hard problem into an exact SDP whose solution has an explicit, interpretable structure (nominal CE-LQR plus strictly causal correction). The explicit controller form and the non-uniqueness result for worst-case distributions provide both computational and theoretical value. Numerical evidence of reduced conservatism relative to DRO, while retaining the regret guarantee, suggests practical utility in robust control design under distributional uncertainty.","major_comments":[{"comment":"The section deriving the SDP reformulation (the main theorem establishing exactness): the multistage extension under common stage-law ambiguity requires explicit verification that the inner supremum over the Gelbrich ball, combined with the regret objective, dualizes to an SDP with no duality gap. The abstract asserts exactness, but the multistage information structure (reuse of the stage law making past disturbances informative) could introduce complications not present in single-stage quadratic cases; the key dualization steps or strong-duality lemma should be highlighted.","section":"SDP reformulation theorem and proof"},{"comment":"The characterization of worst-case distributions: the claim that they are non-unique for the DRRO-optimal policy is load-bearing for the theoretical contribution. The paper should state whether the non-uniqueness is constructive (explicit families of distributions attaining the supremum) or only existential, and confirm that this does not affect the exactness of the SDP solution.","section":"Worst-case distribution analysis"}],"minor_comments":[{"comment":"The motivation that the nominal CE controller is generally not regret-optimal should be illustrated with a low-dimensional numerical example or a short analytic counter-example early in the paper, rather than only asserted via the information-structure argument.","section":"Introduction"},{"comment":"Notation consistency: the abstract uses 'strictly causal empirical-mean correction'; the main text should define this term with an explicit equation (e.g., the form of the correction gain) at first use.","section":"Controller structure"},{"comment":"Numerical results: the reported approach of the correction coefficients to the certainty-equivalent feedforward coefficient is interesting, but the paper should report the number of Monte-Carlo trials, seed variability, and quantitative regret or cost differences (not only qualitative 'substantially less conservative') to support the comparison with DRO.","section":"Numerical experiments"}],"recommendation":"minor_revision","confidential_remarks":"The manuscript appears well-scoped for a math.OC venue. The restriction to linear disturbance-feedback policies is explicitly presented as a tractability device rather than a global optimality claim, which avoids overstatement."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the positive assessment and constructive comments on our manuscript. We address each major comment below and will make the indicated revisions to improve clarity.","responses":[{"response":"We appreciate the referee's suggestion to enhance the presentation of the proof. The existing derivation already verifies strong duality for the multistage setting by exploiting the convexity and compactness of the Gelbrich ball together with the quadratic regret objective under linear disturbance-feedback policies; the common stage-law ambiguity is handled via a reformulation that accounts for the information structure without introducing a duality gap. To address the comment, we will add a dedicated remark immediately following the main theorem that explicitly outlines the key dualization steps and references the strong-duality result used, thereby highlighting the multistage extension relative to the single-stage case.","revision_made":"yes","referee_comment":"[SDP reformulation theorem and proof] The section deriving the SDP reformulation (the main theorem establishing exactness): the multistage extension under common stage-law ambiguity requires explicit verification that the inner supremum over the Gelbrich ball, combined with the regret objective, dualizes to an SDP with no duality gap. The abstract asserts exactness, but the multistage information structure (reuse of the stage law making past disturbances informative) could introduce complications not present in single-stage quadratic cases; the key dualization steps or strong-duality lemma should be highlighted."},{"response":"We thank the referee for this observation. Our characterization is constructive: the manuscript exhibits explicit families of distributions (specific mean and covariance perturbations within the Gelbrich ball) that attain the supremum for the DRRO-optimal policy. This construction is used to establish non-uniqueness while confirming that the attained value matches the SDP optimum. We will revise the relevant section to state explicitly that the non-uniqueness is constructive and to add a sentence confirming that it is compatible with (and does not affect) the exactness of the SDP reformulation.","revision_made":"yes","referee_comment":"[Worst-case distribution analysis] The characterization of worst-case distributions: the claim that they are non-unique for the DRRO-optimal policy is load-bearing for the theoretical contribution. The paper should state whether the non-uniqueness is constructive (explicit families of distributions attaining the supremum) or only existential, and confirm that this does not affect the exactness of the SDP solution."}],"tokens_in":1501,"tokens_out":511,"duration_ms":40588,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's key result is an exact SDP for multistage DRRO-LQR over linear disturbance-feedback policies, where the controller is the nominal CE-LQR law plus a strictly causal empirical-mean correction. They make progress on a setting that was hard before by using the common stage-law ambiguity with Gelbrich balls, which lets them reformulate the problem without losing the regret focus. The work shows why multistage matters here unlike single-stage, and backs it with numerics that DRRO is less conservative than DRO. The non-unique worst-case distributions add some insight. The soft spots are minor but worth noting. The linear policy restriction is necessary for the SDP but might miss better nonlinear options, though they don't overclaim. The ambiguity set is tailored to enable independence with shared parameters, which works but narrows the scope. If the full paper has the derivations, they should hold up given the stress-test, but verification would be key. Readers in robust stochastic control will find this useful for designing less conservative controllers with regret properties. It deserves a serious referee as the tractability is a concrete advance in the area. I recommend putting it through peer review.","headline":"The paper delivers a tractable SDP reformulation for multistage DRRO-LQR under linear disturbance-feedback policies and common stage-law ambiguity, with the solution being nominal CE-LQR plus a causal correction.","tokens_in":2322,"tokens_out":316,"would_cite":false,"duration_ms":36062,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"Standard multistage robust LQR SDP reformulation with no RS-shaped cost or periodicity structure","alignment":"orthogonal","rationale":"The paper's core machinery is an exact SDP reformulation of multistage ex-ante DRRO-LQR under Gelbrich-ball ambiguity on shared stage-law moments, yielding the nominal CE-LQR policy plus a strictly causal empirical-mean correction term (Theorems 1-2, Lemmas 1-2, policy form (35)). This uses conventional quadratic advantage representations, Riccati recursions, and dual LMIs on Bures/Gelbrich distances. None of these elements invoke J-cost forcing, cosh identities, ratio symmetry, φ-ladder spacings, 8-tick periodicity, or parameter-free constant derivations. The domain (math.OC, finite-horizon linear-quadratic control) lies outside the RS forcing chain from a single distinction.","tokens_in":51408,"confidence":"high","tokens_out":203,"duration_ms":10219,"cache_read_input_tokens":32896,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Multistage distributionally robust regret-optimal LQR under common stage-law ambiguity admits an exact SDP reformulation over linear disturbance-feedback policies.","keywords":["distributionally robust optimization","regret optimization","linear quadratic regulator","semidefinite programming","stochastic control","ambiguity sets","Gelbrich distance"],"falsifier":"A concrete instance in which a nonlinear disturbance-feedback policy achieves strictly lower worst-case regret than the SDP-derived linear policy under the same Gelbrich ambiguity set.","tokens_in":2604,"feed_emoji":"","tokens_out":674,"duration_ms":34803,"temperature":0.7,"pith_summary":"The paper establishes that distributionally robust regret optimization for finite-horizon linear quadratic regulators becomes tractable when disturbances are independent but share an unknown stage law whose mean and covariance lie inside a Gelbrich ball. A sympathetic reader would care because this setup captures realistic uncertainty where past realizations inform future decisions, yet standard robust methods often produce overly cautious controllers. By restricting attention to linear disturbance-feedback policies, the multistage problem converts into a semidefinite program whose solution is the nominal certainty-equivalent LQR law plus a strictly causal correction term driven by empirical means. If correct, the resulting policy preserves a regret guarantee while being substantially less conservative than the corresponding distributionally robust optimal controller under identical ambiguity.","feed_headline":"DRRO-LQR solved exactly via SDP as nominal law plus causal correction","feed_subtitle":"Under common stage-law Gelbrich ambiguity the optimal linear policy corrects the certainty-equivalent LQR with past empirical means and is 0","key_machinery":"Linear disturbance-feedback policies together with the Gelbrich-ball ambiguity set, which together convert the multistage regret objective into an exact SDP.","core_discovery":"Over linear disturbance-feedback policies the multistage DRRO-LQR problem with common stage-law ambiguity (Gelbrich ball) admits an exact semidefinite programming reformulation; the optimal controller equals the nominal certainty-equivalent LQR law plus a strictly causal empirical-mean correction. Worst-case distributions realizing the optimal value are nonunique.","pith_inferences":["The same SDP route may apply to infinite-horizon or time-varying versions of the problem if the stage-law ambiguity remains common across periods.","Regret criteria could reduce conservatism in other multistage control settings where past observations inform future decisions under shared distributional uncertainty.","Numerical verification on real plants would test whether the empirical-mean correction measurably improves closed-loop performance over pure certainty-equivalent control."],"forward_implications":["The optimal policy reuses the nominal LQR gain and adds a correction that depends only on past realized disturbances.","Relative to DRO under the identical ambiguity set, the DRRO controller is often substantially less conservative while retaining the regret guarantee.","Worst-case distributions for the DRRO-optimal policy are nonunique.","The correction coefficients in the optimal policy empirically approach the certainty-equivalent feedforward term as horizon length grows."],"fun_headline_variants":["Exact SDP reformulation for multistage DRRO-LQR with causal correction","Optimal DRRO-LQR equals CE-LQR plus empirical-mean causal correction","Multistage DRRO-LQR solved exactly via SDP under common ambiguity set","Gelbrich DRRO yields SDP with nominal LQR law and mean correction"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"That linear disturbance-feedback policies are sufficient to achieve both tractability and optimality under the common stage-law Gelbrich-ball model.","fun_headline_variants_meta":{"raw":{"variants":["Exact SDP reformulation for multistage DRRO-LQR with causal correction","Optimal DRRO-LQR equals CE-LQR plus empirical-mean causal correction","Multistage DRRO-LQR solved exactly via SDP under common ambiguity set","Gelbrich DRRO yields SDP with nominal LQR law and mean correction"]},"model":"grok-4.3","cost_usd":0.008337,"raw_usage":{"total_tokens":3768,"prompt_tokens":651,"num_sources_used":0,"completion_tokens":79,"cost_in_usd_ticks":83374500,"prompt_tokens_details":{"text_tokens":651,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3038,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":651,"tokens_out":79,"duration_ms":31625,"temperature":1.0,"reasoning_tokens":3038,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-10T18:44:45.215300+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A concrete instance in which a nonlinear disturbance-feedback policy achieves strictly lower worst-case regret than the SDP-derived linear policy under the same Gelbrich ambiguity set.","supporting_citations":[],"review_version":1}