{"id":"9cee8e3b-4fe4-4a31-8433-e8e31c998e53","arxiv_id":"2608.13535","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A team-optimal linear controller and communication schedule for partially nested linear-quadratic multi-agent systems can be computed with Riccati-style recursions.","lead":"This paper gives exact formulas for jointly deciding what agents should communicate and how they should control a shared linear system with quadratic costs. It is useful for engineers who design multi-agent control when messages are costly and information flows are partially nested.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: the open-loop claim is internally consistent; Assumption III.4 is restrictive but explicitly scoped and supported by counterexamples.","rationale":"The reader identified Assumption III.4 as the weakest assumption, and I agree that it is the most load-bearing condition for the open-loop result: it is what allows the proof of PN preservation under additional sharing and the SI-CIB property of the strict expansion. The paper explicitly states this assumption, proves its necessity via Lemma III.5, and confines the central theorem to problems satisfying it, so this is a scoped limitation rather than an internal defect. I independently checked the other main pillars: Corollary A.2's extension of Ho-Chu linearity to semidefinite costs and singular Gaussians is plausible and supported by a direct static-team argument; Lemma IV.1's transfer of optimality from the expanded problem is constructive and relies only on the common-information absorption property; Lemma IV.5's pointwise Bellman projection, though terse, is valid under the SI-CIB conditional-Gaussian structure; and the Riccati algebra in Theorem IV.6 is internally consistent, including the pseudo-inverse consistency conditions supplied by Lemma B.1. The numerical experiments are illustrative but not load-bearing. Because the paper's central claim is well supported within its stated assumptions and the main limitation is transparently scoped, I do not find a reason to alter the ACCEPT verdict.","tokens_in":51930,"tokens_out":37968,"duration_ms":406157,"concrete_test":"Independently expand the proof of Lemma IV.5 by writing out the residual-orthogonality and normal-equations-consistency steps for a one-step example with nonzero B_{1,1}, private observation Y_{1,1}, and cross-coupled Q2; verify that the bE* formula in Eq. IV.5 and the F* update in Eq. IV.6 reproduce the exact Bellman minimum. If the expanded proof reveals a dependence of the optimal affine coefficients on the common history beyond the CIB mean, the Riccati recursions would need revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central open-loop claim genuinely depends on Assumption III.4: every nonzero-input action must influence another agent's next-step observation. This assumption is used in Theorem III.6 and Theorem IV.2 to conclude that the acting agent's information is absorbed into common information (I_{i,h} subset of C_{h+1}), and Lemma III.5 shows that when it fails, the optimal strategy can be nonlinear. This is a real limitation of scope, but it is a clearly stated modeling assumption rather than a hidden or circular step, and the paper proves necessity with a counterexample. I also examined the more delicate pointwise Bellman/Riccati step in Lemma IV.5 and Theorem IV.6. The compressed L2 projection argument is valid under the SI-CIB condition: the projection residual is orthogonal to all private-information components because those components are subvectors of the zero-mean Gaussian fluctuation omega_h, and it is independent of future noises. The Kalman recursion in Theorem IV.6 is supported by the strict expansion, which also places the generating private information in the next-step common information, so the known-input update does not improperly treat a private, state-correlated action as exogenous. I find no internal inconsistency in the derivation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper formalizes a joint communication-control strategy optimization (JCCO) problem for multi-agent LQG systems under the common-information-based (CIB) framework. For open-loop communication strategies, it gives structural conditions (Assumptions II.1, III.2, III.4) under which additional sharing preserves partial nestedness and a linear optimal controller exists (Theorem III.6). It then constructs a strict expansion that satisfies the SI-CIB condition, shows the CIB belief is Gaussian, and derives Riccati recursions for the optimal controller (Theorems IV.2, IV.4, IV.6). The same machinery is extended to closed-loop communication strategies under Assumption V.1, yielding a finite-dimensional dynamic program (Section V, Algorithm 1). The paper includes counterexamples showing necessity of the structural assumptions and a numerical study.","tokens_in":52057,"tokens_out":33144,"duration_ms":356253,"significance":"If the main results are correct, the paper is a significant contribution: it connects partial nestedness with the strategy-independent CIB condition and provides a candidate closed-form solution for a class of decentralized LQG problems with output feedback and communication optimization. The appendix contains detailed proofs and explicit counterexamples for the necessity of Assumptions III.2 and III.4, which is a strength. The numerical experiments illustrate the method on four information structures. However, the correctness of the central Riccati recursion depends on a Kalman-filter step that I believe is not justified; this must be repaired before the paper's main computational claim can be accepted.","major_comments":[{"comment":"The recursion for eΘ_{h+1} is derived by the standard known-input Kalman update, but eU_h is not an exogenous input: by Lemma IV.5 it equals bE_h eΘ_h + F_h Ip,h(eS_h-eΘ_h), hence it is correlated with the filter error eS_h-eΘ_h given eC_h. Strict partial nestedness only ensures eU_h (and possibly the relevant private labels) belong to eC_{h+1}; it does not make the prediction E[eS_{h+1}|eC_h,eU_h] equal to A_h eΘ_h+B_h eU_h. A minimal Gaussian counterexample to this formula is X_1∼N(0,1), U_1=X_1, X_2=X_1+U_1+W, Y=[U_1, X_2+V]^T with unit variances: the exact conditional mean of X_2 given Y is 2U_1+0.5(Y_2-U_1), whereas the paper's update gives U_1+0.5(Y_2-U_1). Consequently the matrices {K^j_{h+1}} and the Riccati recursions (IV.5)-(IV.7) are not established; this is load-bearing for the claimed closed-form solution.","section":"§C-6, Eq. (C.6); Theorem IV.6"},{"comment":"The closed-loop DP is presented under the unproved assumption that all displayed Bellman minima are attained by admissible strategies. Since the strategy spaces over continuous states are not compact and the joint minimization over communication actions and control prescriptions can have infima that are not attained, this assumption is not automatic. The paper should either prove attainment under the stated hypotheses or reformulate the claim in terms of epsilon-optimal strategies.","section":"Section V, Appendix D-B (Algorithm 1)"}],"minor_comments":[{"comment":"The condition for adding \\bar U_{i,t} to \\bar C_h uses \\bar I_{i,t}\\subseteq \\bar C_h, but \\bar I_{i,t} is not explicitly defined for the fixed problem D(g^m_{1:H}) after the bar notation is introduced; please define it in the reduced notation.","section":"Eq. (IV.1)"},{"comment":"The notation for the additional-sharing information is inconsistent: Z_h^a, Z^a_h, and \\cup_i Z_{i,h}^a are used interchangeably.","section":"Section II-A"},{"comment":"The line labels 'N' in Figure 2 should be defined in the caption, and it would be useful to report standard deviations over the 10 random seeds in Figures 1 and 2.","section":"Section VI-B, Figures 1 and 2"},{"comment":"The definitions eL^1_h=eL^3_h and eL^2_h=eL^4_h are duplicated; a single definition would reduce confusion.","section":"Theorem IV.6"}],"recommendation":"major_revision","confidential_remarks":"The structural results in Section III and the SI-CIB proof in Section IV are valuable and appear sound. The main risk is the Kalman-filter step in §C-6; if the authors can supply a correct recursion, or show that the known-input formula is valid under additional structure that they prove, the paper could become a strong accept. The closed-loop attainment caveat should also be addressed. I do not see a circularity or authorship issue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Jo: quick read of arXiv:2608.13535. The headline is that the open-loop claim holds up. For a PN JCCO problem satisfying Assumptions II.1, III.2, III.4 and with open-loop communication, they prove existence of a linear optimal control strategy for every fixed schedule and give closed-form Riccati recursions. The key new idea is the strict expansion of the PN information structure that yields strategy-independent CIB beliefs, which then reduces the problem to a centralized LQG with a finite-dimensional sufficient statistic. That is a real advance over [17], which left the private-information component unexplored.\n\nThe paper is also honest. Each structural assumption is tested with a counterexample showing that dropping it causes nonlinearity or non-existence. The counterexamples are constructed carefully, not just asserted. Appendix proofs are full and the central derivation is not circular: linearity comes from PN via Corollary A.2, SI-CIB is proven for the strict expansion, and the Riccati recursion follows from standard centralized LQG once the sufficient statistic is identified.\n\nThe main soft spot is the scope. Assumption III.4—every nonzero input must influence another agent's next observation—excludes a lot of systems, including many with decoupled dynamics or sparse observations. The paper acknowledges this and proves necessity, so it is a scoping assumption rather than a hidden flaw, but it does limit the practical reach. The closed-loop extension in Section V is thinner: it rests on an attainment assumption stated only in Appendix D-B, and the DP in Algorithm 1 is over belief states and message sequences, so the 'more tractable than infinite-dimensional' claim is plausible but not demonstrated with complexity bounds. The numerical experiments are illustrative—no error bars, small scale—but they are not the load-bearing evidence.\n\nI checked the two delicate steps the reader flagged. The pointwise Bellman/L2 projection argument in Lemma IV.5 is compressed but valid under SI-CIB; the projection residual is orthogonal to the private components and to future noises. The Kalman step in Theorem IV.6 is also sound: the strict expansion puts the generating private information in next-step common information, so the known-input update is legitimate.\n\nBottom line: this is a substantive, credibly proven within-subfield contribution. It deserves a serious referee. I wouldn't cite it in my own work in the next year, but I'd want it in the literature.","headline":"Credible Riccati-solvable DP for PN JCCO; the restrictive assumptions are honestly scoped and supported by counterexamples.","tokens_in":52657,"tokens_out":2360,"would_cite":false,"duration_ms":22647,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["93C05","93E20","49N10","90C39","93A14"],"pacs":[],"model":"deepseek-v4-flash","headline":"For partially nested multi-agent linear-quadratic systems, optimal control stays linear under every fixed open-loop communication schedule, and the paper gives closed-form Riccati recursions to compute it.","keywords":["joint communication-control optimization","partially nested information structures","decentralized LQG","Riccati equations","common-information-based approach","output feedback","multi-agent control","communication strategy"],"falsifier":"Build the two-agent, two-step scalar system in Lemma III.5: agent 1 acts first, agent 2 observes only a noisy version of the state after agent 1's action, and agent 1's action influences the state but agent 2's observation matrix is zero for that influence. If communication of agent 1's private observation is allowed at zero communication cost, then a team-optimal strategy must be nonlinear; verifying numerically that the optimal cost over linear strategies is strictly worse than the optimal nonlinear strategy settles whether Assumption III.4 is needed. Equivalently, one could compute the Riccati value under Assumption III.4 on random matrices and check that it coincides with the value of an explicit nonlinear search in small horizons.","tokens_in":51636,"feed_emoji":"🎛️","tokens_out":6844,"duration_ms":64312,"temperature":0.7,"pith_summary":"This paper studies a team of agents controlling a linear system with quadratic costs who may also decide what private information to share with each other at each step, paying a communication cost. The central claim is that when the baseline information structure is partially nested and three structural assumptions hold, fixing any open-loop communication schedule leaves a decentralized linear-quadratic-Gaussian problem whose optimal control strategy is linear and computable by closed-form Riccati recursions; the full communication-control problem is then solved by enumerating the finite set of schedules. The paper further shows that dropping any of the structural assumptions can make the optimal control strategy nonlinear or even nonexistent, so the assumptions carry real weight. For closed-loop communication strategies, an additional assumption on what communication strategies may depend on yields a dynamic program over finite-dimensional Gaussian conditional means rather than infinite-dimensional beliefs. The practical upshot is a tractable, principle-based recipe for co-designing whom to tell what and how to steer.","feed_headline":"Three conditions keep optimal control linear after agents share","feed_subtitle":"Under them, joint communication-control design reduces to Riccati recursions plus a finite schedule search.","key_machinery":"The load-bearing object is the strict expansion of the information structure: starting from the fixed-schedule problem, add to the common information every past control action that has a nonzero input matrix and is already known to a recipient whose information is common, and remove those actions from private information. This expansion turns a partially nested information structure into a strictly partially nested one and makes the common-information-based belief strategy-independent, so the belief is a Gaussian whose mean is a fixed linear function of the common record and whose covariance is deterministic (Lemma IV.3). The Gaussian mean then serves as the state of a Markovian dynamic program, and optimizing over strategies that are linear in the mean and private information reduces to a centralized Kalman filter for the mean and the Riccati recursions of Theorem IV.6 for the gains. The paper also uses the finite set of open-loop schedules to close the loop: enumerate schedules, run the recursion for each, add communication costs, and pick the minimum.","core_discovery":"Let a joint communication-control optimization (JCCO) problem be partially nested, with information evolution satisfying Assumption II.1, useless controls excluded by Assumption III.2, and every effective control observable by at least one other agent at the next time step by Assumption III.4. The paper proves that for any fixed open-loop communication strategy, the induced problem is a decentralized LQG problem with a partially nested information structure, and that a team-optimal control strategy exists that is linear in the available information (Theorem III.6). The proof route expands the information structure so that actions influencing other agents' information are moved into common information; in the expanded problem the common-information-based belief is strategy-independent and Gaussian, so its conditional mean and covariance are finite-dimensional sufficient statistics. A backward Riccati recursion (Theorem IV.6), fed by a centralized Kalman filter for the mean and a convex quadratic minimization for the private gains, computes the optimal linear strategy for each fixed schedule; minimizing over the finite set of schedules solves the original JCCO problem. With closed-loop communication, assuming communication strategies depend only on common information and past messages (Assumption V.1) plus an attainment condition stated in Appendix D-B, the same expansion yields a dynamic program over finite-dimensional Gaussian beliefs.","pith_inferences":["Editorial extension: The schedule enumeration is exponential in the number of agents and horizon; a natural next step is branch-and-bound or beam search over schedules using the Riccati value as a lower bound, which the paper does not investigate.","Editorial extension: Assumption III.4 is a one-step observability condition; one could test whether it can be replaced by a global detectability condition on the pair formed by the state transition and observation maps, which would extend the method to systems where influence takes two or more steps to reach another agent's observations.","Editorial extension: The linearity result suggests a concrete test on random instances: compare the Riccati-computed policy against a nonlinear policy found by numerical search; if the assumptions hold, the linear policy should match or beat it, and violations of III.4 should show a gap."],"forward_implications":["For every fixed open-loop communication schedule, the optimal control strategy is linear and is computed by the closed-form Riccati recursions; the globally optimal schedule is found by finite enumeration.","The assumptions III.2 and III.4 are not technical decorations: dropping either can force nonlinearity or non-existence of the team-optimal control strategy, as shown by explicit two-agent two-step counterexamples.","As a byproduct, the same recursions solve decentralized LQG control with partially nested information structures and output feedback under the common-information-based approach, including settings with singular noise covariances and positive-semidefinite cost matrices.","Under closed-loop communication with Assumption V.1, the dynamic program runs over finite-dimensional Gaussian conditional means and covariances, so the infinite-dimensional belief state of a direct common-information treatment is avoided."],"supporting_citations":[{"why":"Establishes that partially nested information structures admit linear team-optimal strategies; the paper extends this to semidefinite costs and uses it as the anchor for Theorem III.6.","marker":"[15]"},{"why":"The Witsenhausen counterexample is the template for the nonlinearity and non-existence constructions in Lemmas III.1, III.3, and III.5.","marker":"[23]"},{"why":"Supplies the common-information-based framework and the notion of prescriptions used throughout the paper.","marker":"[10]"},{"why":"Supplies the strategy-independent common-information-based belief condition and the Gaussian belief update that Theorem IV.2 and Lemma IV.3 adapt.","marker":"[11]"},{"why":"Shows that linear control strategies with a fixed private-information component reduce to centralized LQG control solved by Riccati equations; the paper removes the fixed-component restriction under partial nestedness.","marker":"[17]"},{"why":"Motivates the dynamic-programming approach and gives the factorized-state case that the paper's method generalizes.","marker":"[16]"},{"why":"Provides the prior formalism for information-sharing control problems whose strict-expansion and strategy-independence technique underlies the closed-loop extension in Section V.","marker":"[14]"}],"fun_headline_variants":["Three conditions keep joint control linear","Linearity preserved under three sharing conditions","Partially nested sharing reduces to Riccati recursions","Joint comms-control: linearity under three rules","Three info-sharing rules ensure linear optimal control"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Assumption III.4: every control action that actually affects the state must show up in at least one other agent's observation at the next time step; the paper demonstrates that without this assumption a partially nested problem can have only nonlinear optimal control strategies, so the linear closed-form solution collapses.","fun_headline_variants_meta":{"raw":{"variants":["Three conditions keep joint control linear","Linearity preserved under three sharing conditions","Partially nested sharing reduces to Riccati recursions","Joint comms-control: linearity under three rules","Three info-sharing rules ensure linear optimal control"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000248,"raw_usage":{"total_tokens":1577,"prompt_tokens":1006,"completion_tokens":571,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":622,"completion_tokens_details":{"reasoning_tokens":503}},"tokens_in":622,"tokens_out":571,"duration_ms":6717,"temperature":1.0,"reasoning_tokens":503,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:45:30.610047+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Build the two-agent, two-step scalar system in Lemma III.5: agent 1 acts first, agent 2 observes only a noisy version of the state after agent 1's action, and agent 1's action influences the state but agent 2's observation matrix is zero for that influence. If communication of agent 1's private observation is allowed at zero communication cost, then a team-optimal strategy must be nonlinear; verifying numerically that the optimal cost over linear strategies is strictly worse than the optimal nonlinear strategy settles whether Assumption III.4 is needed. Equivalently, one could compute the Riccati value under Assumption III.4 on random matrices and check that it coincides with the value of an explicit nonlinear search in small horizons.","supporting_citations":[{"cited_title":"Team decision theory and information structures in optimal control problems – part I,","cited_arxiv_id":null,"evidence_quote":"Establishes that partially nested information structures admit linear team-optimal strategies; the paper extends this to semidefinite costs and uses it as the anchor for Theorem III.6."},{"cited_title":"Decentralized stochastic control with partial history sharing: A common information approach,","cited_arxiv_id":null,"evidence_quote":"Supplies the common-information-based framework and the notion of prescriptions used throughout the paper."},{"cited_title":"Common information based Markov perfect equilibria for linear-gaussian games with asymmetric information,","cited_arxiv_id":null,"evidence_quote":"Supplies the strategy-independent common-information-based belief condition and the Gaussian belief update that Theorem IV.2 and Lemma IV.3 adapt."},{"cited_title":"Sufficient statistics for linear control strategies in decentralized systems with partial history sharing,","cited_arxiv_id":null,"evidence_quote":"Shows that linear control strategies with a fixed private-information component reduce to centralized LQG control solved by Riccati equations; the paper removes the fixed-component restriction under partial nestedness."},{"cited_title":"Optimal decentralized state-feedback control with sparsity and delays,","cited_arxiv_id":null,"evidence_quote":"Motivates the dynamic-programming approach and gives the factorized-state case that the paper's method generalizes."},{"cited_title":"Principled learning-to-communicate with quasi-classical in- formation structures,","cited_arxiv_id":null,"evidence_quote":"Provides the prior formalism for information-sharing control problems whose strict-expansion and strategy-independence technique underlies the closed-loop extension in Section V."}],"review_version":1}