Pith. sign in

REVIEW 2 major objections 5 minor 21 references

Multi-Agent, Multi-Scale Systems with the Koopman Operator

T0 review · 2 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read This paper shows that in a Koopman-lifted multi-agent system, the social optimum and the Nash equilibrium differ by exactly one cross-agent coupling term, making the source of the gap explicit.

desk verdict The Koopman multi-agent lifting is worth having, but the paper calls a decoupled best-response equilibrium 'Nash' and misses the cross-agent adjoint terms, so the central comparison is mislabeled. read the letter →

arxiv 2506.15589 v1 pith:ZYYZKXWM submitted 2025-06-18 math.DS math.OC

classification math.DSmath.OC MSC 37N3591A1093C1037M99
keywords multi-agentsystemsKoopmanoperatoroptimalcontrolgametheoryNashequilibriumtimescaleseparationhierarchicalmixedcomplementarityproblem
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper develops a Koopman-operator formulation of multi-agent nonlinear dynamics in which each agent's state is lifted through its own observables and every other agent's state enters as an input. The structured lift yields a finite-dimensional linear model with explicit cross-agent coupling matrices $K^{xx}_{ij}$. The author then writes the optimality conditions for centralized control (the social optimum) and for general-sum game play (the Nash equilibrium) on the same lifted model. The two conditions are identical except for the term $\sum_{j\ne i}\lambda_{j,t}^T(I-K^{xx}_{jj})K^{xx}_{ji}$, so the gap between the two solution concepts is pinned to the cross-agent couplings. A two-agent numerical demonstration shows the gap is small but nonzero and that agent effort and payoff split asymmetrically.

What carries the argument

The central object is the structured multi-agent Koopman lift of Eq. (3): each agent has its own observable set $\psi_i^x$ and the effect of other agents is folded in as an input through matrices $K^{xx}_{ij}$, with control entering through $K^{xu}_i$. The companion piece is the mixed complementarity problem (MCP) formed from the lifted dynamics and the agents' quadratic costs, whose stationarity conditions are compared with those of the centralized quadratic program to expose the coupling term. The machinery works by making cross-agent terms explicit enough to be analyzed with linear-systems tools such as transient growth bounds, controllability gramians, and SVD-based perturbation maxima.

What would settle it

Build a two-agent system whose coupling is strong and nonlinear enough that no finite Koopman dictionary closes exactly, lift it with a finite set of observables, and compute the difference between the centralized optimum and the Nash equilibrium predicted by Eq. (7) versus Eq. (6). Then solve the true social optimum and Nash equilibrium of the original nonlinear game numerically; if the predicted gap is not close to the true gap, the claim that the coupling term carries the divergence is falsified.

Watch

Extended reading notes

Core claim

The central discovery is a multi-agent generalization of the Koopman lift in which the dynamics of agent $i$ are written as $\psi_i^x(x_{i,t+1}) = K^{xx}_{ii}(\psi_i^x(x_{i,t}) - \sum_{j\ne i}K^{xx}_{ij}\psi_j^x(x_{j,t}) - K^{xu}_i\psi_i^u(u_{i,t})) + \sum_{j\ne i}K^{xx}_{ij}\psi_j^x(x_{j,t}) + K^{xu}_i\psi_i^u(u_{i,t})$. This structure keeps each agent's own dynamics on the diagonal block $K^{xx}_{ii}$ and treats every other agent's lifted state as an additive input. Comparing the Nash first-order condition (Eq. 6) with the centralized optimum condition (Eq. 7), the only difference is the term $\sum_{j\ne i}\lambda_{j,t}^T(I-K^{xx}_{jj})K^{xx}_{ji}$. The paper establishes that, in this lifted linear model, all divergence between decentralized strategic play and a centrally computed optimum is carried by the cross-agent coupling matrices $K^{xx}_{ji}$, and it extends the same comparison to the $\epsilon\to 0$ limit of hierarchical, time-scale-separated multi-agent systems.

Load-bearing premise

The approach assumes that the chosen finite set of lifted coordinates (observables) exactly represents the true nonlinear dynamics over the operating region; when that representation is only approximate, the computed optimum and Nash solutions are solutions of the approximate model, not of the original multi-agent game.

Editorial extensions

If this is right

  • If the finite-dimensional Koopman model is accurate, the same lifted matrices $K^{xx}_{ij}$ serve both centralized and game-theoretic control, making the two solution concepts computationally comparable at the same fidelity.
  • When $(I-K^{xx}_{jj})K^{xx}_{ji}=0$ for all $j\ne i$, the optimum and equilibrium optimality conditions coincide, so agents that do not affect one another in the lifted dynamics will reach the social optimum by playing their own game.
  • In the $\epsilon\to 0$ limit of hierarchical, time-scale-separated systems, the effective slow-scale matrices $B^{xx}_i$ and $B^{xu}_i$ retain the same coupling structure, so the optimum-versus-Nash comparison carries over with a substantially smaller optimization problem.
  • The transient growth metrics and controllability gramians computed from the structured model identify when cross-agent feedback amplifies disturbances and when coupling makes the combined system more controllable than either agent separately.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The norm of $\sum_{j\ne i}\lambda_j^T(I-K^{xx}_{jj})K^{xx}_{ji}$ could serve as a practical diagnostic: when it is small, solving the cheaper centralized optimum and using it as a warm start for the game would save the hour-or-more MCP solve the paper reports.
  • Because the gap term is weighted by the dual variables $\lambda_j$, the optimum-versus-Nash divergence is not a fixed property of the dynamics but depends on the cost structure and constraints, so the same physical coupling can be strategically aligned or misaligned by design.
  • The same lifted model should support Stackelberg (leader-follower) equilibria as mathematical programs with equilibrium constraints, which the author notes but does not develop beyond the Nash case.
  • In the numerical examples, one agent consistently spends more control effort and still does worse than the other; reading which couplings dominate the $K^{xx}_{ij}$ matrices would allow the system designer to rebalance incentives.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper develops a finite-dimensional Koopman-operator representation for multi-agent systems in which each agent's lifted state evolution is written with explicit cross-agent coupling terms (Eq. (3)), and extends this representation to systems with hierarchical control and time-scale separation (Eqs. (12)-(14)). For both settings, the authors formulate a centralized optimal control problem and a purportedly noncooperative game-theoretic problem solved through a mixed complementarity problem (MCP). The central analytical claim is that the centralized social optimum and the Nash equilibrium optimality conditions differ only by the term sum_{j != i} lambda_j^T (I - Kxx_jj) Kxx_ji (Eq. (7) vs. Eq. (6)), which is interpreted as quantifying the coupling-driven divergence between the two solution concepts. A two-agent numerical example with nonlinear oscillators is used to illustrate the difference.

Significance. If the identification of Eq. (6) with the Nash equilibrium were correct, the paper would provide a clean way to use Koopman lifting in multi-agent dynamic games and to quantify the gap between cooperative and noncooperative outcomes. The structured lift with per-agent observables and explicit cross-agent blocks is natural and potentially useful, and the hierarchical/multi-scale extension is a sensible generalization of the author's prior work. The algebraic derivation of the centralized optimality condition (Eq. (7)) is straightforward, and the paper gives a reproducible computational pipeline (Pyomo/IPOPT/PATH) while honestly reporting model errors. However, the central game-theoretic interpretation is, in my assessment, not correct: Eq. (6) corresponds to a coupling-blind, decoupled best-response problem, not to the Nash equilibrium of the coupled dynamic game. The comparison between Eq. (7) and Eq. (6) is therefore a comparison between centralized optimization and a decentralized, coupling-ignoring equilibrium, not between a social optimum and a Nash equilibrium.

major comments (2)
  1. [Sec. II-A, Eqs. (4), (6), (7)] Equation (6), called the 'equilibrium optimality condition', is not the Nash equilibrium condition for the coupled dynamic game defined by (1)-(2) or by its exact Koopman lift (4). In an open-loop Nash equilibrium, each agent i optimizes its own objective with respect to its control sequence subject to the entire coupled dynamics, taking the other agents' control sequences as given. The KKT conditions for agent i's problem must therefore include adjoint equations for all stacked states, and the adjoint equation for psi_x_i,t contains terms of the form sum_{j != i} lambda_{i,j,t}^T (I - Kxx_jj) Kxx_ji, where lambda_{i,j,t} is agent i's multiplier on agent j's dynamics row. Equation (6) contains no such terms; it is the stationarity condition for a single-agent optimal control problem in which the other agents' lifted states are treated as exogenous inputs and only the agent's own row of (3) is a dynamic constraint. Consequently, the MCP solved in Section IV characterizes a fixed point of coupling-blind, decoupled best responses, not a Nash equilibrium of the coupled system. The authors should either re-derive the equilibrium conditions from the stacked dynamics (4) with a separate multiplier vector per agent, or explicitly reframe the contribution as a comparison between centralized optimization and a decoupled decentralized solution and avoid the term 'Nash equilibrium'. This is load-bearing because the paper's central claim that the optimum and equilibrium conditions differ by a single cross-coupling term depends on Eq. (6) being the Nash condition.
  2. [Sec. IV, Tables I-II] The numerical demonstration does not support the game-theoretic conclusions as stated. Because the equilibrium object solved in the MCP is the decoupled fixed point described in the previous comment, the reported optimum/equilibrium differences (average RMS less than 0.005 in Sec. IV-A, roughly 0.02 in Sec. IV-B) measure a difference between centralized and decoupled solutions, not the social-optimum-versus-Nash gap. In addition, the epsilon->0 model used for control in the hierarchical case has mean RMS prediction errors of 0.152 (Agent 1) and 0.175 (Agent 2), which are not negligible relative to the reported policy differences, and no error bars or sensitivity analysis are provided for any learned-model result. The paper's own caveat that the stability analysis 'should be taken with a larger grain of salt' applies equally to the control results. The quantitative claims about agent effort/benefit asymmetries should be re-examined after correcting the equilibrium definition, or softened.
minor comments (5)
  1. [Sec. III-A, text after Eq. (38)] The stated value 'epsilon = 100' is inconsistent with the epsilon -> 0 limit used throughout Sec. II-B and with the identification of y and w as fast-scale variables; if epsilon is meant to be small, the displayed value appears to be a typo (e.g., epsilon = 0.01), and if epsilon = 100 is intended, the time-scale labels 'slow' and 'fast' in the numerical experiment are reversed.
  2. [Sec. II-A, Eq. (5)] The indexing in the quadratic cost approximation is difficult to parse: the notation [psi^i_{k,t}]^T Q_{k,ij} psi^j_{k,t} with i,j in {u,x} should be defined more explicitly, including the dimensions of the Q_{k,ij} blocks and the ordering of the indices, to avoid confusion with the cross-coupling matrices Kxx_{ij}.
  3. [Sec. II-B, text near Eq. (8)] There is a doubled word in 'and and x_{-i}' near Eq. (8), and a similar typo appears in Sec. III-A ('the the k-th variable'); a careful proofread of the manuscript is needed.
  4. [Sec. II-C, Eq. (26)] The Lyapunov equation is written with bare symbols Kii, Kij, and I; these should be decorated (e.g., Kxx_ii, Kxx_ij) to match the notation elsewhere in the paper and to avoid confusion between the identity matrix I and the index i.
  5. [Sec. II-C, Eqs. (27)-(28)] The notation Pmax,i (elsewhere P_{max,i}) and the phrase 'evaluating the fixed k solution over a sufficiently large range of k values' are vague; please state precisely how the maximum over k is computed in the SVD-based procedure.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the optimum-vs-equilibrium comparison is self-contained algebra from the stated Koopman model.

full rationale

The paper's central derivation is Eq. (6) vs Eq. (7): both are stationarity conditions taken from the same multi-agent Koopman model (3)-(4). The difference term sum_{j != i} lambda_j^T (I - Kxx_jj) Kxx_ji is obtained by writing the adjoint of the cross-coupling in the stacked dynamics; it is a direct algebraic consequence of the model, not a fitted parameter or a re-statement of the input. The Koopman lift (3) is an ansatz/approximation, but that is a modeling assumption, not circularity. Self-citations [12], [14], [15], and [18] supply prior component formulations (zero-sum MCP, hierarchical time-scale separation, stability-enforcing KO, domain knowledge) but are not used to conjure the target result; the optimality conditions are derived in this paper. The numerical transient-growth and controllability metrics are computed from the learned matrices; the stability-enforcement used in training does not force the reported log||A||_2 or T_bound values, which can vary independently under stable eigenvalues. The reviewer's concern that Eq. (6) describes a decoupled fixed point rather than the true open-loop Nash condition of the coupled game is a technical correctness issue, not a circularity: even if valid, it would mean the labeled solution concept is mis-specified, not that the claimed derivation reduces to its input by construction.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

No new physical or mathematical entities are postulated. The coupling matrices Kxx_ij are coefficients of the learned linear model, not free-standing inventions. The central claim rests on the modeling fidelity of the Koopman lift and on standard singular perturbation and optimization assumptions.

free parameters (4)
  • Koopman matrix entries (Kxx_ij, Kxu_i, Kyy_i, Kww_i, etc.) = Not reported; learned from 1e4 simulated trajectories
    The Koopman matrices are fit via neural-network observables with a prediction-error plus stability-enforcement loss (Section III-B). All subsequent stability, controllability, and control-policy results depend on these fitted values.
  • Neural network weights for observables ψx, ψy, ψw = Not reported
    12 nonlinear observables for ψx and ψy, 4 for ψw; the network architectures, optimizer, and training hyperparameters are unspecified, making exact reproduction impossible.
  • Number of nonlinear observables (12 for ψx, 12 for ψy, 4 for ψw) = 12, 12, 4
    These counts are chosen by hand and define the dimension of the lifted linear system; they are not derived from any optimality criterion.
  • Stability enforcement penalty weight = Not specified
    The training loss penalizes eigenvalues with magnitude > 1 (Section III-B); the weight of this penalty is not reported, yet it shapes the matrices used in the stability analyses.
assumptions (5)
  • domain assumption A finite-dimensional Koopman-invariant subspace exists for each agent's dynamics spanned by the chosen state-inclusive and neural-network observables.
    Eq. (3) asserts an exact linear relationship ψx_i(x_{i,t+1}) = ...; for generic nonlinear systems this holds only approximately, so the derivation describes the lifted model, not the original system.
  • domain assumption The stability-assuring Koopman formulation of King et al. [15] remains valid when modified with cross-agent coupling terms.
    The structure in (3) is an adaptation of [15]; the paper does not prove the modification preserves the representation properties.
  • standard math The fast dynamics in (9) have a unique, asymptotically stable quasi-steady state as ϵ→0, and the reduced slow dynamics (21)-(23) are valid.
    Standard singular perturbation assumption inherited from [14]; used to collapse the hierarchical control problem to the slow time scale.
  • domain assumption The KKT/MCP optimality conditions of the lifted Koopman problem correctly characterize the optimum and Nash equilibrium of the original nonlinear game.
    The paper computes policies on the approximate lifted model and interprets them as game solutions; the approximation error is not propagated into the optimality conditions.
  • standard math Constraint qualifications hold for the lifted optimization and MCP formulations.
    First-order conditions (6)-(7) are used without verifying constraint qualifications; this is standard practice but an unstated assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Multi-Agent, Multi-Scale Systems with the Koopman Operator." pith.science (2026). https://pith.science/paper/ZYYZKXWM

@misc{pith2026250615589,
  author       = {Pith},
  title        = {Pith review of: Multi-Agent, Multi-Scale Systems with the Koopman Operator},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ZYYZKXWM}},
  note         = {Machine review of arXiv:2506.15589}
}
read the original abstract

The Koopman Operator (KO) takes nonlinear state dynamics and ``lifts'' those dynamics to an infinite-dimensional functional space of observables in which those dynamics are linear. Computational applications typically use a finite-dimensional approximation to the KO. The KO can also be applied to controlled dynamical systems, and the linearity of the KO then facilitates analysis and control calculations. In principle, the potential benefits provided by the KO, and the way that it facilitates the use of game theory via its linearity, would suggest it as a powerful approach for dealing with multi-agent control problems. In practice, though, there has not been much work in this space: most multi-agent KO work has treated those agents as different components of a single system rather than as distinct decision-making entities. This paper develops a KO formulation for multi-agent systems that structures the interactions between decision-making agents and extends this formulation to systems in which the agents have hierarchical control structures and time scale separated dynamics. We solve the multi-agent control problem in both cases as both a centralized optimization and as a general-sum game theory problem. The comparison of the two sets of optimality conditions defining the control solutions illustrates how coupling between agents can create differences between the social optimum and the Nash equilibrium.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

21 extracted references · 5 canonical work pages

  1. [1]

    Applied koopmanism,

    M. Budi ˇsi´c, R. Mohr, and I. Mezi ´c, “Applied koopmanism,” Chaos: An Interdisciplinary Journal of Nonlinear Science, vol. 22, no. 4, p. 047510, 2012

  2. [2]

    The koopman operator: Capabilities and recent advances,

    C. Bakker, A. Bhattacharya, S. Chatterjee, C. J. Perkins, and M. R. Oster, “The koopman operator: Capabilities and recent advances,” 2020 Resilience Week (RWS), pp. 34–40, 2020

  3. [3]

    Linear predictors for nonlinear dynamical sys- tems: Koopman operator meets model predictive control,

    M. Korda and I. Mezi ´c, “Linear predictors for nonlinear dynamical sys- tems: Koopman operator meets model predictive control,” Automatica, vol. 93, pp. 149–160, 2018

  4. [4]

    Linear observer synthesis for nonlin- ear systems using koopman operator framework,

    A. Surana and A. Banaszuk, “Linear observer synthesis for nonlin- ear systems using koopman operator framework,” IFAC-PapersOnLine, vol. 49, no. 18, pp. 716–723, 2016

  5. [5]

    Learning compo- sitional koopman operators for model-based control,

    Y . Li, H. He, J. Wu, D. Katabi, and A. Torralba, “Learning compo- sitional koopman operators for model-based control,” arXiv preprint arXiv:1910.08264, 2019

  6. [6]

    Koopman system approximation-based optimal control of multiple mobile robots,

    Q. Zhao and G. Tao, “Koopman system approximation-based optimal control of multiple mobile robots,” IEEE Transactions on Control Systems Technology, 2025

  7. [7]

    Special section on multidisciplinary design optimization: multidisciplinary design optimization of dynamic engineering systems,

    J. T. Allison and D. R. Herber, “Special section on multidisciplinary design optimization: multidisciplinary design optimization of dynamic engineering systems,” AIAA journal, vol. 52, no. 4, pp. 691–710, 2014

  8. [8]

    Koopman operator methods for global phase space exploration of equivariant dynamical systems,

    S. Sinha, S. P. Nandanoori, and E. Yeung, “Koopman operator methods for global phase space exploration of equivariant dynamical systems,” IFAC-PapersOnLine, vol. 53, no. 2, pp. 1150–1155, 2020

Show all 21 references
  1. [9]

    Integrating machine learning and multiscale mod- eling—perspectives, challenges, and opportunities in the biological, biomedical, and behavioral sciences,

    M. Alber, A. Buganza Tepole, W. R. Cannon, S. De, S. Dura- Bernal, K. Garikipati, G. Karniadakis, W. W. Lytton, P. Perdikaris, L. Petzold, et al. , “Integrating machine learning and multiscale mod- eling—perspectives, challenges, and opportunities in the biological, biomedical...

  2. [10]

    Stackelberg game-theoretic trajectory guid- ance for multi-robot systems with koopman operator,

    Y . Zhao and Q. Zhu, “Stackelberg game-theoretic trajectory guid- ance for multi-robot systems with koopman operator,” arXiv preprint arXiv:2309.16098, 2023

  3. [11]

    Multi-level optimization with the koopman operator for data-driven, domain-aware, and dynamic system security,

    M. R. Oster, E. King, C. Bakker, A. Bhattacharya, S. Chatterjee, and F. Pan, “Multi-level optimization with the koopman operator for data-driven, domain-aware, and dynamic system security,” Reliability Engineering & System Safety , vol. 237, p. 109323, 2023

  4. [12]

    Operator-theoretic methods for differential games,

    C. Bakker, A. Rupe, A. V on Moll, and A. R. Gerlach, “Operator-theoretic methods for differential games,” Journal of Computational Physics , Under Review

  5. [13]

    The path solver: a nommonotone sta- bilization scheme for mixed complementarity problems,

    S. P. Dirkse and M. C. Ferris, “The path solver: a nommonotone sta- bilization scheme for mixed complementarity problems,” Optimization methods and software , vol. 5, no. 2, pp. 123–156, 1995

  6. [14]

    Time scale separation and hierarchical control with the koopman operator,

    C. Bakker, “Time scale separation and hierarchical control with the koopman operator,” 9th IEEE Conference on Control Technology and Applications (CCTA) 2025 , 2025. Under Review

  7. [15]

    Solving the dynamics-aware economic dispatch problem with the koopman operator,

    E. King, C. Bakker, A. Bhattacharya, S. Chatterjee, F. Pan, M. R. Oster, and C. J. Perkins, “Solving the dynamics-aware economic dispatch problem with the koopman operator,” in Proceedings of the twelfth ACM international conference on future energy systems , pp. 137–147, 2021

  8. [16]

    A tutorial review of complementarity models for decision-making in energy markets,

    C. Ruiz, A. J. Conejo, J. D. Fuller, S. A. Gabriel, and B. F. Hobbs, “A tutorial review of complementarity models for decision-making in energy markets,” EURO Journal on Decision Processes , vol. 2, pp. 91– 120, 2014

  9. [17]

    L. N. Trefethen and M. Embree, Spectra and Pseudospectra: The Behavior of Nonnormal Matrices and Operators . Princeton University Press, Princeton, NJ, 2005

  10. [18]

    Deception-based cyber attacks on hierarchical control systems using domain-aware koopman learning,

    C. Bakker, A. August, S. Huang, S. Vasisht, and D. L. Vrabie, “Deception-based cyber attacks on hierarchical control systems using domain-aware koopman learning,” in 2022 Resilience Week (RWS) , pp. 1–8, IEEE, 2022

  11. [19]

    W. E. Hart, C. D. Laird, J.-P. Watson, D. L. Woodruff, G. A. Hackebeil, B. L. Nicholson, J. D. Siirola, et al. , Pyomo-optimization modeling in python, vol. 67. Springer, 2017

  12. [20]

    On the implementation of an interior- point filter line-search algorithm for large-scale nonlinear programming,

    A. W ¨achter and L. T. Biegler, “On the implementation of an interior- point filter line-search algorithm for large-scale nonlinear programming,” Mathematical programming, vol. 106, no. 1, pp. 25–57, 2006

  13. [21]

    Causal discovery in nonlinear dynamical systems using koopman operators,

    A. Rupe, D. DeSantis, C. Bakker, P. Kooloth, and J. Lu, “Causal discovery in nonlinear dynamical systems using koopman operators,” arXiv preprint arXiv:2410.10103 , 2024

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.