{"id":"052d6a59-3e5d-4c9b-bfca-79074f38d173","arxiv_id":"2506.17817","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"For parametric Koopman models, a residual-covariance-weighted closest-point reprojection onto the dictionary manifold gives more reliable predictions and bifurcation diagrams than standard EDMD.","lead":"Koopman-based machine learning models of dynamical systems often drift out of the valid lifted space, which makes their predictions fail. This paper proposes a statistically motivated way to snap each prediction back onto the valid space, and shows the method captures bifurcations in several example systems.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Proposition 4.2 does not establish the maximum-likelihood covariance: the residual affine-in-p and zero-mean assumptions are unproven, and the O(t^2) statement is too weak to identify Q.","rationale":"The reader's conditional accept is appropriate. The practical suggestion to use residual-covariance-weighted reprojection is plausible and the examples are encouraging, but the central theoretical label 'maximum likelihood' is not established: the covariance surrogate is not shown to be the true residual covariance at leading order, and the proof has unstated assumptions. This does not refute the method; it means the paper needs a corrected, sharper analysis or should soften the ML claim. Since the numerical studies provide some evidence but no code/data/error bars, the claim should remain conditional pending validation.","tokens_in":18691,"tokens_out":21116,"duration_ms":217828,"concrete_test":"Check Prop. 4.2 on a scalar parameter-affine system, e.g. dx/dt=x^2+p x with X=L=[-1,1] and dictionary Psi=(1,x,x^2)^T. Build K from (4.4) for t=0.1, then on a dense grid compute R_p(x)=Psi(F_p^t(x))-(K0+pK1)Psi(x). Compare Cov(R_p) with Q(p,p) from (4.10) by evaluating (Cov-Q)/t^2 and taking t to 0 through 0.1, 0.05, 0.025 for several p. If the limit is nonzero for any p, Q is not the leading covariance, so Corollary 4.3's weight is not the ML metric. Also check whether the empirical mean residual is O(t^2) and whether a Q-Q plot of the residual distribution is consistent with Gaussianity; if not, the zero-mean Gaussian model fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Corollary 4.3 is the load-bearing claim. It requires the one-step residual R_p(x)=Psi(F_p^t(x))-(K0+sum p_i K_i)Psi(x) to behave as zero-mean Gaussian noise with covariance Sigma=sum_{i,j} Q_{ij} p_i p_j. The paper's support for this is Proposition 4.2, but that proposition is weaker than it appears: both Cov(R_p) and Q(p,p) are O(t^2), so Cov=Q+O(t^2) is satisfied by any bounded quadratic surrogate and does not imply Q is the leading-order covariance. The proof also substitutes the residual of the jointly learned parametric model by sum bar p_i R_{g_i}^t Psi 'using the parameter-affine structure'; the residuals R_{g_i}^t for the augmented blocks K_i are never defined, and the substitution is not justified because K_i is learned together with K0, not from the individual flow g_i. The zero-mean assertion additionally requires the constant observable to be in V, which is not assumed in Section 2. Thus W=Sigma^{-1} is not demonstrated to be the maximum-likelihood metric; a coordinate projection or any other metric could perform as well, and the numerical examples do not separate these hypotheses.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a geometric reprojection framework for Koopman-based predictions of parametric dynamical systems. The authors consider dictionary-based liftings Ψ:X→R^M, whose image M is invariant under the true Koopman operator but not under a finite-dimensional EDMD approximation. They study weighted closest-point reprojections π_W onto M and argue that the weight matrix W should be the inverse of the covariance of the one-step residual, which gives a maximum-likelihood estimate of the lifted state. For parameter-affine systems, they introduce a quadratic covariance surrogate Q defined by (4.10) and state in Corollary 4.3 that the maximum-likelihood reprojection uses Σ = Σ_{i,j=0}^m Q_{ij} p_i p_j. The paper also proposes an adaptive multistep scheme that propagates the covariance estimate (Section 5.1) and a Riemannian Newton method for computing the reprojection (Section 5.2). Numerical experiments on a pitchfork bifurcation, a Duffing oscillator, and the Lorenz system illustrate the benefits of reprojections over unprojected EDMD predictions.","tokens_in":18944,"tokens_out":7734,"duration_ms":81529,"significance":"If the central claim were established, the paper would give a principled, geometric way to keep Koopman surrogate predictions on the invariant dictionary manifold, with a natural maximum-likelihood interpretation and a clean extension to parameter-affine systems that is relevant for data-driven bifurcation analysis and control. The paper is clearly written and contains several useful ingredients: the parameter-affine regression problem (4.4), the covariance surrogate construction (4.10), the multistep covariance propagation (5.1), the retraction in Lemma 5.5, and a broad numerical study spanning pitchfork, Duffing, and Lorenz dynamics. However, the maximum-likelihood interpretation rests on a number of unproven assumptions: the residual is not shown to be zero-mean unless the constant observable is in the dictionary, the residuals for the individual vector fields g_i are not defined, and the O(t^2) statement in Proposition 4.2 does not identify Q as the leading-order covariance. These gaps are load-bearing for Corollary 4.3, so the paper's main theoretical contribution is not yet established.","major_comments":[{"comment":"The assertion E(R^t_f Ψ) = 0 is not generally true under the assumptions of Section 2. The optimality of \\hat K in (2.3) gives orthogonality of the residual to every dictionary function, but the zero-mean conclusion requires the constant observable to belong to V. This is not assumed in Section 2, and it can fail for dictionaries that omit the constant. In addition, the expectation is written as an unnormalized integral over X, which is not a probability expectation unless the measure is normalized; this matters if the constant is included but the integral is not divided by |X|. Since the zero-mean property is essential for the Gaussian model in Proposition 3.2 and Corollary 4.3, this gap must be fixed, either by adding the constant to the standing assumptions on V or by reworking the residual model.","section":"Section 4.2, equations before (4.8) and Proposition 4.2"},{"comment":"The proof substitutes the covariance of the joint parametric residual by a sum involving residuals R^t_{g_i}Ψ, but these residuals are never defined. The matrices \\hat K_i are learned jointly through (4.4), not as Koopman compressions for the individual flows g_i, and the paper does not introduce any \\hat K_{g_i} with respect to which R^t_{g_i}Ψ could be formed. Without a precise definition of R^t_{g_i} and a supporting argument that the parametric residual decomposes in this way, the key covariance decomposition underlying (4.12) is unsupported. This is load-bearing for the construction of Q in (4.10) and for Corollary 4.3.","section":"Proposition 4.2, equation (4.12) and its proof"},{"comment":"The statement Cov(R) = Q(\\bar p,\\bar p) + O(t^2) is too weak to identify Q as the covariance surrogate needed for the maximum-likelihood projection. Both Cov(R) and Q(\\bar p,\\bar p) are O(t^2) for small t, so the O(t^2) remainder is of the same order as the terms being compared; any bounded quadratic surrogate would satisfy the same asymptotic statement. To support Corollary 4.3, the paper needs either an o(t^2) remainder or an independent argument that Q captures the leading-order covariance. As written, the proof does not establish that W = Σ^{-1} is the maximum-likelihood metric.","section":"Proposition 4.2, equation (4.11)"},{"comment":"The algorithm presented as a Riemannian Newton method does not use the Riemannian Hessian of the objective l_z on the embedded manifold. The ambient Hessian of l_z is W, but the Riemannian Hessian on the submanifold includes a curvature term involving the second fundamental form (equivalently, a second-derivative term (D^2Ψ)^T W(Ψ(x)-z) in coordinates). Equation (5.4) is the Gauss-Newton normal equation for the nonlinear least-squares problem min_x 1/2 ||Ψ(x)-z||^2_W, not the Newton equation, and the quoted quadratic convergence from [1, Theorem 6.3.2] does not apply. The algorithm may still be a reasonable implementation, but it should be relabeled and analyzed as a Gauss-Newton method, or the curvature term must be included if the authors wish to claim a Riemannian Newton method.","section":"Section 5.2, equations (5.3)-(5.5) and Algorithm 1"},{"comment":"The numerical study does not provide train/test separation, error bars, or variability across data realizations, and it does not isolate the effect of the maximum-likelihood weight W = Σ^{-1}. In Figures 5 and 6 the comparison is between geometric reprojection and a specific coordinate reprojection; the improvement could be due to the choice of weight, to the manifold projection itself, or to the particular coordinates selected. To support the claim that the maximum-likelihood metric is the right choice, the authors should add control experiments with, e.g., the identity weight or another fixed positive-definite weight, and should report statistics over multiple training data sets.","section":"Section 4.3, Figures 4-6"}],"minor_comments":[{"comment":"The notation 'YX:' in (4.10) is not defined; if this denotes a pseudoinverse Y X^+, it should be stated explicitly. Moreover, the matrix X = ∫_L (\\bar p\\bar p^T ⊗ \\bar p\\bar p^T) dp is singular for symmetric parameter intervals such as L = [-2,2], so the well-posedness and uniqueness of the solution Q to (4.9) require discussion.","section":"Equation (4.10)"},{"comment":"The expectation E(·) is used as an integral over X without a normalization factor. Since the manuscript later interprets E as a statistical expectation in the Gaussian model, the authors should either normalize the Lebesgue measure on X or explicitly state that E is an unnormalized average that is later reinterpreted probabilistically.","section":"Section 2, around equation (2.3)"},{"comment":"The recursion for η_{k+1} assumes that the model noise μ_k and the propagated error term are independent and Gaussian; this is a modeling assumption that is not stated or justified. Please state the independence assumption explicitly.","section":"Section 5.1, equation (5.1)"},{"comment":"There are several typographical and readability issues: 'to to conduct' in Section 5.1, 'by provided by' in Remark 3.3, and the notation 'V 3ztx1u' in Section 4.3.3, which should be typeset as V_3 \\setminus {x_1}. A careful proofreading pass is recommended.","section":"Throughout"},{"comment":"The retraction R in Lemma 5.5 requires Ψ^{-1} and DΨ(Ψ^{-1}(p))^{-1} to exist globally; Assumption 5.2 guarantees that Ψ is an embedding, but for a general dictionary this is a strong condition. The authors should state clearly that the efficient algorithm in Section 5.2 is only applicable under Assumption 5.2.","section":"Section 5.2, Definition 5.4 and Lemma 5.5"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses an important practical problem in Koopman-based prediction, and the proposed geometric framework is attractive. However, the central maximum-likelihood claim is not yet proven: the residual mean argument requires the constant observable, the individual residual operators are undefined, the O(t^2) statement does not identify the covariance surrogate, and the Riemannian Newton label is not supported by the equations. These issues are substantial but appear fixable within the scope of a revision. I also note that the numerical study is broad but lacks the control experiments needed to isolate the proposed metric, and no code or data are provided for reproducibility. I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the parametric maximum-likelihood reprojection is a real contribution, but it is packaged with a proof that does not quite deliver what the abstract promises. I would send it to review, with the expectation of a major revision.\n\nWhat is new: for parameter-affine systems, the paper proposes the covariance surrogate Q(p,p) and the weight W = Σ^{-1}(p) for closest-point reprojection onto the dictionary manifold. The adaptive multistep triggering in Section 5.1 is a useful practical addition, and the numerical study is genuinely informative: reprojection clearly stabilizes EDMD on the pitchfork, Duffing, and Lorenz examples. The geometric projection outscores the coordinate projection in the case where the dictionary omits a coordinate function, which is the kind of evidence that matters.\n\nSoft spots, roughly in order of importance:\n\n1. Proposition 4.2 does not do the work the paper needs. Both Cov(R) and Q(p,p) are O(t^2), so the statement Cov = Q + O(t^2) is very weak; it does not identify Q as the leading-order covariance. The proof also replaces the joint parametric residual by a sum of individual residuals R^t_{g_i}, but the blocks K_i are fitted jointly in (4.4), not as individual EDMD approximations for g_i. That substitution is not justified on the page.\n\n2. The zero-mean claim for residuals requires the constant observable to be in the dictionary, and that is not assumed in Section 2.\n\n3. Corollary 4.3 depends on postulating zero-mean Gaussian noise. It is fine to propose that as a working model, but it is a postulate; no finite-data or finite-dictionary argument establishes it, and the experiments do not separate the benefit of the Σ^{-1} metric from the effect of using any reasonable reprojection.\n\n4. The method in Section 5.2 is Gauss-Newton, not Riemannian Newton: the curvature term of the embedded manifold is missing from the Hessian. The quadratic convergence claim is therefore not supported by the cited theorem.\n\n5. No code, no data, no train/test separation, and no error bars. Also, in most of the examples the geometric projection and the simple coordinate projection perform almost identically, so the claimed advantage of the maximum-likelihood choice is empirically thin outside Figure 6.\n\nIs the central idea wrong? I do not think so. Reprojections are clearly useful, and the parametric covariance surrogate is worth building on. But the maximum-likelihood interpretation is a modeling assumption, not a demonstrated result. With sharper asymptotics, a clarified definition of what the jointly fitted K_i actually estimate, an explicit constant-observable assumption, and a recalibrated Newton claim, this could be a solid paper. A serious referee should engage with it.","headline":"Genuinely new parametric MLE-reprojection idea for Koopman surrogates, but the flagship covariance result is more heuristic than theorem and the numerical advantage is narrower than the narrative.","tokens_in":19516,"tokens_out":5639,"would_cite":true,"duration_ms":61934,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["37M20","65P30"],"pacs":[],"model":"deepseek-v4-flash","headline":"A maximum-likelihood closest-point reprojection keeps Koopman predictions on the invariant manifold, making parameter-affine surrogates reliable for bifurcation analysis.","keywords":["Koopman operator","Extended Dynamic Mode Decomposition","reprojection","maximum likelihood","parameter-affine systems","bifurcation analysis","Riemannian Newton method"],"falsifier":"Compare, on the pitchfork or Lorenz examples, long-horizon predictions from three reprojections: the proposed inverse-covariance weight, the identity weight, and the coordinate projection, and also measure the empirical mean and distribution of the residuals over many initial states and parameters. If the empirical residual mean is clearly nonzero, or if either simpler projection matches or beats the inverse-covariance weight on prediction error, the maximum-likelihood justification is falsified.","tokens_in":18436,"feed_emoji":"🦋","tokens_out":9153,"duration_ms":86588,"temperature":0.7,"pith_summary":"Koopman-based surrogates such as Extended Dynamic Mode Decomposition learn a linear model in a lifted space, but one-step predictions drift off the nonlinear manifold that corresponds to real states, and once they leave that manifold the model no longer has a physical meaning. This paper claims that the fix is to reproject onto the manifold after each prediction step using a weighted closest-point projection, where the weight estimates the inverse covariance of the prediction error, and that this yields a maximum-likelihood estimate of the true lifted state. For parameter-affine systems, it shows that the error covariance is, to second order in the time step, a quadratic function of the parameters, so the optimal weight can be computed from the same data used to fit the Koopman model. The practical payoff is that reprojected surrogates track attractors and recover bifurcation diagrams in the pitchfork, Duffing, and Lorenz examples, while unprojected predictions fail.","feed_headline":"Weighted reprojection keeps Koopman predictions reliable","feed_subtitle":"A maximum-likelihood projection weight keeps lifted states on the manifold and recovers bifurcations.","key_machinery":"The load-bearing object is the weighted closest-point projection $\\pi_W(z)=\\arg\\min_{q\\in\\mathcal{M}} |q-z|_W^2$ onto the immersed state manifold $\\mathcal{M}=\\Psi(X)$, with $W$ a positive (semi)definite weight matrix. The proposed choice is $W=\\Sigma^{-1}$, where $\\Sigma$ is built from a matrix-valued bilinear form $Q$ that solves the regression (4.9)-(4.10) over the parameter domain; for a fixed parameter the working covariance is $\\Sigma=\\sum_{i,j=0}^m Q_{ij}p_i p_j$. The maximum-likelihood interpretation treats the residual $\\Psi(\\mathrm{Fl}_f^t(x))-(K_0+\\sum_{i=1}^m p_iK_i)\\Psi(x)$ as zero-mean Gaussian noise with that covariance, and the efficient evaluation uses a Riemannian Newton method with the retraction $R_p(\\eta)=\\Psi(\\Psi^{-1}(p)+D\\Psi(\\Psi^{-1}(p))^{-1}\\eta)$.","core_discovery":"The central claim is Corollary 4.3: if $z_+ = (K_0 + \\sum_{i=1}^m p_i K_i)z$ is the one-step lifted prediction for a parameter-affine system, then a maximum-likelihood estimate of the true lifted state is $\\hat z_+ \\in \\pi_{\\Sigma^{-1}}(z_+)$, the weighted closest point on the manifold $\\mathcal{M}=\\Psi(X)$ with weight $\\Sigma^{-1}$, where $\\Sigma = \\sum_{i,j=0}^m Q_{ij} p_i p_j$ and $Q$ is the residual-covariance surrogate defined by the regression (4.10). The argument combines Proposition 3.2, which identifies inverse-covariance weighted closest-point projection with maximum likelihood under Gaussian noise, and Proposition 4.2, which gives the covariance of the parametric residual as $Q(\\bar p,\\bar p)+O(t^2)$. The coordinate projection commonly used in Koopman practice appears as a special singular-weighted case, and the numerical study shows that the geometric reprojection stabilizes long-horizon predictions, recovers the pitchfork bifurcation, and captures the Lorenz attractor, including a case with a dictionary missing the first coordinate function where coordinate reprojection diverges.","pith_inferences":["An implication the paper leaves implicit is that $Q$ itself is a per-parameter uncertainty map: one could use its magnitude to decide where to add training data or which dictionary entries contribute most to prediction error.","The same construction should transfer to parameter-varying and control-affine systems by letting $p$ vary in time, giving a time-varying reprojection metric for Koopman-based model predictive control.","A natural stress test is to replace the Gaussian maximum-likelihood weight with a robust or empirical-covariance weight; if that performs better on the same examples, the residual model is misspecified and the method is best described as a principled heuristic rather than a maximum-likelihood estimator."],"forward_implications":["Reprojecting with the maximum-likelihood weight keeps the surrogate state on $\\mathcal{M}$, so repeated application of the linear model remains meaningful instead of diverging.","The parameter-quadratic covariance surrogate makes a single fitted model usable across a parameter range, allowing bifurcation diagrams and attractor changes to be read off from the reprojected predictions.","The covariance recursion (5.1) lets multi-step predictions skip reprojections until the propagated uncertainty crosses a threshold, cutting the computational cost of the projection step.","The geometric reprojection remains stable when the dictionary omits coordinate functions, a setting in which the standard coordinate reprojection diverges."],"supporting_citations":[{"why":"Introduced the closest-point reprojection framework for autonomous Koopman models that this paper extends to parameter-affine systems.","marker":"[41]"},{"why":"Defined EDMD as dictionary-based lifting and linear regression, producing the manifold $\\Psi(X)$ and the surrogate matrix that the reprojection acts on.","marker":"[43]"},{"why":"Introduced the lifting and coordinate reprojection idea that serves as the baseline method throughout the paper.","marker":"[23]"},{"why":"Developed the Koopman lifting technique whose consistency requirement motivates the reprojection step.","marker":"[24]"},{"why":"Supplies the Riemannian Newton method and convergence theory used to compute the reprojections efficiently.","marker":"[1]"}],"fun_headline_variants":["ML reprojection keeps Koopman states on manifold","Weighted closest-point projection recovers Koopman bifurcations","Koopman reprojection: ML consistency for parametric systems","Maximum-likelihood reprojection for reliable Koopman bifurcation analysis","Geometric reprojection stabilizes Koopman predictions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the error between the true system and the fitted Koopman model is random, centered at zero, and has the bell-shaped spread the paper estimates; if the real error is biased or not bell-shaped, the proposed projection weight is not the maximum-likelihood choice.","fun_headline_variants_meta":{"raw":{"variants":["ML reprojection keeps Koopman states on manifold","Weighted closest-point projection recovers Koopman bifurcations","Koopman reprojection: ML consistency for parametric systems","Maximum-likelihood reprojection for reliable Koopman bifurcation analysis","Geometric reprojection stabilizes Koopman predictions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000329,"raw_usage":{"total_tokens":1831,"prompt_tokens":939,"completion_tokens":892,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":555,"completion_tokens_details":{"reasoning_tokens":808}},"tokens_in":555,"tokens_out":892,"duration_ms":7767,"temperature":1.0,"reasoning_tokens":808,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:01:26.080208+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare, on the pitchfork or Lorenz examples, long-horizon predictions from three reprojections: the proposed inverse-covariance weight, the identity weight, and the coordinate projection, and also measure the empirical mean and distribution of the residuals over many initial states and parameters. If the empirical residual mean is clearly nonzero, or if either simpler projection matches or beats the inverse-covariance weight on prediction error, the maximum-likelihood justification is falsified.","supporting_citations":[{"cited_title":"Reprojection methods for Koopman-based modelling and prediction","cited_arxiv_id":null,"evidence_quote":"Introduced the closest-point reprojection framework for autonomous Koopman models that this paper extends to parameter-affine systems."},{"cited_title":"Williams, Ioannis G Kevrekidis, and Clarence W Rowley","cited_arxiv_id":null,"evidence_quote":"Defined EDMD as dictionary-based lifting and linear regression, producing the manifold $\\Psi(X)$ and the surrogate matrix that the reprojection acts on."},{"cited_title":"Linear identification of nonlinear systems: A lifting technique based on the Koopman operator","cited_arxiv_id":null,"evidence_quote":"Introduced the lifting and coordinate reprojection idea that serves as the baseline method throughout the paper."},{"cited_title":"Koopman-based lifting techniques for nonlinear systems identification.IEEE Transactions on Automatic Control, 65(6):2550–2565, 2019","cited_arxiv_id":null,"evidence_quote":"Developed the Koopman lifting technique whose consistency requirement motivates the reprojection step."},{"cited_title":"Princeton University Press, 2009","cited_arxiv_id":null,"evidence_quote":"Supplies the Riemannian Newton method and convergence theory used to compute the reprojections efficiently."}],"review_version":1}