{"id":"aa3f0a03-688a-47d1-8a9a-cccb4cf22fe5","arxiv_id":"2607.06459","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":4,"one_line_summary":"A practical input-to-state stability certificate for Koopman learning control is derived, separating prediction residuals from selected-channel margins and projection residuals.","lead":"This paper develops a stability certificate for Koopman-operator-based learning control of repetitive nonlinear systems, showing that prediction accuracy alone is insufficient for safe learning. A smart generalist might read it to understand when data-driven control can be formally certified as safe before deployment.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"Lemma 3 perturbation radii may be too conservative for practical certification, but the framework's logic is sound and this is acknowledged in scope.","rationale":"The reader correctly identified the load-bearing concern: Assumption 1's channel perturbation bound (Eq. 17) and the finite-sample transfer in Lemma 3 are the weakest links. The deterministic ISS result (Theorem 1) is a clean, correct comparison-lemma argument — the recursion Ē_k = ΛĒ_{k-1} + r_k with the stated bounds is straightforward and the KL/K function construction is standard. Propositions 1-5 (projection residual, channel margin, request closure) are also correct and well-motivated. The concern is specifically about Lemma 3's perturbation radii: the submultiplicative propagation through matrix powers (Eq. 27) can produce exponentially growing bounds when the lifted matrix Â has norm > 1, which is common for Koopman models of unstable or marginally stable systems. This could make ε_{G,H,β} * Ū_Δ dominate the budget w̄_β, yielding an impractically large certified band. However, several factors mitigate this concern: (1) the paper explicitly acknowledges conservativeness and scopes itself accordingly; (2) the framework allows external audited bounds to replace Lemma 3's radii (stated in the proof of Lemma 3); (3) the numerical experiments show the framework produces non-trivial certificates for moderate-dimensional systems; (4) code and data are shipped for verification. The paper does not overclaim — it presents a certification framework with honest limitations, not a universal stability theorem. The ACCEPT verdict with MODERATE confidence is appropriate. The correctness risk should be noted as 'low for the deterministic result, moderate for the finite-sample extension due to conservativeness of perturbation radii.' No verdict adjustment is needed.","tokens_in":15836,"tokens_out":943,"duration_ms":1017169,"concrete_test":"Recompute the finite-horizon perturbation radii ε_{G,H,β} and ε_{O,H,β} from Lemma 3 (Eqs. 27-30) for the Hard Duffing system with H=70, using the actual identified Â matrix and its one-step error ε_A. Then compare these radii against the values implicitly used in Table 4's budget decomposition (the channel/reset columns). If the Lemma 3 radii exceed the Table 4 values by more than 2x, the reported budgets are using externally audited bounds rather than the Lemma 3 construction, and the tightness of the paper's own finite-sample pipeline needs clarification. If they match within 10%, the conservativeness is as expected and the certificate is non-vacuous for these system dimensions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central ISS certificate (Theorem 1, Eq. 22) is correct given Assumption 1. The finite-sample extension (Theorem 2, Eq. 33) relies on Lemma 3's perturbation radii (Eqs. 27-30), which propagate one-step identification errors through matrix powers using submultiplicative norms. The key concern is whether these radii are tight enough to yield non-vacuous certificates. Specifically, D_A(j) in Eq. 27 accumulates terms involving ||Â||^(j-1-ℓ) * (||Â|| + ε_A)^ℓ, which grows exponentially with horizon H when ||Â|| > 1. Since Koopman lifted matrices often have spectral radii exceeding 1 (the system may be unstable in lifted coordinates even if the physical plant is stable), the stacked radii ε_{G,H,β} and ε_{O,H,β} in Eq. 30 can become extremely large for moderate H, making the certified ultimate band Δ_ISS = w̄_β/(1-λ) impractically large. The numerical experiments (Table 4) show w̄_β values around 0.4-0.6 with bands around 0.8-1.1, which are non-trivial but still relatively wide. However, the paper is transparent about this: it scopes itself to finite-horizon tasks, ships code, and the framework correctly separates the deterministic ISS result (which is clean) from the finite-sample calibration (which is conservative but honest). The reader correctly identified this as the weakest assumption. The concern is real but does not undermine the paper's claims within its declared scope — it limits practical applicability rather than correctness.","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"This paper develops an input-to-state stability (ISS) certificate for Koopman-based learning control of unknown nonlinear repetitive systems operating over finite trial horizons. The central idea is to treat the selected stacked tracking error along the learning-trial axis as the state of a discrete-time system, with Koopman residuals, projection residuals, channel uncertainty, reset mismatch, deployment shifts, and numerical tolerances acting as ISS inputs. The deterministic result (Theorem 1) provides a parameter-free ultimate band derived from independently certifiable quantities. A finite-sample extension (Theorem 2) uses split conformal prediction at the episode level to calibrate the residual budget under exchangeability assumptions. The paper also identifies projection residuals and selected-channel margins as necessary certificate objects, showing that small prediction residuals alone are insufficient for learning stability certification. Numerical experiments on Duffing-type and other nonlinear repetitive systems audit the certificate components.","tokens_in":16079,"tokens_out":1079,"duration_ms":167808,"significance":"The paper makes a clear conceptual contribution by separating model fit from learning certifiability in Koopman-based control. The deterministic ISS certificate (Theorem 1, Eq. 22) is a clean, parameter-free derivation: the budget w_bar is constructed from independently certifiable quantities (residuals, channel margins, projection distances). The projection-residual analysis (Propositions 1-2) correctly identifies a structural obstruction that prediction loss alone cannot capture. The finite-sample extension (Theorem 2) correctly applies split conformal prediction at the episode level under Assumption 2, with the controller frozen before calibration. The paper ships reproducible code and data, and the numerical experiments are structured as certificate audits rather than performance benchmarks, which is appropriate for the claims. The weak-channel rejection mechanism (Proposition 3) is a useful and falsifiable diagnostic.","major_comments":[{"comment":"Lemma 3, Eqs. (27)-(30): The perturbation radii D_A(j) accumulate terms involving ||Â||^(j-1-ℓ) * (||Â|| + ε_A)^ℓ, which grows exponentially with horizon H when ||Â|| > 1. Since Koopman lifted matrices frequently have spectral radii exceeding 1 (the physical plant may be stable while the lifted representation is not), the stacked radii ε_{G,H,β} and ε_{O,H,β} in Eq. (30) can become extremely large for moderate H. The numerical experiments (Table 4) report w_bar values around 0.4-0.6 with bands around 0.8-1.1, which are non-trivial but relatively wide. The paper should explicitly discuss the regime where these radii yield non-vacuous certificates — for instance, characterizing the relationship between H, ||Â||, and the resulting band size, or noting whether the experiments operate in a regime where ||Â|| < 1. This is load-bearing because Theorem 2's practical value depends on these radii.","section":null},{"comment":"Table 6: Several rows report channel and reset terms as exactly 0.000 (e.g., all disturbance and noise sweep rows). Given that Lemma 3's radii are generically nonzero for any finite-data identified model, these zero values require clarification. Are these set to zero because the channel perturbation is negligible relative to the reported precision, or are they excluded from the budget for another reason? If the former, the table should indicate this; if the latter, the relationship between the finite-sample budget (Eq. 33) and the reported numbers needs clarification.","section":null}],"minor_comments":[{"comment":"The paper lists no ad-hoc axioms or invented entities in its formal ledger, which is commendable. However, the relationship between Assumption 0 (Certified nonlinear repetitive task class) and Assumption 1 (Certified deployment event) could be stated more precisely: Assumption 0 defines the task class, while Assumption 1 defines the per-trial event under which the certificate holds. A brief sentence clarifying that Assumption 1 is verified through the finite-sample machinery of Theorem 2 would help the reader.","section":null},{"comment":"Eq. (33) and Remark 3 use slightly different orderings of the budget terms (Δw and Δ_{num,β} appear in different positions). Consistent ordering would improve readability.","section":null},{"comment":"Table 2 lists four systems but Tables 4-9 primarily report results for the Duffing system and its variants. A brief note on which table corresponds to which system in Table 2 would help cross-referencing.","section":null},{"comment":"The CRediT statement and AI declaration are appropriate and transparent.","section":null},{"comment":"Minor typographic issue: 'certifia-bility' in the Highlights section (line break artifact).","section":null}],"recommendation":"minor_revision","confidential_remarks":"The stress-test concern about Lemma 3's perturbation radii being potentially vacuous for moderate H is valid and is the primary practical limitation of the finite-sample result. However, the paper is transparent about its finite-horizon scope and the deterministic ISS result (Theorem 1) stands independently. The concern limits practical applicability rather than correctness, so minor revision is appropriate. The authors should address this in revision but the core contributions are sound."},"author_rebuttal":{"model":"glm-5.2","summary":"The referee recommends minor revision and identifies two major comments: (1) the exponential growth of perturbation radii in Lemma 3 when the lifted matrix norm exceeds 1, and (2) the zero-valued channel and reset terms in Table 6. Both comments are well-taken. We will add an explicit discussion of the non-vacuous certification regime and clarify the table reporting conventions.","responses":[{"response":"The referee is correct that the perturbation radii in Lemma 3 grow exponentially with H when ||Â|| > 1, and that this is a genuine limitation of the bound rather than a presentation artifact. We will add an explicit discussion of this regime. Specifically, we plan to add a remark after Lemma 3 stating three points. First, the bound (30) is conservative because it uses induced-norm submultiplicativity rather than spectral-radius arguments; when the lifted matrix Â has spectral radius below 1 but induced norm above 1, the bound may still grow with H even though the true stacked error does not. Second, in the numerical experiments, the identified lifted matrices for the Duffing, cubic damping, and nonlinear servo systems do have ||Â|| moderately above 1 (typically 1.1-1.5 in the induced 2-norm), but the horizons N = 70-90 are short enough that the stacked radii remain non-vacuous when combined with the other budget terms. Third, the certificate becomes vacuous when the product of ||Â||^H and the one-step error radii exceeds the channel margin, which is a fundamental limitation of any finite-data perturbation bound of this type. We will also note that the radii ε_{G,H,β} and ε_{O,H,β} can be replaced by externally audited bounds (as already mentioned in Lemma 3), which may be tighter than the submultiplicative construction when the lifted spectrum is favorable. We agree this discussion is load-bearing for Theorem 2 and will add it to the revised manuscript.","revision_made":"yes","referee_comment":"Lemma 3, Eqs. (27)-(30): The perturbation radii D_A(j) accumulate terms involving ||Â||^(j-1-ℓ) * (||Â|| + ε_A)^ℓ, which grows exponentially with horizon H when ||Â|| > 1. Since Koopman lifted matrices frequently have spectral radii exceeding 1 (the physical plant may be stable while the lifted representation is not), the stacked radii ε_{G,H,β} and ε_{O,H,β} in Eq. (30) can become extremely large for moderate H. The numerical experiments (Table 4) report w_bar values around 0.4-0.6 with bands around 0.8-1.1, which are non-trivial but relatively wide. The paper should explicitly discuss the regime where these radii yield non-vacuous certificates — for instance, characterizing the relationship between H, ||Â||, and the resulting band size, or noting whether the experiments operate in a regime where ||Â|| < 1. This is load-bearing because Theorem 2's practical value depends on these radii."},{"response":"The referee is correct to flag this. The zero values in the channel and reset columns of Table 6 arise because the sweep experiment was designed to isolate the effect of a single perturbation source (disturbance, noise, or mismatch) on the budget, and in those rows the channel and reset perturbation radii were not separately computed — they were set to zero to isolate the swept variable. This is a reporting choice that should have been stated explicitly. In the revised manuscript, we will add a note to Table 6 clarifying that the sweep rows hold the channel and reset terms at zero to isolate the effect of the swept variable, and that the full budget (including nonzero channel and reset terms) is reported separately in Table 4. We will also verify that the column headers are unambiguous about which terms from Eq. (33) are being reported in each table. If the referee's concern is that the zero values misrepresent the finite-sample budget as defined in Eq. (33), we agree that the current presentation is unclear and will fix it.","revision_made":"yes","referee_comment":"Table 6: Several rows report channel and reset terms as exactly 0.000 (e.g., all disturbance and noise sweep rows). Given that Lemma 3's radii are generically nonzero for any finite-data identified model, these zero values require clarification. Are these set to zero because the channel perturbation is negligible relative to the reported precision, or are they excluded from the budget for another reason? If the former, the table should indicate this; if the latter, the relationship between the finite-sample budget (Eq. 33) and the reported numbers needs clarification."}],"tokens_in":15682,"tokens_out":1015,"duration_ms":161211,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"The main thing to know: this paper formulates Koopman learning control stability as a trial-axis practical ISS problem and decomposes the ultimate band into separately certifiable components — residual, projection, channel margin, reset. That decomposition is the real contribution. The deterministic core (Theorem 1) is a clean parameter-free derivation: the budget w_bar is built from quantities you can actually check before deployment, and the proof is a straightforward scalar comparison recursion from Lemma 2. Propositions 1–3 (projection-residual obstruction, request closure, weak-channel non-certifiability) are correct and well-motivated. The point that a small-residual predictor can still be non-certifiable when the selected channel is weak is a genuinely useful diagnostic distinction. Code and data are shipped, and the numerical experiments audit the specific certificate objects rather than just reporting tracking performance. That matters. The stress-test concern about Lemma 3's perturbation radii is real but does not undermine the paper. The radii in Eqs. 27–30 propagate one-step identification errors through matrix powers using submultiplicative norms, which means they grow exponentially with horizon H when the lifted matrix norm exceeds 1. Since Koopman lifted matrices often have spectral radius above 1, the stacked radii can become large for moderate H, making the certified band wide. The numerical experiments confirm this: bands of 0.8–1.1 in Table 4 are non-vacuous but loose. However, the paper is transparent about this limitation and scopes itself to finite-horizon tasks. The deterministic ISS result stands independently of how tight the finite-sample radii are. I'd push the authors to discuss more concretely when the radii are small enough to be useful — what system classes or horizons yield non-trivial bands — but the framework's logic is sound within its declared scope. The reader's assessment (accept, moderate confidence) is fair. The novelty score of 7 is about right: the decomposition is a new organizational contribution, not a new mathematical technique. This paper is for researchers working on certified learning-based control who need a principled way to separate model fit from deployability. It deserves a serious referee. The main thing a referee should press on is whether the finite-sample radii can be tightened enough for practical use, or whether the framework is primarily conceptual.","headline":"Clean deterministic ISS framework for Koopman learning control; finite-sample radii are conservative but honestly scoped.","tokens_in":16589,"tokens_out":542,"would_cite":true,"duration_ms":279225,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Prediction accuracy does not certify learning stability","keywords":[],"falsifier":"Construct a system where the Koopman predictor has arbitrarily small held-out prediction residuals but the selected finite-horizon channel has a near-zero minimum singular value relative to its perturbation radius. The theory predicts this predictor must be rejected as non-certifiable despite its excellent prediction accuracy.","tokens_in":16043,"feed_emoji":"🛡️","tokens_out":820,"duration_ms":99436,"temperature":0.7,"pith_summary":"This paper proves that a Koopman model with small held-out prediction residuals can still fail to produce a stable learning controller, because stability along the trial axis requires a positive margin on the selected finite-horizon input-output channel and a bounded projection residual measuring the gap between requested output corrections and what the constrained actuators can actually deliver. The authors formulate the stacked tracking error as the state of a discrete-time system indexed by trial number, treat Koopman residuals, reset mismatch, channel uncertainty, projection residuals, deployment shifts, and numerical tolerances as disturbance inputs, and prove that the learning error decays geometrically to a computable ultimate band whose radius is the sum of all these perturbation budgets divided by one minus the learning gain. A finite-sample implementation uses split-conformal calibration on frozen controllers to construct the residual budget with probabilistic coverage, and a rejection-capable certification protocol discards any candidate whose channel margin, projection closure, or certified band fails to meet the requirements.","feed_headline":"Prediction accuracy does not certify learning stability","feed_subtitle":"A Koopman model with small residuals can still fail stability certification if its control channel is weak or actuator constraints block the","key_machinery":"The practical ISS bound (Theorem 1, Eq. 22): ||E_bar_k|| <= lambda^k ||E_bar_0|| + ((1-lambda^k)/(1-lambda)) * w_bar, where w_bar = xi_0 + p_bar + eps_G * U_bar_Delta + eps_O * Z_bar_0. The learning gain lambda controls geometric decay rate; w_bar aggregates residual, projection, channel, and reset budgets into the ultimate band radius w_bar/(1-lambda).","core_discovery":"The central object is the projection residual, defined as the distance from the requested output increment to the constrained reachable set of the learned finite-horizon channel. This residual is a structural actuation limit, not a numerical error, and it enters the ISS budget as an irreducible contribution to the ultimate error band. The paper proves that if the learned channel's minimum singular value minus its certified perturbation radius is nonpositive, no stability certificate can be issued regardless of how small the prediction residuals are, formally separating model fit from learning certifiability.","pith_inferences":[],"forward_implications":["A learned Koopman predictor with excellent prediction accuracy can be rejected from deployment if its selected finite-horizon channel is weak, meaning the actuators cannot reliably produce the output corrections the learning law requests.","Each certification failure can be traced to a specific cause: large residual score implicates the predictor or noise level; nonpositive channel margin implicates data informativeness; large projection residual implicates actuator constraints; dominant shift or numerical terms implicate the implementation protocol.","The certified ultimate band is not zero under realistic conditions because finite samples, calibration scores, and deployment disturbances are nonzero, making this a practical rather than idealized stability statement.","The certification protocol is rejection-capable: it returns 'Not Certified' whenever the contraction condition, channel margin, request closure, or certified band requirement fails, rather than deploying an uncertified controller."],"fun_headline_variants":["Small Koopman residuals do not guarantee learning stability","Projection residuals bound Koopman stability independent of model accuracy","Model fit is decoupled from stability certification in Koopman control","Weak control channels fail stability certificates despite accurate models","Actuation limits, not prediction errors, bound Koopman stability"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The certified deployment event (Assumption 1) requires that the true channel deviates from the learned channel by no more than a certified spectral-norm radius, and that the learned channel's minimum singular value exceeds this radius by a positive margin. If the true channel deviates more than the certified radius, the entire stability budget is invalid.","fun_headline_variants_meta":{"raw":{"variants":["Small Koopman residuals do not guarantee learning stability","Projection residuals bound Koopman stability independent of model accuracy","Model fit is decoupled from stability certification in Koopman control","Weak control channels fail stability certificates despite accurate models","Actuation limits, not prediction errors, bound Koopman stability","Koopman stability requires a positive channel margin, not just accuracy","Structural actuation limits constrain Koopman learning stability"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":1476,"prompt_tokens":527,"completion_tokens":949,"prompt_tokens_details":null},"tokens_in":527,"tokens_out":949,"duration_ms":34683,"temperature":1.0,"reasoning_tokens":932,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-08T04:58:27.486589+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"Construct a system where the Koopman predictor has arbitrarily small held-out prediction residuals but the selected finite-horizon channel has a near-zero minimum singular value relative to its perturbation radius. The theory predicts this predictor must be rejected as non-certifiable despite its excellent prediction accuracy.","supporting_citations":[],"review_version":1}