Pith. sign in

REVIEW 3 major objections 3 minor 29 references

Extrapolating the emergence of Hamiltonian chaos with random-feature Hamiltonian neural networks

T0 review · 3 major / 3 minor · reviewed 2026-08-03 · deepseek-v4-flash

Pith's one-line read A learned Hamiltonian trained on regular orbits predicts the onset of chaos at unseen parameters.

desk verdict The central claim—extrapolating chaos from regular-only training—holds up; the paper is careful and worth a serious referee. read the letter →

arxiv 2607.28977 v1 pith:VIJMWOOM submitted 2026-07-31 nlin.CD cs.LGphysics.comp-ph

classification nlin.CDcs.LGphysics.comp-ph
keywords Hamiltonianneuralnetworksrandomfeaturesparameterextrapolationregular-to-chaotictransitionHénon–HeilesfamilyPoincarésectionsfinite-timeLyapunovexponentstwo-degree-of-freedomHamiltonians
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether a Hamiltonian neural network can predict a dynamical regime it never saw in training: the broad chaotic sea that appears as a control parameter grows past the values where invariant tori dominate. The answer is yes, provided the network is a parameter-aware random-feature Hamiltonian network (RF-HNN) whose only trained part is a linear readout fitted by ridge regression. Trained on vector-field samples from just a few regular control values, the RF-HNN extrapolates to larger, unseen controls and reproduces, across four two-degree-of-freedom families including Hénon–Heiles, the breakup of invariant structures and the growth of chaotic regions, as measured by Poincaré-section occupancy, section geometry, and finite-time Lyapunov exponents. Conventionally trained Hamiltonian networks fitted to the same shells and controls remain too regular out of range. If the claim holds, the route to chaos can be learned from regular dynamics alone, and what governs extrapolation is how the fitted Hamiltonian continues in the control parameter, not Hamiltonian structure by itself.

What carries the argument

The carrying object is the two-layer random-feature Hamiltonian network: the state z and control c pass through two fixed, randomly sampled feature layers with a dedicated control channel, and the Hamiltonian is the linear readout beta^T phi_2(z,c); Hamilton's equations applied to this scalar define the autonomous flow, and the readout beta is the unique ridge-regression solution of the field-matching problem. Because the fit is convex and closed-form, internal hyperparameters such as feature scales, width, activation, and regularization can be tuned by Bayesian optimization on a validation control inside the training interval without touching the extrapolation regime. This low-complexity co

What would settle it

A concrete check: retrain the RF-HNN for the two Hénon–Heiles families after excluding every training orbit whose finite-time Lyapunov exponent at the upper training controls (e.g., µ=0.7 or µ=1.2) exceeds a small threshold, or after shifting the training interval one step below the transition. If the predicted chaotic sea at the largest extrapolation control disappears or degrades sharply, the extrapolation was carried by the mildly unstable minority rather than by continuation of the regular family. Alternatively, vary the control-channel feature scale and check whether the onset control of

Watch

Extended reading notes

Core claim

The central discovery is that a parameter-aware random-feature Hamiltonian network (RF-HNN), trained only at controls whose phase space is dominated by invariant tori and never exposed to the extrapolation regime in training or model selection, generates an autonomous Hamiltonian flow whose long-time behavior at larger, unseen controls matches the true family's development of widespread chaos. In the bounded nonlinear Hénon–Heiles family, the learned occupancy stays within about 0.6 cells of truth across all extrapolation controls and the mean finite-time Lyapunov exponent reaches 0.271 against the true 0.263 at the largest control µ=1.5, while the back-propagation-trained MLP-HNN and its 20

Load-bearing premise

The load-bearing premise is that the training data contain no meaningful seed of chaos: the mildly elevated finite-time instability seen in a minority of orbits at the upper training controls of the two Hénon–Heiles families must not be what teaches the network to produce the extrapolated chaotic sea, because if it is, the claimed qualitatively new extrapolation is closer to interpolation.

Editorial extensions

If this is right

  • Learned Hamiltonian models can act as predictors of qualitatively new dynamical regimes, not just interpolators of regimes seen in training.
  • Parameter continuation, not conservation structure alone, is the decisive factor: two Hamiltonian models that agree inside the training interval can diverge sharply outside it, so extrapolation comparisons must examine the fitted continuation, not just the architecture.
  • Poincaré-section occupancy, Chamfer distance, and finite-time Lyapunov exponents form a workable evaluation package for chaotic extrapolation because they remain meaningful when pointwise trajectories decorrelate.
  • The closed-form ridge fit makes hyperparameter selection over random-feature scales affordable, giving a practical recipe for other order-to-chaos families.
  • The demonstration across Hénon–Heiles, a confined Barbanis-type system, and a coupled Morse pair suggests the phenomenon is generic for near-integrable two-degree-of-freedom Hamiltonians.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: the paper itself flags (Section V, Appendix C) that the mechanism is not yet isolated and that matched ablations are needed; that admitted gap leaves the explanation provisional even if the empirical finding stands.
  • Editorial extension: a sharper test of the mechanism would vary regularization strength directly — e.g., a strongly regularized MLP-HNN or an RF-HNN with much larger control-channel scale — since the Appendix C sweep varied width and activation but not the effective capacity of the backprop-trained fit.
  • Editorial extension: the 'predominantly regular' premise deserves a stress test — retrain after excluding upper-training orbits with mildly elevated finite-time exponents; if the extrapolated chaos collapses, the boundary between extrapolation and interpolation is thinner than the headline suggests.
  • Editorial extension: if the result generalizes, the learned Hamiltonian could serve as a cheap surrogate for mapping the order-to-chaos boundary across a parameter family, since evaluating the surrogate is far cheaper than integrating true chaotic dynamics.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper introduces a parameter-aware random-feature Hamiltonian neural network (RF-HNN) and claims that, when trained only on vector-field data at control-parameter values where invariant tori dominate, it can extrapolate to unseen control values at which a broad chaotic sea develops. The method is tested on four two-degree-of-freedom Hamiltonian families, including two Hénon–Heiles variants. The RF-HNN is compared against conventional MLP-HNN baselines using Poincaré-section occupancy, Chamfer distance, and finite-time Lyapunov exponents. The authors report that the RF-HNN reproduces the growth of chaotic regions quantitatively, whereas MLP-HNNs remain too regular, and conclude that parameter continuation, rather than Hamiltonian structure alone, determines extrapolation success.

Significance. If the claims hold, this is a significant demonstration: a learned Hamiltonian can cross a qualitative dynamical regime boundary—from regular-dominated training data to a chaotic sea—without any data from the target regime. The protocol is carefully designed: training and model selection are restricted to in-band controls, multiple independent diagnostics are used, initial conditions and integrators are shared between truth and models, an interpolation sanity check is included, and baseline architecture sweeps are provided. The paper also explicitly acknowledges its main limitations, including the fact that the mechanism of continuation selection is not yet isolated. These strengths make the result worth taking seriously, provided the sensitivity of the extrapolation to the training-set composition and model-selection choices is established.

major comments (3)
  1. [Section II and Appendix A] The central claim depends on the training data being effectively free of the chaotic mechanism. The paper states that training regimes are only 'predominantly regular' and that a minority of orbits at the upper Hénon–Heiles training controls shows 'mildly elevated finite-time instability.' Although Appendix A reports decaying finite-time instability at T=3840, this does not quantify how close the upper training controls are to the transition, nor does it rule out that the extrapolated chaos is a continuation of a trend already present. A sensitivity analysis that removes the upper training controls (e.g., training only at µ≤0.6 for the canonical family and µ≤1.1 for the bounded family) is needed to support the 'qualitatively new' claim.
  2. [Section V and Appendix C] The contrast with the MLP-HNN is used to conclude that 'what decides parameter extrapolation is not Hamiltonian structure alone but how the fitted Hamiltonian continues in the control parameter.' However, the MLP-HNN baseline is not matched in capacity, regularization, or hyperparameter optimization; Appendix C sweeps only width and activation, not training hyperparameters or early stopping. The conclusion is broader than the evidence. The authors themselves call for matched ablations in Section V. Either the conclusion should be tempered or matched-capacity experiments should be added.
  3. [Section III (model selection)] Hyperparameters are selected using in-band validation RMSE at a single validation control per family. The paper does not report the sensitivity of the extrapolated diagnostics to the choice of validation control or to the number/location of training controls. Given that the mechanism of continuation selection is unknown, this sensitivity is load-bearing for the robustness of the extrapolation claim. A small ablation varying the validation control and training set would substantially strengthen the paper.
minor comments (3)
  1. [General] The data availability statement says code and data are available 'upon request'; for a computational study claiming a novel empirical result, a public repository would improve reproducibility.
  2. [Section III B] In Eq. (7), the summation index and the factor T/Δt are clear, but the notation λ_T for a finite-time exponent averaged over a window could be confused with a true Lyapunov exponent; a brief reminder that this is a finite-time diagnostic would help.
  3. [Table II and Figure 7] The Chamfer distances in Figure 7 span very different magnitudes across families; the text mentions this, but a log scale or normalized version would make the model–truth comparison easier to read.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the extrapolation regime is excluded from training and model selection, and the target chaos metrics are compared post hoc.

full rationale

The derivation chain is self-contained. The only fitted parameters, β, solve Eq. (5), a ridge regression on in-band vector-field targets A_i β = J∇_z H(z_i, c_i) with (z_i, c_i) drawn only at the training controls of Table I. Hyperparameters and the activation are selected by RMSE on a validation control inside the training interval; the paper states explicitly that 'No extrapolation control is used in model selection.' The extrapolated chaos diagnostics—occupancy, Chamfer distance, and finite-time Lyapunov exponents—are evaluated only after the model is fixed and integrated autonomously, so no target-regime quantity enters the fit or the model-selection objective. The MLP-HNN baseline receiving the same data and objective fails to produce the chaotic sea, which shows that the outcome is not forced by Hamiltonian structure or by the training labels alone. The paper's own caveats are limitations, not circularity: Section II describes the training regimes as 'predominantly regular rather than strictly regular' with mildly elevated finite-time instability at some upper training controls, and Section V concedes that the continuation-selection mechanism is 'not yet isolated' and that matched ablations are needed. These are honest interpretational uncertainties about whether the transition is partly interpolative, but they do not make any predicted quantity an input to the fit. The self-citations [14, 16] are background reservoir-computing references and are not load-bearing. Consequently, no circular step is exhibited.

Assumptions & free parameters 1 free parameters · 6 assumptions · 0 invented entities

The central extrapolation claim rests on the selected RF-HNN hyperparameters (in-band fit), the smooth-continuation representational assumption of random features, and the empirical regularity of the training controls. No new physical entities are introduced. The most fragile item is the assumption that the training data are 'predominantly regular' in a meaningful sense and that in-band validation selects the correct out-of-band continuation.

free parameters (1)
  • RF-HNN hyperparameters per family (feature scales s_z, s_c, second-layer scale g_2, width N, ridge α, activation) = Linear HH: N=2000, s_z=0.416, s_c=0.0266, g_2=2.415, log10 α=-10.363; Bounded HH: N=1800, s_z=0.378, s_c=0.0421, g_2=1.0
    Selected by 100-trial Bayesian optimization per activation per family on in-band validation RMSE (Appendix B). These choices control the learned Hamiltonian's continuation in c and are therefore load-bearing for the extrapolation claim, even though they are not fit to extrapolation data.
assumptions (6)
  • standard math Hamilton's equations J∇H define the flow, and the HNN conserves the learned Hamiltonian and phase-space volume by construction.
    Used throughout (Section III) to construct the learned vector field from the scalar Hamiltonian.
  • domain assumption The one-parameter families undergo the generic near-integrable route to chaos (KAM breakup, resonance overlap) as the control increases.
    The task framing (Section II) relies on this standard picture from the cited literature [1,2] to define the extrapolation target.
  • domain assumption Random-feature ridge regression with two fixed hidden layers provides an approximation class whose in-band fit and out-of-band continuation can match the true Hamiltonian's parameter dependence.
    The entire method (Section III, Eq. 3-5) rests on this representational assumption; the paper does not prove it, though the four-family results support it empirically.
  • domain assumption The training controls are 'predominantly regular' and the mildly elevated finite-time instability at the upper training controls does not constitute the chaotic sea being predicted.
    Explicitly stated in Section II: 'None of the training sections contains a broad chaotic sea, although at the upper training controls of the two Hénon–Heiles families a minority of sampled orbits shows mildly elevated finite-time instability.' This is a load-bearing empirical assumption.
  • ad hoc to paper In-band vector-field RMSE at one validation control is a sufficient model-selection criterion for choosing a continuation that extrapolates correctly.
    The hyperparameter search (Appendix B) uses a single held-out control inside the training interval. The paper does not test sensitivity to this choice; the extrapolation success depends on the selected configuration.
  • domain assumption Finite-time Lyapunov exponents with T=240 and d0=10^-5, computed by the two-trajectory Benettin method, are reliable indicators of the long-time regime at the energy shells considered.
    Used for all chaos quantifications (Section III.B); the paper verifies robustness to T, d0, and tangent integration in Appendix A.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Extrapolating the emergence of Hamiltonian chaos with random-feature Hamiltonian neural networks." pith.science (2026). https://pith.science/paper/VIJMWOOM

@misc{pith2026260728977,
  author       = {Pith},
  title        = {Pith review of: Extrapolating the emergence of Hamiltonian chaos with random-feature Hamiltonian neural networks},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VIJMWOOM}},
  note         = {Machine review of arXiv:2607.28977}
}
read the original abstract

Machine learning of Hamiltonian dynamics has driven growing interest in Hamiltonian neural networks (HNNs), which encode Hamilton's equations of motion into the learning architecture. Despite this progress, it remains unknown whether such networks can predict dynamical regimes absent from their training data, in particular the broad chaotic sea that emerges beyond the observed parameter interval. We address this question using a parameter-aware random-feature Hamiltonian neural network (RF-HNN). Trained using data from only a small number of control-parameter values at which invariant tori dominate, the RF-HNN predicts autonomous long-time dynamics at unseen parameter values where mixed phase space develops and chaotic regions expand, with no data from that regime used in training or model selection. The method is demonstrated across four two-degree-of-freedom Hamiltonian families, including the H\'enon-Heiles system. Using Poincar\'e-section geometry and finite-time Lyapunov exponents, we show that the RF-HNN reproduces the breakup of regular structures and the emergence and growth of chaotic regions, whereas conventionally trained HNNs with the same Hamiltonian structure remain too regular. These results show that what decides parameter extrapolation is not Hamiltonian structure alone but how the fitted Hamiltonian continues in the control parameter. To our knowledge, this is the first demonstration that a learned Hamiltonian can qualitatively extrapolate from predominantly regular dynamics into a broad chaotic sea absent from training.

Figures

Figures reproduced from arXiv: 2607.28977 by the authors.

Figure 1
Figure 1. FIG. 1. Problem setting and model. (a) Truth Poincar´e sections of the canonical linear-control H´enon–Heiles family at an [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. FIG. 2. Canonical linear-control H´enon–Heiles family, with no [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. FIG. 3. Canonical linear-control extrapolation. (a) Occupancy (in cells) and mean finite-time exponent [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: FIG. 4. Behavior of the bounded nonlinear-control family in [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: FIG. 5. Bounded nonlinear-control extrapolation. (a) Occupancy (in cells) and mean finite-time exponent [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: FIG. 6. Interpolation counterpart of Fig. 5 for the bounded [PITH_FULL_IMAGE:figures/full_fig_p011_6.png]
Figure 7
Figure 7. Figure 7: FIG. 7. Full comparison for linear H´enon–Heiles, bounded nonlinear-control H´enon–Heiles, the confined Barbanis-type model, [PITH_FULL_IMAGE:figures/full_fig_p012_7.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

29 extracted references · 6 linked inside Pith

  1. [1]

    B. V. Chirikov, A universal instability of many- dimensional oscillator systems, Physics Reports52, 263 (1979)

  2. [2]

    A. J. Lichtenberg and M. A. Lieberman,Regular and Chaotic Dynamics, 2nd ed. (Springer, 1992)

  3. [3]

    H´ enon and C

    M. H´ enon and C. Heiles, The applicability of the third integral of motion: Some numerical experiments, The As- tronomical Journal69, 73 (1964)

  4. [4]

    Barbanis, On the isolating character of the ‘third’ in- tegral in a resonance case, The Astronomical Journal71, 415 (1966)

    B. Barbanis, On the isolating character of the ‘third’ in- tegral in a resonance case, The Astronomical Journal71, 415 (1966)

  5. [5]

    Kryvohuz and J

    M. Kryvohuz and J. Cao, Noise-induced dynamic sym- metry breaking and stochastic transitions in ABA molecules: I. Classification of vibrational modes, The Journal of Physical Chemistry B114, 6549 (2010)

  6. [6]

    Greydanus, M

    S. Greydanus, M. Dzamba, and J. Yosinski, Hamiltonian neural networks, inAdvances in Neural Information Pro- cessing Systems, Vol. 32 (2019) arXiv:1906.01563

  7. [7]

    P. Jin, Z. Zhang, A. Zhu, Y. Tang, and G. E. Karniadakis, SympNets: Intrinsic structure-preserving symplectic net- works for identifying Hamiltonian systems, Neural Net- works132, 166 (2020)

  8. [8]

    Chen and M

    R. Chen and M. Tao, Data-driven prediction of general Hamiltonian dynamics via learning exactly-symplectic maps, inProceedings of the 38th International Confer- ence on Machine Learning, PMLR, Vol. 139 (2021) pp. 1717–1727

Show all 29 references
  1. [9]

    Choudhary, J

    A. Choudhary, J. F. Lindner, E. G. Holliday, S. T. Miller, S. Sinha, and W. L. Ditto, Physics-enhanced neural net- works learn order and chaos, Physical Review E101, 062207 (2020)

  2. [10]

    C.-D. Han, B. Glaz, M. Haile, and Y.-C. Lai, Adaptable Hamiltonian neural networks, Physical Review Research 3, 023156 (2021)

  3. [11]

    G¨ oring, F

    N. G¨ oring, F. Hess, M. Brenner, Z. Monfared, and D. Durstewitz, Out-of-domain generalization in dynam- ical systems reconstruction, inProceedings of the 41st International Conference on Machine Learning (ICML) (2024) arXiv:2402.18377

  4. [12]

    J. Z. Kim, Z. Lu, E. Nozari, G. J. Pappas, and D. S. Bas- sett, Teaching recurrent neural networks to infer global temporal structure from local examples, Nature Machine Intelligence3, 316 (2021)

  5. [13]

    Kong, H.-W

    L.-W. Kong, H.-W. Fan, C. Grebogi, and Y.-C. Lai, Ma- chine learning prediction of critical transition and system collapse, Physical Review Research3, 013090 (2021)

  6. [14]

    Choi and P

    J. Choi and P. Kim, Early warning for critical transitions using machine-based predictability, AIMS Mathematics 7, 20313 (2022)

  7. [15]

    Jaeger and H

    H. Jaeger and H. Haas, Harnessing nonlinearity: Predict- ing chaotic systems and saving energy in wireless com- munication, Science304, 78 (2004)

  8. [16]

    Choi and P

    J. Choi and P. Kim, Homotopy reservoir computing: Har- nessing chaos for computation, Chaos35, 093112 (2025)

  9. [17]

    Panahi and Y.-C

    S. Panahi and Y.-C. Lai, Adaptable reservoir computing: A paradigm for model-free data-driven prediction of crit- ical transitions in nonlinear dynamical systems, Chaos 34, 051501 (2024)

  10. [18]

    Koglmayr and C

    D. Koglmayr and C. R¨ ath, Extrapolating tipping points and simulating non-stationary dynamics of complex sys- tems using efficient machine learning, Scientific Reports 14, 507 (2024)

  11. [19]

    van Tegelen, G

    E. van Tegelen, G. van Voorn, I. Athanasiadis, and P. van Heijster, Neural ordinary differential equations for learn- ing and extrapolating system dynamics across bifurca- tions, Chaos35, 101103 (2025)

  12. [20]

    I. J. S. Shokar, P. H. Haynes, and R. R. Kerswell, Deep learning of the evolution operator enables forecasting of out-of-training dynamics in chaotic systems (2025), arXiv:2502.20603

  13. [21]

    Zhang, H

    H. Zhang, H. Fan, L. Wang, and X. Wang, Learning Hamiltonian dynamics with reservoir computing, Physi- cal Review E104, 024205 (2021)

  14. [22]

    Jakov´ ac, M

    A. Jakov´ ac, M. T. Kurbucz, and P. P´ osfay, Reconstruc- tion of observed mechanical motions with artificial intel- ligence tools, New Journal of Physics24, 073021 (2022)

  15. [23]

    E. L. Bolager, I. Burak, C. Datar, Q. Sun, and F. Diet- rich, Sampling weights of deep neural networks, inAd- vances in Neural Information Processing Systems, Vol. 36 (2023) arXiv:2306.16830

  16. [24]

    Rahma, C

    A. Rahma, C. Datar, and F. Dietrich, Training Hamil- tonian neural networks without backpropagation, in NeurIPS 2024 Workshop on Machine Learning and the Physical Sciences(2024) arXiv:2411.17511

  17. [25]

    Rahma, C

    A. Rahma, C. Datar, A. ˇCukarska, and F. Dietrich, Rapid training of Hamiltonian graph networks using random features, inInternational Conference on Learning Repre- sentations (ICLR)(2026) arXiv:2506.06558

  18. [26]

    P. M. Morse, Diatomic molecules according to the wave mechanics. II. Vibrational levels, Physical Review34, 57 (1929)

  19. [27]

    Akiba, S

    T. Akiba, S. Sano, T. Yanase, T. Ohta, and M. Koyama, Optuna: A next-generation hyperparameter optimiza- tion framework, inProceedings of the 25th ACM SIGKDD International Conference on Knowledge Dis- covery & Data Mining(2019) pp. 2623–2631

  20. [28]

    H. G. Barrow, J. M. Tenenbaum, R. C. Bolles, and H. C. Wolf, Parametric correspondence and chamfer matching: Two new techniques for image matching, inProceedings of the 5th International Joint Conference on Artificial Intelligence(1977) pp. 659–663

  21. [29]

    Benettin, L

    G. Benettin, L. Galgani, A. Giorgilli, and J.-M. Strelcyn, Lyapunov characteristic exponents for smooth dynami- cal systems and for Hamiltonian systems; a method for computing all of them. Part 1: Theory, Meccanica15, 9 (1980)

Pith tools

Reviewed August 3, 2026 · model on record in the stance chip above.