Pith. sign in

REVIEW 3 major objections 5 minor 92 references

This paper claims that machine-learning force fields for molecular dynamics can be reformulated as implicit fixed-point models, and that warm-starting the solver from previous timesteps cuts compute and memory by two- to five-fold while mat

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 12:32 UTC pith:K7CTXREL

load-bearing objection Solid DEQ-force-field paper with a genuinely useful warm-starting contribution, but the headline 2–5x speedup is measured in interaction-layer calls, not wall-clock, so the strong claim needs runtime validation. the 3 major comments →

arxiv 2607.29158 v1 pith:K7CTXREL submitted 2026-07-31 cs.LG cs.AI

Implicit Machine Learning Force Fields Accelerate Molecular Dynamics Simulations

classification cs.LG cs.AI
keywords implicit neural networksmachine learning force fieldsmolecular dynamicsfixed-point iterationwarm-startingdeep equilibrium modelsequivariant graph neural networksimplicit differentiation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to establish that the standard way of computing machine-learning force fields—running a deep network from scratch at every femtosecond timestep—wastes computation that an implicit formulation can recover. It proposes I-MLFFs, where a single interaction layer is iterated to a fixed point that encodes the molecular graph, and where the converged representation from previous timesteps is reused and linearly extrapolated to warm-start the next solve. Because the fixed point changes only slightly between steps, the solver typically converges in one to two layer calls, matching the per-step cost of a one-layer network while retaining the accuracy of explicit models with up to five layers. This is demonstrated across invariant, Cartesian-equivariant, and SO(3)-equivariant architectures, with two- to five-fold reductions in compute and memory, no loss of atomistic resolution or integration timestep, and energy-conserving forces via implicit differentiation. If correct, the approach directly extends the scale and duration of ab-initio-quality molecular simulations within fixed GPU budgets.

Core claim

The central claim is that implicit modeling makes force-field inference amortizable along a molecular dynamics trajectory. Concretely, the paper shows that any explicit graph-neural-network force field can be converted into an implicit model by iterating one learned interaction layer, with input injection and normalization, until it satisfies h* = f(h*, x). The forces are then obtained by implicit differentiation, solving a second fixed-point equation for the adjoint u*, which requires storing only the final application of f. The authors demonstrate that when the solver is warm-started with linear extrapolation h(0)(t+Δt) = 2h*(t) - h*(t-Δt), the fixed point and adjoint converge to a residua

What carries the argument

The load-bearing mechanism is the fixed-point formulation of the latent representation, h* = f(h*, x), together with its implicit derivative. Because the mapping from positions to fixed points is differentiable (by the implicit function theorem), forces can be computed without unrolling the solver, using the adjoint fixed-point equation u* = (∂f/∂h*)ᵀ u* + ∂f_E/∂h*. Temporal amortization is achieved by warm-starting both solves with a linear (Adams-Bashforth-style) extrapolation of the previous two fixed points; training-time Jacobian and iterate-correction regularization keep the solver contractive enough to converge in one to two iterations. The entire argument rests on the smooth evolutio

Load-bearing premise

The central claim collapses if the number of calls to the interaction layer is not a faithful proxy for total cost—i.e., if the extra overhead of the fixed-point solver, residual checks, vector-Jacobian products, and extrapolation steps is large enough to erase the measured 2-5x reduction in layer calls.

What would settle it

Measure end-to-end wall-clock time per MD step for the implicit model versus its explicit counterpart with 3-5 layers, including all solver overhead, on a large system such as the double-walled nanotube, at matched force accuracy. If the implicit model is not faster (or not faster by the claimed factor), the 'compute = layer calls' assumption fails. Alternatively, run the solver with linear warm-start at a deliberately large timestep (e.g., 10 fs) on a high-temperature trajectory: if iteration counts jump far above two, the temporal-continuity premise is violated.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • I-MLFFs reduce the per-step cost of force evaluation to about one interaction-layer call, bringing machine-learned force fields closer to classical force fields in speed without coarse graining or increasing the integration timestep.
  • Memory use for force inference becomes independent of network depth (only the last layer application is stored), enabling larger atomistic systems on fixed-memory GPUs.
  • The efficiency gain is architecture-agnostic, applying to invariant, Cartesian-equivariant, and spherical-tensor equivariant models, and is additive to future architectural and implementation improvements.
  • Implicit models adaptively spend more solver iterations on rare, high-energy, or out-of-distribution conformations while remaining at one iteration near equilibrium, which preserves stability in NVE and NVT simulations.
  • The fixed-point formulation extends the effective range of message passing beyond explicit depth, improving long-range and extrapolative prediction (e.g., cumulenes).

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The warm-starting principle is not limited to force fields: any physics-based ML model whose inference is an iterative solve over a smoothly varying input sequence (e.g., neural wavefunctions, learned self-consistent-field solvers, Hamiltonian networks) could inherit similar amortization, as the paper's I-HNN experiment hints.
  • Because the paper measures compute in interaction-layer calls, the practical wall-clock gain depends on the fixed-point solver overhead being negligible; profiling on large systems would confirm whether the 2-5x claim survives to wall-clock time, especially on GPUs where small kernels have launch overhead.
  • One testable prediction is that the speedup grows as the integration timestep shrinks (fixed points become more similar) and shrinks as temperature rises or timestep increases—a direct consequence of the smoothness assumption the authors make.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes implicit machine learning force fields (I-MLFFs), replacing explicit stacks of graph-neural-network interaction layers with a fixed-point equation h* = f(h*, x). Forces are obtained by implicit differentiation, avoiding unrolled backpropagation, and the fixed-point and adjoint solvers are warm-started from previous MD timesteps, with linear extrapolation from the two previous fixed points. The authors claim that this reduces per-step inference cost to roughly one to two interaction-layer calls, giving a two- to five-fold reduction in compute and memory relative to explicit baselines, while preserving energy conservation. The method is demonstrated on SchNet, PaiNN, and SO3net backbones on MD17 and MD22, with additional stability and generalization experiments.

Significance. If the efficiency claim survives direct runtime measurement, the paper makes a valuable contribution: it couples MLFF inference with temporal coherence of MD trajectories, retains conservative forces through implicit differentiation, and offers architecture-agnostic improvements in interaction-layer call counts and memory. The strengths include extensive benchmarks across three architecture families and two benchmark suites, a clear algorithmic description, memory measurements, NVE/NVT stability tests, and a code release. The main unresolved issue is that the headline '2-5x compute/speedup' claim is supported only by counting interaction-layer calls, not by wall-clock time or FLOPs, so the practical significance is not yet fully established.

major comments (3)
  1. [Sec. II.C & Abstract] The headline 'two- to five-fold reduction in compute' rests on the statement in Sec. II.C that cost is measured as 'the number of calls to the interaction layer f' and that 'other contributions ... are generally negligible and comparable.' This assumption is load-bearing. The implicit pipeline adds residual-norm checks after every iteration, warm-start extrapolation (Eq. 6), storage/retrieval of previous fixed points, and the backward fixed-point solve (Eq. 4), which requires VJPs through f. For GNN layers with edge filters, normalization, and SiLU nonlinearities, reverse-mode cost is not guaranteed to be comparable to forward cost despite the Griewank citation. Since no wall-clock or FLOP comparison is reported, the abstract's 'simulation speedup' claim is not directly supported. Please add GPU wall-clock per MD step and a breakdown of solver/VJP overhead, or temper the claim to 'intera
  2. [Sec. II.C & Fig. 3] The 'matched computational cost' comparison in Fig. 3 requires a precise cost unit. An explicit K-layer model performs K forward calls and K backward VJPs; an implicit model performs I forward iterations and J backward iterations. If one VJP is counted as one f-call, the reported ratios depend on that equivalence. If VJPs are actually two to three times the forward cost, the cost-matched crossover and the claimed speedup change materially. Please specify how VJPs are converted into f-call units for each architecture and show sensitivity of the Fig. 3 results to this multiplier.
  3. [Sec. IV.B & IV.A] The paper states in Sec. IV.B that forces are independent of the warm start because implicit differentiation uses only local derivative information at the fixed point. Strictly, this requires the fixed-point equation h* = f(h*, x) to have a unique solution in the relevant region. Brouwer's theorem (Sec. IV.A) gives existence only, not uniqueness. If multiple fixed points or near-singular I - ∂f/∂h exist, different warm starts can select different branches, and the energy/force may not be a well-defined function of R. Please provide evidence of contraction or uniqueness (e.g., spectral radius estimates) or an empirical test comparing energies and forces obtained from multiple solver initializations along closed conformational loops.
minor comments (5)
  1. [Fig. 3] The placeholder text 'Lorem ipsum' appears in the figure/legend and must be removed before publication.
  2. [SI Sec. S4 vs Fig. 4] SI Sec. S4 refers to 'Fig. 4d' for the fine-tolerance temperature plot, but main-text Fig. 4 has only panels (a)-(c); the relevant panel is (c).
  3. [Fig. 1(d) & Sec. IV.E] Fig. 1(d) says the 2-5x footprint is 'averaged across MD17 and MD22 datasets at 300 K,' while MD17 trajectories are at 500 K and MD22 at 400-500 K (Sec. IV.E). Please clarify the temperature/protocol used for this panel.
  4. [Sec. II.A vs Sec. IV.C] The iterate-correction loss coefficient is reported as 10^3 in Sec. II.A but 10^4 in Sec. IV.C and in the final loss expression. Please unify.
  5. [Algorithm 1] Algorithm 1 has no maximum iteration count. For robustness in production MD, add a cap and define the behavior when the residual is not met within the cap.

Circularity Check

0 steps flagged

No significant circularity: efficiency claims rest on measured iteration counts and external accuracy benchmarks, not on self-referential definitions.

full rationale

I find no circular step. The implicit-force derivation (Eqs. 2-4) is a standard implicit-function-theorem argument and does not assume the efficiency conclusion. Accuracy is benchmarked against external MD17/MD22 DFT references, and the implicit-vs-explicit force-MAE comparisons are independent of the paper's fitted parameters. Hyperparameters (jac coefficient 0.32, itc weight 10^4, solver tolerance 10^-2) are tuned on aspirin and then transferred to other systems; this is a model-selection choice, not a target-equivalent fit. The warm-start speedups (Fig. 2D: 1.12 iterations with linear extrapolation; Fig. 3A: ~1-2 layer calls) are measured iteration counts to a stated residual, not quantities forced by construction, even though Eq. (6) supplies a nearby initial guess. Self-citations are to datasets (MD17/MD22) and to a generalization protocol from Ref. [12] that the paper reproduces rather than assumes. The one caveat that could weaken the headline claim is the cost metric: Sec. II.C defines cost as number of interaction-layer calls and assumes other contributions are negligible, so the 2-5x compute and memory advantage may not translate exactly to wall-clock time if solver/VJP overhead dominates. But this is an external-validity threat, not circularity: the f-call count is a stated, measurable proxy, and no equation in the paper reduces to the fitted benchmarks.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 0 invented entities

The paper introduces no new physical entities. It relies on standard fixed-point theorems, the implicit function theorem, and the empirical smoothness of learned representations over MD trajectories. The main non-standard assumption is that layer-call count is a complete accounting of compute cost; wall-clock time is not measured.

free parameters (4)
  • Forward/backward solver tolerance (epsilon) = 1e-2 (3.2e-3 for drift-free NVE)
    Directly controls the number of layer calls and the strength of energy drift; chosen to balance accuracy and speed.
  • Jacobian regularization coefficient = 0.32
    Grid-searched on aspirin (Sec. II.A) to reduce iterations while preserving force accuracy.
  • Iterate correction coefficient and decay factor = 1e4, gamma=0.4
    Grid-searched on aspirin to ensure convergence within the 10 training iterations without sacrificing accuracy.
  • Force loss weight alpha = 0.95
    Standard weighting between energy and force MSE; could affect accuracy but not the efficiency claim.
axioms (5)
  • standard math Brouwer fixed-point theorem: a continuous self-map on a compact convex set has a fixed point.
    Used in Sec. IV.A to guarantee existence of h* by normalizing the interaction layer f.
  • standard math Implicit Function Theorem: the fixed-point equation h* = f(h*, x) locally defines a differentiable map x -> h*(x).
    Used in Sec. IV.B to derive forces via implicit differentiation.
  • domain assumption The fixed-point trajectory h*(t) is smooth enough along MD trajectories for linear extrapolation to be an accurate warm-start.
    Underlies the central acceleration mechanism (Eq. 6). Empirically tested but not guaranteed for stiff or far-from-equilibrium systems.
  • ad hoc to paper The number of interaction-layer calls is a faithful proxy for computational cost.
    Used to compare implicit vs explicit models in Sec. II.C; overhead of residual checks, VJP, and extrapolation is asserted to be negligible.
  • domain assumption Residual tolerance 10^-2 is sufficient for accurate force prediction.
    Empirically justified in Sec. II.A, but not a rigorous guarantee; the drift in NVE simulations shows tolerance-dependent errors.

pith-pipeline@v1.3.0-daily-deepseek · 29185 in / 10529 out tokens · 107975 ms · 2026-08-03T12:32:31.342849+00:00 · methodology

0 comments
read the original abstract

We introduce implicit machine learning force fields (I-MLFFs), which replace explicit stacks of neural network layers with self-consistent fixed-point equations. In molecular simulations, this formulation enables intermediate representations to be reused across successive timesteps, thereby warm-starting force evaluation. The resulting models effectively combine the computational footprint of a shallow, single-layer MLFF with the representational capacity and accuracy of a deep neural network. Our approach unlocks architecture-agnostic efficiency gains that are inaccessible when force prediction and trajectory integration are considered separately. We demonstrate this across three major classes of graph neural networks: invariant, equivariant Cartesian tensor, and SO(3)-equivariant spherical-tensor architectures. Each yields a two- to five-fold reduction in compute and memory footprint. Crucially, these gains are achieved while retaining full atomistic resolution and the original integration timestep, avoiding spatial or temporal coarse graining. Our contribution therefore advances the scaling frontier of quantum-mechanically faithful molecular simulation, enabling longer trajectories and larger atomistic systems within fixed GPU memory and compute budgets, and thereby opening access to new insights across biomolecular and material systems.

Figures

Figures reproduced from arXiv: 2607.29158 by Johannes Mae{\ss}, Joshua Futterer, J. Thorben Frank, Klaus-Robert M\"uller, Leon Werner, Martin Michajlow, Stefan Chmiela, Winfried Ripken.

Figure 1
Figure 1. Figure 1: FIG. 1 [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: FIG. 2 [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: FIG. 3 [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: a,b visualizes the average iteration count (with and without warmstarts) in (10◦ ) 2 bins of the Ramachandran plot. We find that the simulation becomes highly efficient using warmstarts (Fig. 4a), with only a small fraction of MD steps taking more than one iteration to solve. Both plots highlight a distinctive advantage of implicit mod￾els: the number of solver iterations adapts for chemically rare or out-… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

92 extracted references · 1 canonical work pages

  1. [1]

    Karplus and J

    M. Karplus and J. A. McCammon, Molecular dynam- ics simulations of biomolecules, Nat. Struct. Biol.9, 646 (2002)

  2. [2]

    O. T. Unke, S. Chmiela, H. E. Sauceda, M. Gastegger, I. Poltavsky, K. T. Sch¨ utt, A. Tkatchenko, and K.-R. M¨ uller, Machine learning force fields, Chem. Rev.121, 10142 (2021)

  3. [3]

    Gasteiger, J

    J. Gasteiger, J. Groß, and S. G¨ unnemann, Directional message passing for molecular graphs, inICLR(2020)

  4. [4]

    Sch¨ utt, O

    K. Sch¨ utt, O. Unke, and M. Gastegger, Equivariant mes- sage passing for the prediction of tensorial properties and molecular spectra, inICML(PMLR, 2021) pp. 9377– 9388

  5. [5]

    T. W. Ko, J. A. Finkler, S. Goedecker, and J. Behler, A fourth-generation high-dimensional neural network po- tential with accurate electrostatics including non-local charge transfer, Nat. commun.12, 398 (2021)

  6. [6]

    Gasteiger, F

    J. Gasteiger, F. Becker, and S. G¨ unnemann, Gem- Net: Universal directional graph neural networks for molecules, inNeurIPS, Vol. 34 (2021) pp. 6790–6802

  7. [7]

    Liao and T

    Y.-L. Liao and T. Smidt, Equiformer: Equivariant graph attention transformer for 3D atomistic graphs, inICLR (2023)

  8. [8]

    Y. Wang, S. Li, X. He, M. Li, Z. Wang, N. Zheng, B. Shao, T.-Y. Liu, and T. Wang, ViSNet: an equivariant geometry-enhanced graph neural network with vector- scalar interactive message passing for molecules, arXiv preprint arXiv:2210.16518 (2023)

  9. [9]

    Batatia, D

    I. Batatia, D. P. Kov´ acs, G. Simm, C. Ortner, and G. Cs´ anyi, MACE: Higher order equivariant message passing neural networks for fast and accurate force fields, inNeurIPS, Vol. 35 (2022) pp. 11423–11436

  10. [10]

    Batzner, A

    S. Batzner, A. Musaelian, L. Sun, M. Geiger, J. P. Mailoa, M. Kornbluth, N. Molinari, T. E. Smidt, and B. Kozinsky, E(3)-equivariant graph neural networks for data-efficient and accurate interatomic potentials, Nat. Commun.13, 2453 (2022)

  11. [11]

    Musaelian, S

    A. Musaelian, S. Batzner, A. Johansson, L. Sun, C. J. Owen, M. Kornbluth, and B. Kozinsky, Learning local equivariant representations for large-scale atomistic dy- namics, Nat. Commun.14, 579 (2023)

  12. [12]

    J. T. Frank, O. T. Unke, K.-R. M¨ uller, and S. Chmiela, A Euclidean transformer for fast and stable machine learned force fields, Nat. Commun.15, 6539 (2024)

  13. [13]

    Chmiela, H

    S. Chmiela, H. E. Sauceda, K.-R. M¨ uller, and A. Tkatchenko, Towards exact molecular dynamics simu- lations with machine-learned force fields, Nat. Commun. 9, 3887 (2018)

  14. [14]

    Kabylda, J

    A. Kabylda, J. T. Frank, S. Su´ arez-Dou, A. Khabibrakhmanov, L. Medrano Sandonas, O. T. Unke, S. Chmiela, K.-R. M¨ uller, and A. Tkatchenko, Molecular simulations with a pretrained neural network and universal pairwise force fields, J. Am. Chem. Soc. 147, 33723 (2025)

  15. [15]

    D. P. Kov´ acs, J. H. Moore, N. J. Browning, I. Batatia, J. T. Horton, V. Kapil, W. C. Witt, I.-B. Magd˘ au, D. J. Cole, and G. Cs´ anyi, MACE-OFF23: Transferable ma- chine learning force fields for organic molecules, arXiv preprint arXiv:2312.15211 (2023)

  16. [16]

    M. E. Tuckerman, Ab initio molecular dynamics: basic concepts, current trends and novel applications, J. Phys.: Condens. Matter14, R1297 (2002)

  17. [17]

    W. D. Cornell, P. Cieplak, C. I. Bayly, I. R. Gould, K. M. Merz, D. M. Ferguson, D. C. Spellmeyer, T. Fox, J. W. Caldwell, and P. A. Kollman, A second generation force field for the simulation of proteins, nucleic acids, and organic molecules, J. Am. Chem. Soc.117, 5179 (1995)

  18. [18]

    A. D. MacKerell, D. Bashford, M. Bellott, R. L. Dunbrack, J. D. Evanseck, M. J. Field, S. Fischer, J. Gao, H. Guo, S. Ha, D. Joseph-McCarthy, L. Kuch- nir, K. Kuczera, F. T. K. Lau, C. Mattos, S. Mich- nick, T. Ngo, D. T. Nguyen, B. Prodhom, W. E. Rei- her, B. Roux, M. Schlenkrich, J. C. Smith, R. Stote, J. Straub, M. Watanabe, J. Wiorkiewicz-Kuczera, D. ...

  19. [19]

    J. Wang, S. Olsson, C. Wehmeyer, A. P´ erez, N. E. Charron, G. de Fabritiis, F. No´ e, and C. Clementi, Ma- chine learning of coarse-grained molecular dynamics force fields, ACS Cent. Sci.5, 755 (2019)

  20. [20]

    B. E. Husic, N. E. Charron, D. Lemm, J. Wang, A. P´ erez, M. Majewski, A. Kr¨ amer, Y. Chen, S. Olsson, G. de Fab- ritiis, F. No´ e, and C. Clementi, Coarse graining molecu- lar dynamics with graph neural networks, J. Chem. Phys. 153, 194101 (2020)

  21. [21]

    Majewski, A

    M. Majewski, A. P´ erez, P. Th¨ olke, S. Doerr, N. E. Char- ron, T. Giorgino, B. E. Husic, C. Clementi, F. No´ e, and G. De Fabritiis, Machine learning coarse-grained poten- tials of protein thermodynamics, Nat. Commun.14, 5739 (2023)

  22. [22]

    N. E. Charron, K. Bonneau, A. S. Pasos-Trejo, A. Gul- jas, Y. Chen, F. Musil, J. Venturin, D. Gusew, I. Za- porozhets, A. Kr¨ amer, C. Templeton, A. Kelkar, A. E. P. Durumeric, S. Olsson, A. P´ erez, M. Majewski, B. E. Hu- sic, A. Patel, G. De Fabritiis, F. No´ e, and C. Clementi, Navigating protein landscapes with a machine-learned transferable coarse-gr...

  23. [23]

    A. E. P. Durumeric, Y. Chen, A. S. Pasos-Trejo, F. No´ e, and C. Clementi, Learning data-efficient coarse-grained molecular dynamics from forces and noise, Nat. Com- mun.17, 2493 (2026)

  24. [24]

    T. J. Lane, D. Shukla, K. A. Beauchamp, and V. S. Pande, To milliseconds and beyond: challenges in the simulation of protein folding, Curr. Opin. Struct. Biol. 23, 58 (2013)

  25. [25]

    Simard, M

    P. Simard, M. Ottaway, and D. Ballard, Fixed point anal- ysis for recurrent networks, NeurIPS1, 149 (1988)

  26. [26]

    Miller and M

    J. Miller and M. Hardt, Stable recurrent models, inInter- national Conference on Learning Representations(2019)

  27. [27]

    S. Bai, J. Z. Kolter, and V. Koltun, Deep equilibrium models, inNeurIPS, Vol. 32 (Curran Associates, Inc.,

  28. [28]

    Winston and J

    E. Winston and J. Z. Kolter, Monotone operator equilib- rium networks, inNeurIPS, Vol. 33 (Curran Associates, Inc., 2020) pp. 10718–10728

  29. [29]

    Y. Lu, A. Zhong, Q. Li, and B. Dong, Beyond finite layer neural networks: Bridging deep architectures and numer- ical differential equations, inICML(PMLR, 2018) pp. 3276–3285. 12

  30. [30]

    R. T. Q. Chen, Y. Rubanova, J. Bettencourt, and D. K. Duvenaud, Neural ordinary differential equations, inNeurIPS, Vol. 31 (Curran Associates, Inc., 2018) pp. 6571–6583

  31. [31]

    Haber and L

    E. Haber and L. Ruthotto, Stable architectures for deep neural networks, Inverse Probl.34, 014004 (2017)

  32. [32]

    Ruthotto and E

    L. Ruthotto and E. Haber, Deep neural networks moti- vated by partial differential equations, J. Math. Imaging. Vis.62, 352 (2020)

  33. [33]

    L. L. Schaaf, I. Batatia, J. Tilly, and T. D. Barrett, BoostMD: Accelerated molecular sampling leveraging ml force field features, inNeurIPS 2024 Workshop on Data- driven and Differentiable Simulations, Surrogates, and Solvers(2024)

  34. [34]

    Burger, L

    A. Burger, L. Thiede, A. Aspuru-Guzik, and N. Vijayku- mar, DEQuify your force field: Towards efficient simula- tions using deep equilibrium models, inAI for Accelerated Materials Design Workshop, ICLR(2025)

  35. [35]

    F. L. Thiemann, T. Resch¨ utzegger, M. Esposito, T. Tad- dese, J. D. Olarte-Plata, and F. Martelli, Force-free molecular dynamics through autoregressive equivariant networks, arXiv preprint arXiv:2503.23794 (2025)

  36. [36]

    F. Bigi, S. Chong, A. Kristiadi, and M. Ceriotti, FlashMD: long-stride, universal prediction of molecular dynamics, inNeurIPS(2026)

  37. [37]

    Ripken, M

    W. Ripken, M. Plainer, G. Lied, T. Frank, O. T. Unke, S. Chmiela, F. No´ e, and K.-R. M¨ uller, Learning hamilto- nian flow maps: Mean flow consistency for large-timestep molecular dynamics, arXiv preprint arXiv:2601.22123 (2026)

  38. [38]

    Chmiela, A

    S. Chmiela, A. Tkatchenko, H. E. Sauceda, I. Poltavsky, K. T. Sch¨ utt, and K.-R. M¨ uller, Machine learning of ac- curate energy-conserving molecular force fields, Sci. Adv. 3, e1603015 (2017)

  39. [39]

    W. Hu, M. Shuaibi, A. Das, S. Goyal, A. Sriram, J. Leskovec, D. Parikh, and C. L. Zitnick, ForceNet: A graph neural network for large-scale quantum calcula- tions, arXiv preprint arXiv:2103.01436 (2021)

  40. [40]

    C. L. Zitnick, A. Das, A. Kolluru, J. Lan, M. Shuaibi, A. Sriram, Z. Ulissi, and B. Wood, Spherical channels for modeling atomic interactions, inNeurIPS, Vol. 35 (2022)

  41. [41]

    Gasteiger, M

    J. Gasteiger, M. Shuaibi, A. Sriram, S. G¨ unnemann, Z. Ulissi, C. L. Zitnick, and A. Das, GemNet-OC: Devel- oping graph neural networks for large and diverse molec- ular simulation datasets, Transactions on Machine Learn- ing Research 10.48550/arXiv.2204.02782 (2022)

  42. [42]

    Passaro and C

    S. Passaro and C. L. Zitnick, Reducing SO(3) convolu- tions to SO(2) for efficient equivariant GNNs, inICML, Vol. 202 (PMLR, 2023) pp. 27420–27438

  43. [43]

    Y.-L. Liao, B. M. Wood, A. Das, and T. Smidt, EquiformerV2: Improved equivariant transformer for scaling to higher-degree representations, inICLR(2024)

  44. [44]

    Neumann, J

    M. Neumann, J. Gin, B. Rhodes, S. Bennett, Z. Li, H. Choubisa, A. Hussey, and J. Godwin, Orb: A fast, scalable neural network potential, arXiv preprint arXiv:2410.22570 (2024)

  45. [45]

    Eissler, T

    M. Eissler, T. Korjakow, S. Ganscha, O. T. Unke, K.-R. M¨ uller, and S. Gugler, How simple can you go? an off- the-shelf transformer approach to molecular dynamics, J. Chem. Phys.164(2026)

  46. [46]

    X. Fu, Z. Wu, W. Wang, T. Xie, S. Keten, R. Gomez- Bombarelli, and T. Jaakkola, Forces are not enough: Benchmark and critical evaluation for machine learn- ing force fields with molecular simulations, Trans. Mach. Learn. Res. (2023)

  47. [47]

    F. Bigi, M. F. Langer, and M. Ceriotti, The dark side of the forces: assessing non-conservative force models for atomistic machine learning, inICML, Proceedings of Ma- chine Learning Research, Vol. 267, edited by A. Singh, M. Fazel, D. Hsu, S. Lacoste-Julien, F. Berkenkamp, T. Maharaj, K. Wagstaff, and J. Zhu (PMLR, 2025) pp. 4384–4414

  48. [48]

    J. A. Keith, V. Vassilev-Galindo, B. Cheng, S. Chmiela, M. Gastegger, K.-R. M¨ uller, and A. Tkatchenko, Com- bining machine learning and computational chemistry for predictive insights into chemical systems, Chem. Rev. 121, 9816 (2021)

  49. [49]

    K. T. Sch¨ utt, H. E. Sauceda, P.-J. Kindermans, A. Tkatchenko, and K.-R. M¨ uller, SchNet–a deep learn- ing architecture for molecules and materials, J. Chem. Phys.148(2018)

  50. [50]

    Thomas, T

    N. Thomas, T. Smidt, S. Kearnes, L. Yang, L. Li, K. Kohlhoff, and P. Riley, Tensor field networks: Rotation- and translation-equivariant neural networks for 3D point clouds, arXiv preprint arXiv:1802.08219 (2018)

  51. [51]

    O. T. Unke, S. Chmiela, M. Gastegger, K. T. Sch¨ utt, H. E. Sauceda, and K.-R. M¨ uller, SpookyNet: Learning force fields with electronic degrees of freedom and nonlo- cal effects, Nat. Commun.12, 7273 (2021)

  52. [52]

    X. Fu, B. M. Wood, L. Barroso-Luque, D. S. Levine, M. Gao, M. Dzamba, and C. L. Zitnick, Learning smooth and expressive interatomic potentials for physical prop- erty prediction, inICML, Vol. 267 (2025) pp. 17875– 17893

  53. [53]

    B. M. Wood, M. Dzamba, X. Fu, M. Gao, M. Shuaibi, L. Barroso-Luque, K. Abdelmaqsoud, V. Gharakhanyan, J. R. Kitchin, D. S. Levine, K. Michel, A. Sriram, T. Co- hen, A. Das, A. Rizvi, S. J. Sahoo, Z. W. Ulissi, and C. L. Zitnick, UMA: A family of universal models for atoms, inNeurIPS(2026)

  54. [54]

    L. E. J. Brouwer, ¨Uber abbildung von mannigfaltigkeiten, Mathematische Annalen71, 97 (1911)

  55. [55]

    S. G. Krantz and H. R. Parks,The implicit function the- orem: history, theory, and applications(Springer Science & Business Media, 2002)

  56. [56]

    Sch¨ utt, P.-J

    K. Sch¨ utt, P.-J. Kindermans, H. E. Sauceda Felix, S. Chmiela, A. Tkatchenko, and K.-R. M¨ uller, Schnet: A continuous-filter convolutional neural network for mod- eling quantum interactions, inNeurIPS, Vol. 30 (Curran Associates, Inc., 2017) pp. 991–1001

  57. [57]

    Grisafi, A

    A. Grisafi, A. Fabrizio, B. Meyer, D. M. Wilkins, C. Corminboeuf, and M. Ceriotti, Equivariant graph neural networks for fast electron density estimation of molecules, liquids, and solids, npj Comput. Mater.8, 183 (2022)

  58. [58]

    Esders, T

    M. Esders, T. Schnake, J. Lederer, A. Kabylda, G. Mon- tavon, A. Tkatchenko, and K.-R. M¨ uller, Analyzing atomic interactions in molecules as learned by neural net- works, J. Chem. Theory Comput.21, 714 (2025)

  59. [59]

    Chmiela, V

    S. Chmiela, V. Vassilev-Galindo, O. T. Unke, A. Kabylda, H. E. Sauceda, A. Tkatchenko, and K.-R. M¨ uller, Accurate global machine learning force fields for molecules with hundreds of atoms, Sci. Adv.9, eadf0873 (2023)

  60. [60]

    Kawaguchi, On the theory of implicit deep learning: Global convergence with implicit layers, inICML(2021)

    K. Kawaguchi, On the theory of implicit deep learning: Global convergence with implicit layers, inICML(2021). 13

  61. [61]

    S. Bai, V. Koltun, and Z. Kolter, Stabilizing equilib- rium models by jacobian regularization, inICML, Vol. 139 (PMLR, 2021) pp. 554–565

  62. [62]

    S. Bai, Z. Geng, Y. Savani, and J. Z. Kolter, Deep equi- librium optical flow estimation, inProceedings of the IEEE/CVF conference on computer vision and pattern recognition(2022) pp. 610–620

  63. [63]

    J. S. Spencer, D. Pfau, A. Botev, and W. M. C. Foulkes, Better, faster fermionic neural networks, arXiv preprint arXiv:2011.07125 (2020)

  64. [64]

    Hermann, Z

    J. Hermann, Z. Sch¨ atzle, and F. No´ e, Deep-neural- network solution of the electronic Schr¨ odinger equation, Nat. Chem.12, 891 (2020)

  65. [65]

    K. T. Sch¨ utt, M. Gastegger, A. Tkatchenko, K.-R. M¨ uller, and R. J. Maurer, Unifying machine learning and quantum chemistry with a deep neural network for molec- ular wavefunctions, Nat. Commun.10, 5024 (2019)

  66. [66]

    Song and J

    F. Song and J. Feng, Neural network self-consistent fields for density functional theory, npj Comput. Mater. 10.1038/s41524-026-02110-0 (2026)

  67. [67]

    Zhang, C

    H. Zhang, C. Liu, Z. Wang, X. Wei, S. Liu, N. Zheng, B. Shao, and T.-Y. Liu, Self-consistency training for density-functional-theory Hamiltonian prediction, in ICML, Proceedings of Machine Learning Research, Vol. 235, edited by R. Salakhutdinov, Z. Kolter, K. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkamp (PMLR, 2024) pp. 59329–59357

  68. [68]

    Z. Wang, C. Liu, N. Zou, H. Zhang, X. Wei, L. Huang, L. Wu, and B. Shao, Infusing self-consistency into density functional theory hamiltonian prediction via deep equi- librium models, inNeurIPS, NIPS ’24 (Curran Associates Inc., Red Hook, NY, USA, 2024)

  69. [69]

    Cranmer, S

    M. Cranmer, S. Greydanus, S. Hoyer, P. Battaglia, D. Spergel, and S. Ho, Lagrangian neural networks, arXiv preprint arXiv:2003.04630 (2020)

  70. [70]

    Greydanus, M

    S. Greydanus, M. Dzamba, and J. Yosinski, Hamiltonian neural networks, NeurIPS32(2019)

  71. [71]

    Sohl-Dickstein, E

    J. Sohl-Dickstein, E. Weiss, N. Maheswaranathan, and S. Ganguli, Deep unsupervised learning using nonequi- librium thermodynamics, inICML(PMLR, 2015) pp. 2256–2265

  72. [72]

    Y. Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, Score-based generative modeling through stochastic differential equations, inICLR(2021)

  73. [73]

    A. Q. Nichol and P. Dhariwal, Improved denoising dif- fusion probabilistic models, inICML(PMLR, 2021) pp. 8162–8171

  74. [74]

    Grathwohl, R

    W. Grathwohl, R. T. Chen, J. Bettencourt, I. Sutskever, and D. Duvenaud, FFJORD: Free-form continuous dy- namics for scalable reversible generative models, inICLR (2019)

  75. [75]

    Papamakarios, E

    G. Papamakarios, E. Nalisnick, D. J. Rezende, S. Mo- hamed, and B. Lakshminarayanan, Normalizing flows for probabilistic modeling and inference, J. Mach. Learn. Res.22, 1 (2021)

  76. [76]

    Zhang and R

    B. Zhang and R. Sennrich, Root mean square layer nor- malization, inNeurIPS, Vol. 32 (2019)

  77. [77]

    Y.-L. Liao, A. J. Hoffman, S. C. Shen, A. Duval, S. W. Norwood, and T. Smidt, EquiformerV3: Scaling effi- cient, expressive, and general SE(3)-equivariant graph attention transformers, arXiv preprint arXiv:2604.09130 (2026)

  78. [78]

    Girard, A fast ‘Monte-Carlo cross-validation’ proce- dure for large least squares problems with noisy data, Numer

    A. Girard, A fast ‘Monte-Carlo cross-validation’ proce- dure for large least squares problems with noisy data, Numer. Math.56, 1–23 (1989)

  79. [79]

    Loshchilov and F

    I. Loshchilov and F. Hutter, Decoupled weight decay reg- ularization, inICLR(2019)

  80. [80]

    Sch¨ utt, P

    K. Sch¨ utt, P. Kessel, M. Gastegger, K. A. Nicoli, A. Tkatchenko, and K.-R. M¨ uller, SchNetPack: A deep learning toolbox for atomistic systems, J. Chem. Theory Comput.15, 448 (2018)

Showing first 80 references.