{"id":"f82b9a88-4cb6-4918-8b66-d1a892bcf1fb","arxiv_id":"2412.18633","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"BoostMD accelerates MLFF molecular dynamics by predicting energy changes from previous-step node features and positional displacements, reporting 8x speedup and matching the reference model's sampled free energy surface on an unseen dipeptide.","lead":"BoostMD is a small neural network that reuses hidden features from a previous molecular dynamics step to predict how energy and forces change, so the expensive reference model only needs to run every few steps. It reports up to 8x faster simulations with similar accuracy on organic molecules, including an unseen dipeptide.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Sampling claim rests on an untested architectural variant: Appendix C.1 drops the Kabsch reference frame that §5.1 presents as necessary for rotational equivariance, leaving the tested model potentially non-conservative and biased for Boltzmann sampling.","rationale":"The paper has real independent support: measured speedups on a 3375-atom system, a fairer single-layer MACE baseline, a transferability test on an unseen dipeptide, and a 10 ns stability statement. The issue is not the speed claim but the sampling claim. Section 5.1 derives the rotational reference frame as necessary for rotational equivariance, while Appendix C.1 explicitly says all experiments omitted it. This is an internally acknowledged mismatch, not a disagreement with external consensus. The reader's weakest assumption identifies exactly this point, and I agree it is the most load-bearing soft spot. In addition, the only distributional validation is a visual comparison of two free-energy surfaces, so even a fully equivariant implementation would need a quantitative check. The proposed test is a single experimental control that compares the tested variant, the principled variant, and the ground truth, and would settle whether the omission materially biases sampling or is benign in this regime.","tokens_in":11612,"tokens_out":7275,"duration_ms":72752,"concrete_test":"Re-run the alanine-dipeptide metadynamics experiment of Figure 3 under three conditions: (i) BoostMD exactly as in the paper (no rotational framing), (ii) BoostMD with the full Kabsch reference framing of Eq. (2)/(6), and (iii) the reference MACE-OFF23-M model as ground truth. Compare the resulting (phi, psi) free-energy surfaces with a quantitative divergence (e.g., Jensen-Shannon distance between binned histograms, with bootstrap confidence intervals) and report total angular-momentum and energy-drift time series, especially across the N=10 reference-refresh boundaries. If the framing-on and framing-off surfaces agree within bootstrap noise and drift is negligible, the omission is benign for this system and the paper's claims stand as stated; if they differ, the sampling claim must be restricted to the non-equivariant variant or the method must be retested with framing.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 5.1 derives a per-atom Kabsch rotation Rbar_i as necessary for translational and rotational equivariance of the BoostMD map between a reference configuration and the current configuration. Appendix C.1 then states: \"All experimental results do not use the rotational reference framing during training or inference, due to the associated computational cost.\" Consequently, every number supporting the central claim—Table 1 error/speedups, the 10 ns stability statement, and the Figure 3 free-energy comparison—was produced by a model that is not rotationally equivariant between the stored reference features and the evolving coordinates. Over an N=10 step interval, a small molecule can rotate by a non-negligible angle; for atoms several Å from the center, rotations of a few degrees produce displacements comparable to thermal vibrational amplitudes. Without rotation of the reference features, the predicted energy is not invariant under a global rotation of the current configuration, so the forces can exert a net torque. The method is therefore not a conservative Hamiltonian between refresh steps, and the stationary distribution of this piecewise, state-dependent dynamics is not guaranteed to be the Boltzmann distribution of the reference MLFF. The paper argues by citing thermostat results for systems with energy-conservation violation, but does not report the angular-momentum or energy-drift diagnostics that §7 itself says should be checked. The only sampling evidence, Figure 3, is a qualitative overlay of two 5 ns metadynamics free-energy surfaces; no quantitative distance, error bar, or convergence diagnostic is given. The central claim \"accurately sample the ground-truth Boltzmann distribution\" is thus supported by an untested architectural simplification and a visual comparison.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes BoostMD, a surrogate architecture that reuses node features of a reference machine-learned force field (MACE-OFF23-M) computed at a previous MD time step to predict energy and force changes over a short interval, enabling the reference model to be evaluated only every N steps. The authors report speedups up to about 8.6x on a dipeptide dataset, transferability to an unseen dipeptide, and a qualitative Ramachandran free-energy comparison from metadynamics that they interpret as accurate Boltzmann sampling. The manuscript also discusses architectural design choices, including a Kabsch-based rotational reference frame, readout strategies, and conservation properties.","tokens_in":11906,"tokens_out":4378,"duration_ms":40466,"significance":"If the central sampling claim is established, BoostMD would be a practically valuable method: it could accelerate high-accuracy MLFF molecular dynamics by roughly an order of magnitude with only small bias in the sampled equilibrium ensemble. The paper identifies a promising direction—reusing expressive reference features across time steps—and includes concrete speedup/error measurements, a comparison against a similarly fast MACE baseline, and a transferability test. However, the current evidence does not yet support the full claim: the experiments were run with an architectural variant that omits the rotational reference framing that the paper itself argues is necessary for equivariance and conservation, and the only sampling evidence is a visual free-energy comparison without quantitative error bars or convergence diagnostics. These are load-bearing gaps that can likely be addressed within the manuscript's scope.","major_comments":[{"comment":"Section 5.1 and Appendix A.1 introduce a per-atom Kabsch optimal rotation Rbar_i as necessary for rotational equivariance and for conservation of momentum/angular momentum between BoostMD steps, but Appendix C.1 states: \"All experimental results do not use the rotational reference framing during training or inference, due to the associated computational cost.\" Consequently, every reported result—Table 1 RMSE/speedups, the 10 ns stability statement, and the Figure 3 free-energy comparison—was produced by a model that is not rotationally equivariant between the stored reference features and evolving coordinates. The §5.2 claim that the method is \"energy, momentum and angular momentum conserving\" between boost steps is therefore not supported for the tested model. Please rerun the core experiments with the symmetry-preserving framing, or provide quantitative diagnostics (e.g., angular-momentum drift, energy drift, net torque) demonstrating that the omission is benign over the N=10 interval used in the sampling experiment.","section":"§5.1 and Appendix C.1"},{"comment":"The sampling evidence for the central Boltzmann-distribution claim consists solely of a qualitative comparison of Ramachandran free-energy surfaces from a single 5 ns metadynamics run per method (Figure 3). There is no quantitative distribution discrepancy metric, no error bars, no replicate runs, and no convergence analysis. Please add a quantitative comparison (e.g., Kullback-Leibler or Jensen-Shannon divergence between the two free-energy histograms, with bootstrap or block uncertainty estimates), report the metadynamics bias parameters for both runs, and state whether the runs are statistically converged. Without these, the statement that BoostMD \"accurately samples the ground-truth Boltzmann distribution\" is not demonstrated.","section":"§6.2 / Figure 3"},{"comment":"The conservation/sampling argument relies on the cited result that thermostats can ensure correct sampling despite energy-conservation violation (Ref. [38]), but that result is not directly established for this specific state-dependent, piecewise potential whose energy surface is refreshed every N steps and whose forces are not gradients of a fixed global potential. The paper should either provide a theoretical argument applicable to this setting or add empirical diagnostics, such as comparing sampled distributions for different boost intervals N, measuring energy/angular-momentum jumps at refresh events, and checking stationarity and ergodicity of the trajectory. This is load-bearing for the claim that the sampled distribution is Boltzmann.","section":"§5.2"},{"comment":"The headline \"more than eight times faster\" claim in Section 1 and the Abstract is only supported by the Lref=0 variants in Table 1 (speedups 8.4 and 8.6); the Lref=1 variant, which has better force RMSE, achieves only 2.3x. Since the scalar-feature variant is the one that delivers the advertised speedup, the paper should explicitly qualify the speedup claim with this accuracy/speed trade-off and, ideally, show a broader trade-off curve rather than presenting the 8x figure as the general BoostMD performance.","section":"Table 1 and §1"}],"minor_comments":[{"comment":"The statement that all experimental results omit the rotational reference framing is currently buried in an appendix; this is a major limitation of the experiments as presented and should be stated prominently in the main text, alongside a discussion of its implications.","section":"Appendix C.1"},{"comment":"The paper calls the MACE-OFF model the \"ground truth\" for the Boltzmann distribution, but it is itself an approximate MLFF, not the true quantum-mechanical distribution; this should be stated explicitly to avoid overclaiming.","section":"§6.2 / Figure 3"},{"comment":"The effective speedup formula appears as \"1 + (N − 1)/s n\" and is garbled; please define all symbols (n or N) and present the formula cleanly.","section":"§5.2"},{"comment":"There are several typos: \"transitionally\" should be \"translationally\" (§5.1), \"comapred\" should be \"compared\" (Appendix C.1), \"prallelisation\" should be \"parallelisation\" (Appendix C.1), and \"BostMD\" in the Figure 1 caption should be \"BoostMD\".","section":"Throughout"},{"comment":"The sentence \"MACE-OFF23-M achieves an error of 0.85 meV/atom compared to DFT, with BoostMD models showing similar or lower accuracy relative to MACE-OFF23-M\" is confusing: \"lower accuracy\" is ambiguous; please clarify whether BoostMD errors are lower or higher than the reference model's DFT error.","section":"§6.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript has a clear and potentially useful idea, but the current version contains a significant mismatch between the symmetry-preserving architecture described in the main text and the non-equivariant variant actually used in all experiments. The central sampling claim is also supported only by a qualitative figure. Both issues are fixable with additional experiments and analysis, but they are load-bearing for the paper's main claims. I would encourage the editor to request a revision rather than reject, provided the authors address the rotational-framing gap and provide quantitative ensemble comparisons."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThis paper has a genuine new idea: instead of re-evaluating an expensive equivariant MLFF every step, keep the node features from the last full evaluation and train a small surrogate to predict the energy change from the reference configuration. The 8x speedup is real and measured on a 3,375-atom system; the comparison against a same-speed MACE model is fair and shows the feature-reuse is doing real work. The authors also correctly distinguish their approach from multi-scale reference system propagators like Fu et al. [23]. So the core engineering contribution is solid and worth taking seriously.\n\nThe soft spot is the sampling claim. The paper says BoostMD “accurately samples the ground-truth Boltzmann distribution,” but the evidence is one visual overlay of 5 ns metadynamics free energy surfaces with no error bars, no replicates, and no quantitative distribution distance. And as the paper itself notes in Appendix C.1, all experiments were run without the Kabsch rotational reference framing described in Section 5.1 as necessary for rotational equivariance. That means the tested model is not guaranteed to conserve angular momentum between refresh steps, and the energy landscape depends on the stored reference orientation. The authors acknowledge non-conservation and say one should check for drifts, but no such diagnostics are reported. For a claim about equilibrium sampling, that is a load-bearing gap, not a cosmetic one.\n\nThere are also smaller issues: Table 1 is hard to parse (the variant used for Figure 3 is not identified), and the fastest variants have noticeably higher force RMSE than the Lref=1 ones, so the speed/accuracy trade-off deserves more scrutiny. These are fixable.\n\nIf I were the editor, I would send this to peer review. The idea is novel and the negative result would also be interesting. But I would expect the referee to demand quantitative sampling metrics, drift diagnostics, and either a test with the reference framing or a justification for why omitting it is harmless.\n\nBest.","headline":"A promising surrogate idea with a real 8x speedup, but the Boltzmann-sampling claim rests on a qualitative figure and an untested architectural simplification.","tokens_in":12475,"tokens_out":3095,"would_cite":false,"duration_ms":29289,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"BoostMD accelerates machine-learned force-field molecular dynamics by more than eightfold using previous node features.","keywords":["machine learning force fields","molecular dynamics acceleration","equivariant message passing","surrogate model","Boltzmann sampling","feature reuse","metadynamics"],"falsifier":"Run BoostMD with a larger reference interval N or with a molecular system that rotates substantially between frames, then compare the sampled dihedral free-energy surface to the reference force field; if the free-energy surfaces diverge or the force error grows with the angle of rotation between reference and current frames, the small-rotation assumption fails.","tokens_in":11421,"feed_emoji":"⚡","tokens_out":6216,"duration_ms":51715,"temperature":0.7,"pith_summary":"This paper proposes BoostMD, a surrogate model that makes machine-learning force-field molecular dynamics faster by not recomputing the expensive internal features at every time step. The full reference force field is evaluated every N steps, and a small, fast BoostMD model predicts the energy change for the intervening steps from the previous node features and the atoms' positional changes. The authors claim this yields more than an eightfold speedup over the reference model, transfers to dipeptides not seen in training, and samples the correct equilibrium (Boltzmann) distribution, including under metadynamics. If true, long-timescale simulations of peptides and other organic molecules would become practical at near-quantum accuracy.","feed_headline":"BoostMD runs ML force-field MD eight times faster","feed_subtitle":"Reusing node features from recent steps lets a small model carry most steps while preserving the sampled ensemble.","key_machinery":"The central object is the BoostMD equivariant message-passing layer. For each atom it takes the current positions, the reference positions from the last full evaluation of the reference force field, and the reference node features, forming change vectors that are aligned by the optimal rotation between the reference and current neighborhoods (computed with the Kabsch algorithm). Spherical harmonics of the current displacement and of the change vector are combined with the rotated reference features through learnable tensor products to build a message basis, then a MACE-style product basis creates many-body features, and a small readout predicts the per-atom energy change. A multiplicative factor on the displacement ensures zero energy change when nothing has moved. In the reported experiments a single-layer BoostMD model with a five-angstrom receptive field is used, with the reference model evaluated every ten steps, and the rotational reference framing is omitted during training and inference.","core_discovery":"The discovery is that the internal per-atom features of an equivariant machine-learning force field change slowly during a molecular dynamics trajectory, so a lightweight surrogate conditioned on the features from a recent reference evaluation can accurately predict energy changes and carry most time steps. Running the reference only every N steps and BoostMD in between yields more than eightfold speedup on a large system with over three thousand atoms, with energy and force errors comparable to the reference model's own error against density functional theory. The surrogate reproduces the reference free-energy surface for an unseen dipeptide and remains stable over ten nanoseconds, demonstrating that the accelerated trajectory still samples the correct equilibrium ensemble.","pith_inferences":["If the omission of rotational reference framing is stressed with a larger reference interval or highly rotatable molecules, the reported sampling accuracy may degrade; the paper does not test this.","The method's energy is not exactly conserved at the reference-update steps, and the authors rely on thermostats to maintain correct sampling; combining BoostMD with multiple-time-step integrators or adaptive reference intervals could control drift.","Because the surrogate is conditioned on any machine-learning force field's node features, the same BoostMD recipe could likely be applied to other equivariant force fields, not just the MACE-based reference.","The eightfold speedup may be a conservative figure: the parallel producer/consumer evaluation described but not implemented would remove the serial bottleneck and could deliver the full model speedup."],"forward_implications":["Molecular dynamics with machine-learning force fields can run about eight times faster, bringing microsecond-timescale simulations of organic molecules within practical reach.","The acceleration compounds with enhanced sampling: BoostMD carries metadynamics and reproduces the reference free-energy surface for a molecule outside its training set.","The boost model is small and local, so it should parallelize across GPUs more easily than the full reference model, easing scale-up to large systems.","The surrogate's error is close to the reference model's own error against density functional theory, so the practical accuracy bottleneck becomes the reference force field rather than the acceleration wrapper.","Transferability to unseen dipeptides suggests the scheme may generalize beyond its training molecules, though the paper demonstrates this only for dipeptides."],"supporting_citations":[{"why":"The reference machine-learning force field whose node features BoostMD reuses, whose molecular dynamics trajectories generate the training data, and which serves as the ground truth for energy and force comparisons.","marker":"[1]"},{"why":"The equivariant message-passing architecture whose product basis and tensor contractions form the core of BoostMD layers.","marker":"[11]"},{"why":"The Kabsch algorithm used to compute the optimal rotation aligning reference and current neighborhoods for rotational equivariance.","marker":"[34]"},{"why":"Supplies the dipeptide subset used to seed the molecular dynamics trajectories from which BoostMD is trained.","marker":"[39]"},{"why":"Provides the alanine dipeptide structure used as the unseen test molecule for sampling and free-energy comparisons.","marker":"[40]"},{"why":"Computes the free-energy surfaces from the metadynamics runs used to compare BoostMD with the reference model.","marker":"[42]"},{"why":"Supports the claim that thermostats can maintain correct sampling even when energy is not strictly conserved, underpinning the equilibrium-sampling claim.","marker":"[38]"}],"fun_headline_variants":["ML force-field MD 8x faster via feature reuse","Reusing ML force-field features speeds MD 8x","BoostMD: 8x faster MD without losing accuracy","Feature reuse gives 8x faster molecular dynamics","Old features, new speeds: 8x faster ML MD"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported experiments skip the rotational alignment step that the architecture's equivariance argument requires, so the results depend on molecular rotations between reference updates being small enough that omitting the alignment barely changes the forces and the sampled distribution.","fun_headline_variants_meta":{"raw":{"variants":["ML force-field MD 8x faster via feature reuse","Reusing ML force-field features speeds MD 8x","BoostMD: 8x faster MD without losing accuracy","Feature reuse gives 8x faster molecular dynamics","Old features, new speeds: 8x faster ML MD"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000191,"raw_usage":{"total_tokens":1330,"prompt_tokens":916,"completion_tokens":414,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":532,"completion_tokens_details":{"reasoning_tokens":334}},"tokens_in":532,"tokens_out":414,"duration_ms":4011,"temperature":1.0,"reasoning_tokens":334,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:15:23.190177+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run BoostMD with a larger reference interval N or with a molecular system that rotates substantially between frames, then compare the sampled dihedral free-energy surface to the reference force field; if the free-energy surfaces diverge or the force error grows with the angle of rotation between reference and current frames, the small-rotation assumption fails.","supporting_citations":[{"cited_title":"P., Simm, G","cited_arxiv_id":null,"evidence_quote":"The equivariant message-passing architecture whose product basis and tensor contractions form the core of BoostMD layers."},{"cited_title":"A solution for the best rotation to relate two sets of vectors","cited_arxiv_id":null,"evidence_quote":"The Kabsch algorithm used to compute the optimal rotation aligning reference and current neighborhoods for rotational equivariance."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the dipeptide subset used to seed the molecular dynamics trajectories from which BoostMD is trained."},{"cited_title":"& Rao, F","cited_arxiv_id":null,"evidence_quote":"Provides the alanine dipeptide structure used as the unseen test molecule for sampling and free-energy comparisons."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Computes the free-energy surfaces from the metadynamics runs used to compare BoostMD with the reference model."},{"cited_title":"& Leimkuhler, B","cited_arxiv_id":null,"evidence_quote":"Supports the claim that thermostats can maintain correct sampling even when energy is not strictly conserved, underpinning the equilibrium-sampling claim."}],"review_version":1}