{"id":"a5f1bd8b-154b-48d9-91bf-1ba8bd820c33","arxiv_id":"2506.07579","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A machine learning interatomic potential trained only on total energies and forces learns per-bond energy components that match experimental bond dissociation energies, an ability the authors call emergent.","lead":"An analysis framework called E3D splits the energy predicted by a machine-learned interatomic potential into per-bond pieces, and those pieces line up with measured bond dissociation energies even though the model was never trained on bond energies. The findings suggest that big universal potentials carry chemically meaningful internal representations, which could help scientists decide what training data to add when reaction barrier accuracy stalls.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"D_ij as BDE depends on an untested gauge choice: the per-edge decomposition of Eq. 6 is not identifiable, and no ablation isolates it from the cutoff radius or the forced mu=0, sigma=1 normalization.","rationale":"The reader's weakest_assumption is exactly the gauge / identifiability problem: D_ij is an arbitrary decomposition of E_coh, and the equality at Eq. 6 holds by construction. The stress-test pass finds this to be the single most load-bearing concern. It is concrete, internal to the paper's argument, and testable. Re-analyzing Fig. 2 and the BDE metrics (delta_BDE, sigma_BDE) through gauge ablations would settle whether the central claim holds. The reader's verdict of CONDITIONAL is appropriate; the concern does not by itself warrant REJECT without performing the proposed test, but it does warrant a clear condition on the authors to provide per-bond statistics, error bars, and at least one alternative-decomposition / ablation test. The paper's own limitations (Section 5.4) and its lack of released E3D code further support keeping the verdict conditional rather than upgrading to ACCEPT. The paper's other evidence — the fact that D_ij distributions align with BDE trends across multiple datasets, the self-consistent sum rule, and the entropy analysis — is credible and non-circular, but it does not by itself remove the identifiability concern. No ad hominem; the critique targets the argument's assumption, not the authors. The concrete test proposed (gauge / cutoff / normalization ablations) is specific and feasible with the existing infrastructure described in Appendix A.1.","tokens_in":18699,"tokens_out":1851,"duration_ms":19502,"concrete_test":"Train the same Allegro architecture on the same SPICE 2 data under several gauge variations: (a) per-atom (single-center) energy decomposition, (b) different radial cutoffs (e.g., 3.0, 4.0, 6.5 Å), (c) nonzero per-species shifts mu with E_system as target, and (d) random re-partitioning of a fixed total energy across edges. Then compute the D_ij distributions and their agreement with reference BDEs for each setting. If agreement with BDEs persists only under the specific E3D normalization (mu=0, sigma=1, 5.2 Å cutoff) and changes substantially under any alternative, the claim that the model spontaneously learns BDE is not established; if agreement is invariant across gauges, the concern is resolved.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim — that Allegro spontaneously learns bond dissociation energies — rests on interpreting the edge-wise decomposition of E_coh as physically meaningful bond energies. That interpretation is exactly what is not established. The decomposition in Eq. 5-7 is not identifiable: any function of the local environment can be re-partitioned among edges (and between edge and single-atom terms) without changing total energies or forces. The authors force mu=0, sigma=1 so that Eq. 6 holds by construction, but this only guarantees that the sum of D_ij equals E_coh; it does not guarantee that a particular D_ij equals a bond dissociation energy. The agreement with reference BDEs in Fig. 2 could therefore be a consequence of the imposed normalization and the 5.2 Å cutoff selecting chemically relevant neighborhoods, rather than evidence that the model's internal representation tracks bond energetics. The paper never ablates the gauge: no comparison with alternative decompositions (e.g., per-atom energies, other cutoffs, other sigma/mu settings, or a model trained without the forced normalization), and no per-bond-type quantitative statistics with uncertainties are shown. Section 5.4 limits the claim to Allegro, but even for Allegro the identifiability problem remains. The MatPES-trained result (Fig. S1) is only qualitative and does not resolve the gauge question. An additional concern is that the E_a 'scaling wall' is presented as a causal consequence of representational entropy limits (Section 3.4), but only correlational evidence is given; the causal claim is not load-bearing for the main BDE claim, so the gauge concern is the primary one.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Edge-wise Emergent Decomposition (E3D), an analysis framework that decomposes the total cohesive energy of an Allegro machine-learning interatomic potential into symmetric (D_ij) and asymmetric (A_ij) per-edge components. By setting the per-species shift and scale parameters to 0 and 1 and training on E_coh, the model energy becomes a direct sum of edge energies. The authors report that the resulting D_ij distributions for C-H, N-H, O-H, C-C, C-N, C-O, N-O, and N-N bonds align with literature bond dissociation energies (BDEs) without explicit supervision. They also analyze 2D D-A histograms and their Shannon entropy H2D, observing that H2D decreases with increasing training data and that C-N and C-O bonds in transition states show weaker entropy reduction, coinciding with the observed 'scaling wall' for activation-energy prediction. A hybrid SPICE+MatPES training set is shown to improve E_a predictions. The paper interprets these results as evidence of 'emergent BDE' and as a diagnostic for scaling limitations.","tokens_in":18948,"tokens_out":7439,"duration_ms":83981,"significance":"If the central claim is valid, E3D provides a valuable, interpretable window into the internal representations of equivariant MLIPs and a practical metric for monitoring learning and diagnosing scaling failures. The claim that a model trained only on total energies and forces spontaneously develops quantitatively accurate bond dissociation energies is striking and, if true, would be an important contribution to the emerging understanding of MLIP generalization. The paper also proposes a concrete data-diversity strategy (hybrid training) that appears to improve reactive-property prediction. However, the significance is currently conditional: the identification of D_ij with physical BDEs rests on a gauge choice that is not ablated, and the quantitative evidence is weakened by the absence of per-bond statistics and uncertainties. These issues must be resolved for the results to be fully convincing.","major_comments":[{"comment":"The central identification of D_ij with bond dissociation energy is not identifiable from the training objective. Setting mu=0 and sigma=1 makes Eq. (6) hold by construction, but any zero-sum re-partitioning of edge energies (e.g., adding a flow that cancels in the sum over neighbors) leaves E_coh and the forces invariant while changing D_ij. The paper provides no ablation of this gauge: no comparisons across random seeds, different cutoff radii, alternative decompositions (such as per-atom residuals), or models trained without the forced normalization. In addition, the definition of E_coh requires a choice of one-body reference energies E^(1)_i, which is not specified; if arbitrary offsets are used, the absolute scale of D_ij is correspondingly arbitrary. Without these controls, the agreement with reference BDEs in Fig. 2 could be an artifact of the specific gauge selected by the architecture and initialization rather than evidence that the model internalizes bond energetics.","section":"2.2 (Eqs. 5-7)"},{"comment":"The claim of quantitative agreement is not supported by the reported statistics. Fig. 2 shows distributions of D_ij but no numerical comparison per bond type; Fig. 3 reports only pooled metrics (Delta_BDE and sigma_BDE) averaged over all bond types, with no error bars across training seeds or test-set splits. Pooling can mask systematic deviations that cancel across bond types. The paper should report per-bond-type mean D_ij with standard deviations and confidence intervals against each reference BDE (e.g., a parity plot or a table analogous to Table S2), and ideally repeat training with a few random seeds to establish the stability of the decomposition. Without these, 'quantitatively agree' in the abstract and Section 3.1 is overstated.","section":"3.1 and Fig. 3"},{"comment":"The conclusion that E3D 'identifies the root cause of the scaling wall' is not justified by the evidence presented. The finding is a correlation: H2D for C-N and C-O bonds in transition-state structures plateaus in the same data range where E_a MAE saturates. No causal test or intervention is provided, and the same framework yields an apparent inconsistency: the hybrid dataset improves E_a accuracy while increasing H2D (Section 4.3, Table S4, and Figure S2). This suggests the relationship between representational entropy and reactive accuracy is more complex than a monotonic 'low entropy implies good reactivity' narrative. The authors should either soften the causal language to 'hypothesis' or provide a direct test (e.g., fine-tuning with selected TS structures and showing that the entropy change precedes E_a improvement).","section":"3.4 and 4.3"}],"minor_comments":[{"comment":"The notation in Eq. (5) is unclear: the per-species scale sigma_Zi appears in the outer sum and sigma_ZiZj in the inner sum. Please clarify the relationship between these scales and the statement that both are set to 1. Also specify how the one-body reference energies for E_coh are computed (see also Major Comment 1).","section":"2.2"},{"comment":"The claim of 'robust across diverse training datasets' is based on a single additional dataset (MatPES), with only qualitative agreement (Fig. S1). Please qualify the wording or add more datasets to support the generality claim.","section":"3.2"},{"comment":"The 'sharp decrease in entropy around 10^5 data size' is identified by visual inspection. Please provide a statistical test or at least error bars to support the claim of a discrete transition point rather than a smooth trend.","section":"3.4 / Fig. 5"},{"comment":"The manuscript contains numerous typos and formatting issues, e.g., 'trainig' in Section 4.1, incomplete sentences ('and s.' in Section 3.1), and a missing y-axis label in Figure 2. A careful proofreading pass is needed.","section":"Throughout"},{"comment":"The hyperparameters are said to be 'based on information extracted from arXiv:2502.06073'. Please clarify whether the models used in this paper share the exact architecture of Allegro-FM or whether modifications (e.g., the final edge-energy MLP with a single hidden layer and no nonlinearity) were made. This detail is important for reproducibility.","section":"A.1"},{"comment":"The reference BDEs are a mix of experimental, theoretical, and interpolated values, but Section 2.3 states that reference values are from '[36]' alone. Please be explicit about the source of each value, especially for bond orders where interpolation was used.","section":"Table S2 / Sec. 2.3"}],"recommendation":"major_revision","confidential_remarks":"This is an interesting and timely submission. The central claim is provocative but rests on an untested gauge assumption that must be addressed. The main revisions I would require are: (1) an ablation study of the edge decomposition (seeds, cutoff, reference energies); (2) per-bond-type statistics with uncertainties; (3) softening of the causal claims about the scaling wall. If these are addressed, the paper could be a strong contribution. The paper fits the scope of the journal, though the emphasis on 'emergence' may be better framed as 'interpretable decomposition' unless the gauge issue is resolved."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe thing to know about this paper is that it introduces a genuinely new diagnostic—an edge-wise decomposition of a trained MLIP's energy into symmetric and asymmetric bond energies—and shows that for Allegro, the symmetric part lines up with literature bond dissociation energies without any BDE supervision. I don't know of prior work that does this. The authors also get a concrete, useful result: adding MatPES to SPICE training improves Ea MAE from 0.58 to 0.44 eV.\n\nThe novelty and the qualitative result are real. The BDE trends in Figure 2, including bond-order distinctions for C–C, C–N, and C–O, are the kind of thing that makes you stop and think. And the entropy-based analysis of D–A maps is a reasonable way to quantify representational sharpening.\n\nNow the soft spots. The gauge worry is legitimate. The decomposition in Eq. 6 is enforced by setting mu=0 and sigma=1, which guarantees the sum over D_ij equals E_coh but does not guarantee each D_ij is a bond dissociation energy. The agreement could depend on the 5.2 Å cutoff and on the forced normalization. The paper does not ablate the gauge—no alternative decompositions, no tests at other cutoffs, no comparison of different initializations. And while Figure 2 is suggestive, there are no per-bond-type quantitative statistics with uncertainties; Figure 3 reports pooled metrics with no error bars across seeds. So the 'emergence' claim is credible but undersupported. I'd also tone down the causal language around the Ea scaling wall: the entropy drop near 1e5 structures is correlational, and the hybrid-data result actually increases H2D while improving Ea, which the paper notes but doesn't fully reconcile.\n\nThe paper is honest about its scope—Section 5.4 says only Allegro is tested. No code or analysis scripts are released, so exact reproduction is currently not possible.\n\nWho gets value: anyone working on MLIP interpretability, foundation models, or reaction barrier prediction. It's a solid paper to send to review, with clear requests: per-bond statistics with uncertainties, gauge ablations, and code release. The central idea is worth referee time, not a desk reject.","headline":"A new diagnostic with a plausible but not fully pinned-down claim: per-edge energy decomposition yields BDE-like values, but gauge non-identifiability and missing per-bond statistics leave the emergence claim under-supported.","tokens_in":19595,"tokens_out":3010,"would_cite":true,"duration_ms":35829,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A universal machine-learning interatomic potential trained only on total energies and forces spontaneously learns per-bond dissociation energies that match literature values.","keywords":["machine learning interatomic potentials","bond dissociation energy","emergent ability","energy decomposition","E(3)-equivariant neural networks","Allegro","scaling laws","transition states"],"falsifier":"Train the same Allegro architecture on the same data with several random seeds and with at least one alternative radial cutoff, such as 4.0 Å instead of 5.2 Å, and compare the resulting $D_{ij}$ means to the reference bond dissociation energies; if the agreement scatters by more than the quoted distribution widths, or if an energy-preserving reparametrization of edge energies systematically shifts $D_{ij}$ while leaving total energies unchanged, then the correspondence between $D_{ij}$ and bond dissociation energy is not uniquely learned from the data.","tokens_in":18465,"feed_emoji":"🔗","tokens_out":5205,"duration_ms":62647,"temperature":0.7,"pith_summary":"This paper argues that a machine-learning interatomic potential, trained only on global energies and forces, can spontaneously develop chemically meaningful per-bond energy terms without ever being shown a bond dissociation energy. The authors modify the Allegro architecture so that its output is a sum of per-edge energies, then train it on molecular and bulk datasets. They find that the symmetric part of each edge energy, $D_{ij} = \\varepsilon_{ij} + \\varepsilon_{ji}$, quantitatively matches experimental bond dissociation energies for bonds such as C-C, C-N, C-O, and C-H, including the correct ordering of single, double, and triple bonds. The effect persists even when training only on inorganic bulk materials data, and it sharpens as training data grows. The same decomposition also explains why activation-energy prediction plateaus: internal representations of C-N and C-O bonds in transition states remain diffuse even as stable-structure representations sharpen.","feed_headline":"Neural force field learns real bond energies it was never shown","feed_subtitle":"Fed only total energies and forces, an AI interatomic model assigns per-bond values matching measured dissociation energies.","key_machinery":"The central machinery is the Edge-wise Emergent Decomposition (E3D) framework, built on the Allegro architecture's per-edge energy outputs $\\varepsilon_{ij}$. By fixing the model's scale and shift parameters so that the cohesive energy is exactly the sum of edge energies, the paper defines the symmetric component $D_{ij} = \\varepsilon_{ij} + \\varepsilon_{ji}$ and the asymmetric component $A_{ij} = \\varepsilon_{ij} - \\varepsilon_{ji}$. The sum of $D_{ij}$ over all pairs equals the cohesive energy by construction, and $D_{ij}$ is interpreted as the bond energy, with $A_{ij}$ capturing local-environment asymmetry; two-dimensional $D_{ij}$-$A_{ij}$ histograms and their Shannon entropy serve as the diagnostic that tracks how sharply each bond type is represented as training data grows.","core_discovery":"The central claim is that a universal machine learning interatomic potential trained only on total energies and forces spontaneously acquires per-edge energy components $D_{ij} = \\varepsilon_{ij} + \\varepsilon_{ji}$ whose means agree with literature bond dissociation energies across diverse bond types and bond orders, with no explicit supervision on bond energies. The authors demonstrate this agreement for C-H, N-H, O-H, C-C, C-N, C-O, N-O, and N-N bonds in a model trained on the SPICE 2 molecular dataset, and show that the agreement is qualitatively preserved when training exclusively on the MatPES inorganic bulk dataset. They further claim that data scaling organizes the learned representation, as measured by Shannon entropy of two-dimensional $D_{ij}$-$A_{ij}$ maps, and that the persistently high entropy of C-N and C-O transition-state representations coincides with the observed \"scaling wall\" in activation energy prediction. The paper thus proposes that decomposable learning of local bond energetics is an emergent ability of scaled equivariant networks, not an artifact of explicit BDE supervision.","pith_inferences":["A direct test of the framework's physical identifiability would be to train the same architecture with several random seeds and with different radial cutoffs; if the resulting $D_{ij}$ distributions all overlap within their quoted widths, the per-bond energies are robust, whereas wide scatter would indicate that the agreement with bond dissociation energies is partly a modeling convention.","An energy-preserving reparametrization of the edge energies, such as adding a divergence-free flow to $\\varepsilon_{ij}$, could reveal whether $D_{ij}$ is gauge-invariant; total energies would stay unchanged while per-bond values shift, and a shift would show that the decomposition is not uniquely determined by the training objective.","The entropy transition observed near $10^5$ training structures suggests a concrete, testable prediction: models with insufficient capacity should not exhibit the sharp entropy drop even when given the full dataset, which could be checked by scanning tensor sizes across the data-size axis.","The same decomposition logic could be extended to other additive properties, such as decomposable dipole moments or partial charges, turning the E3D analysis from a bond-energy diagnostic into a general probe of what physics an equivariant potential has learned."],"forward_implications":["If the claim is correct, machine learning interatomic potentials trained purely on energies and forces carry chemically interpretable per-bond information that can be read out directly without retraining or explicit labels.","The persistence of the emergent bond-dissociation energies after inorganic-only training implies that covalent bond energetics is a transferable representation that emerges from local atomic environments rather than from memorized molecular topologies.","The activation-energy scaling wall can be diagnosed and partially circumvented: training on a hybrid SPICE 2 plus MatPES dataset lowers the activation-energy mean absolute error from 0.58 eV to 0.44 eV, and the change is correlated with a reshaping of the $D_{ij}$-$A_{ij}$ maps for transition states.","The Shannon entropy of the $D_{ij}$-$A_{ij}$ maps provides a data-driven signal for when a bond type's internal representation is still diffuse, offering a way to spot dataset limitations for reactive structures before running expensive reference calculations.","The emergent per-bond energies could serve as collective variables in enhanced-sampling simulations and as real-time bond-dissociation progress monitors in molecular dynamics, as the paper suggests in its outlook."],"supporting_citations":[{"why":"Supplies the Allegro architecture whose per-edge energy structure is modified to obtain the $D_{ij}$ and $A_{ij}$ decomposition.","marker":"[34]"},{"why":"SPICE 2 is the primary molecular training dataset of roughly two million conformations used to train the main models.","marker":"[14]"},{"why":"Lange's Handbook of Chemistry provides the reference bond dissociation energies against which learned $D_{ij}$ values are compared.","marker":"[36]"},{"why":"MatPES is the inorganic bulk dataset used to test whether emergent bond energies transfer from non-molecular training data and to build the hybrid training set.","marker":"[20]"},{"why":"Transition1x supplies the benchmark reaction and transition-state structures used to evaluate activation-energy scaling and to map representation entropy.","marker":"[33]"},{"why":"NequIP is the training framework used to train the modified Allegro models with the standardized normalization parameters.","marker":"[10]"},{"why":"The Atomic Cluster Expansion formalism underlies Allegro's ability to express many-body interactions as a sum over edge energies.","marker":"[9]"},{"why":"Prior hybrid-data training of Allegro-FM motivates the paper's data-diversity strategy and serves as a baseline in the hybrid dataset comparison.","marker":"[31]"}],"fun_headline_variants":["AI force field infers bond energies it was never trained on","ML interatomic potential spontaneously learns bond strengths","Emergent bond energies from an energy-only neural network","Neural network discovers bond energies without supervision"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that $D_{ij}$ is a physically meaningful bond energy assumes that the per-edge decomposition of the cohesive energy is identifiable, meaning that the specific value the network assigns to a particular C-C or C-N pair is the bond's true dissociation energy rather than an artifact of the chosen normalization, cutoff radius, or network initialization.","fun_headline_variants_meta":{"raw":{"variants":["AI force field infers bond energies it was never trained on","ML interatomic potential spontaneously learns bond strengths","Emergent bond energies from an energy-only neural network","Neural network discovers bond energies without supervision"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000185,"raw_usage":{"total_tokens":1326,"prompt_tokens":955,"completion_tokens":371,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":571,"completion_tokens_details":{"reasoning_tokens":310}},"tokens_in":571,"tokens_out":371,"duration_ms":5024,"temperature":1.0,"reasoning_tokens":310,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:31:18.331171+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same Allegro architecture on the same data with several random seeds and with at least one alternative radial cutoff, such as 4.0 Å instead of 5.2 Å, and compare the resulting $D_{ij}$ means to the reference bond dissociation energies; if the agreement scatters by more than the quoted distribution widths, or if an energy-preserving reparametrization of edge energies systematically shifts $D_{ij}$ while leaving total energies unchanged, then the correspondence between $D_{ij}$ and bond dissociation energy is not uniquely learned from the data.","supporting_citations":[{"cited_title":"Learning local equivariant representations for large-scale atomistic dynamics.Nat","cited_arxiv_id":null,"evidence_quote":"Supplies the Allegro architecture whose per-edge energy structure is modified to obtain the $D_{ij}$ and $A_{ij}$ decomposition."},{"cited_title":"Nutmeg and SPICE: Models and data for biomolecular machine learning.J","cited_arxiv_id":null,"evidence_quote":"SPICE 2 is the primary molecular training dataset of roughly two million conformations used to train the main models."},{"cited_title":"McGraw-Hill Education, Columbus, OH, 17 edition, October 2016","cited_arxiv_id":null,"evidence_quote":"Lange's Handbook of Chemistry provides the reference bond dissociation energies against which learned $D_{ij}$ values are compared."},{"cited_title":"A foundational potential energy surface dataset for materials.arXiv [cond-mat.mtrl-sci], March 2025","cited_arxiv_id":null,"evidence_quote":"MatPES is the inorganic bulk dataset used to test whether emergent bond energies transfer from non-molecular training data and to build the hybrid training set."},{"cited_title":"Transition1x - a dataset for building generalizable reactive machine learning potentials.Sci","cited_arxiv_id":null,"evidence_quote":"Transition1x supplies the benchmark reaction and transition-state structures used to evaluate activation-energy scaling and to map representation entropy."},{"cited_title":"E(3)-equivariant graph neural networks for data- efficient and accurate interatomic potentials.Nat","cited_arxiv_id":null,"evidence_quote":"NequIP is the training framework used to train the modified Allegro models with the standardized normalization parameters."},{"cited_title":"Atomic cluster expansion for accurate and transferable interatomic potentials.Phys","cited_arxiv_id":null,"evidence_quote":"The Atomic Cluster Expansion formalism underlies Allegro's ability to express many-body interactions as a sum over edge energies."},{"cited_title":"Allegro-FM: Towards equivariant foundation model for exascale molecular dynamics simulations.arXiv [physics.comp-ph], February 2025","cited_arxiv_id":null,"evidence_quote":"Prior hybrid-data training of Allegro-FM motivates the paper's data-diversity strategy and serves as a baseline in the hybrid dataset comparison."}],"review_version":1}