REVIEW 2 major objections 2 minor
Approximate label symmetries improve data efficiency
T0 review · 2 major / 2 minor · reviewed 2026-06-29 · grok-4.3
Pith's one-line read Approximate label symmetries improve machine learning scaling laws for electron densities and molecular energies.
desk verdict Approximate label symmetries improve scaling for these quantum ML models, with the Hessian correction addressing the main error on convex wells. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Label symmetries applied to augment training data, with a Hessian correction for approximate cases in convex potential wells.
What would settle it
Check whether learning curves using approximate label symmetries plateau exactly at heights predicted by the measured degree of symmetry breaking, and whether the Hessian correction removes the dominant error term in convex wells of the water potential energy surface.
Extended reading notes
Core claim
Exploiting exact as well as approximate label symmetries can benefit scaling laws. ML models of the s, p, d orbital densities of the hydrogen atom, the three vibrational normal modes of the water molecule, and its full 3D potential energy hypersurface exhibit superior learning curves. When label symmetries are not exact the same principles govern learning behavior up to convergence floors set by the degree of approximation. For convex wells a Hessian-based correction suppresses the leading symmetry-breaking error in augmented labels.
Load-bearing premise
Scaling principles continue to govern learning behavior when label symmetries are approximate, with performance floors set solely by the degree of approximation.
Editorial extensions
If this is right
- ML models for electron density and potential energies achieve improved generalization efficiency.
- Learning curves follow the same scaling principles for approximate symmetries until limited by approximation degree.
- Hessian correction suppresses leading symmetry-breaking error for convex wells in molecular potential energy surfaces.
Reading between the lines
- The method may extend to other quantum chemistry tasks where near-symmetries appear in molecular properties.
- It could lower data requirements for training models on systems with partial symmetry.
- Similar augmentation might apply to learning curves in other domains with approximate invariances.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript claims that exploiting exact and approximate label symmetries improves data scaling in machine learning models for physical systems. It illustrates this with s/p/d orbital densities of the hydrogen atom, the three vibrational normal modes of water, and the full 3D potential energy surface of water. Resulting models show superior learning curves. For approximate symmetries the same scaling principles are said to hold up to convergence floors set by the degree of approximation; a Hessian-based correction is proposed to suppress the leading symmetry-breaking error for convex wells in the PES.
Significance. If the quantitative results hold, the work provides a practical route to data-efficient ML for quantum chemistry by relaxing the requirement for exact symmetries while retaining scaling benefits. The concrete examples (H-atom orbitals, water modes, PES) and the explicit Hessian correction for convex wells are strengths; the absence of free parameters or ad-hoc axioms in the core argument is also positive.
major comments (2)
- [Abstract, §3–4] Abstract and §3–4 (results on learning curves): the central empirical claims of superior learning curves and the quantitative effect of the Hessian correction are asserted without reported error bars, training-set sizes, number of independent runs, or exclusion criteria for the augmented labels. This information is load-bearing for the scaling-law assertions and must be supplied before the claims can be evaluated.
- [§4.2] §4.2 (Hessian correction): the statement that the correction 'suppresses the leading symmetry-breaking error' for convex wells is presented without an explicit derivation showing that higher-order terms remain negligible across the tested range of displacements; a short expansion or numerical check of the neglected terms is needed to support the claim.
minor comments (2)
- [Figures 2–3] Figure 2 and 3 captions should state the precise definition of the 'augmented label' set and the metric used for the learning curves (e.g., MAE on density or energy).
- [Introduction] The introduction should define 'label symmetry' at first use rather than relying on the later technical sections.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback. The two major comments identify important omissions in statistical reporting and justification of the Hessian correction; both are addressable by revision.
read point-by-point responses
-
Referee: [Abstract, §3–4] Abstract and §3–4 (results on learning curves): the central empirical claims of superior learning curves and the quantitative effect of the Hessian correction are asserted without reported error bars, training-set sizes, number of independent runs, or exclusion criteria for the augmented labels. This information is load-bearing for the scaling-law assertions and must be supplied before the claims can be evaluated.
Authors: We agree that these experimental details are required to evaluate the scaling claims. The revised manuscript now reports error bars obtained from ten independent training runs per data point, lists the precise training-set cardinalities used for each learning curve, and states the exclusion criterion applied to augmented labels (augmented labels were retained only when the symmetry-breaking residual lay below a threshold set by the norm of the Hessian at the reference geometry). These additions appear in §§3–4 together with a brief methods paragraph; the abstract has been updated to note the statistical controls. revision: yes
-
Referee: [§4.2] §4.2 (Hessian correction): the statement that the correction 'suppresses the leading symmetry-breaking error' for convex wells is presented without an explicit derivation showing that higher-order terms remain negligible across the tested range of displacements; a short expansion or numerical check of the neglected terms is needed to support the claim.
Authors: We accept that an explicit expansion strengthens the claim. The revised §4.2 now contains a short Taylor expansion of the potential about the equilibrium geometry, demonstrating that the leading symmetry-breaking term is quadratic in the displacement vector and is exactly cancelled by the Hessian correction, while cubic and higher contributions scale as O(‖δ‖³). For the displacement magnitudes employed in the water PES experiments (‖δ‖ ≤ 0.1 Å), a supplementary numerical check shows that the neglected terms remain below 5 % of the quadratic residual. This material has been inserted as a new paragraph with an accompanying figure panel. revision: yes
Circularity Check
No significant circularity in derivation chain
full rationale
The paper reports empirical results on concrete systems (H atom orbital densities, water vibrational modes, and 3D PES) showing improved learning curves when exploiting exact and approximate label symmetries. Claims about scaling behavior and Hessian corrections for convex wells are presented as observations from these examples, with performance floors tied to approximation degree. No load-bearing step reduces by construction to fitted inputs, self-citations, or renamed known results; the central claims remain independent of the enumerated circularity patterns.
Assumptions & free parameters
assumptions (1)
- domain assumption Approximate label symmetries still improve scaling laws up to a convergence floor determined by the degree of approximation
Cite this review
Pith. "Pith review of Approximate label symmetries improve data efficiency." pith.science (2026). https://pith.science/paper/YYREAJT2
@misc{pith2026260528238,
author = {Pith},
title = {Pith review of: Approximate label symmetries improve data efficiency},
year = {2026},
howpublished = {\url{https://pith.science/paper/YYREAJT2}},
note = {Machine review of arXiv:2605.28238}
}
read the original abstract
Enforcing feature symmetries in machine learning (ML) models is a common strategy to mitigate data scarcity. Confirming expectations from statistical learning theory, we show that exact, as well as approximate, label symmetries can also improve data efficiency. We illustrate the idea for the s, p, d orbital densities of the electron in the hydrogen atom, for the three vibrational normal modes of the water molecule, and for its full 3D potential energy hypersurface. Resulting ML models of electron density and potential energies exhibit superior learning curves, demonstrating improved generalization efficiency. We observe that learning curves similarly improve even when label symmetries are not exact - up to the convergence floors set by the degree to which the symmetry is approximate. Further improvements are obtained for approximate label symmetries in the molecular potential energy surface, using a Hessian-based correction that suppresses the leading order term in the error.
Figures
Reviewed June 29, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.