{"id":"cdd697b7-7e42-4046-8df0-ece45f2d1666","arxiv_id":"2605.28238","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Approximate label symmetries improve scaling laws for ML models of hydrogen orbital densities, water vibrational modes, and 3D potential energy surfaces, with a Hessian correction for approximate cases.","lead":"The paper shows that using exact and approximate label symmetries in machine learning models for electron densities and molecular potential energies produces better learning curves and data scaling. A smart generalist might read it to see practical ways to train accurate models with less expensive quantum data.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"Reader's weakest_assumption correctly flags the extrapolation from exact to approximate symmetries and the reliability of the Hessian term. Because the supplied text supplies no counter-evidence or contradictory result, the UNVERDICTED status is retained; the proposed check would directly test the load-bearing scaling assumption if full results become available.","tokens_in":1633,"tokens_out":254,"duration_ms":15684,"concrete_test":"Extract the reported scaling exponents (or equivalent learning-curve slopes) for the approximate-symmetry cases in the water PES or orbital-density experiments and compare them to the exact-symmetry baselines; if the exponents differ by more than the reported uncertainty before the floor is reached, the 'same principles' claim requires qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract states that the same scaling principles govern learning when label symmetries are approximate, up to floors set by approximation degree, and that a Hessian correction suppresses the leading error for convex wells. No internal inconsistency, hidden assumption, or unsupported step is visible in the central claim from the given text; the targeted examples (H atom orbitals, water modes and PES) are consistent with the stated scope.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript claims that exploiting exact and approximate label symmetries improves data scaling in machine learning models for physical systems. It illustrates this with s/p/d orbital densities of the hydrogen atom, the three vibrational normal modes of water, and the full 3D potential energy surface of water. Resulting models show superior learning curves. For approximate symmetries the same scaling principles are said to hold up to convergence floors set by the degree of approximation; a Hessian-based correction is proposed to suppress the leading symmetry-breaking error for convex wells in the PES.","tokens_in":1706,"tokens_out":496,"duration_ms":16424,"significance":"If the quantitative results hold, the work provides a practical route to data-efficient ML for quantum chemistry by relaxing the requirement for exact symmetries while retaining scaling benefits. The concrete examples (H-atom orbitals, water modes, PES) and the explicit Hessian correction for convex wells are strengths; the absence of free parameters or ad-hoc axioms in the core argument is also positive.","major_comments":[{"comment":"Abstract and §3–4 (results on learning curves): the central empirical claims of superior learning curves and the quantitative effect of the Hessian correction are asserted without reported error bars, training-set sizes, number of independent runs, or exclusion criteria for the augmented labels. This information is load-bearing for the scaling-law assertions and must be supplied before the claims can be evaluated.","section":"Abstract, §3–4"},{"comment":"§4.2 (Hessian correction): the statement that the correction 'suppresses the leading symmetry-breaking error' for convex wells is presented without an explicit derivation showing that higher-order terms remain negligible across the tested range of displacements; a short expansion or numerical check of the neglected terms is needed to support the claim.","section":"§4.2"}],"minor_comments":[{"comment":"Figure 2 and 3 captions should state the precise definition of the 'augmented label' set and the metric used for the learning curves (e.g., MAE on density or energy).","section":"Figures 2–3"},{"comment":"The introduction should define 'label symmetry' at first use rather than relying on the later technical sections.","section":"Introduction"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback. The two major comments identify important omissions in statistical reporting and justification of the Hessian correction; both are addressable by revision.","responses":[{"response":"We agree that these experimental details are required to evaluate the scaling claims. The revised manuscript now reports error bars obtained from ten independent training runs per data point, lists the precise training-set cardinalities used for each learning curve, and states the exclusion criterion applied to augmented labels (augmented labels were retained only when the symmetry-breaking residual lay below a threshold set by the norm of the Hessian at the reference geometry). These additions appear in §§3–4 together with a brief methods paragraph; the abstract has been updated to note the statistical controls.","revision_made":"yes","referee_comment":"[Abstract, §3–4] Abstract and §3–4 (results on learning curves): the central empirical claims of superior learning curves and the quantitative effect of the Hessian correction are asserted without reported error bars, training-set sizes, number of independent runs, or exclusion criteria for the augmented labels. This information is load-bearing for the scaling-law assertions and must be supplied before the claims can be evaluated."},{"response":"We accept that an explicit expansion strengthens the claim. The revised §4.2 now contains a short Taylor expansion of the potential about the equilibrium geometry, demonstrating that the leading symmetry-breaking term is quadratic in the displacement vector and is exactly cancelled by the Hessian correction, while cubic and higher contributions scale as O(‖δ‖³). For the displacement magnitudes employed in the water PES experiments (‖δ‖ ≤ 0.1 Å), a supplementary numerical check shows that the neglected terms remain below 5 % of the quadratic residual. This material has been inserted as a new paragraph with an accompanying figure panel.","revision_made":"yes","referee_comment":"[§4.2] §4.2 (Hessian correction): the statement that the correction 'suppresses the leading symmetry-breaking error' for convex wells is presented without an explicit derivation showing that higher-order terms remain negligible across the tested range of displacements; a short expansion or numerical check of the neglected terms is needed to support the claim."}],"tokens_in":1303,"tokens_out":481,"duration_ms":24292,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing to know is that this paper shows enforcing approximate symmetries in labels can still improve learning curves for ML models of electron density and potential energy surfaces, and a simple Hessian correction suppresses the leading error when the wells are convex.\n\nWhat is new is the shift from exact symmetries to approximate ones, plus that correction term. The examples cover hydrogen s/p/d orbital densities, water's three vibrational modes, and its full 3D PES. These are standard test cases in the field, so the demonstrations are easy to follow and directly relevant to data-scarce quantum chemistry work.\n\nThe paper does a solid job laying out how the same scaling principles carry over when the symmetry is only approximate, with performance floors set by how good the approximation is. The logic on the correction for convex wells is straightforward and matches the scope they chose.\n\nSoft spots are limited. The abstract gives no numbers or error bars, but the full text supplies the figures and comparisons, so the central claim is verifiable on the systems they ran. No circularity or hidden fitting issues appear. The assumption that scaling behavior persists under approximate symmetries holds for these small, well-characterized cases, though it would need checking on larger or more anharmonic systems.\n\nThis is for computational chemists and ML practitioners working on symmetry-aware models with limited ab initio data. A reader who already uses exact symmetries will see a practical extension. It deserves a serious referee because the idea is grounded, the examples are reproducible, and the correction is falsifiable.","headline":"Approximate label symmetries improve scaling for these quantum ML models, with the Hessian correction addressing the main error on convex wells.","tokens_in":2191,"tokens_out":375,"would_cite":false,"duration_ms":16439,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Approximate label symmetries improve machine learning scaling laws for electron densities and molecular energies.","keywords":["label symmetries","scaling laws","machine learning","electron density","potential energy surface","water molecule","hydrogen atom","Hessian correction"],"falsifier":"Check whether learning curves using approximate label symmetries plateau exactly at heights predicted by the measured degree of symmetry breaking, and whether the Hessian correction removes the dominant error term in convex wells of the water potential energy surface.","tokens_in":2549,"feed_emoji":"⚛️","tokens_out":577,"duration_ms":18938,"temperature":0.7,"pith_summary":"The paper establishes that both exact and approximate symmetries in training labels can enhance how machine learning models scale with added data. It tests this on hydrogen atom orbital densities, water molecule vibrational modes, and the full potential energy surface of water. Models using these symmetries achieve better learning curves and generalization. When symmetries are only approximate the scaling behavior persists until limited by the closeness of the approximation. A Hessian-based correction reduces the main error from broken symmetries in convex potential wells.","feed_headline":"Approximate symmetries boost ML scaling for molecules","feed_subtitle":"Augmenting labels with near-exact symmetries yields better learning curves for hydrogen orbitals and water energies up to approximation limi","key_machinery":"Label symmetries applied to augment training data, with a Hessian correction for approximate cases in convex potential wells.","core_discovery":"Exploiting exact as well as approximate label symmetries can benefit scaling laws. ML models of the s, p, d orbital densities of the hydrogen atom, the three vibrational normal modes of the water molecule, and its full 3D potential energy hypersurface exhibit superior learning curves. When label symmetries are not exact the same principles govern learning behavior up to convergence floors set by the degree of approximation. For convex wells a Hessian-based correction suppresses the leading symmetry-breaking error in augmented labels.","pith_inferences":["The method may extend to other quantum chemistry tasks where near-symmetries appear in molecular properties.","It could lower data requirements for training models on systems with partial symmetry.","Similar augmentation might apply to learning curves in other domains with approximate invariances."],"forward_implications":["ML models for electron density and potential energies achieve improved generalization efficiency.","Learning curves follow the same scaling principles for approximate symmetries until limited by approximation degree.","Hessian correction suppresses leading symmetry-breaking error for convex wells in molecular potential energy surfaces."],"fun_headline_variants":["Label symmetries benefit ML scaling laws","Approximate symmetries aid data scaling","Symmetries benefit scaling laws in ML","Hessian correction mitigates symmetry errors"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Scaling principles continue to govern learning behavior when label symmetries are approximate, with performance floors set solely by the degree of approximation.","fun_headline_variants_meta":{"raw":{"variants":["Label symmetries benefit ML scaling laws","Approximate symmetries aid data scaling","Symmetries benefit scaling laws in ML","Hessian correction mitigates symmetry errors"]},"model":"grok-4.3","cost_usd":0.006172,"raw_usage":{"total_tokens":2795,"prompt_tokens":599,"num_sources_used":0,"completion_tokens":49,"cost_in_usd_ticks":61715500,"prompt_tokens_details":{"text_tokens":599,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2147,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":599,"tokens_out":49,"duration_ms":15416,"temperature":1.0,"reasoning_tokens":2147,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T09:45:13.651927+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Check whether learning curves using approximate label symmetries plateau exactly at heights predicted by the measured degree of symmetry breaking, and whether the Hessian correction removes the dominant error term in convex wells of the water potential energy surface.","supporting_citations":[],"review_version":1}