REVIEW 4 major objections 7 minor 22 references
Edge many-body products and radial rotary attention push SO(2) interatomic potentials past prior Matbench leaders.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 10:07 UTC pith:GB27VOGS
load-bearing objection Solid SO(2) methods paper with real operators and ablations; the Matbench SOTA is a 0.001 CPS edge and should not be the main reason you care. the 4 major comments →
Edge Cluster Expansion with Radial Rotary Attention for Interatomic Potentials
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Conventional SO(2) Linear is weaker than Clebsch–Gordan tensor products mainly because of design choices (path-wise radial weights, real versus complex weights, edge nonlinearities), not because of SO(2) equivariance itself; Edge Complex Product Basis via generalized asymmetric contraction raises effective body order on edges, and Radial Rotary Complex Attention improves extrapolation over prior attention-vector schemes, yielding TECE-OAM-RRA-1.0 with state-of-the-art overall Matbench Discovery performance after training on OMat24, sAlex, and MPTrj.
What carries the argument
Edge Cluster Expansion (ECE): generalized asymmetric contraction that builds higher-order product bases on edge features with SO(2) coupling coefficients, paired with Radial Rotary Complex Attention whose logits are the real part of a complex QK product scaled by a radial phase and shifted by a radial bias.
Load-bearing premise
Modules are kept or discarded mainly by whether they keep force error below a fixed high-temperature molecular threshold, treating that single check as a reliable stand-in for how well a universal materials model will extrapolate.
What would settle it
Train the same TECE stack with ECE and RRA ablated (or replaced by standard symmetric contraction and Equiformer-style attention vectors) on the same OMat24/sAlex/MPTrj mix and re-score Matbench Discovery; if CPS, F1, and kappa_SRME no longer lead, the claimed gains do not come from those blocks.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript systematically analyzes SO(2) equivariant operators for machine learning interatomic potentials, contrasts uuSO2Linear with uvSO2Linear and real versus complex SO(2) weights, and proposes two constructions of Wigner-D matrices (direct Cartesian and recursive Clebsch–Gordan). It introduces Edge Cluster Expansion (ECE) via generalized asymmetric contraction on edges and Radial Rotary Complex Attention (RRA), plus ACE improvements (group linear, GLU-style nonlinear coefficients, residual/LayerNorm choices). Controlled molecular ablations (3BPA/AcAc, M-fold symmetry, Pozdnyakov body-order graphs) support the design choices. Models trained on OMat24, sAlex, and MPTrj yield TECE-OAM-RRA-1.0, reported as state-of-the-art on Matbench Discovery (CPS 0.908).
Significance. If the methodological claims hold, the paper is a useful consolidation of SO(2) practice for MLIPs: the O(2) analysis of complex weights (Eq. 14), the uu/uv path distinction, the recursive Wigner-D construction with favorable scaling (Fig. 1), and the completeness arguments for ECE (M-fold math and Table 3 body-order tests) are concrete contributions. RRA’s radial phase/bias design is a clear inductive-bias improvement over attention-vector baselines on 3BPA (Table 2). Code and model releases strengthen reproducibility. The Matbench ranking is of practical interest to the materials community, though the reported margin is very small. Overall significance is primarily architectural and theoretical rather than a decisive leap in materials accuracy.
major comments (4)
- Table 4 and Appendix Table 7: TECE-OAM-RRA-1.0 is ranked SOTA by CPS 0.908 vs EquFlashV2 0.907. The two models share F1=0.929 and RMSD=0.058; the only reported difference is κSRME 0.093 vs 0.094. No multi-seed variance, bootstrap, or uncertainty on CPS is given. A 0.001 CPS edge is not, by itself, a robust SOTA claim. Either quantify ranking uncertainty, report multiple independent runs, or temper the abstract/introduction language to “competitive / among top models” unless stronger evidence is added.
- Section 3.6: architecture modules are retained only if 3BPA 1200K force RMSE stays ≤65 meV/Å. This single molecular high-T threshold is treated as the gate for a universal materials model. The manuscript does not show that modules ranked by this gate also rank by Matbench F1/κSRME/RMSD, nor that ECE/RRA (vs ablated variants) improve Matbench metrics. Without materials-side ablations or a demonstrated correlation between the 3BPA gate and Matbench, the causal link from ECE/RRA to the Matbench ranking remains under-supported.
- Section 3.11 states Adam is best for extrapolation and that Muon/SOAP can compromise it; Appendix A reports the released TECE-OAM-RRA-1.0 was trained with Muon for speed. For a paper whose narrative emphasizes extrapolation (RRA, real weights, 3BPA high-T), this is a load-bearing inconsistency. Please either retrain/fine-tune a key checkpoint with AdamW and report Matbench/3BPA deltas, or clearly qualify that the SOTA model may not reflect the authors’ preferred extrapolation recipe and discuss the risk.
- Tables 1–3 and the M-fold derivation establish ECE/RRA on small molecular tasks, but the large-scale claim (Abstract, §3.12) attributes Matbench SOTA to “these advances” without component ablations at OAM scale (ECE on/off, RRA vs EquiformerV3 attention, uu vs uv under matched params). At minimum, a smaller matched-budget materials ablation or intermediate TACE-OAM-L-style comparison that isolates ECE/RRA would make the contribution narrative falsifiable rather than confounded with data, width, DeNS, and optimizer.
minor comments (7)
- Abstract and §1: “propose direct Cartesian construction and recursive Clebsch-Gordan construction” — spacing/typos (“proposedirect”, “Complex At- tention”) should be cleaned throughout the arXiv text.
- §3.7: sentence fragment “two key factors that are critical to the a (Joshi et al., 2023) identifies” needs repair.
- Eq. (35): θ_{m,h}(r) is written with an m index but the surrounding text often uses θ_h; clarify whether phase is m-dependent and how it is shared across heads/orders.
- Table 5: TECE OOMs earlier than EquFlashV2; the discussion correctly notes fusion advantages of SO(3), but a brief note on peak memory vs parameter count (222M vs 44.9M) would help readers interpret speed fairly.
- §3.5 / Table 2: “w1 w1 w1 w1 w2” column headers are hard to parse; a compact legend (w1 | w1=w2 | w1+iw2) would improve readability.
- §4.1 diatomic discussion is valuable; consider moving a short quantitative diatomic table into the main text or appendix so the “poor diatomic / high Matbench” trade-off is documented rather than only narrated.
- References and arXiv dates in the bibliography include 2026 entries; ensure consistency of citation keys and that concurrent work (EquiformerV3, EquFlashV2, DPA4) is fairly scoped in related work.
Circularity Check
No significant circularity: SOTA and ablation claims are external-benchmark evaluations of proposed operators, not tautologies of fitted inputs or self-citation uniqueness.
full rationale
The paper proposes architectural operators (Cartesian/recursive Wigner-D construction, Edge Cluster Expansion via generalized asymmetric contraction, Radial Rotary Complex Attention, and ACE group-linear/GLU-style improvements) and evaluates them on held-out molecular sets (3BPA temperatures/dihedrals, AcAc), geometric completeness tests (M-fold SO(2) structures; Pozdnyakov k-body counterexamples), and the external Matbench Discovery leaderboard after training on OMat24/sAlex/MPTrj. None of these results is obtained by fitting a parameter to a quantity and then reporting that same quantity as a prediction, nor by defining an operator in terms of the metric it is said to derive. Self-citations to prior TACE Cartesian work supply background ACE/Cartesian operators and residual/norm design context; they do not import a uniqueness theorem that forces the Matbench ranking or the ECE/RRA claims. The 3BPA-1200K 65 meV/Å module gate and the Muon-vs-Adam optimizer choice are design/selection decisions that may affect causal attribution of SOTA gains, but they are not circular reductions of outputs to inputs. Score 0 with empty steps is therefore the correct finding.
Axiom & Free-Parameter Ledger
free parameters (5)
- 3BPA-1200K force RMSE exclusion threshold =
65 meV/Å
- Interaction / product channel widths and Lmax/mmax =
64 / 256 / 4 / 2
- RRA temperature bounds and radial phase form =
τ∈[0.25,4.0]
- Muon learning-rate schedule and DeNS/stochastic-depth rates =
see Table 6
- Cutoff radius =
6 Å
axioms (5)
- standard math SO(2) complex multiplication with m3=m1±m2 is equivariant under planar rotations (Eqs. 1–4).
- standard math Wigner-D matrices implement SO(3) action on spherical tensors and can be obtained from Cartesian rotations via ICTD or recursive CG coupling (Eqs. 10, 13).
- domain assumption Local edge frames from bond directions make global SO(3) features into local SO(2) features without breaking equivariance when D-matrices are applied correctly.
- domain assumption OMat24, sAlex, and MPTrj labels are sufficiently consistent DFT targets for ranking universal MLIPs on Matbench Discovery.
- ad hoc to paper Real-valued SO(2) weights (w1) are preferred because complex weights break O(2) reflection equivariance unless w2=0 (Eq. 14).
invented entities (4)
-
Edge Cluster Expansion (ECE) / Edge Complex Product Basis via Generalized Asymmetric Contraction
no independent evidence
-
Radial Rotary Complex Attention (RRA)
no independent evidence
-
uuSO2Linear vs uvSO2Linear distinction
no independent evidence
-
Recursive Clebsch-Gordan Wigner-D construction (TACE-Recursive)
independent evidence
read the original abstract
In this paper, we provide a systematic investigation of SO(2) theory to machine learning interatomic potentials (MLIPs) and identify the limitations of conventional SO(2) Linear architectures relative to SO(3) Clebsch-Gordan Tensor Products (CGTP). Building on these insights, we propose direct Cartesian construction and recursive Clebsch-Gordan construction of Wigner D-matrices and introduce two novel interaction building blocks. First, we propose the Edge Complex Product Basis based on Generalized Asymmetric Contraction, a new formulation for many-body expansion that directly constructs higher-order interactions on edges through complex-valued equivariant multiplications. Second, we introduce Radial Rotary Complex Attention(RRA), which enhances extrapolation performance and surpasses existing attention vector formulations. We also introduce several improvements to the Atomic Cluster Expansion module. Building on these advances, we train our models on OMat24, sAlex, and MPTrj, and introduce TECE-OAM-RRA-1.0, which achieve state-of-the-art (SOTA) performance on the Matbench Discovery.
Figures
Reference graph
Works this paper leans on
-
[1]
M., Dzamba, M., Gao, M., Rizvi, A., Uyttendaele, M., Zit- nick, C
Barros-Luque, L., Shuaibi, M., Fu, X., Wood, B. M., Dzamba, M., Gao, M., Rizvi, A., Uyttendaele, M., Zit- nick, C. L., and Ulissi, Z. W. The open materials 2024 (omat24) inorganic materials dataset and models.Nature Computational Science, pp. 1–11,
2024
-
[2]
URL https://arxiv.org/abs/ 2601.16195. Dauphin, Y . N., Fan, A., Auli, M., and Grangier, D. Lan- guage modeling with gated convolutional networks. In Precup, D. and Teh, Y . W. (eds.),Proceedings of the 34th International Conference on Machine Learning, volume 70 ofProceedings of Machine Learning Re- search, pp. 933–941. PMLR, 06–11 Aug
-
[3]
Harari, G., Zimmermann, Y ., Kulseng, O
URL https: //arxiv.org/abs/2607.03433. Harari, G., Zimmermann, Y ., Kulseng, O. T., Zichi, L., Tan, C. W., Descoteaux, M. L., and Kozinsky, B. Beyond adam: Soap and muon for faster, label-efficient training of machine learning interatomic potentials,
-
[4]
He, K., Zhang, X., Ren, S., and Sun, J
URL https://arxiv.org/abs/2607.02499. He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learn- ing for image recognition. InProceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778,
-
[5]
Huang, G., Sun, Y ., Liu, Z., Sedra, D., and Weinberger, K
URL https://arxiv.org/abs/2603.08630. Huang, G., Sun, Y ., Liu, Z., Sedra, D., and Weinberger, K. Deep networks with stochastic depth,
-
[6]
Huang, L., Huang, C., Wang, Z., Du, Y ., Wang, C., Lu, H., Li, Y ., Liu, X., Jiang, A., and Zhang, J
URL https://arxiv.org/abs/1603.09382. Huang, L., Huang, C., Wang, Z., Du, Y ., Wang, C., Lu, H., Li, Y ., Liu, X., Jiang, A., and Zhang, J. E2former- v2: On-the-fly equivariant attention with linear activation memory.arXiv preprint arXiv:2601.16622,
-
[7]
URL https://arxi v.org/abs/2401.04088. Joshi, C. K., Bodnar, C., Mathis, S. V ., Cohen, T., and Lio, P. On the expressive power of geometric graph neural networks
-
[8]
Kondor, R., Lin, Z., and Trivedi, S
URL https://arxiv.org/ab s/1412.6980. Kondor, R., Lin, Z., and Trivedi, S. Clebsch–gordan nets: a fully fourier space spherical convolutional neural network. In Bengio, S., Wallach, H., Larochelle, H., Grauman, K., Cesa-Bianchi, N., and Garnett, R. (eds.),Advances in Neural Information Processing Systems, volume
-
[9]
Li, T., Li, W., Peng, A., Xue, J., Zhang, L., Zhang, D., and Wang, H
URL https://openreview.net /forum?id=wiQe95BPaB. Li, T., Li, W., Peng, A., Xue, J., Zhang, L., Zhang, D., and Wang, H. Dpa4: Pushing the accuracy-cost frontier of interatomic potentials with emfa so(2) convolution, 2026a. URLhttps://arxiv.org/abs/2606.02419. Li, Y ., Huang, L., Ding, Z., Wei, X., Wang, C., Yang, H., Wang, Z., Liu, C., Shi, Y ., Jin, P., Q...
-
[10]
URL https: //arxiv.org/abs/2403.09549. Liao, Y .-L., Hoffman, A. J., Shen, S. C., Duval, A., Nor- wood, S. W., and Smidt, T. Equiformerv3: Scaling ef- ficient, expressive, and general se (3)-equivariant graph attention transformers.arXiv preprint arXiv:2604.09130,
-
[11]
URL https://arxiv.org/abs/ 2502.16982. Loshchilov, I. and Hutter, F. Decoupled weight decay regu- larization,
-
[12]
Luo, S., Chen, T., and Krishnapriyan, A
URL https://arxiv.org/abs/ 1711.05101. Luo, S., Chen, T., and Krishnapriyan, A. S. Enabling effi- cient equivariant operations in the fourier basis via gaunt tensor products. InThe Twelfth International Confer- ence on Learning Representations,
-
[13]
Orb-v3: atomistic simulation at scale.arXiv preprint arXiv:2504.06231,
Rhodes, B., Vandenhaute, S., ˇSimkus, V ., Gin, J., Godwin, J., Duignan, T., and Neumann, M. Orb-v3: atomistic simulation at scale.arXiv preprint arXiv:2504.06231,
-
[14]
URL https://arxiv.org/ab s/2104.09864. Team, K., Chen, G., Zhang, Y ., Su, J., Xu, W., Pan, S., Wang, Y ., Wang, Y ., Chen, G., Yin, B., et al. Attention residuals.arXiv preprint arXiv:2603.15031,
-
[15]
Thomas, N., Smidt, T., Kearnes, S., Yang, L., Li, L., Kohlhoff, K., and Riley, P. Tensor field networks: Rotation-and translation-equivariant neural networks for 3d point clouds.arXiv preprint arXiv:1802.08219,
-
[16]
Unke, O. T. and Maennel, H. E3x: E(3)-equivariant deep learning made easy.arXiv preprint arXiv:2401.07595,
-
[17]
URLhttps://arxiv.org/abs/2409.11321. Warford, T., Thiemann, F. L., and Cs´anyi, G. Better without u: Impact of selective hubbard u correction on founda- tional mlips,
-
[18]
15 Weiler, M., Geiger, M., Welling, M., Boomsma, W., and Cohen, T
URL https://arxiv.org/ab s/2601.21056. 15 Weiler, M., Geiger, M., Welling, M., Boomsma, W., and Cohen, T. S. 3d steerable cnns: Learning rotationally equivariant features in volumetric data. In Bengio, S., Wallach, H., Larochelle, H., Grauman, K., Cesa-Bianchi, N., and Garnett, R. (eds.),Advances in Neural Informa- tion Processing Systems, volume
-
[19]
Xie, Y ., Daigavane, A., Kotak, M., and Smidt, T
URL https://arxiv.org/abs/ 2506.23971. Xie, Y ., Daigavane, A., Kotak, M., and Smidt, T. Asymp- totically fast clebsch-gordan tensor products with vector spherical harmonics,
-
[20]
URL https://arxiv. org/abs/2602.21466. Xu, Z., Wu, C., Xie, W., and Hu, P. A cartesian-3j frame- work for machine learning interatomic potentials, 2026a. URLhttps://arxiv.org/abs/2512.16882. Xu, Z., Xie, W., and Hu, P. Spectral/spatial tensor atomic cluster expansion with universal embeddings in cartesian space.arXiv preprint arXiv:2509.14961, 2026b. Yu, ...
-
[21]
Training Details All TECE models were trained in torch.float32 precision using 32 NVIDIA H20 GPUs
16 A. Training Details All TECE models were trained in torch.float32 precision using 32 NVIDIA H20 GPUs. During pretraining, we employed DeNS (Liao et al., 2024)and stochastic depth (Huang et al.,
2024
-
[22]
DeNS was mainly used with the direct model to accelerate convergence, rather than to improve the final accuracy
as additional regularization techniques. DeNS was mainly used with the direct model to accelerate convergence, rather than to improve the final accuracy. With appropriately chosen hyperparameters, the direct and conservative models achieved comparable performance upon convergence. The DeNS hyperparameters were kept consistent with those used in Equiformer...
2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.