Pith. sign in

REVIEW 7 cited by

Does equivariance matter at scale?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.23179 v2 pith:7GWX2CYT submitted 2024-10-30 cs.LG

classification cs.LG
keywords computenon-equivarianttrainingdataequivariantmodelssizearchitectures
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Given large datasets and sufficient compute, is it beneficial to design neural architectures for the structure and symmetries of each problem? Or is it more efficient to learn them from data? We study empirically how equivariant and non-equivariant networks scale with compute and training samples. Focusing on a benchmark problem of rigid-body interactions and on general-purpose transformer architectures, we perform a series of experiments, varying the model size, training steps, and dataset size. We find evidence for three conclusions. First, equivariance improves data efficiency, but training non-equivariant models with data augmentation can close this gap given sufficient epochs. Second, scaling with compute follows a power law, with equivariant models outperforming non-equivariant ones at each tested compute budget. Finally, the optimal allocation of a compute budget onto model size and training duration differs between equivariant and non-equivariant models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CP$^2$: Leveraging Geometry for Conformal Prediction via Canonicalization

    stat.ML 2025-06 conditional novelty 7.0 of 10

    Canonicalizing inputs before conformal prediction preserves coverage and shrinks prediction sets under rotation shifts, without retraining the underlying model.

  2. Wall Shear Stress Estimation in Abdominal Aortic Aneurysms: Towards Generalisable Neural Surrogate Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A neural network trained on 100 aneurysm patients estimates transient wall shear stress on 3D vessel surfaces and generalizes, without retraining, to new patients, new branches, and different mesh resolutions.

  3. Distributed Equivariant Graph Neural Networks for Large-Scale Electronic Structure Prediction

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A distributed equivariant GNN with a neighbor-minimizing graph partitioner scales electronic-structure (Hamiltonian) prediction to 512 GPUs and 190,000 atoms, with an 87% weak-scaling efficiency.

  4. Machine Learning Interatomic Potentials: library for efficient training, model development and simulation of molecular systems

    physics.chem-ph 2025-05 conditional novelty 6.0 of 10

    InstaDeep's mlip library ports MACE, NequIP, and ViSNet to JAX with a JAX-MD backend, ships SPICE2-trained organics models, reports faster MD steps than its own Torch routes, and proposes a faster gated MACE variant i...

  5. EQ-VAE: Equivariance Regularized Latent Space for Improved Generative Image Modeling

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Enforcing scale and rotation equivariance in pretrained image autoencoders via a reconstruction loss on transformed latents speeds up and improves latent generative models.

  6. Flopping for FLOPs: Leveraging equivariance for computational efficiency

    cs.CV 2025-02 conditional novelty 5.0 of 10

    Flopping-equivariant versions of ResMLP, ViT and ConvNeXt use half the FLOPs of the originals and match their ImageNet-1K accuracy at large model sizes.

  7. Energy & Force Regression on DFT Trajectories is Not Enough for Universal Machine Learning Interatomic Potentials

    cond-mat.mtrl-sci 2025-02 conditional novelty 4.0 of 10

    A perspective arguing that current MLIP training on DFT data is insufficient, and proposing CCSD(T)-quality data, metrology, and efficient inference as research priorities.

Pith tools