REVIEW 7 cited by
Does equivariance matter at scale?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Given large datasets and sufficient compute, is it beneficial to design neural architectures for the structure and symmetries of each problem? Or is it more efficient to learn them from data? We study empirically how equivariant and non-equivariant networks scale with compute and training samples. Focusing on a benchmark problem of rigid-body interactions and on general-purpose transformer architectures, we perform a series of experiments, varying the model size, training steps, and dataset size. We find evidence for three conclusions. First, equivariance improves data efficiency, but training non-equivariant models with data augmentation can close this gap given sufficient epochs. Second, scaling with compute follows a power law, with equivariant models outperforming non-equivariant ones at each tested compute budget. Finally, the optimal allocation of a compute budget onto model size and training duration differs between equivariant and non-equivariant models.
Forward citations
Cited by 7 Pith papers
-
CP$^2$: Leveraging Geometry for Conformal Prediction via Canonicalization
Canonicalizing inputs before conformal prediction preserves coverage and shrinks prediction sets under rotation shifts, without retraining the underlying model.
-
Wall Shear Stress Estimation in Abdominal Aortic Aneurysms: Towards Generalisable Neural Surrogate Models
A neural network trained on 100 aneurysm patients estimates transient wall shear stress on 3D vessel surfaces and generalizes, without retraining, to new patients, new branches, and different mesh resolutions.
-
Distributed Equivariant Graph Neural Networks for Large-Scale Electronic Structure Prediction
A distributed equivariant GNN with a neighbor-minimizing graph partitioner scales electronic-structure (Hamiltonian) prediction to 512 GPUs and 190,000 atoms, with an 87% weak-scaling efficiency.
-
Machine Learning Interatomic Potentials: library for efficient training, model development and simulation of molecular systems
InstaDeep's mlip library ports MACE, NequIP, and ViSNet to JAX with a JAX-MD backend, ships SPICE2-trained organics models, reports faster MD steps than its own Torch routes, and proposes a faster gated MACE variant i...
-
EQ-VAE: Equivariance Regularized Latent Space for Improved Generative Image Modeling
Enforcing scale and rotation equivariance in pretrained image autoencoders via a reconstruction loss on transformed latents speeds up and improves latent generative models.
-
Flopping for FLOPs: Leveraging equivariance for computational efficiency
Flopping-equivariant versions of ResMLP, ViT and ConvNeXt use half the FLOPs of the originals and match their ImageNet-1K accuracy at large model sizes.
-
Energy & Force Regression on DFT Trajectories is Not Enough for Universal Machine Learning Interatomic Potentials
A perspective arguing that current MLIP training on DFT data is insufficient, and proposing CCSD(T)-quality data, metrology, and efficient inference as research priorities.
Discussion (0). Continue with ORCID to comment.