REVIEW 6 cited by
Fast Training of Convolutional Networks through FFTs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Convolutional networks are one of the most widely employed architectures in computer vision and machine learning. In order to leverage their ability to learn complex functions, large amounts of data are required for training. Training a large convolutional network to produce state-of-the-art results can take weeks, even when using modern GPUs. Producing labels using a trained network can also be costly when dealing with web-scale datasets. In this work, we present a simple algorithm which accelerates training and inference by a significant factor, and can yield improvements of over an order of magnitude compared to existing state-of-the-art implementations. This is done by computing convolutions as pointwise products in the Fourier domain while reusing the same transformed feature map many times. The algorithm is implemented on a GPU architecture and addresses a number of related challenges.
Forward citations
Cited by 6 Pith papers
-
Flash EQ-Linear: Accelerating Equivariant Linear Layers via Group-wise Discrete Fourier Transform
EQ-Linear can be computed exactly as pointwise multiplications in the Fourier domain along the group dimension, cutting FLOPs from NDC to ~2NDC/T and yielding up to ~2× wall-clock speedups for p4 equivariant transformers.
-
Simple Graph Contrastive Learning via Fractional-order Neural Diffusion Networks
FD-GCL is an augmentation-free, negative-free graph contrastive learner that uses two fractional-diffusion encoders with different orders to generate local and global views, reporting state-of-the-art node classificat...
-
Fourier analysis of the physics of transfer learning for data-driven subgrid-scale models of ocean turbulence
A CNN subgrid model trained on one quasi-geostrophic regime underestimates out-of-distribution activation and output spectra, and retraining only the first hidden layer with target data corrects the spectral bias.
-
Symmetry group factorization reveals the structure-function relation in the neural connectome of Caenorhabditis elegans
Symmetry groups of the C. elegans locomotion circuits factorize into subgroups whose neuron sets match known functional classes, and finer imprimitive blocks form circulant 'filter' matrices.
-
Demystifying the 7-D Convolution Loop Nest for Data and Instruction Streaming in Reconfigurable AI Accelerators
A fold-based weight-stationary mapping of the 7D convolution loop nest is proposed for the MAVeC accelerator, with simulated VGG-16 performance claims of 12.7 KIPS and over 90 percent utilization.
-
FPGA-based Acceleration for Convolutional Neural Networks: A Comprehensive Review
A comprehensive review of FPGA-based CNN accelerators that consolidates evaluation metrics, acceleration methods, parallel computing strategies, and toolflows, concluding that dynamic parallelism is the most promising...
Discussion (0). Continue with ORCID to comment.