REVIEW 3 cited by
Lookahead: A Far-Sighted Alternative of Magnitude-based Pruning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Magnitude-based pruning is one of the simplest methods for pruning neural networks. Despite its simplicity, magnitude-based pruning and its variants demonstrated remarkable performances for pruning modern architectures. Based on the observation that magnitude-based pruning indeed minimizes the Frobenius distortion of a linear operator corresponding to a single layer, we develop a simple pruning method, coined lookahead pruning, by extending the single layer optimization to a multi-layer optimization. Our experimental results demonstrate that the proposed method consistently outperforms magnitude-based pruning on various networks, including VGG and ResNet, particularly in the high-sparsity regime. See https://github.com/alinlab/lookahead_pruning for codes.
Forward citations
Cited by 3 Pith papers
-
Post-Training Pruning for Diffusion Transformers
DiT-Pruning keeps CLIP and FID nearly unchanged at 50% sparsity on FLUX and PixArt by an energy-motivated squared-weight metric plus clustering-aware granularity, beating Wanda and magnitude baselines.
-
Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study
Structural pruning with finetuning plus hidden-state distillation recovers most performance in multimodal LLMs, with 5% of training data sufficient at moderate compression levels.
-
Pruning-based Data Selection and Network Fusion for Efficient Deep Learning
PruneFuse combines pruning at initialization with weight fusion and knowledge distillation to make active learning data selection cheaper and to initialize the final model.
Discussion (0). Continue with ORCID to comment.