Pith. sign in

REVIEW 3 cited by

Lookahead: A Far-Sighted Alternative of Magnitude-based Pruning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.04809 v1 pith:UKXSQACJ submitted 2020-02-12 cs.LG stat.ML

classification cs.LGstat.ML
keywords pruningmagnitude-basedlookaheadlayermethodnetworksoptimizationsingle
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Magnitude-based pruning is one of the simplest methods for pruning neural networks. Despite its simplicity, magnitude-based pruning and its variants demonstrated remarkable performances for pruning modern architectures. Based on the observation that magnitude-based pruning indeed minimizes the Frobenius distortion of a linear operator corresponding to a single layer, we develop a simple pruning method, coined lookahead pruning, by extending the single layer optimization to a multi-layer optimization. Our experimental results demonstrate that the proposed method consistently outperforms magnitude-based pruning on various networks, including VGG and ResNet, particularly in the high-sparsity regime. See https://github.com/alinlab/lookahead_pruning for codes.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Post-Training Pruning for Diffusion Transformers

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    DiT-Pruning keeps CLIP and FID nearly unchanged at 50% sparsity on FLUX and PixArt by an energy-motivated squared-weight metric plus clustering-aware granularity, beating Wanda and magnitude baselines.

  2. Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Structural pruning with finetuning plus hidden-state distillation recovers most performance in multimodal LLMs, with 5% of training data sufficient at moderate compression levels.

  3. Pruning-based Data Selection and Network Fusion for Efficient Deep Learning

    cs.LG 2025-01 conditional novelty 4.0 of 10

    PruneFuse combines pruning at initialization with weight fusion and knowledge distillation to make active learning data selection cheaper and to initialize the final model.

Pith tools