Pith. sign in

REVIEW 3 cited by

Compression-aware Training of Neural Networks using Frank-Wolfe

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.11921 v2 pith:F5JS22K6 submitted 2022-05-24 cs.LG math.OC

classification cs.LGmath.OC
keywords trainingapproachescompression-awareconvergencedecompositiondenseexistingfrank-wolfe
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Many existing Neural Network pruning approaches rely on either retraining or inducing a strong bias in order to converge to a sparse solution throughout training. A third paradigm, 'compression-aware' training, aims to obtain state-of-the-art dense models that are robust to a wide range of compression ratios using a single dense training run while also avoiding retraining. We propose a framework centered around a versatile family of norm constraints and the Stochastic Frank-Wolfe (SFW) algorithm that encourage convergence to well-performing solutions while inducing robustness towards convolutional filter pruning and low-rank matrix decomposition. Our method is able to outperform existing compression-aware approaches and, in the case of low-rank matrix decomposition, it also requires significantly less computational resources than approaches based on nuclear-norm regularization. Our findings indicate that dynamically adjusting the learning rate of SFW, as suggested by Pokutta et al. (2020), is crucial for convergence and robustness of SFW-trained models and we establish a theoretical foundation for that practice.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SLORR: Simple and Efficient In-Training Low-Rank Regularization

    cs.LG 2026-07 accept novelty 6.0 of 10

    A stateless, SVD-free regularizer approximates polar factors to induce low-rank weight structure during training, enabling better post-training compression of vision models and LLMs at under 8% overhead.

  2. prunAdag: an adaptive pruning-aware gradient method

    math.OC 2025-02 conditional novelty 6.0 of 10

    prunAdag separates parameters into optimisable and decreasable sets, updates them with Adagrad-like rules, and provably drives the average gradient norm to zero at rate O(log(k)/sqrt(k+1)).

  3. Compression Aware Certified Training

    cs.LG 2025-06 conditional novelty 5.0 of 10

    CACTUS trains a single network on pruned and weight-perturbed copies of itself, beating prior certified-training baselines on compressed MNIST and CIFAR-10 models.

Pith tools