Pith. sign in

REVIEW 3 cited by

An Overview of Neural Network Compression

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.03669 v2 pith:6RL4HVPB submitted 2020-06-05 cs.LG stat.ML

classification cs.LGstat.ML
keywords networksneuralcitetdeepfootnoteoverviewcitepcompression
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Overparameterized networks trained to convergence have shown impressive performance in domains such as computer vision and natural language processing. Pushing state of the art on salient tasks within these domains corresponds to these models becoming larger and more difficult for machine learning practitioners to use given the increasing memory and storage requirements, not to mention the larger carbon footprint. Thus, in recent years there has been a resurgence in model compression techniques, particularly for deep convolutional neural networks and self-attention based networks such as the Transformer. Hence, this paper provides a timely overview of both old and current compression techniques for deep neural networks, including pruning, quantization, tensor decomposition, knowledge distillation and combinations thereof. We assume a basic familiarity with deep learning architectures\footnote{For an introduction to deep learning, see ~\citet{goodfellow2016deep}}, namely, Recurrent Neural Networks~\citep[(RNNs)][]{rumelhart1985learning,hochreiter1997long}, Convolutional Neural Networks~\citep{fukushima1980neocognitron}~\footnote{For an up to date overview see~\citet{khan2019survey}} and Self-Attention based networks~\citep{vaswani2017attention}\footnote{For a general overview of self-attention networks, see ~\citet{chaudhari2019attentive}.},\footnote{For more detail and their use in natural language processing, see~\citet{hu2019introductory}}. Most of the papers discussed are proposed in the context of at least one of these DNN architectures.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Rateless Joint Source-Channel Coding, and a Blueprint for 6G Semantic Communications System Design

    cs.IT 2025-02 conditional novelty 6.0 of 10

    Rateless joint source-channel coding lets a network puncture a coded image to match channel capacity, and a new autoencoder code (RLACS) shows graceful image-quality degradation without channel state information at th...

  2. Beyond Backbone Backpropagation: A Decoupled Strategy for Efficient Transfer Learning

    cs.LG 2026-06 conditional novelty 5.0 of 10

    A decoupled transfer-learning method that freezes the backbone, retunes only normalization statistics, and reweights ambiguous samples matches standard last-layer fine-tuning on most medical benchmarks at 10-20x lower...

  3. INSIGHT: A Survey of In-Network Systems for Intelligent, High-Efficiency AI and Topology Optimization

    cs.NI 2025-05 conditional

    A survey of in-network AI computing that catalogs architectures, model-compression methods, aggregation frameworks, and applications, but offers no new experimental results.

Pith tools