Pith. sign in

REVIEW 8 cited by

A Survey of Model Compression and Acceleration for Deep Neural Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1710.09282 v9 pith:QBKK4ERH submitted 2017-10-23 cs.LG cs.CV

classification cs.LGcs.CV
keywords networksdeepmodelneuralperformancerecenttechniquesacceleration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep neural networks (DNNs) have recently achieved great success in many visual recognition tasks. However, existing deep neural network models are computationally expensive and memory intensive, hindering their deployment in devices with low memory resources or in applications with strict latency requirements. Therefore, a natural thought is to perform model compression and acceleration in deep networks without significantly decreasing the model performance. During the past five years, tremendous progress has been made in this area. In this paper, we review the recent techniques for compacting and accelerating DNN models. In general, these techniques are divided into four categories: parameter pruning and quantization, low-rank factorization, transferred/compact convolutional filters, and knowledge distillation. Methods of parameter pruning and quantization are described first, after that the other techniques are introduced. For each category, we also provide insightful analysis about the performance, related applications, advantages, and drawbacks. Then we go through some very recent successful methods, for example, dynamic capacity networks and stochastic depths networks. After that, we survey the evaluation matrices, the main datasets used for evaluating the model performance, and recent benchmark efforts. Finally, we conclude this paper, discuss remaining the challenges and possible directions for future work.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MixQuant: Adaptive Mixed-Precision Quantization for Large Language Models

    cs.LG 2026-07 conditional novelty 7.0 of 10

    A single calibration that marginalizes layer distortion over random quantized upstream contexts yields budget-agnostic bit allocations that beat FP16-scored adaptive baselines across Llama-3.2-3B, Llama-2-7B, and Mistral-7B.

  2. Contrastive Predictive Coding with Compression for Enhanced Channel State Feedback in Wireless Networks

    cs.IT 2026-06 conditional novelty 6.0 of 10

    Integrating CPC into 3GPP CSI compression yields age-aware latent prediction at fixed 64-bit overhead, with CPC-before exceeding 90% SGCS and 32× lighter decoder compute than the baseline.

  3. LACE-SVD: Loss-Aware SVD with Cumulative Error Correction for LLM Compression

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Loss-aware rank allocation plus residual-stream output correction yields substantially lower WikiText-2 perplexity than prior SVD LLM compressors at 60% compression.

  4. CoT-X: An Adaptive Framework for Cross-Model Chain-of-Thought Transfer and Optimization

    cs.AI 2025-11 conditional novelty 5.0 of 10

    An adaptive summarization framework compresses chain-of-thought traces and transfers them across model families, claiming up to 40.5% accuracy gains over truncation on medical QA and 84% fewer configuration evaluation...

  5. UnIT: Scalable Unstructured Inference-Time Pruning for MAC-efficient Neural Inference on MCUs

    cs.LG 2025-07 conditional novelty 5.0 of 10

    UnIT enables unstructured, input-aware pruning of individual MACs on MCUs without retraining, reporting up to 82% MAC reduction and up to 84% energy savings at 0.48 to 7% accuracy loss.

  6. Smooth Model Compression without Fine-Tuning

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Training a ResNet-18 with smoothness penalties on weights, then compressing via truncated SVD, keeps 91% CIFAR-10 accuracy at 70% sparsity with no post-compression fine-tuning.

  7. Modality Alignment with Multi-scale Bilateral Attention for Multimodal Recommendation

    cs.IR 2025-09 conditional novelty 4.0 of 10

    MambaRec improves multimodal recommendation accuracy on Baby, Sports, and Clothing datasets through local dilated-attention alignment and global MMD/contrastive alignment.

  8. Foundation Models for Autonomous Driving Perception: A Survey Through Core Capabilities

    cs.RO 2025-09 conditional novelty 4.0 of 10

    Foundation-model perception for autonomous driving is surveyed through four capability lenses: generalized knowledge, spatial understanding, multi-sensor robustness, and temporal understanding.

Pith tools