Pith. sign in

REVIEW 2 cited by

Efficient Transformer-based Large Scale Language Representations using Hardware-friendly Block Structured Pruning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2009.08065 v4 pith:4UUZJ4UH submitted 2020-09-17 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords languagemodelsaccuracydistilbertpre-trainedpruningblockcompression
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pre-trained large-scale language models have increasingly demonstrated high accuracy on many natural language processing (NLP) tasks. However, the limited weight storage and computational speed on hardware platforms have impeded the popularity of pre-trained models, especially in the era of edge computing. In this work, we propose an efficient transformer-based large-scale language representation using hardware-friendly block structure pruning. We incorporate the reweighted group Lasso into block-structured pruning for optimization. Besides the significantly reduced weight storage and computation, the proposed approach achieves high compression rates. Experimental results on different models (BERT, RoBERTa, and DistilBERT) on the General Language Understanding Evaluation (GLUE) benchmark tasks show that we achieve up to 5.0x with zero or minor accuracy degradation on certain task(s). Our proposed method is also orthogonal to existing compact pre-trained language models such as DistilBERT using knowledge distillation, since a further 1.79x average compression rate can be achieved on top of DistilBERT with zero or minor accuracy degradation. It is suitable to deploy the final compressed model on resource-constrained edge devices.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. F$^3$OCUS -- Federated Finetuning of Vision-Language Foundation Models with Optimal Client Layer Updating Strategy via Multi-objective Meta-Heuristics

    cs.CV 2024-11 conditional novelty 6.0 of 10

    F3OCUS combines per-client LNTK layer importance scores with server-side meta-heuristic optimization of layer diversity to improve federated fine-tuning of vision-language models for medical tasks, and releases the 70...

  2. Systolic Arrays and Structured Pruning Co-design for Efficient Transformers in Edge Systems

    cs.AR 2024-11 conditional novelty 4.0 of 10

    Matching the size of pruned weight blocks to a systolic array lets edge transformer inference skip zero tiles, yielding up to 44% encoder speedup with roughly 1.4% WER degradation.

Pith tools