Pith. sign in

REVIEW 2 cited by

The Need for Speed: Pruning Transformers with One Recipe

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.17921 v1 pith:OOV5EWQR submitted 2024-03-26 cs.LG

classification cs.LG
keywords textbfoptinre-trainingtextitframeworktaskstransformerwithout
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We introduce the $\textbf{O}$ne-shot $\textbf{P}$runing $\textbf{T}$echnique for $\textbf{I}$nterchangeable $\textbf{N}$etworks ($\textbf{OPTIN}$) framework as a tool to increase the efficiency of pre-trained transformer architectures $\textit{without requiring re-training}$. Recent works have explored improving transformer efficiency, however often incur computationally expensive re-training procedures or depend on architecture-specific characteristics, thus impeding practical wide-scale adoption. To address these shortcomings, the OPTIN framework leverages intermediate feature distillation, capturing the long-range dependencies of model parameters (coined $\textit{trajectory}$), to produce state-of-the-art results on natural language, image classification, transfer learning, and semantic segmentation tasks $\textit{without re-training}$. Given a FLOP constraint, the OPTIN framework will compress the network while maintaining competitive accuracy performance and improved throughput. Particularly, we show a $\leq 2$% accuracy degradation from NLP baselines and a $0.5$% improvement from state-of-the-art methods on image classification at competitive FLOPs reductions. We further demonstrate the generalization of tasks and architecture with comparative performance using Mask2Former for semantic segmentation and cnn-style networks. OPTIN presents one of the first one-shot efficient frameworks for compressing transformer architectures that generalizes well across different class domains, in particular: natural language and image-related tasks, without $\textit{re-training}$.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Data-to-Model Distillation: Data-Efficient Learning Framework

    cs.CV 2024-11 conditional novelty 6.0 of 10

    D2M distills a dataset's knowledge into the parameters of a pre-trained GAN, enabling flexible, architecture-general synthetic training data with state-of-the-art classification accuracy.

  2. Adaptive Pruning of Pretrained Transformer via Differential Inclusions

    cs.LG 2025-01 conditional novelty 5.0 of 10

    A single differential-inclusion search over masks produces a whole family of pruned transformers at different sparsity levels from one pretrained model.

Pith tools