REVIEW 1 cited by
AdapterDrop: On the Efficiency of Adapters in Transformers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Massively pre-trained transformer models are computationally expensive to fine-tune, slow for inference, and have large storage requirements. Recent approaches tackle these shortcomings by training smaller models, dynamically reducing the model size, and by training light-weight adapters. In this paper, we propose AdapterDrop, removing adapters from lower transformer layers during training and inference, which incorporates concepts from all three directions. We show that AdapterDrop can dynamically reduce the computational overhead when performing inference over multiple tasks simultaneously, with minimal decrease in task performances. We further prune adapters from AdapterFusion, which improves the inference efficiency while maintaining the task performances entirely.
Forward citations
Cited by 1 Pith paper
-
SSH: Sparse Spectrum Adaptation via Discrete Hartley Transformation
SSH fine-tunes large models by learning sparse Hartley-spectrum coefficients selected by energy of the pretrained weights, matching or beating LoRA and FourierFT with fewer parameters.
Discussion (0). Continue with ORCID to comment.