Pith. sign in

REVIEW 1 cited by

$\rm SP^3$: Enhancing Structured Pruning via PCA Projection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.16475 v3 pith:GOAJLNDG submitted 2023-08-31 cs.CL cs.AI

classification cs.CLcs.AI
keywords pruningstructuredaccuracycompressdimensioneffectivemethodsmodel
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Structured pruning is a widely used technique for reducing the size of pre-trained language models (PLMs), but current methods often overlook the potential of compressing the hidden dimension (d) in PLMs, a dimension critical to model size and efficiency. This paper introduces a novel structured pruning approach, Structured Pruning with PCA Projection (SP3), targeting the effective reduction of d by projecting features into a space defined by principal components before masking. Extensive experiments on benchmarks (GLUE and SQuAD) show that SP3 can reduce d by 70%, compress 94% of the BERTbase model, maintain over 96% accuracy, and outperform other methods that compress d by 6% in accuracy at the same compression ratio. SP3 has also proven effective with other models, including OPT and Llama. Our data and code are available at an anonymous repo.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. P$^2$ Law: Scaling Law for Post-Training After Model Pruning

    cs.AI 2024-11 conditional novelty 6.0 of 10

    Post-training loss of a pruned LLM follows a power-law curve fixed by pre-pruning model size, pruning rate, number of tokens, and the model's original loss.

Pith tools