Pith. sign in

REVIEW 1 cited by

AVSS: Layer Importance Evaluation in Large Language Models via Activation Variance-Sparsity Analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.02117 v1 pith:RZ6PM6PH submitted 2024-11-04 cs.CL

classification cs.CL
keywords activationlanguagelayersmodelavssimportancelargelayer
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The evaluation of layer importance in deep learning has been an active area of research, with significant implications for model optimization and interpretability. Recently, large language models (LLMs) have gained prominence across various domains, yet limited studies have explored the functional importance and performance contributions of individual layers within LLMs, especially from the perspective of activation distribution. In this work, we propose the Activation Variance-Sparsity Score (AVSS), a novel metric combining normalized activation variance and sparsity to assess each layer's contribution to model performance. By identifying and removing approximately the lowest 25% of layers based on AVSS, we achieve over 90% of original model performance across tasks such as question answering, language modeling, and sentiment classification, indicating that these layers may be non-essential. Our approach provides a systematic method for identifying less critical layers, contributing to efficient large language model architectures.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EM-MIAs: Enhancing Membership Inference Attacks in Large Language Models through Ensemble Modeling

    cs.RO 2024-12 reject novelty 2.0 of 10

    Feeding four standard membership-inference scores into an XGBoost classifier yields higher AUC-ROC than the individual attacks on seven datasets for LLMs from 160M to 12B parameters.

Pith tools