Pith. sign in

REVIEW 5 cited by

What Would Elsa Do? Freezing Layers During Transformer Fine-Tuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.03090 v1 pith:IQZWZ7AY submitted 2019-11-08 cs.CL

classification cs.CL
keywords layersfinalfine-tunedlanguagemodelsneedtasksacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pretrained transformer-based language models have achieved state of the art across countless tasks in natural language processing. These models are highly expressive, comprising at least a hundred million parameters and a dozen layers. Recent evidence suggests that only a few of the final layers need to be fine-tuned for high quality on downstream tasks. Naturally, a subsequent research question is, "how many of the last layers do we need to fine-tune?" In this paper, we precisely answer this question. We examine two recent pretrained language models, BERT and RoBERTa, across standard tasks in textual entailment, semantic similarity, sentiment analysis, and linguistic acceptability. We vary the number of final layers that are fine-tuned, then study the resulting change in task-specific effectiveness. We show that only a fourth of the final layers need to be fine-tuned to achieve 90% of the original quality. Surprisingly, we also find that fine-tuning all layers does not always help.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Vision Transformer Finetuning Benefits from Non-Smooth Components

    cs.LG 2026-02 conditional novelty 6.0 of 10

    For vision transformers, components with higher input-output sensitivity (attention and feedforward layers) yield better and more stable fine-tuning accuracy than smoother LayerNorm components.

  2. MoEcho: Exploiting Side-Channel Attacks to Compromise User Privacy in Mixture-of-Experts LLMs

    cs.CR 2025-08 reject novelty 6.0 of 10

    MoEcho claims to compromise user privacy in MoE LLMs and VLMs via four CPU and GPU side channels, but the provided manuscript body contains no supporting content.

  3. Tiny Reward Models

    cs.CL 2025-07 conditional novelty 6.0 of 10

    TinyRM shows that 400M-parameter bidirectional masked language models, tuned with FLAN-style prompting, DoRA, and layer freezing, outperform a 70B reward model on RewardBench reasoning and come close on safety.

  4. On Finetuning Tabular Foundation Models

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Full finetuning of TabPFNv2 outperforms in-context learning and partial finetuning on medium tabular datasets, and its gains come from sharper query-key attention that better reflects target similarity.

  5. DiffoRA: Enabling Parameter-Efficient Fine-Tuning via Differential Module Selection

    cs.CV 2025-02 reject novelty 4.0 of 10

    DiffoRA selects a subset of modules for LoRA fine-tuning using a learned binary mask, reporting improved accuracy over standard LoRA on GLUE and SQuAD.

Pith tools