Pith. sign in

REVIEW 2 cited by

What Happens To BERT Embeddings During Fine-tuning?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.14448 v1 pith:X2Y4XEEX submitted 2020-04-29 cs.CL

classification cs.CL
keywords fine-tuningmodelbertfindrepresentationsaffectsanalysislinguistic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While there has been much recent work studying how linguistic information is encoded in pre-trained sentence representations, comparatively little is understood about how these models change when adapted to solve downstream tasks. Using a suite of analysis techniques (probing classifiers, Representational Similarity Analysis, and model ablations), we investigate how fine-tuning affects the representations of the BERT model. We find that while fine-tuning necessarily makes significant changes, it does not lead to catastrophic forgetting of linguistic phenomena. We instead find that fine-tuning primarily affects the top layers of BERT, but with noteworthy variation across tasks. In particular, dependency parsing reconfigures most of the model, whereas SQuAD and MNLI appear to involve much shallower processing. Finally, we also find that fine-tuning has a weaker effect on representations of out-of-domain sentences, suggesting room for improvement in model generalization.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Pre-training Time Series Models with Stock Data Customization

    cs.CE 2025-06 conditional novelty 6.0 of 10

    Stock Specialized Pre-trained Transformer (SSPT) uses stock code classification, sector classification, and moving average prediction as pre-training tasks, and reports improved stock selection returns and Sharpe rati...

  2. A Framework for Deductive Semantic Content Analysis at Scale in Science Education Using Text Embeddings

    cs.CL 2025-08 conditional novelty 4.0 of 10

    A few-shot text embedding classification framework achieves high agreement with human coders (Cohen's Kappa 0.74-0.83) on a simulated exhaustive coding task over 2,899 physics education survey responses.

Pith tools