Pith. sign in

REVIEW 1 cited by

Pretrained Language Model Embryology: The Birth of ALBERT

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.02480 v2 pith:BK4JYX4X submitted 2020-10-06 cs.CL

classification cs.CL
keywords modelpretrainedknowledgelanguagepretrainingduringalbertdifferent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While behaviors of pretrained language models (LMs) have been thoroughly examined, what happened during pretraining is rarely studied. We thus investigate the developmental process from a set of randomly initialized parameters to a totipotent language model, which we refer to as the embryology of a pretrained language model. Our results show that ALBERT learns to reconstruct and predict tokens of different parts of speech (POS) in different learning speeds during pretraining. We also find that linguistic knowledge and world knowledge do not generally improve as pretraining proceeds, nor do downstream tasks' performance. These findings suggest that knowledge of a pretrained model varies during pretraining, and having more pretrain steps does not necessarily provide a model with more comprehensive knowledge. We will provide source codes and pretrained models to reproduce our results at https://github.com/d223302/albert-embryology.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fairness Dynamics During Training

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Gender bias in Pythia-6.9b grows sharply after about 80k training steps even as general performance improves, and stopping earlier could trade 1.7% LAMBADA accuracy for a large fairness gain.

Pith tools