Pith. sign in

REVIEW 1 cited by

Pre-training LLMs using human-like development data corpus

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.04666 v4 pith:ZZC2RY6F submitted 2023-11-08 cs.CL cs.AI

classification cs.CLcs.AI
keywords llmspre-traininglanguagenumberseenstricttasktokens
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pre-trained Large Language Models (LLMs) have shown success in a diverse set of language inference and understanding tasks. The pre-training stage of LLMs looks at a large corpus of raw textual data. The BabyLM shared task compares LLM pre-training to human language acquisition, where the number of tokens seen by 13-year-old kids is magnitudes smaller than the number of tokens seen by LLMs. In this work, we pre-train and evaluate LLMs on their ability to learn contextual word representations using roughly the same number of tokens as seen by children. We provide a strong set of baselines; with different architectures, evaluation of changes in performance across epochs, and reported pre-training metrics for the strict small and strict tracks of the task. We also try to loosely replicate the RoBERTa baseline given by the task organizers to observe the training robustness to hyperparameter selection and replicability. We provide the submission details to the strict and strict-small tracks in this report.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The potential -- and the pitfalls -- of using pre-trained language models as cognitive science theories

    cs.CL 2025-01 accept novelty 4.0 of 10

    Pretrained language models can serve as credible cognitive science theories only if researchers validate linking hypotheses and avoid pitfalls of commission and omission.

Pith tools