REVIEW 4 cited by
Datasheet for the Pile
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This datasheet describes the Pile, a 825 GiB dataset of human-authored text compiled by EleutherAI for use in large-scale language modeling. The Pile is comprised of 22 different text sources, ranging from original scrapes done for this project, to text data made available by the data owners, to third-party scrapes available online.
Forward citations
Cited by 4 Pith papers
-
Causal Estimation of Tokenisation Bias
Using regression discontinuity, the paper shows that adding a subword to a tokenizer's vocabulary can raise the model's probability for that string by up to about 17 times in small models.
-
Language Models Represent and Transform Concepts with Shared Geometry
Contextual displacements of concepts in LLMs form semantically organized vector fields whose relational geometry is shared across models and predicts held-out displacements above chance.
-
Natural Context Drift Undermines the Natural Language Understanding of Large Language Models
QA accuracy of open-weight LLMs drops as Wikipedia passages semantically drift from training-time content, while human accuracy stays flat.
-
How Quantization Impacts Privacy Risk on LLMs for Code?
Quantizing code LLMs reduces membership inference effectiveness, with 4-bit compression giving larger privacy and performance drops than 8-bit.
Discussion (0). Continue with ORCID to comment.