Pith. sign in

Paper Citation Record · LEDGER

Drop Dropout on Single-Epoch Language Model Pretraining

As of 10 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 1 inbound Pith citation observation for arXiv:2505.24788.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24788 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:18:56.233445Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T05:38:12.882914Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T05:41:02.181700Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ec892921-0533-4864-8a44-1ec55ded0f7b · outbound

This paper cites InProceedings of the 2021 Confer- ence on Empirical Methods in Natural Language Pro- cessing, pages 5484–5495, Online and Punta Cana, Dominican Republic.

Drop Dropout on Single-Epoch Language Model Pretraining InProceedings of the 2021 Confer- ence on Empirical Methods in Natural Language Pro- cessing, pages 5484–5495, Online and Punta Cana, Dominican Republic

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:58.303745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:18:54.922866Z digest=sha256:701338f20d69cb3f21ca7ca8fe28e7675572a84fa6e5ed5d3d7708c3a9dc2397

Observation d0fc1467-7b46-4b9d-8151-76d043095499 · outbound

This paper cites Association for Computational Linguistics.

Drop Dropout on Single-Epoch Language Model Pretraining Association for Computational Linguistics

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:57.917281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:18:55.307375Z digest=sha256:27548ebc6d94a1942660c70e17318368e1f182e1b538205a8eca86c337a5f3bd

Observation 8b9aed26-2221-4d6c-8020-ac3c1db84b21 · outbound

This paper cites an unresolved cited work.

Drop Dropout on Single-Epoch Language Model Pretraining Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:18:57.569992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:18:55.555633Z digest=sha256:d80a3ca5c88997179e8a5066800aa70a373cebe3800b78561ed542311dfb90cd

Observation 9c845021-90ce-4769-b02e-54015936c838 · outbound

This paper cites InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online.

Drop Dropout on Single-Epoch Language Model Pretraining InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:57.370618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:18:55.697998Z digest=sha256:47f5e46766a44aee2c97b2d1efc2abbc2ca9ed9eb1e04e3fcc37eba38e4603f9

Observation 60612f92-7e0d-45af-a219-1073cb261708 · outbound

This paper cites An adapted version of the official evaluation script was used to obtain the dev-slice results reported in this work.

Drop Dropout on Single-Epoch Language Model Pretraining An adapted version of the official evaluation script was used to obtain the dev-slice results reported in this work

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:56.523414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:18:56.183766Z digest=sha256:6130a69822b2ff063912aa512d2e96c447b1e9f8d4d8069e3477a5de3c6f6e72

Observation 97e2a51e-cf1b-4b21-9de7-94b0fdac0da8 · outbound

This paper cites When MLP dropout is used, p= 0.1.

Drop Dropout on Single-Epoch Language Model Pretraining When MLP dropout is used, p= 0.1

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:57.207353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:18:55.880575Z digest=sha256:ed72ac97d8e3fc7e651c3a5fe61ccafc107ae4493ebf727798d30c2975cff9e4

Observation 61ab845e-e8a5-4776-80d5-d0e8cf75666c · outbound

This paper cites Batching was done sequentially with the Pytorch Data Loader, sequence lengths are capped at 512 tokens.

Drop Dropout on Single-Epoch Language Model Pretraining Batching was done sequentially with the Pytorch Data Loader, sequence lengths are capped at 512 tokens

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:56.970449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:18:55.962071Z digest=sha256:a75009e35f011dc3c20d26dc494a1bb89e09c38f4c7addf9dfce25d4bebc0d26

Observation 20f42294-0e67-46a4-8297-ab6a97aa1076 · outbound

This paper cites Optimization was donewith regulariza- tionusing AdamW (Loshchilov and Hutter,.

Drop Dropout on Single-Epoch Language Model Pretraining Optimization was donewith regulariza- tionusing AdamW (Loshchilov and Hutter,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:56.712805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:18:56.060676Z digest=sha256:f110331abc181e7b65e5b1a577ce823889ee9dbfaf4dad11f38aef9e2e4d2f86

Observation 6d2db1f3-db5b-493c-9e77-02120fb5abc2 · outbound

This paper cites Batch size was set to 128, and dropout rate was set to 0.15 regardless of whether pretrain- ing the BERT model used dropout consistent with previous approaches.

Drop Dropout on Single-Epoch Language Model Pretraining Batch size was set to 128, and dropout rate was set to 0.15 regardless of whether pretrain- ing the BERT model used dropout consistent with previous approaches

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:56.414228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:18:56.233445Z digest=sha256:3a1a59d83ccec8f384ff32db337f6098aa029cd23f644ce46235d0d2e83b75a1

Observation c64bc1a1-f446-49a1-8444-2caa456e0456 · outbound

This paper cites Improving neural networks by preventing co-adaptation of feature detectors.

Drop Dropout on Single-Epoch Language Model Pretraining Improving neural networks by preventing co-adaptation of feature detectors

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:55.233415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:55.233415Z digest=sha256:db548f3378eb09783c26e61bd635d32d6b52738ace1c477c40c1273669b7ad04

Observation 24b0913e-3420-446b-9874-bdd1cad2b9f2 · outbound

This paper cites Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al.

Drop Dropout on Single-Epoch Language Model Pretraining Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:57.722046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:18:55.402456Z digest=sha256:0ff6ec623b083ba55ac32b4fad08beb111055474e02fb10b64336208a8f38913

Observation 8daa68c7-a906-4998-9fb2-58914a0172b7 · outbound

This paper cites InProceedings of the 8th Workshop on Cognitive Modeling and Com- putational Linguistics (CMCL 2018), pages 10–18, Salt Lake City, Utah.

Drop Dropout on Single-Epoch Language Model Pretraining InProceedings of the 8th Workshop on Cognitive Modeling and Com- putational Linguistics (CMCL 2018), pages 10–18, Salt Lake City, Utah

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:58.117199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:18:55.139696Z digest=sha256:48e3e0211ef1a3ee89dae59a90c3a3ade25a9b98dd9fb7dab7007f24433df24a

Observation a843f7c7-d1b9-4eec-8491-b4946b423fa7 · outbound

This paper cites an unresolved cited work.

Drop Dropout on Single-Epoch Language Model Pretraining Unresolved cited work

Reference 2019

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:18:58.481051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:18:54.555900Z digest=sha256:ab994a7229350d0e8cb640bf4ced41260b4b4eef7873737c841a50e8937cc2a5

Observation 856c6ebb-6a4b-4186-ae65-14f8d09836ee · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Drop Dropout on Single-Epoch Language Model Pretraining The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:54.703706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:54.703706Z digest=sha256:631c7bf319ebe93b3f2d425957e60113475773404691c0f6e5c3e6d217f37c3f

Observation 5d8ec641-5d65-4c91-a58c-9eddb65b716e · outbound

This paper cites InPro- ceedings of the 2021 Conference on Empirical Meth- ods in Natural Language Processing, pages 6491– 6506, Online and Punta Cana, Dominican Republic.

Drop Dropout on Single-Epoch Language Model Pretraining InPro- ceedings of the 2021 Conference on Empirical Meth- ods in Natural Language Processing, pages 6491– 6506, Online and Punta Cana, Dominican Republic

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:58.666912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:18:54.415114Z digest=sha256:607af98786d36b199db188b970b7cf38d21d23d75ccc5a54779b5433c2080eba

Observation 3fe71e64-6957-4a44-bb46-2f501760d989 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Drop Dropout on Single-Epoch Language Model Pretraining LLaMA: Open and Efficient Foundation Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:55.499541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:55.499541Z digest=sha256:18007007f19bbc59adbcf5262dbb43e99157e27cdce67457769be9be5a0e99e8

Observation 3f795873-d19c-46f8-9103-0ee8f35cf793 · outbound

This paper cites ReFT: Representation Finetuning for Language Models.

Drop Dropout on Single-Epoch Language Model Pretraining ReFT: Representation Finetuning for Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:55.792408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:55.792408Z digest=sha256:15fd7098f1fd4dc757ec74264d2e92fa359c5afc8ebca26fb00c92460c13bff1

Pith citing papers

Observation 09c0af84-9f63-4978-80fb-afcb9cce6080 · inbound

Language models recognize dropout and Gaussian noise applied to their activations cites this paper.

Language models recognize dropout and Gaussian noise applied to their activations Drop Dropout on Single-Epoch Language Model Pretraining

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:41:02.183113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T05:38:12.882914Z digest=sha256:06d496855ef30bc3813b6464dc961d81542e846d0200d08eb91634b32d86c8c9