Pith. sign in

Paper Citation Record · LEDGER

Drop Dropout on Single-Epoch Language Model Pretraining

As of 22 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 1 inbound Pith citation observation for arXiv:2505.24788.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.24788 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:18:56.233445Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T05:38:12.882914Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T05:41:02.181700Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ec892921-0533-4864-8a44-1ec55ded0f7b · outbound

This paper cites InProceedings of the 2021 Confer- ence on Empirical Methods in Natural Language Pro- cessing, pages 5484–5495, Online and Punta Cana, Dominican Republic.

Drop Dropout on Single-Epoch Language Model Pretraining InProceedings of the 2021 Confer- ence on Empirical Methods in Natural Language Pro- cessing, pages 5484–5495, Online and Punta Cana, Dominican Republic

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:58.303745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:18:54.922866Z digest=sha256:91d29a5945a74be8427ae0705805a208e3416c2186856b2ac2a137f716e50b0f

Observation d0fc1467-7b46-4b9d-8151-76d043095499 · outbound

This paper cites Association for Computational Linguistics.

Drop Dropout on Single-Epoch Language Model Pretraining Association for Computational Linguistics

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:57.917281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:18:55.307375Z digest=sha256:6522c00cef57fefa0fb7458140b764e5dfdc62de11bf4c6965de0c19725f5d82

Observation 8b9aed26-2221-4d6c-8020-ac3c1db84b21 · outbound

This paper cites an unresolved cited work.

Drop Dropout on Single-Epoch Language Model Pretraining Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:18:57.569992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:18:55.555633Z digest=sha256:9e86cf5ab60ef947f4c5e804f0a5aa794c6317e390eeb0217daa7dfb91470040

Observation 9c845021-90ce-4769-b02e-54015936c838 · outbound

This paper cites InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online.

Drop Dropout on Single-Epoch Language Model Pretraining InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:57.370618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:18:55.697998Z digest=sha256:96b487b04b7883e5f147093d678e3ce684f0bd23a012762faedad890bfb09bb7

Observation 60612f92-7e0d-45af-a219-1073cb261708 · outbound

This paper cites An adapted version of the official evaluation script was used to obtain the dev-slice results reported in this work.

Drop Dropout on Single-Epoch Language Model Pretraining An adapted version of the official evaluation script was used to obtain the dev-slice results reported in this work

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:56.523414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:18:56.183766Z digest=sha256:28f726ed381e880e23682efabdb1f6c98abc472ec3e2be2cf232389a39db335a

Observation 97e2a51e-cf1b-4b21-9de7-94b0fdac0da8 · outbound

This paper cites When MLP dropout is used, p= 0.1.

Drop Dropout on Single-Epoch Language Model Pretraining When MLP dropout is used, p= 0.1

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:57.207353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:18:55.880575Z digest=sha256:dba8ebf0bc25bf2a797a338e1f6bcfad534b9367ed50c4987c7c4ee3b796ed35

Observation 61ab845e-e8a5-4776-80d5-d0e8cf75666c · outbound

This paper cites Batching was done sequentially with the Pytorch Data Loader, sequence lengths are capped at 512 tokens.

Drop Dropout on Single-Epoch Language Model Pretraining Batching was done sequentially with the Pytorch Data Loader, sequence lengths are capped at 512 tokens

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:56.970449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:18:55.962071Z digest=sha256:417475f466490df515486ee244b56682a319747d4194a05a44d6cb72fd6f969c

Observation 20f42294-0e67-46a4-8297-ab6a97aa1076 · outbound

This paper cites Optimization was donewith regulariza- tionusing AdamW (Loshchilov and Hutter,.

Drop Dropout on Single-Epoch Language Model Pretraining Optimization was donewith regulariza- tionusing AdamW (Loshchilov and Hutter,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:56.712805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:18:56.060676Z digest=sha256:7e97603240aac30321aa86801348cfd736d7ce76c0888729672d55b3a8c39080

Observation 6d2db1f3-db5b-493c-9e77-02120fb5abc2 · outbound

This paper cites Batch size was set to 128, and dropout rate was set to 0.15 regardless of whether pretrain- ing the BERT model used dropout consistent with previous approaches.

Drop Dropout on Single-Epoch Language Model Pretraining Batch size was set to 128, and dropout rate was set to 0.15 regardless of whether pretrain- ing the BERT model used dropout consistent with previous approaches

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:56.414228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:18:56.233445Z digest=sha256:06a89b149c25579a99c45f142f7d1351634bfb285b90e8115cd7089ed1d662b6

Observation c64bc1a1-f446-49a1-8444-2caa456e0456 · outbound

This paper cites Improving neural networks by preventing co-adaptation of feature detectors.

Drop Dropout on Single-Epoch Language Model Pretraining Improving neural networks by preventing co-adaptation of feature detectors

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:55.233415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:55.233415Z digest=sha256:a5456ff64e8183ef11268197665071e40045da8b1a00ba532f87f9c074733511

Observation 24b0913e-3420-446b-9874-bdd1cad2b9f2 · outbound

This paper cites Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al.

Drop Dropout on Single-Epoch Language Model Pretraining Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:57.722046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:18:55.402456Z digest=sha256:ba5de3feeb13b9b25ac39ce6d391cbdee610e75295351674efeb0b70e10ffa85

Observation 8daa68c7-a906-4998-9fb2-58914a0172b7 · outbound

This paper cites InProceedings of the 8th Workshop on Cognitive Modeling and Com- putational Linguistics (CMCL 2018), pages 10–18, Salt Lake City, Utah.

Drop Dropout on Single-Epoch Language Model Pretraining InProceedings of the 8th Workshop on Cognitive Modeling and Com- putational Linguistics (CMCL 2018), pages 10–18, Salt Lake City, Utah

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:58.117199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:18:55.139696Z digest=sha256:2c1e0dec873701eb99c6e4a597411143f55ba5f99060a743edfa0b01f750a4f4

Observation a843f7c7-d1b9-4eec-8491-b4946b423fa7 · outbound

This paper cites an unresolved cited work.

Drop Dropout on Single-Epoch Language Model Pretraining Unresolved cited work

Reference 2019

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:18:58.481051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:18:54.555900Z digest=sha256:07c89cf121b3b78964a0aa04ae28eb0b440a7180053e253c6a206d1e76a61428

Observation 856c6ebb-6a4b-4186-ae65-14f8d09836ee · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Drop Dropout on Single-Epoch Language Model Pretraining The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:54.703706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:54.703706Z digest=sha256:424e712a4789cc274b1aaf8c68200bc5d289dbaa0ba6d8f697f77cebaedf6245

Observation 5d8ec641-5d65-4c91-a58c-9eddb65b716e · outbound

This paper cites InPro- ceedings of the 2021 Conference on Empirical Meth- ods in Natural Language Processing, pages 6491– 6506, Online and Punta Cana, Dominican Republic.

Drop Dropout on Single-Epoch Language Model Pretraining InPro- ceedings of the 2021 Conference on Empirical Meth- ods in Natural Language Processing, pages 6491– 6506, Online and Punta Cana, Dominican Republic

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:18:58.666912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:18:54.415114Z digest=sha256:eaef1c597cbf90e066bea06a09907b98bf2ed0153470eaccaa62dd7475dd713a

Observation 3fe71e64-6957-4a44-bb46-2f501760d989 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Drop Dropout on Single-Epoch Language Model Pretraining LLaMA: Open and Efficient Foundation Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:55.499541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:55.499541Z digest=sha256:2d3061f46668b0cea37d3a312a77c04d52948b1c7c85b51ac424af9bbe434a2e

Observation 3f795873-d19c-46f8-9103-0ee8f35cf793 · outbound

This paper cites ReFT: Representation Finetuning for Language Models.

Drop Dropout on Single-Epoch Language Model Pretraining ReFT: Representation Finetuning for Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:55.792408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:55.792408Z digest=sha256:6ba550974041ac3b9ef6b617a0c7e1bded5a861138dff5678a52163e63417a87

Pith citing papers

Observation 09c0af84-9f63-4978-80fb-afcb9cce6080 · inbound

Language models recognize dropout and Gaussian noise applied to their activations cites this paper.

Language models recognize dropout and Gaussian noise applied to their activations Drop Dropout on Single-Epoch Language Model Pretraining

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:41:02.183113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T05:38:12.882914Z digest=sha256:27eef70f5bf115894483d04d44886a3d1654e0f339a46dc845559a7ea1811620