Pith. sign in

Paper Citation Record · LEDGER

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data

As of 17 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2605.10129.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.10129 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T03:38:39.738401Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T05:33:53.927154Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T05:33:56.202284Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact3
  • verified fuzzy50
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch8

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d287a732-f633-40f0-bdb6-6b9b80251e28 · outbound

This paper cites 2025 , eprint =.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2025 , eprint =

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.235488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:66b18ff905d6413753f19a61c6447e29c01c80325f8dc8ddf484059650491d51

Observation 0f259c7e-1e20-426c-a65a-596c85b38de7 · outbound

This paper cites 2020 , eprint =.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2020 , eprint =

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.199058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:11a9e2c23318eae61de9b55235c7cafb0e8aa807071eccf60e5fcea282416205

Observation f9adb72c-2e3d-4924-a6b8-36338f9189dc · outbound

This paper cites 2023 , eprint =.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2023 , eprint =

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.250841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:8c4bfe9849fcb2f45950435c74ff94d86177fac40eabc1bd8295859b2cb0a738

Observation d70c2b7c-5b01-4234-989b-5dc30519babd · outbound

This paper cites 2023 , eprint=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2023 , eprint=

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.269897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:bb21d21fbed99cbcb331cd4f5dc5ee8a56d69786e5db899fd9dd452ded3cff41

Observation ed854f06-72b2-4435-8c32-7352400f8dff · outbound

This paper cites 2016 , eprint=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2016 , eprint=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.218638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:ad1f2ab888622eb919940394bd980107d693de950df741ac518bfcfcf25cf408

Observation 492389c6-3cdf-4966-b087-780a0dd86bdb · outbound

This paper cites 2023 , eprint=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2023 , eprint=

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.207303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:cc9204bea338d402839679d800f99278432d47a7dbd95f774e31d0e648c91c31

Observation cda71a16-1d14-48d0-abb4-c7ab090cec4e · outbound

This paper cites Journal of machine learning research , volume=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data Journal of machine learning research , volume=

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.180878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:b7cfc135c7d00baad241706bb7da6b67646da2a83ce48cacf4895935b0dee474

Observation e1e2027e-3cc9-4d7c-b501-64368c06d202 · outbound

This paper cites 2024 , eprint=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2024 , eprint=

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.191776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:a3d146adde1dc22cb575bbccd87860cf6bb255dbe7190e29449d6eec6c97b387

Observation 23329a6c-f76c-4ccf-8c70-bc261f15e40e · outbound

This paper cites 2024 , eprint=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2024 , eprint=

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.170042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:3f0fa97a53cf201cbc7d41bf464e295ed8c2b477f6dfb6909492b53297bf28be

Observation b083a3ea-d17a-4322-b4a4-218451aee362 · outbound

This paper cites 2019 , eprint=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2019 , eprint=

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.354528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:39f28ee6b632845ef63bf512eed46ea40a1a4d8631f13c1e7f01aa85bdb816d7

Observation 42705cf6-c977-4fa7-90a3-f7a193ee31eb · outbound

This paper cites Proceedings of the 26th annual international conference on machine learning , pages=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data Proceedings of the 26th annual international conference on machine learning , pages=

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.223503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:a23c890af4555553bd8e8a0047ca2b109d107283b87a0815966268ac2fe758a4

Observation 8801d9af-1640-48ab-9a77-54508cb3cb07 · outbound

This paper cites 2023 , eprint=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2023 , eprint=

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.227068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:717645962449b50b01b786becf8d569171984bd665b23bf6902ed89c65176644

Observation 17be6ef5-e260-4893-9a65-ecb99ea4d1a9 · outbound

This paper cites 2024 , eprint=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2024 , eprint=

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.195417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:730066cfbb9003ee6de940066f12b943373a7952149e37a4be4078159af86544

Observation 8eced79e-cccc-4ff8-b469-05e0efe860c1 · outbound

This paper cites 2024 , eprint=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2024 , eprint=

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.211449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:f41f9679f1951d1921734fd02c43354b60f6b8939817b8df530d2f6ea2bb048a

Observation 2cf2127d-59fe-421d-ba7a-84a4d3443ced · outbound

This paper cites an unresolved cited work.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-12T18:46:46.166265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:e8b037c0dab93ea6f62dc1d8bb8a70d8ebb3b0a7c213886336f77449bc7d5d70

Observation 818872a7-d3eb-4101-802f-0326c69cc1df · outbound

This paper cites 2022 , eprint=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2022 , eprint=

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.184085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:eceacf868b2f475ce2d13aea15d0b4869200087b45195e9f8d441aec1d144663

Observation 77783ced-bd0d-459a-95fc-263252d0dd7c · outbound

This paper cites 2023 , eprint=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2023 , eprint=

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.297101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:823e8328d6fc320f151dac918182a2c6163cd33ba72534cedba6186ee586700a

Observation aa669298-61bf-4744-8685-1e8843bf5820 · outbound

This paper cites 2019 , eprint=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2019 , eprint=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.285995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:f68f608a4375af889a1bceac7bd867261ed84bc5e205071ba19fab94074f4017

Observation 729c4318-0ba5-45a6-9d26-c0fa65bbcf1b · outbound

This paper cites 2019 , eprint=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2019 , eprint=

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.262744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:b26090d82a27868c99239705a8c4e37eac0a6a1153aebf3098afb3c60dd0af4f

Observation ea8cd691-c147-4be3-9c6e-1cd2a79187c3 · outbound

This paper cites 2021 , eprint=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2021 , eprint=

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.215510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:12f62d18d4f27030156097d5b98000eaed657cf4825533122f94fd15a0587bcd

Observation 151dd78a-33bb-4d9c-8229-fa6b5c263b65 · outbound

This paper cites 2022 , eprint=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2022 , eprint=

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.187708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:326062b7cfe128fca0a15e6da2eb36cadb3cbac13edf498db2990a1acee63e76

Observation 83e16d8b-c117-4cbc-b555-b4dab99d6b0e · outbound

This paper cites 2019 , eprint=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2019 , eprint=

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.331131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:3ba09f8d61bfc54bab136d6b6c9105c322731d64291454b4238de3e4829355ab

Observation 0a8fc2e3-8365-454a-a081-3f61172562cd · outbound

This paper cites 2020 , eprint =.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2020 , eprint =

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.174153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:e7bfbd2fcb899818bccca607438f994143e494448e88bb919e5d7c234e57ac02

Observation bab011ff-3097-4ebf-875a-92c60c30327a · outbound

This paper cites 2019 , eprint =.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2019 , eprint =

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.277811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:b3a5a5397b7303989795b89504aa2ab7c1393d78c80ad651430edbaca7b99018

Observation 7a7996d2-c3f0-46d2-9a89-3c701e2a78ce · outbound

This paper cites 2026 , eprint =.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2026 , eprint =

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.281584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:1af6ba76e4dbf86079de92a9b9840e86d43dbbb51ba5f693221c165264c65c45

Observation eee5a271-277a-45df-96d3-25c7bc40e090 · outbound

This paper cites 2023 , eprint =.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2023 , eprint =

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.202583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:acd258b24f32423ff6e5b56b21e72f5a489578fcadf9ae76517d16f2af2e1ded

Observation c7faa5e0-1d93-4e99-934e-81db9b9eb64e · outbound

This paper cites 2023 , eprint=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2023 , eprint=

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.266309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:69473cdb79a7c5ee477c5ccedfc1ae4acadb11dfc427920ea1479b66a9c45795

Observation 8adb8273-210c-4925-a467-a2cfa94f928c · outbound

This paper cites 2026 , eprint=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2026 , eprint=

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.254198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:9be246d52de55e81fe99ad2c3395f44c87a7ddf32700bd49e1da066e448e8786

Observation a8a15e44-01f1-4ff9-b07c-deb67889f734 · outbound

This paper cites 2025 , eprint =.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2025 , eprint =

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.300681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:34aa959e8ac4a62c419aebc75e2485bce0b86c09f4d249b8c0f74074206cd605

Observation e830adf2-4980-4f5d-85c2-ec17dd9d9d43 · outbound

This paper cites 2025 , eprint =.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2025 , eprint =

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.289798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:f5fae85be824556c0b4262df7ceea08b6a7d381ddba7fb535c297d94e6cc20fb

Observation 3b04163b-91ee-4411-baa7-d9cb8037a424 · outbound

This paper cites 2026 , eprint =.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2026 , eprint =

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.293490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:319a1fad9092421a2f7a2ba44c6ae6aafe0f2d638fa3b200a1e2c057aa28e8cf

Observation a89f1628-0e46-41c9-839c-bd4fd3848a24 · outbound

This paper cites 2025 , eprint =.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2025 , eprint =

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.361687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:13cbd6394f702ed75cd730f63ccf0dbe06fc756298cf18138820e8c44c13755c

Observation 2cadef74-2d1d-473a-ac86-086e18866327 · outbound

This paper cites 2025 , eprint=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2025 , eprint=

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.346507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:13227064550f8ff2d297a3fd3d4759ee55f71ce0eaccb672c438a1b3c01afa15

Observation 386ee380-9276-4481-a806-2fe871be326d · outbound

This paper cites 2021 , eprint=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2021 , eprint=

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.239097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:edfeef6a61bfdb4bfa1a6a5dc3fa4c057046ad695a2141461a0baa4c378ead59

Observation 98f4ac27-a15c-43ca-b9bf-329a73edf8b8 · outbound

This paper cites 2023 , eprint=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2023 , eprint=

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.315702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:99c27cc99dc3093774bfbf2287804d7bac9b0fa9b4f878473dfb46acf760943f

Observation a1dde3cc-463c-4e10-b40e-47f365573326 · outbound

This paper cites CharBERT: Character-aware Pre-trained Language Model , url=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data CharBERT: Character-aware Pre-trained Language Model , url=

Reference 36

Resolution
verified exact
doi, observed 2026-05-12T03:41:18.830806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:5d164b75c307659a2fccc7221ea21f20a224d0a18d2b8f09e144e9b97416ced4

Observation f474ad41-ac19-4da7-b030-f79182ae960c · outbound

This paper cites 2021 , eprint=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2021 , eprint=

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.308062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:62aa4baf3f2bda610308fc6263de9eee562e7ae7f05618f067be77fe93a5dfef

Observation c9e6492b-11b5-407d-8de9-2d6863e5f9ec · outbound

This paper cites 2026 , eprint=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2026 , eprint=

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.350569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:ea00b6c21bb11849e15244f85ea69066c0e106be6fdd00180bbe7b596b35064d

Observation ed5f108a-2c36-483c-b909-90f0b6ec5655 · outbound

This paper cites Proceedings of the fifth annual workshop on Computational learning theory , pages=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data Proceedings of the fifth annual workshop on Computational learning theory , pages=

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.327322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:5e99c3eb98549cedc50582c44df359c8646e0f783691c42c3222c6bc1f1dffc6

Observation c4e001ab-636e-48ba-b855-595511b6f612 · outbound

This paper cites echo state.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data echo state

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.358329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:71a5b91b72376c74752b74c1a072275aba2022d35c6e35915225502e7401e436

Observation bc94a105-5148-4b79-934d-b73e858c5e5e · outbound

This paper cites Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , pages=

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.304091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:7c40bc9d4b98dd7d5618d0bbb36ad377aea2044695dfadfcf4d3dd5200f5afe4

Observation 99ebdd60-b139-45b0-b786-47c73094e184 · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data DataComp-LM: In search of the next generation of training sets for language models

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T22:58:17.761776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:e5d81e1291ed5053ba129b976a3f7b980817f66bb26e5be5de9e5bb29e0b453a

Observation 12747c7f-6f4d-43f5-9509-588a0e47b182 · outbound

This paper cites Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:11:25.676262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:b289f5e743eb5f56cc8cb2bb7ead581b69a624641810732eba6a046696283224

Observation 978ced0b-b564-494d-bc32-4d9e41158f5d · outbound

This paper cites Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:11:25.667213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:02c0691eca544236d1e85660929dc360f09a1095783e55239ab175d2b0ff80a1

Observation 415dda88-a562-4c8b-923e-7f92fc27e751 · outbound

This paper cites 2021 , eprint=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2021 , eprint=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.342644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:2cc4a32449aaee4901e276e416699ddca1bc55a4c9593d9d10480e921d58cc62

Observation bbac2115-6314-4e2b-9496-5cf873fe255e · outbound

This paper cites 2018 , eprint =.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2018 , eprint =

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.243722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:fbf9c71cd45b1db3266bd9bbdedc37c809b66db373a825e7656a222b1da1ae9e

Observation 7aef18a7-ab57-4e18-89dc-bdae0e7560fe · outbound

This paper cites 2018 , eprint =.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2018 , eprint =

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.322396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:6952f2c084a2f2ac5acdf8e3291724704aa43750de5321454607a18b4bf3325d

Observation e1fc255c-527b-42af-a14e-540b78c0ce87 · outbound

This paper cites 2023 , eprint =.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2023 , eprint =

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.257814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:6b49b5bb0abbfa85a06489102c37045495fe3548c8dc5e469516a63798908922

Observation f8716b3e-fecc-4f62-8bef-5d1f815cf7fb · outbound

This paper cites 2024 , eprint =.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2024 , eprint =

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.319101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:ba36cd4aeba6f416831fdecbb11908350057a7e3dc6bef4162d99ace5ed01bf9

Observation f7c53a5a-a6e7-4cac-91bc-17cc28b58910 · outbound

This paper cites What's In My Big Data?.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data What's In My Big Data?

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:11:25.726433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:92b3770eb06914df72f5fa84bd0236e215c27a4507ba01089fcbb78942735f48

Observation 0d35ce35-d6a4-488b-9bc9-9718ee2571b9 · outbound

This paper cites Physics of Language Models: Part 3.1, Knowledge Storage and Extraction.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data Physics of Language Models: Part 3.1, Knowledge Storage and Extraction

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:11:25.715025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:dfec6382ea2168aa3c51ce9d8a06398564bcec7f440aec3b7ee5d1cf1ff287a7

Observation ab0cf40e-d3e9-400d-8816-7ab6d0e3cd69 · outbound

This paper cites Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:11:25.645042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:39863a7b5bf146f230f59de7ce3eb1627fae2c6466c7931fad747f4c92e4dd96

Observation f8e2c283-87ae-4273-9e98-9c468ff3cbf6 · outbound

This paper cites IEEE transactions on neural networks and learning systems , volume =.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data IEEE transactions on neural networks and learning systems , volume =

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.247430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:8d82d6a88cebf6740e5c8d5fccb28c15f959dba93d8fe4b833de0b19937f1199

Observation d321721b-aada-470a-9cee-26bc89547039 · outbound

This paper cites A Survey on Data Selection for Language Models.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data A Survey on Data Selection for Language Models

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:11:25.683736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:058677fc6e7330c977b63de976cbaf4268113d73b08a6618ec44ba430095e990

Observation 06f56ca4-35e0-4986-a6f7-3aed6d7455d1 · outbound

This paper cites FastText.zip: Compressing text classification models.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data FastText.zip: Compressing text classification models

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:11:25.705591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:793c5970a35b22303b0192ad0fe5ac01c5bbefc2bd4ae5b0c72e7b5dce479c08

Observation 496b429d-ddda-4fc9-8345-f0245b0a4046 · outbound

This paper cites 2019 , eprint =.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2019 , eprint =

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.334728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:9550dcfdde502d8b583662829ff183dfe4f8c9b677ba1a2775615e917221791a

Observation 2dbfe5d2-7f6c-4204-8f81-f212719b6796 · outbound

This paper cites Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data Understanding and Mitigating the Label Noise in Pre-training on Downstream Tasks

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:25.660670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:1fcb10dfed64af004ebd4545eaa4266ca1c208c618c39410dd9fe19c000ab561

Observation bd638b21-2036-4d50-a4a9-d7d365f2c294 · outbound

This paper cites 2017 , eprint =.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2017 , eprint =

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.338810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:73df7b26c749f2c48e62aa60f90c7fcf1d7cf6b44c2957a4ba0dcc88d96bb4e2

Observation 7a065798-f2c1-4324-aacd-0ad011cb591a · outbound

This paper cites 2017 , eprint =.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2017 , eprint =

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.311794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:a8772b44e79eba133d4583af956a1ad6f88a1af5f7a6926790ca96a74e6708f8

Observation 39960009-aeb8-413c-8ac2-00a2f6c3188d · outbound

This paper cites 2023 , eprint =.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data 2023 , eprint =

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.231678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:8a8418de091d489dd2ed861e799d10b5a481ef277db0a3af43f8ba0ba5dec4b6

Observation c22bcc74-a081-485b-b914-a2ce1f2d7e9e · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T18:46:46.274025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:77c013692aab2a8aa02b050151375ed9364f9b089c7e354cea92c1e6a0e0a42f

Observation 62030f58-7b87-407a-b593-e6793d47850c · outbound

This paper cites Transformers Pretrained on Procedural Data Contain Modular Structures for Algorithmic Reasoning.

Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data Transformers Pretrained on Procedural Data Contain Modular Structures for Algorithmic Reasoning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:11:25.693145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-12T03:38:39.738401Z digest=sha256:adadaff5d3eda5af7e0cf261f27d89d9547099f7e7dd6d005139514b8455f15a

Pith citing papers

Observation 5ca3996f-5305-417f-86a2-e3e2e34cfb33 · inbound

Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility cites this paper.

Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-05T05:33:56.215173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-05T05:33:53.927154Z digest=sha256:dcbb31f3b5bb80acbcff4b9cf8b88b523b68057eee25373bc37d37cf873cbf96