Pith. sign in

Paper Citation Record · LEDGER

Do we really have to filter out random noise in pre-training data for language models?

As of 9 August 2026, this Paper Citation Record lists 100 of 138 outbound references and 2 inbound Pith citation observations for arXiv:2502.06604.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06604 v2

Coverage vector

measured 100 of 138 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:04:29.760015Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:53:41.436077Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:53:27.573309Z

Reference resolution

100 of 138 outbound references displayed

  • verified exact3
  • verified fuzzy17
  • unresolved80
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e60751f7-b8b0-4b28-a56d-96ef3d09d876 · outbound

This paper cites A pretrainer’s guide to training data: Measuring the effects of data age, domain coverage, quality, & toxicity,.

Do we really have to filter out random noise in pre-training data for language models? A pretrainer’s guide to training data: Measuring the effects of data age, domain coverage, quality, & toxicity,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.387312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.387312Z digest=sha256:7b3edc9b49caf8b39927a0bfe89caf3fbe1c1192039d713ae7fa8a60c5ed88e0

Observation e6135b2d-7382-4223-809c-3bf0fb78439a · outbound

This paper cites What’s in my big data?.

Do we really have to filter out random noise in pre-training data for language models? What’s in my big data?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.392350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.392350Z digest=sha256:904f408ca788e3fc18daf6583eaf0a27d9ef468aa75eafa7e85649af5039eb4f

Observation 2aa6417b-9931-4b3e-9186-7d3f83ea71c0 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Do we really have to filter out random noise in pre-training data for language models? LLaMA: Open and Efficient Foundation Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.396346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.396346Z digest=sha256:4a54d8e7b9a8941fe2dd9ffd0c66a80939b63276c870631264e585517981930f

Observation 38acb914-94a6-434e-8d8a-2e3fee02c94b · outbound

This paper cites Physics of language models: Part 3.1, knowledge storage and extraction,.

Do we really have to filter out random noise in pre-training data for language models? Physics of language models: Part 3.1, knowledge storage and extraction,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.400803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.400803Z digest=sha256:1693361cec9f4bed3a2e495c921520807113a3bdc3d78a724a7a0d7f53ec033b

Observation 7fea4bdf-ed93-45c8-9def-120f76dce541 · outbound

This paper cites Data selection for language models via importance resampling,.

Do we really have to filter out random noise in pre-training data for language models? Data selection for language models via importance resampling,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.404634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.404634Z digest=sha256:8ba403ec333212b753694e37e1b04df1fc3f0b5a7e3638b894c2ea71ad65961a

Observation 79e2dc12-4a69-4fca-8669-c043fca3a086 · outbound

This paper cites Ai models collapse when trained on recursively generated data,.

Do we really have to filter out random noise in pre-training data for language models? Ai models collapse when trained on recursively generated data,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.408327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.408327Z digest=sha256:4899350f025822ff4af06a7c1ae4d7d9fe0d59aaa67c34a1c0f2e6ce4f5f631b

Observation 22e7b213-867c-4393-9522-fd02e345e742 · outbound

This paper cites How bad is training on synthetic data? a statistical analysis of language model collapse,.

Do we really have to filter out random noise in pre-training data for language models? How bad is training on synthetic data? a statistical analysis of language model collapse,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.412362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.412362Z digest=sha256:b006584d885010b7851600f85423941da3c5c1fe27e18050e0c0c6145df8be99

Observation fa364b09-c266-4501-9349-a65b3182dada · outbound

This paper cites Leveraging Web-Crawled Data for High-Quality Fine-Tuning.

Do we really have to filter out random noise in pre-training data for language models? Leveraging Web-Crawled Data for High-Quality Fine-Tuning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-08T15:04:30.283408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:04:29.416371Z digest=sha256:837f7746fbcb9d8a01636f612b3ec12e403afbd64852000305bb811f0b27e157

Observation 75887f41-308c-4dc8-af72-e977df203bb1 · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing,.

Do we really have to filter out random noise in pre-training data for language models? Wavlm: Large-scale self-supervised pre-training for full stack speech processing,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.420304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.420304Z digest=sha256:5271a8aeccd0efd0859f5dd695335cf14888f5cdd7a0313f83dbce227c5f00e3

Observation 47cbc124-37d8-43cc-8496-f966bf5ad458 · outbound

This paper cites Noise-aware learning from web-crawled image-text data for image captioning,.

Do we really have to filter out random noise in pre-training data for language models? Noise-aware learning from web-crawled image-text data for image captioning,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.423875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.423875Z digest=sha256:5e1c1273aabd65480394e4fc0d0be2c2f35750fb4f8d1bbd8218a87dce403376

Observation 3946fc89-0778-4629-a98b-5f1a212527aa · outbound

This paper cites A survey on data selection for language models,.

Do we really have to filter out random noise in pre-training data for language models? A survey on data selection for language models,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.427445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.427445Z digest=sha256:55112fb6d4a9d1ad78ef07757c20ce2430b30447fff7d8e5fc6bab6815e8115b

Observation 92057b1e-53da-4d4e-94c9-6cebc1575b13 · outbound

This paper cites Dolma: an open corpus of three trillion tokens for language model pretraining research,.

Do we really have to filter out random noise in pre-training data for language models? Dolma: an open corpus of three trillion tokens for language model pretraining research,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.431091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.431091Z digest=sha256:1f71892225fe68d1ceda6370ef14e9d587c575f8ca32a62542a35346e98aec6d

Observation f2a68c19-904d-4a4d-9cd9-641711725e38 · outbound

This paper cites Openwebtext corpus,.

Do we really have to filter out random noise in pre-training data for language models? Openwebtext corpus,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.434719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.434719Z digest=sha256:dff0cd0c5418bb06b125ebdfd85301ffa4f0bee221a74208c5215aba76cc84d6

Observation 7b87a0bf-91bc-4dde-83b0-b4b8762b67b9 · outbound

This paper cites ATRI: Mitigating Multilingual Audio Text Retrieval Inconsistencies by Reducing Data Distribution Errors.

Do we really have to filter out random noise in pre-training data for language models? ATRI: Mitigating Multilingual Audio Text Retrieval Inconsistencies by Reducing Data Distribution Errors

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.438363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.438363Z digest=sha256:7bad9d9e2add9248e8607c4790453e433f90ff1bac435ff2886e6b1bab4fa140

Observation eeeb5c56-95d2-4e74-ba1a-1502d5f79c9d · outbound

This paper cites VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model.

Do we really have to filter out random noise in pre-training data for language models? VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.442501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.442501Z digest=sha256:e62fd1218777a3ea32b354fc932ec2179170a347d982f089135b5c1ae28080dd

Observation fe997328-c058-4a0c-b19e-22884066b795 · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

Do we really have to filter out random noise in pre-training data for language models? UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.446787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.446787Z digest=sha256:5c2e1a7be564606882f97e844635add00f0266a15120addd507f2cd2246be9d2

Observation 7e65718f-1a0d-4617-aef4-9ae9ac8b0aa1 · outbound

This paper cites Understanding and mitigating the label noise in pre-training on downstream tasks,.

Do we really have to filter out random noise in pre-training data for language models? Understanding and mitigating the label noise in pre-training on downstream tasks,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.450755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.450755Z digest=sha256:00cc1a73950732dec648ce846ef5f4b4a215a93c42a579da19855fde1b85fb42

Observation 00bdec02-4ec0-4253-b5e5-da38957682c8 · outbound

This paper cites Is out-of-distribution detection learnable?.

Do we really have to filter out random noise in pre-training data for language models? Is out-of-distribution detection learnable?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.454312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.454312Z digest=sha256:35aaa66b56baedcaa107248b3dbb8f163db16e1566d21835739c34fa7ad0ba73

Observation 8f696481-e67a-4059-80f9-c963ad7e472a · outbound

This paper cites Defweb: Defending user privacy against cache-based website fingerprinting attacks with intelligent noise injection,.

Do we really have to filter out random noise in pre-training data for language models? Defweb: Defending user privacy against cache-based website fingerprinting attacks with intelligent noise injection,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.458560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.458560Z digest=sha256:ea8508e608504d5c08da03d3a1bdca97f3ce7cb4d5ba0a32cf99526bbfd8583e

Observation fb20402c-8f72-47da-ab12-a1fd82f11a32 · outbound

This paper cites Language models are unsupervised multitask learners,.

Do we really have to filter out random noise in pre-training data for language models? Language models are unsupervised multitask learners,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.461908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.461908Z digest=sha256:297926fca9869175a25c5954a86d3e6015345e357d267ddf05f3758fe2755022

Observation 7ef48a16-b7d7-4e12-9042-1baab41f0a3f · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Do we really have to filter out random noise in pre-training data for language models? The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.465197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.465197Z digest=sha256:824b22b2b0be53e73fd043af18d995a6f4f20fda049bf230b08fa9051fe02898

Observation 7b6f107f-4494-4862-8644-2aedd2d01f2c · outbound

This paper cites Shalev-Shwartz and S.

Do we really have to filter out random noise in pre-training data for language models? Shalev-Shwartz and S

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.468714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.468714Z digest=sha256:9fecb0448ef3374bbc91dc703a32003d7cef97ae93e5809c61f80abd87f8998e

Observation c076fbe9-7998-4093-afc9-25de527a1229 · outbound

This paper cites A theory of learning from different domains,.

Do we really have to filter out random noise in pre-training data for language models? A theory of learning from different domains,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.471874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.471874Z digest=sha256:948ae7875899cef63c95c4b4e74762d6a2e57d1f385d2fbc2d073c5f43b41ac4

Observation c316f708-bb04-49f2-8cb9-6feebcd9af46 · outbound

This paper cites A mathematical exploration of why language models help solve downstream tasks,.

Do we really have to filter out random noise in pre-training data for language models? A mathematical exploration of why language models help solve downstream tasks,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.475290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.475290Z digest=sha256:f492e3887aeaec101eb53127f007085892a8a9c61d38e0e0f1fca42df8d31cfb

Observation fefa69fb-4802-412d-a70d-7a67c616aae2 · outbound

This paper cites Why do pretrained language models help in downstream tasks? an analysis of head and prompt tuning,.

Do we really have to filter out random noise in pre-training data for language models? Why do pretrained language models help in downstream tasks? an analysis of head and prompt tuning,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.478461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.478461Z digest=sha256:a8c494a063d44f8763507886f3ecbef63f1cda5173e3bab7548b19e92914133b

Observation b270e9a9-3d55-46d3-88ac-99334c3d53a7 · outbound

This paper cites Same pre-training loss, better downstream: Implicit bias matters for language models,.

Do we really have to filter out random noise in pre-training data for language models? Same pre-training loss, better downstream: Implicit bias matters for language models,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.482138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.482138Z digest=sha256:873334596528b1d8adc68863f269c95df892c08782f483ab536345a8f6450621

Observation 94bbd6ee-7abd-4753-b62f-edf8feb6b7a2 · outbound

This paper cites Revisiting discriminative vs. generative classifiers: Theory and implications,.

Do we really have to filter out random noise in pre-training data for language models? Revisiting discriminative vs. generative classifiers: Theory and implications,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.485597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.485597Z digest=sha256:c6a882be6c93318cdcaf6f8cdee7e3f5fe31d277e5b7c2abcdb77a881298ac5f

Observation 00b13d9d-13bd-43d5-83df-b2b70bebe7d0 · outbound

This paper cites How multilingual is multilingual BERT?.

Do we really have to filter out random noise in pre-training data for language models? How multilingual is multilingual BERT?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.489312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.489312Z digest=sha256:a9a2dd628e560993aa752918d55f9921ecd35443d25b14dbff4ba481d0fd3147

Observation 866bd5a5-8db5-4929-95bb-cc9f993183dd · outbound

This paper cites Finding universal grammatical relations in multilingual BERT,.

Do we really have to filter out random noise in pre-training data for language models? Finding universal grammatical relations in multilingual BERT,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.492969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.492969Z digest=sha256:ffc6ef645fb43ea946ea518e9afcf59e97dcd5e53564b35d48ea6b2b8899e2e6

Observation 629bf092-6af7-4abd-9925-3e39f4c1718a · outbound

This paper cites Zeronlg: Aligning and autoencoding domains for zero-shot multimodal and multilingual natural language generation,.

Do we really have to filter out random noise in pre-training data for language models? Zeronlg: Aligning and autoencoding domains for zero-shot multimodal and multilingual natural language generation,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.496487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.496487Z digest=sha256:df4f32e19ec8b41c939fb7487b3518433a298093e4739ac1a4dbe92a3fa3e4f4

Observation 252ec1b0-aecb-4c11-b068-5b2db3f1d3d5 · outbound

This paper cites VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning.

Do we really have to filter out random noise in pre-training data for language models? VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.499875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.499875Z digest=sha256:77609f494537e81a9e0c7e616c5dcae37c5e946b71ba1bb7683a468faef9eeaf

Observation 5bea82d2-a5dc-4f64-9633-aade6bdb05c6 · outbound

This paper cites ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language Modeling.

Do we really have to filter out random noise in pre-training data for language models? ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language Modeling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.503731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.503731Z digest=sha256:9a13ba4eadbad4ece3a16d62a6342f67ce0801ab330d2b3f8684fc61c9bdfce1

Observation a1ab164d-cf59-4c48-9c5a-0deb307b7619 · outbound

This paper cites Stochastic collapse: How gradient noise attracts SGD dynamics towards simpler subnetworks,.

Do we really have to filter out random noise in pre-training data for language models? Stochastic collapse: How gradient noise attracts SGD dynamics towards simpler subnetworks,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.507769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.507769Z digest=sha256:7f9b58253d2cea067b35c86182256e235d8678409f726240ee3a6bbf1bca4590

Observation eba9f5e8-8580-4847-bf29-5fcc0d4d2285 · outbound

This paper cites Positive-negative momentum: Manipulating stochastic gradient noise to improve generalization,.

Do we really have to filter out random noise in pre-training data for language models? Positive-negative momentum: Manipulating stochastic gradient noise to improve generalization,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.511466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.511466Z digest=sha256:7fcf8845ef587fb1e5f8c8349e6d965d1f211affa22bdb79acaf8988ba3d0e5a

Observation 584760e0-207c-484d-a224-1ca056c494fe · outbound

This paper cites Noise stability regularization for improving BERT fine-tuning,.

Do we really have to filter out random noise in pre-training data for language models? Noise stability regularization for improving BERT fine-tuning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.514949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.514949Z digest=sha256:b3973f25cfe20863619a2be7eb5ba0b2ace307f3c6b914778e19a688b1958336

Observation 23805558-8f6d-4dc5-b82c-fd51353079f7 · outbound

This paper cites A diffusion theory for deep learning dynamics: Stochastic gradient descent exponentially favors flat minima,.

Do we really have to filter out random noise in pre-training data for language models? A diffusion theory for deep learning dynamics: Stochastic gradient descent exponentially favors flat minima,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.518831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.518831Z digest=sha256:5b6e8f546a027d59b4cab433d559404850c53e3c6d3ce2db08e7cdefe822c714

Observation bfa49cda-dd51-47e8-b1ea-6a75032f69c1 · outbound

This paper cites Unveiling the structure of wide flat minima in neural networks,.

Do we really have to filter out random noise in pre-training data for language models? Unveiling the structure of wide flat minima in neural networks,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.522524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.522524Z digest=sha256:7f2fd9aa35136c4779c28d3db797153669150d21c22fb9d85e1407818e9e8e59

Observation 9a9478f6-edb8-4bdc-a2b4-c2d552f370ec · outbound

This paper cites The Llama 3 Herd of Models.

Do we really have to filter out random noise in pre-training data for language models? The Llama 3 Herd of Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.530022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.530022Z digest=sha256:3eb13c64883c76e43ab7da05057b13b106c5dadaff2dca09524e06fa4887a5fc

Observation a210335c-813c-4deb-b5e8-0c2af119d242 · outbound

This paper cites Le Gall, Measure theory, probability, and stochastic processes.

Do we really have to filter out random noise in pre-training data for language models? Le Gall, Measure theory, probability, and stochastic processes

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.533826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.533826Z digest=sha256:d0f1791329d83dd39932079b0afa7676db712f38ef8e9d156f31d9e79e8fe1a3

Observation 0088e24e-719c-4f9c-b46e-38fe8f375052 · outbound

This paper cites Structured pruning of large language models,.

Do we really have to filter out random noise in pre-training data for language models? Structured pruning of large language models,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.537245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.537245Z digest=sha256:0c4b9eaeb89e123f12d3f095f137922c0b41e62079632d59685d9b477cc6825c

Observation 79431427-9586-4017-af28-f6cc6264c527 · outbound

This paper cites The optimal BERT surgeon: Scalable and accurate second-order pruning for large language models,.

Do we really have to filter out random noise in pre-training data for language models? The optimal BERT surgeon: Scalable and accurate second-order pruning for large language models,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.540868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.540868Z digest=sha256:c419778edffcced61bc5f39453fc4aec297b77dee25873bdd3fab040d90c7d46

Observation fdad228a-4fb6-4cf7-8a4f-f0d2fc69244b · outbound

This paper cites Plug-and- play: An efficient post-training pruning method for large language models,.

Do we really have to filter out random noise in pre-training data for language models? Plug-and- play: An efficient post-training pruning method for large language models,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.544412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.544412Z digest=sha256:f1ef3de5e3c30cc4ccc9d8fc45cdfac101b17e4dd85fd000a69f7f822db760bd

Observation 68b6c3f1-818d-4220-8f1c-77cb64c3c97f · outbound

This paper cites LRQuant: Learnable and robust post-training quantization for large language models,.

Do we really have to filter out random noise in pre-training data for language models? LRQuant: Learnable and robust post-training quantization for large language models,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.547924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.547924Z digest=sha256:98e582bec1dada784729d24172f2c2a9b09503cb172f9f7d1c18d125d428edf5

Observation e3b849ba-d072-4778-99f5-86790833e8b4 · outbound

This paper cites A comprehensive evaluation of quantization strategies for large language models,.

Do we really have to filter out random noise in pre-training data for language models? A comprehensive evaluation of quantization strategies for large language models,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.551345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.551345Z digest=sha256:88355d89a24b02f50544aec7b7d038ec9a0486fd59af305788c73b5130d5d174

Observation 5eb38568-ab4d-4861-8590-ba97b05a7fee · outbound

This paper cites IntactKV: Improving large language model quantization by keeping pivot tokens intact,.

Do we really have to filter out random noise in pre-training data for language models? IntactKV: Improving large language model quantization by keeping pivot tokens intact,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.554973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.554973Z digest=sha256:9d294382eaeb0b7d1b19c5a37a6da6eec433b2c5d78f8c76ec2a40dbf4cf8969

Observation 1382e693-6788-4267-8b0f-73b44233d5a4 · outbound

This paper cites Cost-effective distillation of large language models,.

Do we really have to filter out random noise in pre-training data for language models? Cost-effective distillation of large language models,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.558646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.558646Z digest=sha256:adfea7fb91a0a3f9a099bace7ec778e778d44a543c876066581bdbb831fa8726

Observation 45dbc5fa-08e0-4ea5-b570-82f66ece80ba · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Do we really have to filter out random noise in pre-training data for language models? Distilling the Knowledge in a Neural Network

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.562282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.562282Z digest=sha256:f89115bd0b1ede34863dd2404984f0dc41024d355fe6cbaf97ed38254b94597b

Observation e3dd482d-64b7-4cd8-a400-c81c0bd365be · outbound

This paper cites mGPT: Few-shot learners go multilingual,.

Do we really have to filter out random noise in pre-training data for language models? mGPT: Few-shot learners go multilingual,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.566153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.566153Z digest=sha256:974c13f10b01f8251244412adcc789ddece6d70b65c17ccd79af099220dab3e4

Observation 8fbaf7c4-ea48-4cfa-9835-59c80f03d751 · outbound

This paper cites Toward understanding generative data augmentation,.

Do we really have to filter out random noise in pre-training data for language models? Toward understanding generative data augmentation,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.570112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.570112Z digest=sha256:aa42a1470be99c3912a4c9684391b55582937e4f50fb3a35f541804114e8ee2f

Observation 29e76057-9c23-4af4-b06d-b8432af52151 · outbound

This paper cites Smoothness, low noise and fast rates,.

Do we really have to filter out random noise in pre-training data for language models? Smoothness, low noise and fast rates,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.573984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.573984Z digest=sha256:807c9a9947086273fde5f4d2c52c86aa0c864df13c8a29cd2f3efb8aa1bf2daf

Observation 5d3296b2-cd1c-42e4-aa74-7ea80d1f121a · outbound

This paper cites Diffsound: Discrete diffusion model for text-to-sound generation,.

Do we really have to filter out random noise in pre-training data for language models? Diffsound: Discrete diffusion model for text-to-sound generation,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.577927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.577927Z digest=sha256:06923e9b81efbfe59274a7005a7837c078d350e186f68aafab092c4ff4d0ef07

Observation fe0d3fa3-778b-42b8-b5d6-0736991d1b37 · outbound

This paper cites Towards explainable joint models via information theory for multiple intent detection and slot filling,.

Do we really have to filter out random noise in pre-training data for language models? Towards explainable joint models via information theory for multiple intent detection and slot filling,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.581544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.581544Z digest=sha256:688c3c011ac6b3a26d08162fc719452fd0a040fb2897fc3a94041a79a4138584

Observation c0b1eddb-9c65-4d41-9060-d87339143893 · outbound

This paper cites Kdpror: A knowledge-decoupling probabilistic framework for video-text retrieval,.

Do we really have to filter out random noise in pre-training data for language models? Kdpror: A knowledge-decoupling probabilistic framework for video-text retrieval,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.585134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.585134Z digest=sha256:0a1f537a6bcb1612c8f12c1aa2cfb7bbc646580b38330bc1ceea43f1aa5a7af8

Observation 62813d17-e566-4673-a8e8-1f57d73c2d0d · outbound

This paper cites Gpa: global and prototype alignment for audio-text retrieval,.

Do we really have to filter out random noise in pre-training data for language models? Gpa: global and prototype alignment for audio-text retrieval,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.588827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.588827Z digest=sha256:1a10221ef502f26528aba85b7c89fb9f6e7bb484cfad9efa74baccffdf9a2311

Observation 8f816ae0-4ba1-4cc6-bf05-757899058524 · outbound

This paper cites Preparing lessons for progressive training on language models,.

Do we really have to filter out random noise in pre-training data for language models? Preparing lessons for progressive training on language models,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.592013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.592013Z digest=sha256:5d676f2e59dc47fdaf16d3fdd6d65609fb11c6d8bd2626075f16a05249282ee5

Observation ae724c0e-fb80-42b1-920b-9625418b0727 · outbound

This paper cites Do as We Do, Not as You Think: the Conformity of Large Language Models.

Do we really have to filter out random noise in pre-training data for language models? Do as We Do, Not as You Think: the Conformity of Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.595184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.595184Z digest=sha256:489f74db7ec0f7e5eec82bcbc984de333cf7f718ac94af36d29e9cea75322ea9

Observation e962eadd-59e1-41c6-acdc-935f0a37f0c1 · outbound

This paper cites Rmt: Retentive networks meet vision transformers,.

Do we really have to filter out random noise in pre-training data for language models? Rmt: Retentive networks meet vision transformers,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.598822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.598822Z digest=sha256:c986ff58344e467a2f9aad084e879d76cc2c7ed079725d65e2e45debe0278bdd

Observation 3c947bd0-3178-46e6-9632-5cd030dc9219 · outbound

This paper cites Reusing pretrained models by multi-linear operators for efficient training,.

Do we really have to filter out random noise in pre-training data for language models? Reusing pretrained models by multi-linear operators for efficient training,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.602320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.602320Z digest=sha256:0039d9e278449a9bba39225762197a67f54a4d6b7ef2f56b73b38805d2e3dd5d

Observation baeeb66d-2163-40d3-937e-2d8e8f4bc873 · outbound

This paper cites Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens.

Do we really have to filter out random noise in pre-training data for language models? Semantic Equitable Clustering: A Simple and Effective Strategy for Clustering Vision Tokens

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.606779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.606779Z digest=sha256:21316eeb4a3ae50a2780098076e118cfa2aabef6e450c806c84dc800fd9f2124

Observation f19448a6-bf5b-49ad-854e-7502baf40aed · outbound

This paper cites Pcad: Towards asr- robust spoken language understanding via prototype calibration and asymmetric decoupling,.

Do we really have to filter out random noise in pre-training data for language models? Pcad: Towards asr- robust spoken language understanding via prototype calibration and asymmetric decoupling,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.610887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.610887Z digest=sha256:6f922a61e6fa9710c436bb06903764e21c5f3d31ac4b41e8191f5e445889f5ac

Observation 26d058df-78ad-4d0e-a689-e49ae0853d94 · outbound

This paper cites Macsc: Towards multimodal- augmented pre-trained language models via conceptual prototypes and self-balancing calibra- tion,.

Do we really have to filter out random noise in pre-training data for language models? Macsc: Towards multimodal- augmented pre-trained language models via conceptual prototypes and self-balancing calibra- tion,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.614649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.614649Z digest=sha256:26ca4de341b1a7f73fc1ce27b914fb4e574f2f1045a481af6849148ab478bf7e

Observation 8d564f31-603c-406e-b171-5f5efd47bbc6 · outbound

This paper cites HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec.

Do we really have to filter out random noise in pre-training data for language models? HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.618212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.618212Z digest=sha256:1791a3beb8126d023be70d941c915ec3372f44fa2a8d3f6f96296f3bb4e4b1af

Observation 247aa966-32d4-440b-945a-b96d11c7df17 · outbound

This paper cites UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner.

Do we really have to filter out random noise in pre-training data for language models? UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.622125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.622125Z digest=sha256:e9857f598c471380e1be7344c32aaac7983b5e88d902984d106e9f3ff6bc5868

Observation 0fa6f811-b05d-4edc-bc88-d3cdb4d87d0d · outbound

This paper cites FedZKP: Federated Model Ownership Verification with Zero-knowledge Proof.

Do we really have to filter out random noise in pre-training data for language models? FedZKP: Federated Model Ownership Verification with Zero-knowledge Proof

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.625853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.625853Z digest=sha256:62a2469cf81a66d9495b4bdcdf591d069b5de3a980129a6d0577e9f8d84a05dd

Observation 24cbec1a-f8e5-4c3d-90fc-2d2cd14ac512 · outbound

This paper cites FedSOV: Federated Model Secure Ownership Verification with Unforgeable Signature.

Do we really have to filter out random noise in pre-training data for language models? FedSOV: Federated Model Secure Ownership Verification with Unforgeable Signature

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-08-08T15:04:30.120370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:04:29.629607Z digest=sha256:1fd00a3d54c17cb6bcdc2182ca916794e780adf63c7a4bec4eb8099fa81c3aa7

Observation c78fe4ee-1fad-4b44-9d0c-300efe993bd4 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

Do we really have to filter out random noise in pre-training data for language models? Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.633527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.633527Z digest=sha256:ee1c51ea5fd29e4a30e5202a3b3478fe018c9fda6f9203f0f400608f21f5c696

Observation 53abca1b-6c74-4c3c-ba55-f33eee8ce64b · outbound

This paper cites Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws.

Do we really have to filter out random noise in pre-training data for language models? Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.637076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.637076Z digest=sha256:6464de2069474335af57876f76c3fb34b8b46050acde5abf316e806ba5bbfc2b

Observation ef3112b6-b093-4aff-a072-3fedc2da4819 · outbound

This paper cites Parameterized algorithms and complexity for the traveling purchaser problem and its variants,.

Do we really have to filter out random noise in pre-training data for language models? Parameterized algorithms and complexity for the traveling purchaser problem and its variants,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.640544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.640544Z digest=sha256:49b695233154bf094aead8412897d7ad507bbb1f5250260ec66c0576aee0e43c

Observation 8aa7ae6b-ecb4-44ca-847b-8749c533de5a · outbound

This paper cites Dataset pruning: Reducing training data by examining generalization influence,.

Do we really have to filter out random noise in pre-training data for language models? Dataset pruning: Reducing training data by examining generalization influence,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.644119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.644119Z digest=sha256:27a2f71afe294e14bd169ec39232db5e843396dbcd11d8638fd645833fd9acb6

Observation da8125c3-4813-4f7e-9005-a336803ab07e · outbound

This paper cites On training data influence of GPT models,.

Do we really have to filter out random noise in pre-training data for language models? On training data influence of GPT models,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.648416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.648416Z digest=sha256:a10e288571c9543d2573212f7eaabb3a1d25e896536bc67a012229743154f304

Observation 2388707d-a097-4d8c-8237-5941474258e6 · outbound

This paper cites From quantity to quality: Boosting LLM performance with self-guided data selection for instruction tuning,.

Do we really have to filter out random noise in pre-training data for language models? From quantity to quality: Boosting LLM performance with self-guided data selection for instruction tuning,

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.652719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.652719Z digest=sha256:69f2439537fbeb8759e06d0038b2ffadf60be6b35f05a3bee3ac141e94d29767

Observation f5423c0c-be98-4c67-8dd4-13da3ec33721 · outbound

This paper cites Doremi: Optimizing data mixtures speeds up language model pretraining,.

Do we really have to filter out random noise in pre-training data for language models? Doremi: Optimizing data mixtures speeds up language model pretraining,

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.656553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.656553Z digest=sha256:49bca8aca798a6bd5c0c67a5a590ad26bf38d3bcfcf144e6849d4995a43b8c5b

Observation 4b190eac-44c4-4306-b116-456537a6340d · outbound

This paper cites Beyond Scale: The Diversity Coefficient as a Data Quality Metric for Variability in Natural Language Data.

Do we really have to filter out random noise in pre-training data for language models? Beyond Scale: The Diversity Coefficient as a Data Quality Metric for Variability in Natural Language Data

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.660572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.660572Z digest=sha256:bfa9b9e6c1599749a4afdb95e81fba8f13afc5dbb2127c4a69e33985b80f61f2

Observation edb1e5e7-3608-478f-8b80-6b770f8fada6 · outbound

This paper cites Talking Nonsense: Probing Large Language Models' Understanding of Adversarial Gibberish Inputs.

Do we really have to filter out random noise in pre-training data for language models? Talking Nonsense: Probing Large Language Models' Understanding of Adversarial Gibberish Inputs

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.664691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.664691Z digest=sha256:c69285eaa30ff3022e5b65805f78b8b2c4b490c0bfd52dd987cc311d16e7d6d2

Observation 28c16ea2-b981-4783-87fd-ba387ec0ee1f · outbound

This paper cites A comprehensive survey on source-free domain adaptation,.

Do we really have to filter out random noise in pre-training data for language models? A comprehensive survey on source-free domain adaptation,

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.668537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.668537Z digest=sha256:15eaa54114b7c8975253811012a4eae2e648cea3dcd1b554551446462fd7e81d

Observation ee691dd7-34a9-4f4e-a8b8-42c26fdd5635 · outbound

This paper cites Cross-domain mutual information adversarial maximiza- tion,.

Do we really have to filter out random noise in pre-training data for language models? Cross-domain mutual information adversarial maximiza- tion,

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.672375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.672375Z digest=sha256:dff587b047d1f542222afc0e7634ec860bcfaa7671e56f5f22d3eca23f1e15cb

Observation 16e8b51b-0a27-4e39-9c34-53260530560e · outbound

This paper cites Domain adaptive land-cover classification via local consistency and global diversity,.

Do we really have to filter out random noise in pre-training data for language models? Domain adaptive land-cover classification via local consistency and global diversity,

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.676159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.676159Z digest=sha256:e2d0f1955c3cee8080ba73e85571ec8c23b92641aea8a9aade1747ff6501f456

Observation 0ee60054-d64d-4308-957b-6be7389cadb3 · outbound

This paper cites Diffusion-based probabilistic uncertainty estimation for active domain adaptation,.

Do we really have to filter out random noise in pre-training data for language models? Diffusion-based probabilistic uncertainty estimation for active domain adaptation,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:04:30.843335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:04:29.679733Z digest=sha256:86e62b166ae74b1bb118fef9831cef8cd1790def6c6c553723709bb55513db87

Observation d46f6f92-ff54-4a9f-a061-090ee4327ae1 · outbound

This paper cites Online adaptive fault diagnosis with test-time domain adaptation,.

Do we really have to filter out random noise in pre-training data for language models? Online adaptive fault diagnosis with test-time domain adaptation,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:04:30.832007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:04:29.683655Z digest=sha256:a19401150ebaa1762cf94a16110f9df21789a3787d36cf18e36b45b8c0dfa156

Observation 93f0ad21-3f65-4d91-99ec-f10a5728d2fe · outbound

This paper cites Divergence-agnostic unsupervised domain adaptation by adversarial attacks,.

Do we really have to filter out random noise in pre-training data for language models? Divergence-agnostic unsupervised domain adaptation by adversarial attacks,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:04:30.820172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:04:29.687543Z digest=sha256:06c9f3b3d65eb64913a72f074ae3c041cab582c2ae96d8bc049d2e779c68a9f5

Observation 30f9f738-75a4-4e6a-9435-96a8c1fcd894 · outbound

This paper cites Imbalanced open set domain adaptation via moving-threshold estimation and gradual alignment,.

Do we really have to filter out random noise in pre-training data for language models? Imbalanced open set domain adaptation via moving-threshold estimation and gradual alignment,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:04:30.809300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:04:29.691420Z digest=sha256:1e150d9fc9781ecd52c0bcad557e6a8e9b4ecb03575d3e9054a5bdc7ac719147

Observation a5e4410d-7ab1-421c-b67f-402a748c8713 · outbound

This paper cites Learning from noisy labels with deep neural networks: A survey,.

Do we really have to filter out random noise in pre-training data for language models? Learning from noisy labels with deep neural networks: A survey,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:04:30.798241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:04:29.695000Z digest=sha256:45744e26634a7c2b3d47c35c9bfbe7d46e826fb46fbb8822711577a5d312fcb2

Observation 80dbfc26-d9e2-40fb-845e-ee3750565c80 · outbound

This paper cites Does label smoothing mitigate label noise?.

Do we really have to filter out random noise in pre-training data for language models? Does label smoothing mitigate label noise?

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:04:30.787591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:04:29.698504Z digest=sha256:e79743cd25d6edf3a8ce410e40882ebe660885a000b738e3c99c1da77de9e680

Observation ccfdff97-884e-4f24-94b3-bd5932737885 · outbound

This paper cites SmoothGrad: removing noise by adding noise.

Do we really have to filter out random noise in pre-training data for language models? SmoothGrad: removing noise by adding noise

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.702487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.702487Z digest=sha256:cea28f091f6a0b064e87060080fca2b26c0dfc55497b8d1dcf064342ccee8748

Observation ef34113c-8311-4996-963e-8a343280bf3d · outbound

This paper cites Pure noise to the rescue of insufficient data: Improving imbalanced classification by training on random noise images,.

Do we really have to filter out random noise in pre-training data for language models? Pure noise to the rescue of insufficient data: Improving imbalanced classification by training on random noise images,

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:04:30.776292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:04:29.706481Z digest=sha256:6bf936b5974352e94b5fe80db872e9a377a0ba4f8e4de10a8bc55cc4258f9fae

Observation 772f9bf0-235a-43d4-b1a4-3991426f7380 · outbound

This paper cites LLM-PCGC: Large Language Model-based Point Cloud Geometry Compression.

Do we really have to filter out random noise in pre-training data for language models? LLM-PCGC: Large Language Model-based Point Cloud Geometry Compression

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-08-08T15:04:30.053906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:04:29.710090Z digest=sha256:74c505b8d6a078bbd1c2b371d9e9b11f30d9d36e46969ea93fdb36a1014ea62d

Observation f06efcc4-fd35-448f-91a4-da2750cd68f1 · outbound

This paper cites Vision Transformer with Sparse Scan Prior.

Do we really have to filter out random noise in pre-training data for language models? Vision Transformer with Sparse Scan Prior

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.713823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.713823Z digest=sha256:38265507ace098fcf71dd3005bf3b63a4992955925473a801b5b5ef8bc8dd0ca

Observation 6be1d87c-bd45-45bb-82b3-b55d9898feed · outbound

This paper cites Semigmmpoint: Semi-supervised point cloud segmentation based on gaussian mixture models,.

Do we really have to filter out random noise in pre-training data for language models? Semigmmpoint: Semi-supervised point cloud segmentation based on gaussian mixture models,

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:04:30.765395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:04:29.717670Z digest=sha256:901fa09b89a815b9e3444f75ccf1b0c7cceccebc3a0d8bf0568ef86cf9ebb66f

Observation f86614ab-2c94-482c-bac0-a2e9655febc5 · outbound

This paper cites Composed fine-tuning: Freezing pre-trained denoising autoencoders for improved generalization,.

Do we really have to filter out random noise in pre-training data for language models? Composed fine-tuning: Freezing pre-trained denoising autoencoders for improved generalization,

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:04:30.755594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:04:29.721617Z digest=sha256:87deb69cab3408518d055aa5395a2758120c41efbe3cd8a66aeac6467926112c

Observation 9ecbcaf8-cfc0-46ba-bc3a-b6ec955b262c · outbound

This paper cites Game on tree: Visual hallucination mitigation via coarse-to-fine view tree and game theory,.

Do we really have to filter out random noise in pre-training data for language models? Game on tree: Visual hallucination mitigation via coarse-to-fine view tree and game theory,

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:04:30.745462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:04:29.725287Z digest=sha256:a609ec3a356012020f2f979ecb5ba417c1083309cd28c395c9f1602e88137f11

Observation b5edfe61-66e0-4e72-bde9-2e8fed3ccf39 · outbound

This paper cites Not all texts are the same: Dynamically querying texts for scene text detection,.

Do we really have to filter out random noise in pre-training data for language models? Not all texts are the same: Dynamically querying texts for scene text detection,

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:04:30.734208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:04:29.729029Z digest=sha256:4a63d093286ff03ee8e13bcdf4b1d6bb3251cdc9419ee4475b266b1c432b8325

Observation da431da8-f5cb-471d-8978-956b43446485 · outbound

This paper cites Towards multimodal-augmented pre-trained language models via self-balanced expectation-maximization iteration,.

Do we really have to filter out random noise in pre-training data for language models? Towards multimodal-augmented pre-trained language models via self-balanced expectation-maximization iteration,

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:04:30.722785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:04:29.732662Z digest=sha256:63651132f9c006aaf4476e5de13ff075f7ed2b430f72f2671d9149706a6490da

Observation 92a36130-2c45-45bc-ac40-cf254b3a3596 · outbound

This paper cites VASparse: Towards Efficient Visual Hallucination Mitigation via Visual-Aware Token Sparsification.

Do we really have to filter out random noise in pre-training data for language models? VASparse: Towards Efficient Visual Hallucination Mitigation via Visual-Aware Token Sparsification

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.736634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.736634Z digest=sha256:f1e97498c82c2e00a6fc73224af8fe37d6d8a92ba1fa664e13b4b0f3e221509c

Observation 615d0e26-bab0-47f2-b697-8d9426c1aadc · outbound

This paper cites Improving pretrained language model fine-tuning with noise stability regularization,.

Do we really have to filter out random noise in pre-training data for language models? Improving pretrained language model fine-tuning with noise stability regularization,

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:04:30.711292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:04:29.740630Z digest=sha256:432e7ec0e362de54d3ebe4ffc0b543d56abccefa413c33fd6bd98f27e9a865a6

Observation 2708e4a6-1e28-4ab6-8f69-aa3d9401753d · outbound

This paper cites SMART: Robust and efficient fine-tuning for pre-trained natural language models through principled regularized optimization,.

Do we really have to filter out random noise in pre-training data for language models? SMART: Robust and efficient fine-tuning for pre-trained natural language models through principled regularized optimization,

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:04:30.699686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:04:29.744313Z digest=sha256:279f5f9e33dd44efb13260ff0d8839a70a7cc33ce936fdb467d562a8c7e7db87

Observation 77a5dca0-0548-4d5b-a001-9143f37c2c51 · outbound

This paper cites Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling.

Do we really have to filter out random noise in pre-training data for language models? Tip-Adapter: Training-free CLIP-Adapter for Better Vision-Language Modeling

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.747280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.747280Z digest=sha256:e321744799ce7d68a15bbb66ace840d15be1bf60d6670f246f69aa281ff44a89

Observation 3c737a18-7d15-4104-b4cd-42b466f07a3a · outbound

This paper cites Learning to prompt for vision-language models,.

Do we really have to filter out random noise in pre-training data for language models? Learning to prompt for vision-language models,

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:04:30.688043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:04:29.750651Z digest=sha256:18000860c0847e099db9ba9526d7786e4fc3ccd005e5ea0906d61b800da5a8a7

Observation 9098e3a3-4c5a-4533-ac00-d9aa85b754c1 · outbound

This paper cites Conditional prompt learning for vision-language models,.

Do we really have to filter out random noise in pre-training data for language models? Conditional prompt learning for vision-language models,

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:04:30.676767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:04:29.753787Z digest=sha256:5820c1620998111ffee3dee049b80bd44c21b6a8da714023759eadef35cdcd9e

Observation c7f8ce62-0baf-414d-ad87-05dcaeffe454 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Do we really have to filter out random noise in pre-training data for language models? Learning transferable visual models from natural language supervision,

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-08T15:04:29.756870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:04:29.756870Z digest=sha256:32cce9306b7f4e5b517cd22c9872499e1ed7e9fe175249885bf9adc3cf7596ae

Observation 380e6b3e-f672-4e7d-b65c-cd6002930eed · outbound

This paper cites LoRA: Low-rank adaptation of large language models,.

Do we really have to filter out random noise in pre-training data for language models? LoRA: Low-rank adaptation of large language models,

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T15:04:30.657663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T15:04:29.760015Z digest=sha256:bb28392d131251431a464d77ef9c0982ac29880783cc416095ad1dd3ecd0c2b5

Pith citing papers

Observation 0386a878-954f-4cf9-93d5-6959bef397b7 · inbound

HAD: Hybrid Architecture Distillation Outperforms Teacher in Genomic Sequence Modeling cites this paper.

HAD: Hybrid Architecture Distillation Outperforms Teacher in Genomic Sequence Modeling Do we really have to filter out random noise in pre-training data for language models?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:41.436077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:41.436077Z digest=sha256:2eac64f22cd97012e9485384fffad06c07f1766c62c0940001ac4744c5c7f92f

Observation d129c012-6cc2-400f-b314-b3ba900413d5 · inbound

Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation cites this paper.

Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation Do we really have to filter out random noise in pre-training data for language models?

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:53:27.615962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T00:53:24.931453Z digest=sha256:8707ef1d370f8cb9e5bdd780be0f2ea2be3a19a538083ae6ee8b9e330f5c07bb