Pith. sign in

Paper Citation Record · LEDGER

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs

As of 21 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 3 inbound Pith citation observations for arXiv:2505.02009.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.02009 v3

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:09:26.124106Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:13:05.224677Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:26:57.033442Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d4925574-3d7e-4987-9230-c4838ca40804 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:09:25.966440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:09:25.966440Z digest=sha256:c15d4c8716d71a5abcfc94cc321a6273be5bdd48c1b5737179ca1451d62a892a

Observation 05346a7b-52ee-4d55-b419-aa249f7f858e · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T04:09:25.994378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:09:25.994378Z digest=sha256:e6f302143f71b9fea578cc122c3acb432d77f153f358b9ceb4df49f4f379f2fd

Observation 3b647bc3-f328-4af7-9a8e-9a8141c1cba4 · outbound

This paper cites Illicit darkweb clas- sification via natural-language processing: Classifying illicit content of webpages based on textual information.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Illicit darkweb clas- sification via natural-language processing: Classifying illicit content of webpages based on textual information

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:09:26.660624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:09:26.001429Z digest=sha256:ae616397e0f7f4c452e91a1411331e14bcb657d59930088c6c90daf93bbc6fca

Observation f6130c86-cf74-411b-9289-054db11cb7ac · outbound

This paper cites HateBERT: Retrain- ing BERT for abusive language detection in English.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs HateBERT: Retrain- ing BERT for abusive language detection in English

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:09:26.641747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:09:26.006613Z digest=sha256:dbabcac2072c023714bf0952a23095705e9818c1a651873247cd514f6f4c0b34

Observation 345d4b0e-9013-4a49-8489-fb6b75a02e73 · outbound

This paper cites an unresolved cited work.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:09:26.602751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:09:26.015939Z digest=sha256:a9bc06777352c0f870b4fb361ad2a6506ba300ad9ce6a6c19a66bd2b728a2d96

Observation b5088d15-c80a-4efe-bc58-362ccaa25ea7 · outbound

This paper cites The Llama 3 Herd of Models.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T04:09:26.024807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:09:26.024807Z digest=sha256:8402e0c9d764194d4b38e5bf503f9c8be2574962efa134eb33480a466885308d

Observation 699ffc2a-35f1-4296-8b27-78c90485a0d9 · outbound

This paper cites an unresolved cited work.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:09:26.568895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:09:26.029462Z digest=sha256:684f226dbbc1d61f853c6c840bf7f8ca6c36973ca9ad2ac823967dc095237ae4

Observation a215cf9b-a5b3-4ed5-a531-9849e4c774f4 · outbound

This paper cites [Gemma, 2024] Riviere et al.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs [Gemma, 2024] Riviere et al

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:09:26.552504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:09:26.033791Z digest=sha256:f2b502a0337a0f9feb9baa28d5b085193a98d0819c199b681ff5841483ec49ac

Observation dc7efd2c-c8cd-438f-a6dc-3f25025189ba · outbound

This paper cites A survey on automated fact-checking.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs A survey on automated fact-checking

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T04:09:26.038331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:09:26.038331Z digest=sha256:3353b5e0ddf1a691eff1921910b0981657c08997b7d087ad14537902776d8462

Observation 7742323f-cfe7-45af-b66b-25d55f925ff4 · outbound

This paper cites Llama guard: Llm-based input-output safeguard for human-ai conversations,.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Llama guard: Llm-based input-output safeguard for human-ai conversations,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:09:26.524585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:09:26.043728Z digest=sha256:7b6f336b13fbd1113f539d78ede604676be97abe4045730547612ceb95f475b9

Observation e413571b-7068-4038-84e7-d96ac011f2d8 · outbound

This paper cites Perplexed by quality: A perplexity-based method for adult and harmful content detection in multilingual heterogeneous web data,.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Perplexed by quality: A perplexity-based method for adult and harmful content detection in multilingual heterogeneous web data,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:09:26.506568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:09:26.048719Z digest=sha256:56c1f444bbfe1ef80b5891aafa4d5bac47e8197d44bc7f1eecffa6525e827dbe

Observation d25c268e-e2b4-4d1a-8ebd-aa952b333597 · outbound

This paper cites A pretrainer’s guide to training data: Measuring the effects of data age, domain coverage, quality, & toxicity.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs A pretrainer’s guide to training data: Measuring the effects of data age, domain coverage, quality, & toxicity

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:09:26.486439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:09:26.054262Z digest=sha256:d190920784b7fed80df3579e006698e69374ba601e8fa20d05a98642b8759293

Observation 5a8a4753-8141-4d5b-8b30-b5d5b6844b03 · outbound

This paper cites [Loshchilov and Hutter, 2019] Ilya Loshchilov and Frank Hutter.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs [Loshchilov and Hutter, 2019] Ilya Loshchilov and Frank Hutter

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:09:26.468220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:09:26.059337Z digest=sha256:622a8a385dc8a5d8631fc77436588f68e727671162b3198b1c2ce12b90ecc828

Observation c2b61ccc-d146-422d-9398-fd8766926cba · outbound

This paper cites [Markov et al., 2023] Todor Markov, Chong Zhang, Sand- hini Agarwal, Florentine Eloundou Nekoul, Theodore Lee, Steven Adler, Angela Jiang, and Lilian Weng.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs [Markov et al., 2023] Todor Markov, Chong Zhang, Sand- hini Agarwal, Florentine Eloundou Nekoul, Theodore Lee, Steven Adler, Angela Jiang, and Lilian Weng

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:09:26.432060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:09:26.069506Z digest=sha256:2e30023df56315d9145e55f2c89661f482aa37217242b8894a2aa9d161d11821

Observation 51e11f62-3248-474f-be49-7c79497feb97 · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vision with open, customizable models.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Llama 3.2: Revolutionizing edge ai and vision with open, customizable models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:09:26.412569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:09:26.074840Z digest=sha256:af4158d8166f96bcf87633d3c6bca55167e66a8a3dc3c52676e7fd07a9116187

Observation a8065ae7-ba00-41c9-8ae8-fb8d563ae6cd · outbound

This paper cites Mistral-7b-v0.3.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Mistral-7b-v0.3

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:09:26.396544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:09:26.079677Z digest=sha256:b8220944a0156be019a35a7e39df0296b8032b6468488b6c20b362535608445c

Observation 22d0562e-55fb-4037-a5d2-c1152b222359 · outbound

This paper cites Ollama - get up and running with large language models.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Ollama - get up and running with large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:09:26.379148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:09:26.084577Z digest=sha256:cf6165ed1218efe4c8d540514f730758e15216ebd999d7c1739c01b8ba915e76

Observation 1fc408a3-7f32-4ab1-8109-72b8d66f5fbc · outbound

This paper cites Data, data everywhere: A guide for pretraining dataset construction.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Data, data everywhere: A guide for pretraining dataset construction

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:09:26.361071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:09:26.089270Z digest=sha256:022902a23feb3aa0c27bb0ce4f6f6399035e229742d6ad0ff6dd70fbcc9524b7

Observation 4635636d-4916-40c5-bc2a-c5e95dae7338 · outbound

This paper cites [Penedo et al., 2024] Guilherme Penedo, Hynek Kydl ´ıˇcek, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro V on Werra, and Thomas Wolf.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs [Penedo et al., 2024] Guilherme Penedo, Hynek Kydl ´ıˇcek, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro V on Werra, and Thomas Wolf

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:09:26.343610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:09:26.094105Z digest=sha256:c7a986465056bbdce02deb91c56d479154c35211f7d99a906ec9f9740256ab9d

Observation d53124c4-172b-4a91-b5ec-749c2587d68c · outbound

This paper cites an unresolved cited work.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:09:26.324416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:09:26.098958Z digest=sha256:aeadffb9ac8677d0dd52e236bf8ca2e2f3863ff043a5f46aef42684205d88fec

Observation 45f3d535-32df-4deb-826b-cef15983f794 · outbound

This paper cites Common crawl – building an open web-scale crawl using hadoop,.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Common crawl – building an open web-scale crawl using hadoop,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:09:26.305902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:09:26.103868Z digest=sha256:7c92b2d12de1c96ea7bcd16bb57ad15fa6e497445f78aaafd74c9ed5b7c9f1e6

Observation 5a2cb45f-01e1-43db-bdae-8e61c63d6046 · outbound

This paper cites [Truic˘a and Apostol, 2023] Ciprian-Octavian Truic ˘a and Elena-Simona Apostol.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs [Truic˘a and Apostol, 2023] Ciprian-Octavian Truic ˘a and Elena-Simona Apostol

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:09:26.269867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:09:26.113990Z digest=sha256:4a1ea19d2f22a3e93432ac1d10aa2c1d40011ade806104dfe5e56b1744ea095b

Observation 4e8a4e84-1bf9-4403-8acb-ee9130a06e32 · outbound

This paper cites Gomez, Łukasz Kaiser, and Illia Polosukhin.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Gomez, Łukasz Kaiser, and Illia Polosukhin

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:09:26.250492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:09:26.119077Z digest=sha256:6a7e22bb743b561eaf94b97f37756185c2d29b813446f55fcc88cee821711848

Observation 0340ade0-f2c9-44f9-ab90-2b20cd62a7fc · outbound

This paper cites Problematic webpage identification: A trilogy of hatespeech, search engines and GPT.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Problematic webpage identification: A trilogy of hatespeech, search engines and GPT

Reference 2010

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:09:26.288078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:09:26.108871Z digest=sha256:efb16c1dfb66f87d70e24f04a42cf4cb324b7e3eb03dfb32eb63971752dd2524

Observation c2cdef70-35aa-41d9-8261-c059ce7e61ac · outbound

This paper cites [Wei et al., 2022] Jason Wei, Xuezhi Wang, Dale Schuur- mans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs [Wei et al., 2022] Jason Wei, Xuezhi Wang, Dale Schuur- mans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:09:26.233298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:09:26.124106Z digest=sha256:81d78ca012e252bca13c5567966d64affcfd84a0f76d0fd9c2b77190ce39d011

Observation 6be51992-5f56-430c-b6b1-998cad748a1d · outbound

This paper cites What’s in the box? an analysis of undesirable content in the Common Crawl corpus.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs What’s in the box? an analysis of undesirable content in the Common Crawl corpus

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:09:26.449968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:09:26.064197Z digest=sha256:ea80d4d65a5a531dc9e3aafb19f93a284db3e742d7b104530ccfa267fa7bde7d

Observation 59f5f5cf-870c-4cf2-858a-e0daee52ffc1 · outbound

This paper cites Language models are few-shot learn- ers.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Language models are few-shot learn- ers

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-16T04:09:25.988757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:09:25.988757Z digest=sha256:af1dbb3517d2efc7446b8468d4922307ea1e79bd724939b2da67d373aa47d6b8

Observation da30a9a9-46b4-42e5-bc0f-8e9dc8f1b8af · outbound

This paper cites an unresolved cited work.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Unresolved cited work

Reference 2021

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:09:26.621617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:09:26.011611Z digest=sha256:3141afc3017f3e77572045d73c976b18616951f3d59d538150125cf0a1a3f59b

Observation 2686cbdf-9c57-41ab-8cde-05653776f2ab · outbound

This paper cites GPT-4 Technical Report.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs GPT-4 Technical Report

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-16T04:09:25.977828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:09:25.977828Z digest=sha256:7b2807761d59eb7a152cbb82581dbd008bb5d1e41f7077adfc3c2fff4a27190e

Observation ea172620-5d56-45af-95b4-9d8c8b7f8108 · outbound

This paper cites Peters, and Ar- man Cohan.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Peters, and Ar- man Cohan

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:09:26.690671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:09:25.983406Z digest=sha256:6028c2710536bd33257a329fb0a90171fc78578913afaecccdca9686902c9c06

Observation 4faf27d3-69a6-4065-acf3-ca02546ee48e · outbound

This paper cites Suicidal ideation detection on social me- dia: A review of machine learning methods,.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Suicidal ideation detection on social me- dia: A review of machine learning methods,

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:09:26.708853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:09:25.972649Z digest=sha256:99dc8e99888d981580d3e04238eb8335ff1a06310259cd0846182a4a1d03d262

Observation 7de96306-a4bf-4e67-bb0a-fed1115cc866 · outbound

This paper cites Documenting large webtext corpora: A case study on the colossal clean crawled corpus.

Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Documenting large webtext corpora: A case study on the colossal clean crawled corpus

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:09:26.585711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T04:09:26.020508Z digest=sha256:912d65617342a60fac53c5634b8863889b2a24687b5810d9a4a6502be0d8be41

Pith citing papers

Observation 803bed5d-fdd5-462f-86f8-adeff192f10d · inbound

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM cites this paper.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs

Reference 212

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:05.224677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:05.224677Z digest=sha256:f54ae04dd3d404b312f3ad9c964662bd1e21ce8b43b4baaa37550c12501d2f93

Observation bb034246-6af4-4a59-b1e4-b960e64aae42 · inbound

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection cites this paper.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.878381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.878381Z digest=sha256:6b0ebe761fe4fc05d57ac04d954ab753bd45e4b5acde4b46d944a4b376899f48

Observation ea22ca27-f0da-4ec6-8833-486a85e51208 · inbound

Epistemic Injustice in Language Models: An Audit of Pretraining Filters and Guardrails cites this paper.

Epistemic Injustice in Language Models: An Audit of Pretraining Filters and Guardrails Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:26:57.034737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T02:06:31.341519Z digest=sha256:79789d98d826974891493e0860d893dd9ce1456761287721fe49b88336d5e11f