Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:09:26.124106Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 3 inbound Pith citation observations for arXiv:2505.02009.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T04:09:26.124106Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T23:13:05.224677Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T12:26:57.033442Z
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d4925574-3d7e-4987-9230-c4838ca40804 · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05346a7b-52ee-4d55-b419-aa249f7f858e · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b647bc3-f328-4af7-9a8e-9a8141c1cba4 · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Illicit darkweb clas- sification via natural-language processing: Classifying illicit content of webpages based on textual information
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f6130c86-cf74-411b-9289-054db11cb7ac · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs HateBERT: Retrain- ing BERT for abusive language detection in English
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 345d4b0e-9013-4a49-8489-fb6b75a02e73 · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b5088d15-c80a-4efe-bc58-362ccaa25ea7 · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs The Llama 3 Herd of Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 699ffc2a-35f1-4296-8b27-78c90485a0d9 · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a215cf9b-a5b3-4ed5-a531-9849e4c774f4 · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs [Gemma, 2024] Riviere et al
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation dc7efd2c-c8cd-438f-a6dc-3f25025189ba · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs A survey on automated fact-checking
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7742323f-cfe7-45af-b66b-25d55f925ff4 · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Llama guard: Llm-based input-output safeguard for human-ai conversations,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e413571b-7068-4038-84e7-d96ac011f2d8 · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Perplexed by quality: A perplexity-based method for adult and harmful content detection in multilingual heterogeneous web data,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d25c268e-e2b4-4d1a-8ebd-aa952b333597 · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs A pretrainer’s guide to training data: Measuring the effects of data age, domain coverage, quality, & toxicity
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5a8a4753-8141-4d5b-8b30-b5d5b6844b03 · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs [Loshchilov and Hutter, 2019] Ilya Loshchilov and Frank Hutter
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c2b61ccc-d146-422d-9398-fd8766926cba · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs [Markov et al., 2023] Todor Markov, Chong Zhang, Sand- hini Agarwal, Florentine Eloundou Nekoul, Theodore Lee, Steven Adler, Angela Jiang, and Lilian Weng
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 51e11f62-3248-474f-be49-7c79497feb97 · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Llama 3.2: Revolutionizing edge ai and vision with open, customizable models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a8065ae7-ba00-41c9-8ae8-fb8d563ae6cd · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Mistral-7b-v0.3
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 22d0562e-55fb-4037-a5d2-c1152b222359 · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Ollama - get up and running with large language models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1fc408a3-7f32-4ab1-8109-72b8d66f5fbc · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Data, data everywhere: A guide for pretraining dataset construction
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4635636d-4916-40c5-bc2a-c5e95dae7338 · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs [Penedo et al., 2024] Guilherme Penedo, Hynek Kydl ´ıˇcek, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro V on Werra, and Thomas Wolf
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d53124c4-172b-4a91-b5ec-749c2587d68c · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 45f3d535-32df-4deb-826b-cef15983f794 · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Common crawl – building an open web-scale crawl using hadoop,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5a2cb45f-01e1-43db-bdae-8e61c63d6046 · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs [Truic˘a and Apostol, 2023] Ciprian-Octavian Truic ˘a and Elena-Simona Apostol
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4e8a4e84-1bf9-4403-8acb-ee9130a06e32 · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Gomez, Łukasz Kaiser, and Illia Polosukhin
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0340ade0-f2c9-44f9-ab90-2b20cd62a7fc · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Problematic webpage identification: A trilogy of hatespeech, search engines and GPT
Reference 2010
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c2cdef70-35aa-41d9-8261-c059ce7e61ac · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs [Wei et al., 2022] Jason Wei, Xuezhi Wang, Dale Schuur- mans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H
Reference 2017
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6be51992-5f56-430c-b6b1-998cad748a1d · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs What’s in the box? an analysis of undesirable content in the Common Crawl corpus
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 59f5f5cf-870c-4cf2-858a-e0daee52ffc1 · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Language models are few-shot learn- ers
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da30a9a9-46b4-42e5-bc0f-8e9dc8f1b8af · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Unresolved cited work
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2686cbdf-9c57-41ab-8cde-05653776f2ab · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs GPT-4 Technical Report
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea172620-5d56-45af-95b4-9d8c8b7f8108 · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Peters, and Ar- man Cohan
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4faf27d3-69a6-4065-acf3-ca02546ee48e · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Suicidal ideation detection on social me- dia: A review of machine learning methods,
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7de96306-a4bf-4e67-bb0a-fed1115cc866 · outbound
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 803bed5d-fdd5-462f-86f8-adeff192f10d · inbound
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs
Reference 212
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb034246-6af4-4a59-b1e4-b960e64aae42 · inbound
Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea22ca27-f0da-4ec6-8833-486a85e51208 · inbound
Epistemic Injustice in Language Models: An Audit of Pretraining Filters and Guardrails Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.