Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:46:42.375506Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 100 of 118 outbound references and 22 inbound Pith citation observations for arXiv:2506.20920.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:46:42.375506Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:16:45.263479Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
100 of 118 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 5ec22fa5-a012-41c0-82da-f189097500a9 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Yi: Open Foundation Models by 01.AI
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93645589-9dc1-4d2a-8128-c1a94669f570 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6894d93a-aa44-48df-8279-75e4b1f5b8f6 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Llama 3 model card
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b2b22c6-7c9f-43bf-b997-1ec8c9115267 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language A survey on data selection for language models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 767dab6a-c430-4bea-a8fa-31aa9e767aec · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Open llm turkish leaderboard v0.2
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1f05eab-e33f-40fb-977c-3aac94923944 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language A l G hafa evaluation benchmark for A rabic language models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c00582aa-05a4-47a7-96da-83d765789020 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language 101 billion arabic words dataset, 2024
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20cce32f-fab4-43bc-8d2e-76412bc7fc45 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language On the cross-lingual transferability of monolingual representations
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67eaca72-07bb-4902-b2db-fa169d3747a0 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language A call for more rigor in unsupervised cross-lingual learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fd5889bf-fa8d-4af2-aa2b-5ae7d4a83975 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Japanese massive multitask language understanding benchmark, 2023
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b0d82cf-0d66-40f5-8a0e-da42bf5d4814 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language The belebele benchmark: a parallel reading comprehension dataset in 122 language variants
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d29a952-39ae-463d-87d5-d28ef38c1316 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Building Machine Translation Systems for the Next Thousand Languages
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aee1a052-c635-4462-b70b-a26749d0af24 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3578c5c-d4c7-49a8-a653-2ca0dfa9d86e · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language On the resemblance and containment of documents
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ead3e28e-4569-40f6-878e-666442a5ef75 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language An open dataset and model for language identification
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f6c3cb5-c38d-4cce-b6f4-f6f4b634ed5a · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51c0b978-de55-4141-810a-32a5f05bf112 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language PTT5: Pretraining and validating the T5 model on Brazilian Portuguese data
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d97dbaa2-3c1b-4e30-bb53-124de1eeb969 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Language ID in the wild: Unexpected challenges on the path to a thousand-language web text corpus
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f48cd5d1-bcbe-49cb-8ec9-074466a7393e · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a73efc9f-d303-4a64-a261-dd8ea894f2dd · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Command r+
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 204dfc02-b923-4ba9-85cb-9b4525f83964 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Unsupervised cross-lingual representation learning at scale
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ebfebb77-afd4-4762-8ae6-1ec465b2ef68 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Neural learning for question answering in italian
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df62415c-318d-40c7-80cc-884a41aa911f · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Dataset for the First Evaluation on Chinese Machine Reading Comprehension
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 504fbcb6-49c0-4587-9919-3ba7ae8ddc77 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Daniels and William Bright (eds.)
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75a59b1b-7199-4d01-ba28-fac54f1e3df1 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language A New Massive Multilingual Dataset for High-Performance Language Technologies
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed081d84-e88a-4f48-92f3-42000a1ef7c8 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language BERTje: A Dutch BERT Model
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 488d3324-dc9d-45c4-a85d-3f57bcfdbcf5 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 579cc685-a676-4f08-aa87-8aeca05345ef · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language RobBERT: a Dutch RoBERTa-based Language Model
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48891c63-e204-485a-b9db-06320ce3391e · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language FQuAD: French Question Answering Dataset
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf5d4220-13a8-4211-ae7c-eab29ffc1d12 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Sailor2: Sailing in South-East Asia with Inclusive Multilingual LLMs
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96fee737-29a8-477b-b0ed-85d088c6a9fc · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Chinese Tiny LLM: Pretraining a Chinese-Centric Large Language Model
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3edfc52-178b-4997-99d4-9ae23425345f · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Eberhard, Gary F
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 086ea4a7-24ab-48bc-a1e0-83a70fe57691 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language SberQuAD – Russian Reading Comprehension Dataset: Description and Analysis, pp.\ 3–15
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f6ba979-b408-4848-b114-102f4590aae4 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Arabicweb24: Creating a high quality arabic web-only pre-training dataset, 2024
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44f6f449-cb49-424c-803c-bad6e90723f9 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Guerreiro, António Loison, Duarte M
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be88cc1e-59c7-4bed-bca5-0a003eddfb9e · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language MERA: A Comprehensive LLM Evaluation in Russian
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fd9627c-c59c-4bd2-9299-0f896ee29e4e · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Open llm leaderboard v2
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb661be8-d0a6-4bce-a8a6-859bc397a9bf · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Gemma: Open Models Based on Gemini Research and Technology
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 981672ec-3d2a-4080-855f-8cb4f19284c7 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language The Llama 3 Herd of Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94c3acd5-d761-4e6b-9598-f18da8e52e00 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Studying Large Language Model Generalization with Influence Functions
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7072c36e-25f9-4ad5-9295-6e4aa128b307 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language OLMES: A Standard for Language Model Evaluations
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a33ca50-0ec2-4830-a8ab-8355fa9c5c2c · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language EXAMS: A Multi-Subject High School Examinations Dataset for Cross-Lingual and Multilingual Question Answering
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1d7babeb-2c21-4c95-8506-9831e9a7cab5 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Measuring massive multitask language understanding
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6cfbcb3-9399-4019-bdd0-090d0eb25307 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Khmer natural language processing tookit
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59ff659f-ad21-44bc-867b-0135918901f7 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language spaCy: Industrial-strength Natural Language Processing in Python
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 561fdd31-a08e-42b2-bf8f-c9cb73492142 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac51f328-e38f-40e5-8db0-f5462da4da08 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Mistral 7B
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4ff6bd0-52bb-4dfb-9c64-fbd5f186fe61 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Mixtral of Experts
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bff3b4a-7fb8-4da2-8b86-0d413dfd1696 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language The state and fate of linguistic diversity and inclusion in the NLP world
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf778262-2509-4a1f-8333-9248fed2b284 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language FastText.zip: Compressing text classification models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d33314e-4afd-42de-954f-636bcf10589b · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language G lot LID : Language identification for low-resource languages
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f91fd40f-c527-4989-b2c1-955c2585825c · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Glot CC : An open broad-coverage commoncrawl corpus and pipeline for minority languages
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab275403-84c5-447e-86d0-96b6802e314c · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian Languages
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf5468c1-6200-469f-b7b6-e5dccd617073 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Large language models only pass primary school exams in I ndonesia: A comprehensive test on I ndo MMLU
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a77dfd9-3344-4fa0-ad00-4b36543393b8 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96434442-1e71-403d-96da-6cea0a979b97 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language The IndicNLP Library
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aecd3d84-6991-4604-9374-a45c68a30b04 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language JGLUE : J apanese general language understanding evaluation
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ee3b8ed-aa0d-41e0-91be-f4186e87631a · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedback
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaeafdc2-d5d4-495e-8f1b-8115b58b28c5 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language F lau BERT : Unsupervised language model pre-training for F rench
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 957c7581-a3b7-4084-a9cc-2950589ca17d · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Open-arabic-llm-leaderboard-v1
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98be8b18-486a-4cbf-ac8a-6ea4421f473f · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Deduplicating Training Data Makes Language Models Better
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6b25d1c-e7e3-4731-87d8-d9ad2d2f05ab · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Kiwipiepy: Kiwi package for python, 2024
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 072cbc6c-d94e-493c-9e08-bb7b96b8a997 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language MLQA: Evaluating Cross-lingual Extractive Question Answering
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82379ae4-9e75-4a1f-b422-34a70798f7d2 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language CMMLU: Measuring massive multitask language understanding in Chinese
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6e28839-9292-470b-9c28-2b6c7cf4ae35 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language DataComp-LM: In search of the next generation of training sets for language models
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4208a9a-d1bd-463b-bb92-2efd48be185b · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Common sense beyond E nglish: Evaluating and improving multilingual language models for commonsense reasoning
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a01ec591-bec5-493d-bd80-54c62c0db555 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Few-shot Learning with Multilingual Language Models
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adc5788a-159a-4393-8e78-f5702c2bec77 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language FinGPT: Large Generative Models for a Small Language
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffa8195f-6c42-430e-a18d-f5c38ae41b92 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Quantifying Variance in Evaluation Benchmarks
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4be25e36-faac-4254-9da6-e62f2ffad73a · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language C amem BERT : a tasty F rench language model
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30ff154b-ccfc-425f-815b-7b6da12380b0 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2027824-04d2-46e6-8767-f458b66c54f0 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Mnbvc: Massive never-ending bt vast chinese corpus
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 996645bd-7901-40f4-a1f0-8f1ca508ca5d · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Neural A rabic question answering
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8441a717-8366-4f55-8bd0-2c7e762cf8be · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Crosslingual generalization through multitask finetuning, 2022
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f82ef352-d1ca-4cac-a91d-b5d765dd15d9 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Rossi, and Thien Huu Nguyen
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4a84145-1fcf-45ba-930a-8d8230b3334f · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language No Language Left Behind: Scaling Human-Centered Machine Translation
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b02b7449-4b62-4a97-bf29-1576201f80f5 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Omnia russica
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 967d7d78-2977-47f4-8d9b-6bc072bb1a8d · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Botok: State-of-the-art tokenizers for tibetan language, 2025
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41d85cdc-552c-43ca-8bde-c85e358c797f · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Building pre-train llm dataset for the indic languages: A case study on hindi
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13ff38a0-299a-4cff-8429-f354bf0330b1 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Hellaswag-th, 2023
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3636b3d9-ad11-4cf7-aa9b-70f4aad44a20 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1421b3f6-cc9a-4d8e-9540-d4986d14c6f1 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language The fineweb datasets: Decanting the web for the finest text data at scale
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6c30583-bdce-4a57-86ea-01fa20c397b2 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Laonlp: Lao language natural language processing, July 2022
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d4910a25-953d-42a0-a5cc-a60dfbd80149 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language P y T hai NLP : T hai natural language processing in P ython, June 2024
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 792e8fff-e19d-49a1-8874-bc9c2e996040 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Typhoon: Thai Large Language Models
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00ab1a9a-a5b7-47ee-a3ee-117d04220b57 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Pllum: A family of polish large language models
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 800b7eb2-19b6-4ae1-bbfa-70d87904a68d · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Chinesesquad
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ae140c9-8575-42bd-a403-2f666001c830 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language XCOPA: A Multilingual Dataset for Causal Commonsense Reasoning
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 213d288c-62ad-464c-a708-530c35a4b8d7 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Stanza: A Python Natural Language Processing Toolkit for Many Human Languages
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16b60db9-7c02-4cb9-90ab-2fdbfc3e98e6 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Scaling Language Models: Methods, Analysis & Insights from Training Gopher
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b23d45d5-4bb6-4328-afa5-fa2979fd8043 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 078640bb-dca3-4481-b332-812e91aec898 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Impact of Pretraining Term Frequencies on Few-Shot Reasoning
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f195de96-efca-4b7d-8ce1-43649ab4f262 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language How Much Knowledge Can You Pack Into the Parameters of a Language Model?
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9647ceb-cb4a-488a-a35e-0db91dae8727 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b72f10a-e44f-434e-9e6c-32f17e657e5c · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Pyidaungsu: Python library for myanmar language, 2024
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 63725564-5a04-489c-a39b-9f21105c3bc5 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Compact language detector v3
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f76eddb-8eb6-4b46-a375-2dc8251a7cbd · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Mintaka: A Complex, Natural, and Multilingual Dataset for End-to-End Question Answering
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f7c332a-cbf0-48f2-8dd8-067bbe507e12 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language INDIC QA BENCHMARK: A Multilingual Benchmark to Evaluate Question Answering capability of LLMs for Indic Languages
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 982ab6e6-25d6-4030-8a74-e6e7278fcfaa · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c20a06fc-aaf4-4213-a34f-b98e46657df4 · outbound
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Thquad: Turkish historic question answering dataset for reading comprehension
Reference 102
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09e337de-4ea6-40da-aabe-c6305c9f5365 · inbound
From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c6bb196-3a1f-4b67-ae83-0f7f9f57eabe · inbound
Observation of momentum dependent charge density wave gap in EuTe4 FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67cc670d-1c49-4e0f-9050-c5fdcc28d126 · inbound
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7b73d5d6-df5b-429e-b414-072673213de2 · inbound
Entropy2Vec: Crosslingual Language Modeling Entropy as End-to-End Learnable Language Representations FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45b04929-9776-4355-bbde-389fd154ecf0 · inbound
The Effect of Scripts and Formats on LLM Numeracy FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e263f088-3f4e-4306-8aae-d48749d9a534 · inbound
A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Reference 167
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79a8722b-9b1e-435f-9f17-7bc8c9414d93 · inbound
OpenLID-v3: Improving the Precision of Closely Related Language Identification -- An Experience Report FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 763e30af-8044-4a58-8526-9f21ffcb16c1 · inbound
Scaling Laws for Mixture Pretraining Under Data Constraints FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3a3d4225-0be4-4675-bc7f-70b881a3c0a0 · inbound
Scaling Laws for Mixture Pretraining Under Data Constraints FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 10a7a53c-b215-4103-8017-0c99cb34c20e · inbound
Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 009b4acd-a370-4dba-a981-20466c2442fe · inbound
Granite Embedding Multilingual R2 Models FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9c65265e-e2be-427e-a18f-c02d71c8ba7e · inbound
A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAM$\Delta$ Integration into Upcycled MoE FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 55dab8fb-b24c-476f-a89e-a007b018028b · inbound
SomaliWeb v1: A Quality-Filtered Somali Web Corpus with a Matched Tokenizer and a Public Language-Identification Benchmark FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed65c45e-2831-430a-8d11-502aa052c297 · inbound
Multilingual Knowledge Transfer under Data Constraints via Lexical Interventions FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cf0129b7-a60b-40ca-9fb4-83adcd620143 · inbound
Mimir: Large-scale Multilingual Concept Modeling FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 57f65719-6f44-406a-861c-fcf0faa9aabd · inbound
Ishigaki-IDS: An Open-Weight Verifier-Aware Model for Information Delivery Specification Drafting in Building Information Modeling FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation df6a9db1-36aa-41d6-b236-56f1861d6a01 · inbound
Grouped Query Experts: Mixture-of-Experts on GQA Self-Attention FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d2bd763f-f994-4834-9947-29e54e4eb04e · inbound
moBERTo: A Modern Encoder for Portuguese via Continued Pretraining of ModernBERT FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1b5b1bdb-58fb-408b-9c29-2bf74907bb15 · inbound
LangMAP: A Language-Adaptive Approach to Tokenization FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3e45960-722e-4f56-bb55-437ffa4ab47a · inbound
SARA: Unlocking Multilingual Knowledge in Mixture-of-Experts via Semantically Anchored Routing Alignment FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 65447489-0aee-4602-976e-e5023ff38cb3 · inbound
MultiHashFormer: Hash-based Generative Language Models FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a12afaeb-34ce-473a-91f9-9ad1b79e3171 · inbound
In-Place Tokenizer Expansion for Pre-trained LLMs FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.