Pith. sign in

Paper Citation Record · LEDGER

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

As of 7 August 2026, this Paper Citation Record lists 100 of 118 outbound references and 22 inbound Pith citation observations for arXiv:2506.20920.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20920 v1

Coverage vector

measured 100 of 118 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:46:42.375506Z

measured 122 of 122 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:16:45.263479Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 118 outbound references displayed

  • verified exact8
  • verified fuzzy2
  • unresolved89
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 5ec22fa5-a012-41c0-82da-f189097500a9 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Yi: Open Foundation Models by 01.AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.435145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.435145Z digest=sha256:9d2c474bc428530e59c96856c0b25555201dc413eb275186cd943dec55f1d055

Observation 93645589-9dc1-4d2a-8128-c1a94669f570 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.489072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.489072Z digest=sha256:22a42b1ec7bf5dfc8bfe9618b5ad488eda362ac46efc58588decbd8b4f4b46ba

Observation 6894d93a-aa44-48df-8279-75e4b1f5b8f6 · outbound

This paper cites Llama 3 model card.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Llama 3 model card

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.557062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.557062Z digest=sha256:dc2a463a14686899aab282e881a3353f0ca54d926b32c31fb18e973826d19dc6

Observation 9b2b22c6-7c9f-43bf-b997-1ec8c9115267 · outbound

This paper cites A survey on data selection for language models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language A survey on data selection for language models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.625972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.625972Z digest=sha256:b075b540226ff59b7aa41077ea08ef84f839c468c243b70fc68ea1060e83e88e

Observation 767dab6a-c430-4bea-a8fa-31aa9e767aec · outbound

This paper cites Open llm turkish leaderboard v0.2.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Open llm turkish leaderboard v0.2

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.702254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.702254Z digest=sha256:5ed7f72c8309177e763b86be94dac677f7f49a3c0747850d5d6082c225986fac

Observation a1f05eab-e33f-40fb-977c-3aac94923944 · outbound

This paper cites A l G hafa evaluation benchmark for A rabic language models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language A l G hafa evaluation benchmark for A rabic language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.797026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.797026Z digest=sha256:7ed0579fbc5b78099b9c514be61cd6a3f6e51e492daa3aca82a055a0a32195dd

Observation c00582aa-05a4-47a7-96da-83d765789020 · outbound

This paper cites 101 billion arabic words dataset, 2024.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language 101 billion arabic words dataset, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.843819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.843819Z digest=sha256:3363ca043dfcb9d0f08d33b242abe8657e76505287a9514486665abdd572351f

Observation 20cce32f-fab4-43bc-8d2e-76412bc7fc45 · outbound

This paper cites On the cross-lingual transferability of monolingual representations.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language On the cross-lingual transferability of monolingual representations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.881227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.881227Z digest=sha256:253447ff3f75fa6a3a1188ddbaf689e184f9540878400c9f88d39de5d94cdd4d

Observation 67eaca72-07bb-4902-b2db-fa169d3747a0 · outbound

This paper cites A call for more rigor in unsupervised cross-lingual learning.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language A call for more rigor in unsupervised cross-lingual learning

Reference 9

Resolution
verified exact
doi, observed 2026-08-06T22:46:43.869917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:46:33.948108Z digest=sha256:59b2bdf2caefa2620adad3e135ade26916051ad90d86ff4c7ad4d8fc41543036

Observation fd5889bf-fa8d-4af2-aa2b-5ae7d4a83975 · outbound

This paper cites Japanese massive multitask language understanding benchmark, 2023.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Japanese massive multitask language understanding benchmark, 2023

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.995458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.995458Z digest=sha256:b4393928aac48abea6c3b82685d101520425c991a7bc8f4ce61e910c6ac172b0

Observation 9b0d82cf-0d66-40f5-8a0e-da42bf5d4814 · outbound

This paper cites The belebele benchmark: a parallel reading comprehension dataset in 122 language variants.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language The belebele benchmark: a parallel reading comprehension dataset in 122 language variants

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:34.075955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.075955Z digest=sha256:2b79ff3bc0c01f0094c7578415093b90554e58e0ec4215797f6bb8682355ae6f

Observation 2d29a952-39ae-463d-87d5-d28ef38c1316 · outbound

This paper cites Building Machine Translation Systems for the Next Thousand Languages.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Building Machine Translation Systems for the Next Thousand Languages

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:34.150087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.150087Z digest=sha256:ad4647bcdb0fd97f247a96f24aac9d2f4fb66bed3c31cef388cff4a32caa215c

Observation aee1a052-c635-4462-b70b-a26749d0af24 · outbound

This paper cites Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:34.210423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.210423Z digest=sha256:540c66d180a3803e6db79cd42810c4925ace2fced9ed53e0dadc62c44e606bc1

Observation c3578c5c-d4c7-49a8-a653-2ca0dfa9d86e · outbound

This paper cites On the resemblance and containment of documents.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language On the resemblance and containment of documents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:34.319292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.319292Z digest=sha256:345a3ed1743fc3e22b6d87c6e494a8615ae6476c2120a4e70ce2f57804356ea2

Observation ead3e28e-4569-40f6-878e-666442a5ef75 · outbound

This paper cites An open dataset and model for language identification.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language An open dataset and model for language identification

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:34.474853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.474853Z digest=sha256:d587208e9f252099aab037e65cb45cbb3264f112ad4c791a0bcb21d487390c96

Observation 6f6c3cb5-c38d-4cce-b6f4-f6f4b634ed5a · outbound

This paper cites An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT).

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:34.583082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.583082Z digest=sha256:61598311fd56b6d0a987c8f7ea9a1847e32df03427bae3d66ed5c5ce51bb9900

Observation 51c0b978-de55-4141-810a-32a5f05bf112 · outbound

This paper cites PTT5: Pretraining and validating the T5 model on Brazilian Portuguese data.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language PTT5: Pretraining and validating the T5 model on Brazilian Portuguese data

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:34.736671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.736671Z digest=sha256:63035089be8212c17863bfec8b4b87dd7685b53225ed277c6a3f3feaad2d120a

Observation d97dbaa2-3c1b-4e30-bb53-124de1eeb969 · outbound

This paper cites Language ID in the wild: Unexpected challenges on the path to a thousand-language web text corpus.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Language ID in the wild: Unexpected challenges on the path to a thousand-language web text corpus

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:34.831748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.831748Z digest=sha256:f94399fec150cab3e709f1190e1e487a5deb4747f2bde74386b44b27a1f9261e

Observation f48cd5d1-bcbe-49cb-8ec9-074466a7393e · outbound

This paper cites Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki

Reference 19

Resolution
malformed identifier
no resolver link, observed 2026-08-06T22:46:34.988316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.988316Z digest=sha256:e5f78e5779d738ee5a3fcca580cc3f26e6fe3397e39e675853b27e8af3b81f05

Observation a73efc9f-d303-4a64-a261-dd8ea894f2dd · outbound

This paper cites Command r+.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Command r+

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:35.082235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:35.082235Z digest=sha256:385902bf1d89f7296723f4bdec8d01901b49db0b5a8251081aca5908825c0777

Observation 204dfc02-b923-4ba9-85cb-9b4525f83964 · outbound

This paper cites Unsupervised cross-lingual representation learning at scale.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Unsupervised cross-lingual representation learning at scale

Reference 21

Resolution
verified exact
doi, observed 2026-08-06T22:46:43.742969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:46:35.190771Z digest=sha256:e22aa1c0d549fd35758c269e9a37a0631b2af7242c517c3a50e3e02fe56f2ba1

Observation ebfebb77-afd4-4762-8ae6-1ec465b2ef68 · outbound

This paper cites Neural learning for question answering in italian.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Neural learning for question answering in italian

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:35.333775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:35.333775Z digest=sha256:abb2b70b35121676041bec9900d4701a49fa56cf3c2ab182f9fce7ce2fedca64

Observation df62415c-318d-40c7-80cc-884a41aa911f · outbound

This paper cites Dataset for the First Evaluation on Chinese Machine Reading Comprehension.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Dataset for the First Evaluation on Chinese Machine Reading Comprehension

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:46:45.068840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:46:35.445116Z digest=sha256:4e4b694df24e6a7a56449b91beeb4fc0c8c36c942a2dd4de78d79b38b63a2887

Observation 504fbcb6-49c0-4587-9919-3ba7ae8ddc77 · outbound

This paper cites Daniels and William Bright (eds.).

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Daniels and William Bright (eds.)

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:35.543260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:35.543260Z digest=sha256:7972bdd18f75d4045761096b1de47c5e8c0c31b0876ccb195bd20886367f38d9

Observation 75a59b1b-7199-4d01-ba28-fac54f1e3df1 · outbound

This paper cites A New Massive Multilingual Dataset for High-Performance Language Technologies.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language A New Massive Multilingual Dataset for High-Performance Language Technologies

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:46:44.888181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:46:35.727398Z digest=sha256:30f739d9653f01b6f978eaa726d1ad999e07590444c029767e18e09f514bffcd

Observation ed081d84-e88a-4f48-92f3-42000a1ef7c8 · outbound

This paper cites BERTje: A Dutch BERT Model.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language BERTje: A Dutch BERT Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:35.817846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:35.817846Z digest=sha256:6a845069b437f9da7754c36e4c456a35e7a5322ff5bd49640a3a9be62093c4df

Observation 488d3324-dc9d-45c4-a85d-3f57bcfdbcf5 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:35.894130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:35.894130Z digest=sha256:0335b5c9f0311553fcbf59d8b720f3d6542fb7036b369ff6c860d97ac2ea2123

Observation 579cc685-a676-4f08-aa87-8aeca05345ef · outbound

This paper cites RobBERT: a Dutch RoBERTa-based Language Model.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language RobBERT: a Dutch RoBERTa-based Language Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:35.985633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:35.985633Z digest=sha256:e0a335264e88d4d57c5370dd90e16baae7743987db780431ebdb53ee7960b284

Observation 48891c63-e204-485a-b9db-06320ce3391e · outbound

This paper cites FQuAD: French Question Answering Dataset.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language FQuAD: French Question Answering Dataset

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:46:44.741197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:46:36.078826Z digest=sha256:b0870f1bab88c2309632374cbe43f6d69207f5b779d688d1fdbc4b7d691df69f

Observation bf5d4220-13a8-4211-ae7c-eab29ffc1d12 · outbound

This paper cites Sailor2: Sailing in South-East Asia with Inclusive Multilingual LLMs.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Sailor2: Sailing in South-East Asia with Inclusive Multilingual LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.172908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.172908Z digest=sha256:47479dc6d4d1fc7fc92427c1f7290e46ff3dea3bff2abeeb5e6df5388f7f45cd

Observation 96fee737-29a8-477b-b0ed-85d088c6a9fc · outbound

This paper cites Chinese Tiny LLM: Pretraining a Chinese-Centric Large Language Model.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Chinese Tiny LLM: Pretraining a Chinese-Centric Large Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.260283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.260283Z digest=sha256:0ddce6d185487a401e7941d43a17f50597b6ba9b50809b5e47bcf3b7a3a80e7b

Observation c3edfc52-178b-4997-99d4-9ae23425345f · outbound

This paper cites Eberhard, Gary F.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Eberhard, Gary F

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.357550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.357550Z digest=sha256:802172d0e75c8eda44f54aced8c378a872116e556ba4e26f48e7d5c97718b34d

Observation 086ea4a7-24ab-48bc-a1e0-83a70fe57691 · outbound

This paper cites SberQuAD – Russian Reading Comprehension Dataset: Description and Analysis, pp.\ 3–15.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language SberQuAD – Russian Reading Comprehension Dataset: Description and Analysis, pp.\ 3–15

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.484014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.484014Z digest=sha256:692b7f4e7a4e958ddacc5e55b11369d1f895574c65a70d1c7d577828127a6811

Observation 4f6ba979-b408-4848-b114-102f4590aae4 · outbound

This paper cites Arabicweb24: Creating a high quality arabic web-only pre-training dataset, 2024.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Arabicweb24: Creating a high quality arabic web-only pre-training dataset, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.563221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.563221Z digest=sha256:d18776682ba543d7f95a3203dcad24ee2c2df904163af9090bcd93e03a2f65a8

Observation 44f6f449-cb49-424c-803c-bad6e90723f9 · outbound

This paper cites Guerreiro, António Loison, Duarte M.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Guerreiro, António Loison, Duarte M

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.638830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.638830Z digest=sha256:deec4b0fd11bd8e23a79908773041b11045cfa6900e5164b2647d768d1fe93cd

Observation be88cc1e-59c7-4bed-bca5-0a003eddfb9e · outbound

This paper cites MERA: A Comprehensive LLM Evaluation in Russian.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language MERA: A Comprehensive LLM Evaluation in Russian

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.710892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.710892Z digest=sha256:be526f8c3715fdcce145bcdcff0f3b691fd282fc29cc0669a886f8835a030716

Observation 4fd9627c-c59c-4bd2-9299-0f896ee29e4e · outbound

This paper cites Open llm leaderboard v2.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Open llm leaderboard v2

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.843320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.843320Z digest=sha256:670ec616699aaf778184194f95cdd4b7c5142826ddcaa269b0949fd32dff7e80

Observation fb661be8-d0a6-4bce-a8a6-859bc397a9bf · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Gemma: Open Models Based on Gemini Research and Technology

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.956706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.956706Z digest=sha256:3f8438b61eb0a7d61269c42bd298600265e8c5a0ed194eb166e2d53e530e6b15

Observation 981672ec-3d2a-4080-855f-8cb4f19284c7 · outbound

This paper cites The Llama 3 Herd of Models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language The Llama 3 Herd of Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:37.083695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:37.083695Z digest=sha256:53dc3e01bc8f979184e92eaade22c506911657b354234dbca5df0d49100138ae

Observation 94c3acd5-d761-4e6b-9598-f18da8e52e00 · outbound

This paper cites Studying Large Language Model Generalization with Influence Functions.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Studying Large Language Model Generalization with Influence Functions

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:37.222530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:37.222530Z digest=sha256:344b7ad028b17f0d6330bf24622186eccaaf7eb052885cac62783ba95cc11542

Observation 7072c36e-25f9-4ad5-9295-6e4aa128b307 · outbound

This paper cites OLMES: A Standard for Language Model Evaluations.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language OLMES: A Standard for Language Model Evaluations

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:37.365791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:37.365791Z digest=sha256:be23da788293e9ee853f1dd6702b4ebe021248508fe9287d18b9492921e16636

Observation 9a33ca50-0ec2-4830-a8ab-8355fa9c5c2c · outbound

This paper cites EXAMS: A Multi-Subject High School Examinations Dataset for Cross-Lingual and Multilingual Question Answering.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language EXAMS: A Multi-Subject High School Examinations Dataset for Cross-Lingual and Multilingual Question Answering

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:46:44.529872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:46:37.521847Z digest=sha256:034206e80855ef12add8ea5d3ba0b82e92c6a8e7f2b47b164293c9aabad1b7ad

Observation 1d7babeb-2c21-4c95-8506-9831e9a7cab5 · outbound

This paper cites Measuring massive multitask language understanding.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Measuring massive multitask language understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:37.613762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:37.613762Z digest=sha256:d2861daaf1f82683dfc471bf262ed102ca67c311355fb102c81c5b5c2fb064f2

Observation e6cfbcb3-9399-4019-bdd0-090d0eb25307 · outbound

This paper cites Khmer natural language processing tookit.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Khmer natural language processing tookit

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:37.677450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:37.677450Z digest=sha256:26c09f46621a5f878a564b09bdb84fcbf383d549b140b9b9748038e5013b02b3

Observation 59ff659f-ad21-44bc-867b-0135918901f7 · outbound

This paper cites spaCy: Industrial-strength Natural Language Processing in Python.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language spaCy: Industrial-strength Natural Language Processing in Python

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:37.821687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:37.821687Z digest=sha256:e1de226b431990147d8b7f5834eb0fe402359cc107f4b4e2534e2fc903c7bd8c

Observation 561fdd31-a08e-42b2-bf8f-c9cb73492142 · outbound

This paper cites C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:37.938316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:37.938316Z digest=sha256:5589602bdd7ae119e35a1260704b30b99dbc32bb81c7f9b054b318c9d7c8f3fd

Observation ac51f328-e38f-40e5-8db0-f5462da4da08 · outbound

This paper cites Mistral 7B.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Mistral 7B

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:38.104108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:38.104108Z digest=sha256:9dbde09cc1445a1ab62558b26081d7b273f634c1bab600c0e6f964ab51df4cfc

Observation f4ff6bd0-52bb-4dfb-9c64-fbd5f186fe61 · outbound

This paper cites Mixtral of Experts.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Mixtral of Experts

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:38.254676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:38.254676Z digest=sha256:c0b2a9224a1229c68cc7dc3eaa2bf07579112744ca9754a1a5b7d2466daeb059

Observation 1bff3b4a-7fb8-4da2-8b86-0d413dfd1696 · outbound

This paper cites The state and fate of linguistic diversity and inclusion in the NLP world.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language The state and fate of linguistic diversity and inclusion in the NLP world

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:38.411289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:38.411289Z digest=sha256:06ddae62a89f52634028cea82fe31a1a4cc594256a2d6788137bafd6a8e47d8b

Observation bf778262-2509-4a1f-8333-9248fed2b284 · outbound

This paper cites FastText.zip: Compressing text classification models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language FastText.zip: Compressing text classification models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:38.542891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:38.542891Z digest=sha256:a51386263b6c536de6aba0f2365cdb97ff71ee78fc41466aad9aba37629f40df

Observation 0d33314e-4afd-42de-954f-636bcf10589b · outbound

This paper cites G lot LID : Language identification for low-resource languages.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language G lot LID : Language identification for low-resource languages

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:38.664452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:38.664452Z digest=sha256:6ae553a216912f2b273c2b274c28d013ce550fb1f1f7e31d93a90a3236076591

Observation f91fd40f-c527-4989-b2c1-955c2585825c · outbound

This paper cites Glot CC : An open broad-coverage commoncrawl corpus and pipeline for minority languages.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Glot CC : An open broad-coverage commoncrawl corpus and pipeline for minority languages

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:38.729060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:38.729060Z digest=sha256:bee7a55daec7d63009ec26b6f531fa6c21c9e1b052d0f98b7fd3262aee66ca9d

Observation ab275403-84c5-447e-86d0-96b6802e314c · outbound

This paper cites IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian Languages.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian Languages

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:38.868060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:38.868060Z digest=sha256:4968e9b5f74516d076667f3e0740247ea979dedbdda4b1b3d381dac895046d30

Observation bf5468c1-6200-469f-b7b6-e5dccd617073 · outbound

This paper cites Large language models only pass primary school exams in I ndonesia: A comprehensive test on I ndo MMLU.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Large language models only pass primary school exams in I ndonesia: A comprehensive test on I ndo MMLU

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:38.953519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:38.953519Z digest=sha256:2d9314d5693a29b46029e80ec7964939c0a5aa37e6bb0c2f9fafd6bd557de671

Observation 7a77dfd9-3344-4fa0-ad00-4b36543393b8 · outbound

This paper cites ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.027217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.027217Z digest=sha256:fd3893b62b360797afb9144240b5bb24d2f61ac2e04e620ac4de834c96db6622

Observation 96434442-1e71-403d-96da-6cea0a979b97 · outbound

This paper cites The IndicNLP Library.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language The IndicNLP Library

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.099607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.099607Z digest=sha256:cc32b6a94e75e46b9e7005884e44268d7456ec142faed092bf8647e54e688650

Observation aecd3d84-6991-4604-9374-a45c68a30b04 · outbound

This paper cites JGLUE : J apanese general language understanding evaluation.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language JGLUE : J apanese general language understanding evaluation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.169580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.169580Z digest=sha256:4cc846dd44df46e1b9be3d263eec99b434102658847e0d8e2b63dad4119068d2

Observation 8ee3b8ed-aa0d-41e0-91be-f4186e87631a · outbound

This paper cites Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedback.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedback

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.239094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.239094Z digest=sha256:c1410ea110fad886abd39d4299796eb4d371e01d369dfe6d62ae1e3631b16efc

Observation eaeafdc2-d5d4-495e-8f1b-8115b58b28c5 · outbound

This paper cites F lau BERT : Unsupervised language model pre-training for F rench.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language F lau BERT : Unsupervised language model pre-training for F rench

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.343322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.343322Z digest=sha256:d04a3d7954a9a4e50e7da22974247057a949aa8353c3411a3a60fa692cbb2261

Observation 957c7581-a3b7-4084-a9cc-2950589ca17d · outbound

This paper cites Open-arabic-llm-leaderboard-v1.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Open-arabic-llm-leaderboard-v1

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.454235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.454235Z digest=sha256:2bde525e7ba414ab720b0d769cbe62d9420af6e79fbccaa424670451f57f3d56

Observation 98be8b18-486a-4cbf-ac8a-6ea4421f473f · outbound

This paper cites Deduplicating Training Data Makes Language Models Better.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Deduplicating Training Data Makes Language Models Better

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.550459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.550459Z digest=sha256:86086b267da9900a1dc29c6cc78a1532828aaa5f2386b48b4ef3d1ef6f785f41

Observation b6b25d1c-e7e3-4731-87d8-d9ad2d2f05ab · outbound

This paper cites Kiwipiepy: Kiwi package for python, 2024.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Kiwipiepy: Kiwi package for python, 2024

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.641312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.641312Z digest=sha256:f17b7fb17cb904cabafa2fcb0150133a01abad243112fdc87710683907067cdb

Observation 072cbc6c-d94e-493c-9e08-bb7b96b8a997 · outbound

This paper cites MLQA: Evaluating Cross-lingual Extractive Question Answering.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language MLQA: Evaluating Cross-lingual Extractive Question Answering

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.718750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.718750Z digest=sha256:e4a1403ae1eed9a0bb78b941b9c0c402d02fd4bb304a018dce45dfab73b8d27f

Observation 82379ae4-9e75-4a1f-b422-34a70798f7d2 · outbound

This paper cites CMMLU: Measuring massive multitask language understanding in Chinese.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language CMMLU: Measuring massive multitask language understanding in Chinese

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.789508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.789508Z digest=sha256:45b0dd3255424c89ee9821061c1622a9f6950c87456cd655f6df957b8aef2399

Observation c6e28839-9292-470b-9c28-2b6c7cf4ae35 · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language DataComp-LM: In search of the next generation of training sets for language models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.887384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.887384Z digest=sha256:1d8b73cf19579506a72754fc59e1415f139a47c82021ef8700028c2a9fb0b6b6

Observation d4208a9a-d1bd-463b-bb92-2efd48be185b · outbound

This paper cites Common sense beyond E nglish: Evaluating and improving multilingual language models for commonsense reasoning.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Common sense beyond E nglish: Evaluating and improving multilingual language models for commonsense reasoning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.950982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.950982Z digest=sha256:343aa3aa54da924c8694d81fe7e752b67f0fdaef668beb456822212953ef28df

Observation a01ec591-bec5-493d-bd80-54c62c0db555 · outbound

This paper cites Few-shot Learning with Multilingual Language Models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Few-shot Learning with Multilingual Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.120607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.120607Z digest=sha256:2ab0d1d9cd5b07458b14de70e7a11c05bee3b753a12916897dc3d0261e0e551e

Observation adc5788a-159a-4393-8e78-f5702c2bec77 · outbound

This paper cites FinGPT: Large Generative Models for a Small Language.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language FinGPT: Large Generative Models for a Small Language

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.228960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.228960Z digest=sha256:bf18591f06b607d2ee736ac5cfda1ae7e8b0050485b42487983531f77e0ab30b

Observation ffa8195f-6c42-430e-a18d-f5c38ae41b92 · outbound

This paper cites Quantifying Variance in Evaluation Benchmarks.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Quantifying Variance in Evaluation Benchmarks

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.311443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.311443Z digest=sha256:a32db50b18d2fd89be7ea1e1e4246ff0d27e07899b56cae8c3973f25b947946f

Observation 4be25e36-faac-4254-9da6-e62f2ffad73a · outbound

This paper cites C amem BERT : a tasty F rench language model.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language C amem BERT : a tasty F rench language model

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.376511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.376511Z digest=sha256:753a35859b8544c3011f75efcd698fb66e25a0e3a6dc16c28e1a732934398edc

Observation 30ff154b-ccfc-425f-815b-7b6da12380b0 · outbound

This paper cites Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.446794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.446794Z digest=sha256:6813083e35427b284e24d59d0b9ad4a1b639447aedc715f374a761a537b4795a

Observation f2027824-04d2-46e6-8767-f458b66c54f0 · outbound

This paper cites Mnbvc: Massive never-ending bt vast chinese corpus.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Mnbvc: Massive never-ending bt vast chinese corpus

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.530378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.530378Z digest=sha256:11215c66d0ffba4ec121d77139b7878d5c964c15199d07ea7b3cec2fa57ae727

Observation 996645bd-7901-40f4-a1f0-8f1ca508ca5d · outbound

This paper cites Neural A rabic question answering.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Neural A rabic question answering

Reference 75

Resolution
verified exact
doi, observed 2026-08-06T22:46:43.644012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:46:40.612521Z digest=sha256:9fe3bb24e6f297ed66a45af6933f8af55df7e0bbcf3bc14877d286f4c06324aa

Observation 8441a717-8366-4f55-8bd0-2c7e762cf8be · outbound

This paper cites Crosslingual generalization through multitask finetuning, 2022.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Crosslingual generalization through multitask finetuning, 2022

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.708439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.708439Z digest=sha256:0236267374ca41716d7d78a31fa89a1ee14d14030e204d81d6f0b9ee1f019d4f

Observation f82ef352-d1ca-4cac-a91d-b5d765dd15d9 · outbound

This paper cites Rossi, and Thien Huu Nguyen.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Rossi, and Thien Huu Nguyen

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.789698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.789698Z digest=sha256:b482d66f7f95cb7a9454d9bda96b78ca8cd19d28947a27a0de32b8b77d6b7e67

Observation a4a84145-1fcf-45ba-930a-8d8230b3334f · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.884821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.884821Z digest=sha256:f6b753b68bc10b71c80c8ba16113d2b53976266bccabbb17666213de3b7d103f

Observation b02b7449-4b62-4a97-bf29-1576201f80f5 · outbound

This paper cites Omnia russica.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Omnia russica

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.940992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.940992Z digest=sha256:6755a1a1fa69b64fca84a71b643c3cbe3d88f6f906dc1ba43b432a392f8d06d2

Observation 967d7d78-2977-47f4-8d9b-6bc072bb1a8d · outbound

This paper cites Botok: State-of-the-art tokenizers for tibetan language, 2025.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Botok: State-of-the-art tokenizers for tibetan language, 2025

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.003899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.003899Z digest=sha256:df86855d4aaaa898591841c32dfefa3f4da7bcc09e61a2999a2fb5f347e2cde0

Observation 41d85cdc-552c-43ca-8bde-c85e358c797f · outbound

This paper cites Building pre-train llm dataset for the indic languages: A case study on hindi.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Building pre-train llm dataset for the indic languages: A case study on hindi

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.104136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.104136Z digest=sha256:da77735488a9c7f6d7ff52ab3c1ba00ff0163ac834e9d14054b93b3901d97697

Observation 13ff38a0-299a-4cff-8429-f354bf0330b1 · outbound

This paper cites Hellaswag-th, 2023.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Hellaswag-th, 2023

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.162030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.162030Z digest=sha256:596b42b655ce673471023eac8474bdda614ff66f4a66d52e9a7285bc3e9b1b82

Observation 3636b3d9-ad11-4cf7-aa9b-70f4aad44a20 · outbound

This paper cites The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.208321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.208321Z digest=sha256:60f217128f68a78982daa62f47697879d4c5f4e509ce4d2bb387293db6b1f4da

Observation 1421b3f6-cc9a-4d8e-9540-d4986d14c6f1 · outbound

This paper cites The fineweb datasets: Decanting the web for the finest text data at scale.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language The fineweb datasets: Decanting the web for the finest text data at scale

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.292670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.292670Z digest=sha256:93e453e8d3de73478fbce52717f5e4a08ca1651172131b6753ba64097d5de983

Observation f6c30583-bdce-4a57-86ea-01fa20c397b2 · outbound

This paper cites Laonlp: Lao language natural language processing, July 2022.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Laonlp: Lao language natural language processing, July 2022

Reference 85

Resolution
verified exact
doi, observed 2026-08-06T22:46:43.524600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:46:41.328021Z digest=sha256:74683e46c859a2e6e876ade741e81895894cf3b02bfa8e73953dfaffcdf8fbaa

Observation d4910a25-953d-42a0-a5cc-a60dfbd80149 · outbound

This paper cites P y T hai NLP : T hai natural language processing in P ython, June 2024.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language P y T hai NLP : T hai natural language processing in P ython, June 2024

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.368981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.368981Z digest=sha256:bd5a41d6ccc9ba902122c9a3dc8140d82e3868645dd27a7a3ed529a39141fd0f

Observation 792e8fff-e19d-49a1-8874-bc9c2e996040 · outbound

This paper cites Typhoon: Thai Large Language Models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Typhoon: Thai Large Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.432866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.432866Z digest=sha256:5fee15730d61303e81ead5bb0ba01ca43f95e34708876897b8852100437b147b

Observation 00ab1a9a-a5b7-47ee-a3ee-117d04220b57 · outbound

This paper cites Pllum: A family of polish large language models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Pllum: A family of polish large language models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.515692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.515692Z digest=sha256:cc3272175c89cf591f09e72a4ec4048dce4a40046a903d7d29e451ad1d083e45

Observation 800b7eb2-19b6-4ae1-bbfa-70d87904a68d · outbound

This paper cites Chinesesquad.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Chinesesquad

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.584393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.584393Z digest=sha256:fe69e3e8c0b0941cd22db5b035fe2b1f858e53bade7b90ab81fd3a66d5431e6b

Observation 8ae140c9-8575-42bd-a403-2f666001c830 · outbound

This paper cites XCOPA: A Multilingual Dataset for Causal Commonsense Reasoning.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language XCOPA: A Multilingual Dataset for Causal Commonsense Reasoning

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.630958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.630958Z digest=sha256:8f44c00a8af7a7c9d730b013229979e205a4126ac325fa6983b1257e70dd8ee8

Observation 213d288c-62ad-464c-a708-530c35a4b8d7 · outbound

This paper cites Stanza: A Python Natural Language Processing Toolkit for Many Human Languages.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Stanza: A Python Natural Language Processing Toolkit for Many Human Languages

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.693007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.693007Z digest=sha256:8bce6cd0cba4db137c4074d84c85544b241499a5f6dd40674e3369942e59065a

Observation 16b60db9-7c02-4cb9-90ab-2fdbfc3e98e6 · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.766456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.766456Z digest=sha256:e481d967dd1886b3ffc1ee40c5a769acc2aed3622aca1602005bf5177dccf07a

Observation b23d45d5-4bb6-4328-afa5-fa2979fd8043 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.821123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.821123Z digest=sha256:8c6eb0cfc79312230dc96a37633a342f9aaef93568544591a42072332505a4bd

Observation 078640bb-dca3-4481-b332-812e91aec898 · outbound

This paper cites Impact of Pretraining Term Frequencies on Few-Shot Reasoning.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Impact of Pretraining Term Frequencies on Few-Shot Reasoning

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.895184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.895184Z digest=sha256:6aa169fc3d77d8830f1a9616c9ac86bc285f5a8ac649778fe617f078ceae4408

Observation f195de96-efca-4b7d-8ce1-43649ab4f262 · outbound

This paper cites How Much Knowledge Can You Pack Into the Parameters of a Language Model?.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language How Much Knowledge Can You Pack Into the Parameters of a Language Model?

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.932613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.932613Z digest=sha256:b8c4396f98bbf624859eade1346089695448d55a1f32f8282b015825f56e157c

Observation e9647ceb-cb4a-488a-a35e-0db91dae8727 · outbound

This paper cites How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.999731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.999731Z digest=sha256:d23b317b98b54312c2e3599615789cf5151db3dcecddb48357f4dd3f6bfbd155

Observation 4b72f10a-e44f-434e-9e6c-32f17e657e5c · outbound

This paper cites Pyidaungsu: Python library for myanmar language, 2024.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Pyidaungsu: Python library for myanmar language, 2024

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:46:45.888138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:46:42.080031Z digest=sha256:f0eb44e57786ba5691f99417750ce5477f9bbedd2ea0ce0dc1713b413e57f253

Observation 63725564-5a04-489c-a39b-9f21105c3bc5 · outbound

This paper cites Compact language detector v3.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Compact language detector v3

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:46:45.879187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T22:46:42.103931Z digest=sha256:b888977c645081574e9ed0f34fbaf4ff25ee55ae939c030e5d5c50c0d01c4bd0

Observation 8f76eddb-8eb6-4b46-a375-2dc8251a7cbd · outbound

This paper cites Mintaka: A Complex, Natural, and Multilingual Dataset for End-to-End Question Answering.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Mintaka: A Complex, Natural, and Multilingual Dataset for End-to-End Question Answering

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:42.192719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:42.192719Z digest=sha256:05d96affd77a6480d0ff60367b8d2c7ffa107b454850480afe8bd7eda57cebd3

Observation 6f7c332a-cbf0-48f2-8dd8-067bbe507e12 · outbound

This paper cites INDIC QA BENCHMARK: A Multilingual Benchmark to Evaluate Question Answering capability of LLMs for Indic Languages.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language INDIC QA BENCHMARK: A Multilingual Benchmark to Evaluate Question Answering capability of LLMs for Indic Languages

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:42.248951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:42.248951Z digest=sha256:c3ce7f863f5035484869d022e61329dc3a70cb8d2faac76cc5385be310830829

Observation 982ab6e6-25d6-4030-8a74-e6e7278fcfaa · outbound

This paper cites Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:42.309351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:42.309351Z digest=sha256:047ce6fb2599367ab988a9274bf6c0e554992e81e633a97049da5d040f7b4c96

Observation c20a06fc-aaf4-4213-a34f-b98e46657df4 · outbound

This paper cites Thquad: Turkish historic question answering dataset for reading comprehension.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Thquad: Turkish historic question answering dataset for reading comprehension

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:42.375506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:42.375506Z digest=sha256:14c3c1c7e3b4ab2386a19825defa63d26fd44f16d54ddef0186bf5d255d9a13f

Pith citing papers

Observation 09e337de-4ea6-40da-aabe-c6305c9f5365 · inbound

From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation cites this paper.

From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:16:45.263479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:16:45.263479Z digest=sha256:91eb2f0421bdf224c7e488645e7d1471c090ac5f19356c622f977eec9c1137e3

Observation 0c6bb196-3a1f-4b67-ae83-0f7f9f57eabe · inbound

Observation of momentum dependent charge density wave gap in EuTe4 cites this paper.

Observation of momentum dependent charge density wave gap in EuTe4 FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T22:45:15.379631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:45:15.379631Z digest=sha256:4fe53fe723df1ea96c639228c0dc12c0c7e921a74b3e8f310ea6740ddb4ac20c

Observation 67cc670d-1c49-4e0f-9050-c5fdcc28d126 · inbound

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models cites this paper.

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:50:08.479194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T17:50:08.399160Z digest=sha256:a09f6a5c6934214ec730c0330c5311c56d664443a68d9aa1acb5ad0923a0b02a

Observation 7b73d5d6-df5b-429e-b414-072673213de2 · inbound

Entropy2Vec: Crosslingual Language Modeling Entropy as End-to-End Learnable Language Representations cites this paper.

Entropy2Vec: Crosslingual Language Modeling Entropy as End-to-End Learnable Language Representations FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T05:42:50.078953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:42:50.078953Z digest=sha256:7355a9bfeda87d6cb9e7b133c7a2ff82654ccb04223be4d4cbd3f71d396e610b

Observation 45b04929-9776-4355-bbde-389fd154ecf0 · inbound

The Effect of Scripts and Formats on LLM Numeracy cites this paper.

The Effect of Scripts and Formats on LLM Numeracy FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T08:58:50.692729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:58:50.692729Z digest=sha256:4dbba667f54b6d15f4193b4ade5b89fd05728743280b250af7fc983c462c9f31

Observation e263f088-3f4e-4306-8aae-d48749d9a534 · inbound

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data cites this paper.

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 167

Resolution
unresolved
no resolver link, observed 2026-08-03T08:15:27.454474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:15:27.454474Z digest=sha256:1eacadfed6d2078cc46b237e4e7939c64503cc76e09b173cf3223137c57f8441

Observation 79a8722b-9b1e-435f-9f17-7bc8c9414d93 · inbound

OpenLID-v3: Improving the Precision of Closely Related Language Identification -- An Experience Report cites this paper.

OpenLID-v3: Improving the Precision of Closely Related Language Identification -- An Experience Report FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T23:37:59.719719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T23:37:59.719719Z digest=sha256:8f86a9e9989c17015eb3e2abc91084ebc7f5b7a8c6586d653465aca96fd7a75a

Observation 763e30af-8044-4a58-8526-9f21ffcb16c1 · inbound

Scaling Laws for Mixture Pretraining Under Data Constraints cites this paper.

Scaling Laws for Mixture Pretraining Under Data Constraints FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:48:00.949304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T21:44:31.429223Z digest=sha256:a54e01311d1e61e2489e1de611b6aa4ed72de4ca2407c8ef08bf136e55a13e9d

Observation 3a3d4225-0be4-4675-bc7f-70b881a3c0a0 · inbound

Scaling Laws for Mixture Pretraining Under Data Constraints cites this paper.

Scaling Laws for Mixture Pretraining Under Data Constraints FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:37:39.604606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T16:36:30.007014Z digest=sha256:631a8a84dcb46c570163f42eb1e529e5a3292bbf1936c83a72802626fc9242d6

Observation 10a7a53c-b215-4103-8017-0c99cb34c20e · inbound

Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings cites this paper.

Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:19:27.448259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T20:17:26.661595Z digest=sha256:ab428d03626d7eac97aa38be47e4c07646281097a8d8138a0b9eff836b41a40c

Observation 009b4acd-a370-4dba-a981-20466c2442fe · inbound

Granite Embedding Multilingual R2 Models cites this paper.

Granite Embedding Multilingual R2 Models FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:17:34.932496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T18:16:49.148303Z digest=sha256:963d23a394e28d765d3d5253d271472838809162c9b33b24ae45f1fea16c30ef

Observation 9c65265e-e2be-427e-a18f-c02d71c8ba7e · inbound

A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAM$\Delta$ Integration into Upcycled MoE cites this paper.

A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAM$\Delta$ Integration into Upcycled MoE FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:13:13.616472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T11:09:22.027588Z digest=sha256:71d674f101f06a6417db5a86a4fe7817e10b3619b5f732f7df3e97970347b457

Observation 55dab8fb-b24c-476f-a89e-a007b018028b · inbound

SomaliWeb v1: A Quality-Filtered Somali Web Corpus with a Matched Tokenizer and a Public Language-Identification Benchmark cites this paper.

SomaliWeb v1: A Quality-Filtered Somali Web Corpus with a Matched Tokenizer and a Public Language-Identification Benchmark FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:38:12.634240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T10:34:49.783942Z digest=sha256:e2393170232b104f90e1a789c2fe063d9e9255bd2085ed881a3a3c0d0e66b489

Observation ed65c45e-2831-430a-8d11-502aa052c297 · inbound

Multilingual Knowledge Transfer under Data Constraints via Lexical Interventions cites this paper.

Multilingual Knowledge Transfer under Data Constraints via Lexical Interventions FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:15:20.257197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-25T04:11:06.312503Z digest=sha256:f2f300c37f005b71b98c5275ae681bdc964158f32799703addcc274baaef46f6

Observation cf0129b7-a60b-40ca-9fb4-83adcd620143 · inbound

Mimir: Large-scale Multilingual Concept Modeling cites this paper.

Mimir: Large-scale Multilingual Concept Modeling FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T11:24:38.681356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T11:08:44.027943Z digest=sha256:85073fc5006d4ee660cf9d4ce4c9e378c310cc068b5c65cd3419342225be282f

Observation 57f65719-6f44-406a-861c-fcf0faa9aabd · inbound

Ishigaki-IDS: An Open-Weight Verifier-Aware Model for Information Delivery Specification Drafting in Building Information Modeling cites this paper.

Ishigaki-IDS: An Open-Weight Verifier-Aware Model for Information Delivery Specification Drafting in Building Information Modeling FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 25

Resolution
malformed identifier
arxiv_id, observed 2026-07-02T23:17:29.057127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T18:24:53.592902Z digest=sha256:d7ea4733d39f87fa0a49c3f97e6ec6af901e39281283025b22f9b10ff88289de

Observation df6a9db1-36aa-41d6-b236-56f1861d6a01 · inbound

Grouped Query Experts: Mixture-of-Experts on GQA Self-Attention cites this paper.

Grouped Query Experts: Mixture-of-Experts on GQA Self-Attention FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:49:29.893255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T17:40:06.036019Z digest=sha256:7ac52b4ef593fcf2b7fa2846e4692adb667c13c8881dc8ac2bcd3475cd2809b2

Observation d2bd763f-f994-4834-9947-29e54e4eb04e · inbound

moBERTo: A Modern Encoder for Portuguese via Continued Pretraining of ModernBERT cites this paper.

moBERTo: A Modern Encoder for Portuguese via Continued Pretraining of ModernBERT FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:19:44.124861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T10:07:55.521194Z digest=sha256:2e776fa1ed42638d709d9d25a931f01f80e8f28c5ec9b8b8dcf350c81c2c3fcc

Observation 1b5b1bdb-58fb-408b-9c29-2bf74907bb15 · inbound

LangMAP: A Language-Adaptive Approach to Tokenization cites this paper.

LangMAP: A Language-Adaptive Approach to Tokenization FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:49:46.653379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T08:26:33.185340Z digest=sha256:77f6240a3956753504b8052da5bb986cb5906961584c8cdf669f93c4e1300ff2

Observation d3e45960-722e-4f56-bb55-437ffa4ab47a · inbound

SARA: Unlocking Multilingual Knowledge in Mixture-of-Experts via Semantically Anchored Routing Alignment cites this paper.

SARA: Unlocking Multilingual Knowledge in Mixture-of-Experts via Semantically Anchored Routing Alignment FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:00:09.022875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-25T20:48:04.003465Z digest=sha256:32e4e15acc710e0f8b55edb5a2ebb3cf5fa8c912940d53533290075c3d276240

Observation 65447489-0aee-4602-976e-e5023ff38cb3 · inbound

MultiHashFormer: Hash-based Generative Language Models cites this paper.

MultiHashFormer: Hash-based Generative Language Models FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-29T04:23:05.508816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T04:13:05.082903Z digest=sha256:2d1a19c0bc921a0b5c75cacb6ecafa0bfa9477f3cc81cf89abd409e23be74158

Observation a12afaeb-34ce-473a-91f9-9ad1b79e3171 · inbound

In-Place Tokenizer Expansion for Pre-trained LLMs cites this paper.

In-Place Tokenizer Expansion for Pre-trained LLMs FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:16.706422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:16.706422Z digest=sha256:18233fbed682dd4ba038668c5cd36ffb4f9f51dc6c1962d3e2b085652a5ab09d