Pith. sign in

Paper Citation Record · LEDGER

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

As of 18 August 2026, this Paper Citation Record lists 100 of 118 outbound references and 24 inbound Pith citation observations for arXiv:2506.20920.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20920 v1

Coverage vector

measured 100 of 118 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:46:42.375506Z

measured 124 of 124 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:18:26.734265Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 118 outbound references displayed

  • verified exact8
  • verified fuzzy2
  • unresolved89
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 5ec22fa5-a012-41c0-82da-f189097500a9 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Yi: Open Foundation Models by 01.AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.435145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.435145Z digest=sha256:7ebedd45ca60d3f047ebaec402785107d9c1cd0892c2ed790d49c4722a03e1ae

Observation 93645589-9dc1-4d2a-8128-c1a94669f570 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.489072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.489072Z digest=sha256:b6aaca39bbc702f43c00ec3e3e611d99cc6baedc19dcc25bd79d0f60b1915183

Observation 6894d93a-aa44-48df-8279-75e4b1f5b8f6 · outbound

This paper cites Llama 3 model card.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Llama 3 model card

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.557062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.557062Z digest=sha256:0ce919bcb9e441d87549734a75a28d8b1f4913bdea3cc6e45dbdd5beaecea393

Observation 9b2b22c6-7c9f-43bf-b997-1ec8c9115267 · outbound

This paper cites A survey on data selection for language models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language A survey on data selection for language models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.625972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.625972Z digest=sha256:d1dfac074e5551e3fc5b345269d70bfd5cc6f7a65b81f865f7a6a719162d7bcb

Observation 767dab6a-c430-4bea-a8fa-31aa9e767aec · outbound

This paper cites Open llm turkish leaderboard v0.2.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Open llm turkish leaderboard v0.2

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.702254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.702254Z digest=sha256:bf5f5796b657a73891b5b784982ce272ba517b0ffe54f082956d607b3a3abbe8

Observation a1f05eab-e33f-40fb-977c-3aac94923944 · outbound

This paper cites A l G hafa evaluation benchmark for A rabic language models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language A l G hafa evaluation benchmark for A rabic language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.797026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.797026Z digest=sha256:5e1c2b17be2c0bd5b221dfc3831dfb91129c08969740cf2771b7777b7f196445

Observation c00582aa-05a4-47a7-96da-83d765789020 · outbound

This paper cites 101 billion arabic words dataset, 2024.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language 101 billion arabic words dataset, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.843819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.843819Z digest=sha256:424ef0134dff1a13cb6c6f474931c69aa1d0050ce825ef3e0f283f22d7645072

Observation 20cce32f-fab4-43bc-8d2e-76412bc7fc45 · outbound

This paper cites On the cross-lingual transferability of monolingual representations.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language On the cross-lingual transferability of monolingual representations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.881227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.881227Z digest=sha256:e4391ea595b3d09bf41be3b99aaa20fcc19021708d94063ce72fe1eaabfe9fca

Observation 67eaca72-07bb-4902-b2db-fa169d3747a0 · outbound

This paper cites A call for more rigor in unsupervised cross-lingual learning.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language A call for more rigor in unsupervised cross-lingual learning

Reference 9

Resolution
verified exact
doi, observed 2026-08-06T22:46:43.869917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:46:33.948108Z digest=sha256:0ad81b3fd18153b5da2a145eeef2a21715ae216f3a00ac91f66ceba65a14d62b

Observation fd5889bf-fa8d-4af2-aa2b-5ae7d4a83975 · outbound

This paper cites Japanese massive multitask language understanding benchmark, 2023.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Japanese massive multitask language understanding benchmark, 2023

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:33.995458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:33.995458Z digest=sha256:dce3fb5d3788093a50e775fbebf60f4a95ab1ccc7ddef0115f0862bf1bd92fd2

Observation 9b0d82cf-0d66-40f5-8a0e-da42bf5d4814 · outbound

This paper cites The belebele benchmark: a parallel reading comprehension dataset in 122 language variants.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language The belebele benchmark: a parallel reading comprehension dataset in 122 language variants

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:34.075955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.075955Z digest=sha256:a00728d1ad59b00af34d2093bb9b4b69cddc8c828e609cdd672d7fa73aa70d20

Observation 2d29a952-39ae-463d-87d5-d28ef38c1316 · outbound

This paper cites Building Machine Translation Systems for the Next Thousand Languages.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Building Machine Translation Systems for the Next Thousand Languages

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:34.150087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.150087Z digest=sha256:b60efc5fb5e004315f978edc40adc96fda72e37dc7cafb9e41e5272217d29646

Observation aee1a052-c635-4462-b70b-a26749d0af24 · outbound

This paper cites Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:34.210423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.210423Z digest=sha256:56e1a01073a3b563423ee29de51224d383eee5c03218baed7d8e8e14913ac8b1

Observation c3578c5c-d4c7-49a8-a653-2ca0dfa9d86e · outbound

This paper cites On the resemblance and containment of documents.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language On the resemblance and containment of documents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:34.319292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.319292Z digest=sha256:62cc1ec5d600a193fc34440c10be420e493525a39e6e61c84a0cc105c640de74

Observation ead3e28e-4569-40f6-878e-666442a5ef75 · outbound

This paper cites An open dataset and model for language identification.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language An open dataset and model for language identification

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:34.474853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.474853Z digest=sha256:a50d65265f63cccac1d59929b993646d65452e3b8ad523a63ba5adfb28dc0369

Observation 6f6c3cb5-c38d-4cce-b6f4-f6f4b634ed5a · outbound

This paper cites An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT).

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:34.583082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.583082Z digest=sha256:e3344abe751e925c9c60eec83103b5c1b4ad671aeced98a184a910fd7e8a39ed

Observation 51c0b978-de55-4141-810a-32a5f05bf112 · outbound

This paper cites PTT5: Pretraining and validating the T5 model on Brazilian Portuguese data.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language PTT5: Pretraining and validating the T5 model on Brazilian Portuguese data

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:34.736671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.736671Z digest=sha256:0e30e4d9a879c2ac5043693cb259e175f9a3a1acea92752fb1f156983bc10dc6

Observation d97dbaa2-3c1b-4e30-bb53-124de1eeb969 · outbound

This paper cites Language ID in the wild: Unexpected challenges on the path to a thousand-language web text corpus.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Language ID in the wild: Unexpected challenges on the path to a thousand-language web text corpus

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:34.831748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.831748Z digest=sha256:3e522e3862f81b4d04a2ac0f1a72e3afc5f1a3aff2071d5c86305c10d4c3744b

Observation f48cd5d1-bcbe-49cb-8ec9-074466a7393e · outbound

This paper cites Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki

Reference 19

Resolution
malformed identifier
no resolver link, observed 2026-08-06T22:46:34.988316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:34.988316Z digest=sha256:249990f451e6248ee2b01bf2b1a72dca3ed8bca903c2086a923a7eea1eb514f1

Observation a73efc9f-d303-4a64-a261-dd8ea894f2dd · outbound

This paper cites Command r+.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Command r+

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:35.082235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:35.082235Z digest=sha256:c0605e2bb1e46ab046681c41754f8e09a0b0722187ada5fec793dc3801269429

Observation 204dfc02-b923-4ba9-85cb-9b4525f83964 · outbound

This paper cites Unsupervised cross-lingual representation learning at scale.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Unsupervised cross-lingual representation learning at scale

Reference 21

Resolution
verified exact
doi, observed 2026-08-06T22:46:43.742969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:46:35.190771Z digest=sha256:930fd643fc782c04e2ed688cf14eaf328eca1dc49b3f3172efccf7b707e7a6ac

Observation ebfebb77-afd4-4762-8ae6-1ec465b2ef68 · outbound

This paper cites Neural learning for question answering in italian.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Neural learning for question answering in italian

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:35.333775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:35.333775Z digest=sha256:8f79e25cb34d35dfcfda863d3f15cfbd228adc5655d2f512f5b264d4551bd707

Observation df62415c-318d-40c7-80cc-884a41aa911f · outbound

This paper cites Dataset for the First Evaluation on Chinese Machine Reading Comprehension.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Dataset for the First Evaluation on Chinese Machine Reading Comprehension

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:46:45.068840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:46:35.445116Z digest=sha256:9573c8ac413935d88f77134e0747349851f66cbe79b261ca6286ce9cf34bf8a7

Observation 504fbcb6-49c0-4587-9919-3ba7ae8ddc77 · outbound

This paper cites Daniels and William Bright (eds.).

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Daniels and William Bright (eds.)

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:35.543260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:35.543260Z digest=sha256:d2bf36a0d27a2f1aae486e5a385a01f03c426134ef13ba35e666ff326a106d73

Observation 75a59b1b-7199-4d01-ba28-fac54f1e3df1 · outbound

This paper cites A New Massive Multilingual Dataset for High-Performance Language Technologies.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language A New Massive Multilingual Dataset for High-Performance Language Technologies

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:46:44.888181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:46:35.727398Z digest=sha256:9ee398234e3e6ae888d69b9d7962708813422f59e218b80f6264ce3a3550feb7

Observation ed081d84-e88a-4f48-92f3-42000a1ef7c8 · outbound

This paper cites BERTje: A Dutch BERT Model.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language BERTje: A Dutch BERT Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:35.817846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:35.817846Z digest=sha256:55e3209b146c0ddcc735dc267958862e2b843d07986477a02d4137758d30e6e5

Observation 488d3324-dc9d-45c4-a85d-3f57bcfdbcf5 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:35.894130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:35.894130Z digest=sha256:315e466bcf7b20016fbb59e03097f96d4c9a5605ea040c83faae21668db42b85

Observation 579cc685-a676-4f08-aa87-8aeca05345ef · outbound

This paper cites RobBERT: a Dutch RoBERTa-based Language Model.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language RobBERT: a Dutch RoBERTa-based Language Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:35.985633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:35.985633Z digest=sha256:a787f913e13c7b37e7c7567800e5660d959578cb1b41015bef210b6930b61084

Observation 48891c63-e204-485a-b9db-06320ce3391e · outbound

This paper cites FQuAD: French Question Answering Dataset.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language FQuAD: French Question Answering Dataset

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:46:44.741197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:46:36.078826Z digest=sha256:5c6b15bf8c8a4737a88cb65e9b9193fd5cecbf0ffad05ec3226c0ec264a97663

Observation bf5d4220-13a8-4211-ae7c-eab29ffc1d12 · outbound

This paper cites Sailor2: Sailing in South-East Asia with Inclusive Multilingual LLMs.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Sailor2: Sailing in South-East Asia with Inclusive Multilingual LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.172908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.172908Z digest=sha256:a3a4d0353cc75bdb4aaf3205ceb77fa68f9b4316d77a2f85d15e57dcb4fcce54

Observation 96fee737-29a8-477b-b0ed-85d088c6a9fc · outbound

This paper cites Chinese Tiny LLM: Pretraining a Chinese-Centric Large Language Model.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Chinese Tiny LLM: Pretraining a Chinese-Centric Large Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.260283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.260283Z digest=sha256:0e01dd1c712b4ccaeb6e536dcf50d8352ee1cd5fbe4c14ac1e4c050c0527f361

Observation c3edfc52-178b-4997-99d4-9ae23425345f · outbound

This paper cites Eberhard, Gary F.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Eberhard, Gary F

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.357550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.357550Z digest=sha256:6d5c26c6c0008a4479ae5c1d48b5c43f83d96d2d129b093af2714c0697defcfa

Observation 086ea4a7-24ab-48bc-a1e0-83a70fe57691 · outbound

This paper cites SberQuAD – Russian Reading Comprehension Dataset: Description and Analysis, pp.\ 3–15.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language SberQuAD – Russian Reading Comprehension Dataset: Description and Analysis, pp.\ 3–15

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.484014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.484014Z digest=sha256:a4fee2490865fe7d5cf1d5e7e27d235a116a791158ef3ee69ec902948f4f443a

Observation 4f6ba979-b408-4848-b114-102f4590aae4 · outbound

This paper cites Arabicweb24: Creating a high quality arabic web-only pre-training dataset, 2024.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Arabicweb24: Creating a high quality arabic web-only pre-training dataset, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.563221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.563221Z digest=sha256:d656c55af1b81d80fb8a2e9eb1e3d6f3b5aa5fbd6c24784a3b24a94cbe5752de

Observation 44f6f449-cb49-424c-803c-bad6e90723f9 · outbound

This paper cites Guerreiro, António Loison, Duarte M.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Guerreiro, António Loison, Duarte M

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.638830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.638830Z digest=sha256:7e012647d5f8d80a6e8a2a57e86b06de32790433d1a8041176afe8d434aa5b3c

Observation be88cc1e-59c7-4bed-bca5-0a003eddfb9e · outbound

This paper cites MERA: A Comprehensive LLM Evaluation in Russian.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language MERA: A Comprehensive LLM Evaluation in Russian

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.710892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.710892Z digest=sha256:a8c25f5dd5f4b6311ae51289247bf1e7ca00b69042de96ae0a54b74ad458bcc1

Observation 4fd9627c-c59c-4bd2-9299-0f896ee29e4e · outbound

This paper cites Open llm leaderboard v2.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Open llm leaderboard v2

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.843320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.843320Z digest=sha256:e14242545cc91a9a96fbf1d2d5eedc6b0408c1a1208945b9549e03ee4c76376f

Observation fb661be8-d0a6-4bce-a8a6-859bc397a9bf · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Gemma: Open Models Based on Gemini Research and Technology

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:36.956706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:36.956706Z digest=sha256:587409297df9ce000b89ff5fc976f9b44d36dff3258a4835452ab5efc6122c4d

Observation 981672ec-3d2a-4080-855f-8cb4f19284c7 · outbound

This paper cites The Llama 3 Herd of Models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language The Llama 3 Herd of Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:37.083695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:37.083695Z digest=sha256:5db425d8a5f1c2099a038bf4a5013216de5b2e6e47646005b63757af26e1691d

Observation 94c3acd5-d761-4e6b-9598-f18da8e52e00 · outbound

This paper cites Studying Large Language Model Generalization with Influence Functions.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Studying Large Language Model Generalization with Influence Functions

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:37.222530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:37.222530Z digest=sha256:c0d76cabe2d2a78f65c77ed6f965564f4007efb435d3406ef8514e41693a3aee

Observation 7072c36e-25f9-4ad5-9295-6e4aa128b307 · outbound

This paper cites OLMES: A Standard for Language Model Evaluations.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language OLMES: A Standard for Language Model Evaluations

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:37.365791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:37.365791Z digest=sha256:9a9907604db9e4fb2699a3e258d0ebe23873fd62f531a3b70ea6a1f6cec9d43f

Observation 9a33ca50-0ec2-4830-a8ab-8355fa9c5c2c · outbound

This paper cites EXAMS: A Multi-Subject High School Examinations Dataset for Cross-Lingual and Multilingual Question Answering.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language EXAMS: A Multi-Subject High School Examinations Dataset for Cross-Lingual and Multilingual Question Answering

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:46:44.529872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:46:37.521847Z digest=sha256:c121903bcd0415842282725fd456f6bd5577df2c3eb67d4f2d36529698f4c1f5

Observation 1d7babeb-2c21-4c95-8506-9831e9a7cab5 · outbound

This paper cites Measuring massive multitask language understanding.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Measuring massive multitask language understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:37.613762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:37.613762Z digest=sha256:6654481d347ed69f316466418e42c4fbc47c01e1bfd262d3b1e5be2d84bc9cc3

Observation e6cfbcb3-9399-4019-bdd0-090d0eb25307 · outbound

This paper cites Khmer natural language processing tookit.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Khmer natural language processing tookit

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:37.677450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:37.677450Z digest=sha256:a89060ca3f95a1c867ede615558c17ea1c695bbf37fcb920475e6d833d2edac9

Observation 59ff659f-ad21-44bc-867b-0135918901f7 · outbound

This paper cites spaCy: Industrial-strength Natural Language Processing in Python.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language spaCy: Industrial-strength Natural Language Processing in Python

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:37.821687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:37.821687Z digest=sha256:f687814468640f03eaf1717fba3dff167a34a472c14d7b925d75a26710c403b7

Observation 561fdd31-a08e-42b2-bf8f-c9cb73492142 · outbound

This paper cites C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:37.938316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:37.938316Z digest=sha256:bc9d128f03918b84f46fffecccf6d12f7324eac78996610842dfa388d8fbb067

Observation ac51f328-e38f-40e5-8db0-f5462da4da08 · outbound

This paper cites Mistral 7B.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Mistral 7B

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:38.104108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:38.104108Z digest=sha256:7a2bec0b0e651107d99800e31e987da417d5e79c22374f742c78e5c82d465710

Observation f4ff6bd0-52bb-4dfb-9c64-fbd5f186fe61 · outbound

This paper cites Mixtral of Experts.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Mixtral of Experts

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:38.254676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:38.254676Z digest=sha256:ce202447b6fafdc904c3e8ef40fb0efcab2549a041e43dfba84d07514db84793

Observation 1bff3b4a-7fb8-4da2-8b86-0d413dfd1696 · outbound

This paper cites The state and fate of linguistic diversity and inclusion in the NLP world.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language The state and fate of linguistic diversity and inclusion in the NLP world

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:38.411289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:38.411289Z digest=sha256:43c1b078a7c43372d898e60ef4902fe625cc1486883673339fdc834ca25bd339

Observation bf778262-2509-4a1f-8333-9248fed2b284 · outbound

This paper cites FastText.zip: Compressing text classification models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language FastText.zip: Compressing text classification models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:38.542891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:38.542891Z digest=sha256:3a304a810fdc4079143cd5a5dd32164019851552e0ffb7ea4c40224b8a68efbc

Observation 0d33314e-4afd-42de-954f-636bcf10589b · outbound

This paper cites G lot LID : Language identification for low-resource languages.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language G lot LID : Language identification for low-resource languages

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:38.664452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:38.664452Z digest=sha256:2d0200e39353a7b22c3797f3d73b86e7f09cd0d6c4d59629c1fc67f9c71bea8f

Observation f91fd40f-c527-4989-b2c1-955c2585825c · outbound

This paper cites Glot CC : An open broad-coverage commoncrawl corpus and pipeline for minority languages.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Glot CC : An open broad-coverage commoncrawl corpus and pipeline for minority languages

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:38.729060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:38.729060Z digest=sha256:8ccf90e0f3959654b1a34a3380a7f6bb49173500742e1fd77b21c3a3530b8d26

Observation ab275403-84c5-447e-86d0-96b6802e314c · outbound

This paper cites IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian Languages.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language IndicLLMSuite: A Blueprint for Creating Pre-training and Fine-Tuning Datasets for Indian Languages

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:38.868060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:38.868060Z digest=sha256:5c7b847d928995d780c2f1e503eed3a746509fde3a6220f4ba92a8d994025f11

Observation bf5468c1-6200-469f-b7b6-e5dccd617073 · outbound

This paper cites Large language models only pass primary school exams in I ndonesia: A comprehensive test on I ndo MMLU.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Large language models only pass primary school exams in I ndonesia: A comprehensive test on I ndo MMLU

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:38.953519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:38.953519Z digest=sha256:5d6340a02f744b209f90bfd376273ff909aec785d834ed326b78edf95c82fbd6

Observation 7a77dfd9-3344-4fa0-ad00-4b36543393b8 · outbound

This paper cites ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.027217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.027217Z digest=sha256:d73577e9bb7c7e24c8ec6cca89e4cc8ddb8abf430c50ecd32476934c5a537597

Observation 96434442-1e71-403d-96da-6cea0a979b97 · outbound

This paper cites The IndicNLP Library.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language The IndicNLP Library

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.099607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.099607Z digest=sha256:c3f41a68adfc0129f849899f683139c9ba9c91d0ae5735ca9258a3652825f3e7

Observation aecd3d84-6991-4604-9374-a45c68a30b04 · outbound

This paper cites JGLUE : J apanese general language understanding evaluation.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language JGLUE : J apanese general language understanding evaluation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.169580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.169580Z digest=sha256:d990ab689354b46622504ce4ca83be8d931252e423942db385f41515621562c7

Observation 8ee3b8ed-aa0d-41e0-91be-f4186e87631a · outbound

This paper cites Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedback.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Okapi: Instruction-tuned Large Language Models in Multiple Languages with Reinforcement Learning from Human Feedback

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.239094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.239094Z digest=sha256:7f90424e96b66a93881ac45eb49c27a6cd043cdcf4e32ba1f30ac45f533127a6

Observation eaeafdc2-d5d4-495e-8f1b-8115b58b28c5 · outbound

This paper cites F lau BERT : Unsupervised language model pre-training for F rench.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language F lau BERT : Unsupervised language model pre-training for F rench

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.343322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.343322Z digest=sha256:e676f8ddea4f3baaf77fb07ff4a3181b20352dd3d7d38a9f47be1a921d3f4932

Observation 957c7581-a3b7-4084-a9cc-2950589ca17d · outbound

This paper cites Open-arabic-llm-leaderboard-v1.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Open-arabic-llm-leaderboard-v1

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.454235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.454235Z digest=sha256:703f63d9dc1a3903b6de8226ac2326c3d224e3f63b4d4e68b6c2a584d9107a65

Observation 98be8b18-486a-4cbf-ac8a-6ea4421f473f · outbound

This paper cites Deduplicating Training Data Makes Language Models Better.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Deduplicating Training Data Makes Language Models Better

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.550459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.550459Z digest=sha256:5d1f5db049341d16b49e7fd43e341d1ffd6e939c38f6bad93d82b72649595179

Observation b6b25d1c-e7e3-4731-87d8-d9ad2d2f05ab · outbound

This paper cites Kiwipiepy: Kiwi package for python, 2024.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Kiwipiepy: Kiwi package for python, 2024

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.641312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.641312Z digest=sha256:966acba893384604f3625a13ed40542476767c9c3ebed9da5a29688417916c7d

Observation 072cbc6c-d94e-493c-9e08-bb7b96b8a997 · outbound

This paper cites MLQA: Evaluating Cross-lingual Extractive Question Answering.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language MLQA: Evaluating Cross-lingual Extractive Question Answering

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.718750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.718750Z digest=sha256:95e762084827c017b85bc30b6f954f9249b0a8d1368daf10015aad9f5880b64a

Observation 82379ae4-9e75-4a1f-b422-34a70798f7d2 · outbound

This paper cites CMMLU: Measuring massive multitask language understanding in Chinese.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language CMMLU: Measuring massive multitask language understanding in Chinese

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.789508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.789508Z digest=sha256:7a523fb8ad96ccf79ec1650c246acbc44b28d6207ebb8a28e8dede43a644d8ee

Observation c6e28839-9292-470b-9c28-2b6c7cf4ae35 · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language DataComp-LM: In search of the next generation of training sets for language models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.887384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.887384Z digest=sha256:5e5c639a0d795ec39e691adcc9a1d3bddedf37c73da5ee7e4cd7e04c087c7713

Observation d4208a9a-d1bd-463b-bb92-2efd48be185b · outbound

This paper cites Common sense beyond E nglish: Evaluating and improving multilingual language models for commonsense reasoning.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Common sense beyond E nglish: Evaluating and improving multilingual language models for commonsense reasoning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:39.950982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:39.950982Z digest=sha256:7ba3b6adefe131b84e1edb16fa532d77afeecd26c80be71abd31d284aa43a860

Observation a01ec591-bec5-493d-bd80-54c62c0db555 · outbound

This paper cites Few-shot Learning with Multilingual Language Models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Few-shot Learning with Multilingual Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.120607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.120607Z digest=sha256:5dd9291a8ccec7c9b4c9dcb32808d4d24ae2ac491b618559d0ad7b0f43d2134f

Observation adc5788a-159a-4393-8e78-f5702c2bec77 · outbound

This paper cites FinGPT: Large Generative Models for a Small Language.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language FinGPT: Large Generative Models for a Small Language

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.228960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.228960Z digest=sha256:b318bc6e1c95ee58e77e33f2a40fb2c080b69be34d1b5f635ea2285c6ac5b27a

Observation ffa8195f-6c42-430e-a18d-f5c38ae41b92 · outbound

This paper cites Quantifying Variance in Evaluation Benchmarks.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Quantifying Variance in Evaluation Benchmarks

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.311443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.311443Z digest=sha256:e5b807e7cbd92308623123b586874cac4e7abf976ef73a02513c1d4765ed0df4

Observation 4be25e36-faac-4254-9da6-e62f2ffad73a · outbound

This paper cites C amem BERT : a tasty F rench language model.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language C amem BERT : a tasty F rench language model

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.376511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.376511Z digest=sha256:74b3665547921800637584d89e9039484c2428526bba57f1c4b9e4af96d53b51

Observation 30ff154b-ccfc-425f-815b-7b6da12380b0 · outbound

This paper cites Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.446794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.446794Z digest=sha256:30c97c1d826e512cadb03d39d75a907dcc1a1ced01cab261f0b94ac0a3ddcde5

Observation f2027824-04d2-46e6-8767-f458b66c54f0 · outbound

This paper cites Mnbvc: Massive never-ending bt vast chinese corpus.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Mnbvc: Massive never-ending bt vast chinese corpus

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.530378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.530378Z digest=sha256:7b95aa3f908f6418291e56f653d2d154ee124a2d6442206ab29ef4d44673cf16

Observation 996645bd-7901-40f4-a1f0-8f1ca508ca5d · outbound

This paper cites Neural A rabic question answering.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Neural A rabic question answering

Reference 75

Resolution
verified exact
doi, observed 2026-08-06T22:46:43.644012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:46:40.612521Z digest=sha256:360c2d30ca18a6b7aabf89f76a77f8ca1a05a49a811dae486cd606f753687458

Observation 8441a717-8366-4f55-8bd0-2c7e762cf8be · outbound

This paper cites Crosslingual generalization through multitask finetuning, 2022.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Crosslingual generalization through multitask finetuning, 2022

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.708439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.708439Z digest=sha256:4ff2c95be4d3a02be4a22dfbe281a5916376512576ef7372e6179931ff915803

Observation f82ef352-d1ca-4cac-a91d-b5d765dd15d9 · outbound

This paper cites Rossi, and Thien Huu Nguyen.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Rossi, and Thien Huu Nguyen

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.789698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.789698Z digest=sha256:d9088fff4b78c120c531e29adba4024d2c20f812548701ae9d0d73bf86c5348a

Observation a4a84145-1fcf-45ba-930a-8d8230b3334f · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.884821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.884821Z digest=sha256:6339ec0043cd5e94edb33a0fed4cdcd700250d71e617d3979e3e56ea6386d6d8

Observation b02b7449-4b62-4a97-bf29-1576201f80f5 · outbound

This paper cites Omnia russica.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Omnia russica

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.940992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.940992Z digest=sha256:f4a9c427783647ae29922a52a81ba2c8eb18c6787ab8b494677cb302aa77adb7

Observation 967d7d78-2977-47f4-8d9b-6bc072bb1a8d · outbound

This paper cites Botok: State-of-the-art tokenizers for tibetan language, 2025.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Botok: State-of-the-art tokenizers for tibetan language, 2025

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.003899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.003899Z digest=sha256:654a97cb209f2e76f5dd2992c17c24f64ec70a012d724c3afd072b3a0bbcd6df

Observation 41d85cdc-552c-43ca-8bde-c85e358c797f · outbound

This paper cites Building pre-train llm dataset for the indic languages: A case study on hindi.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Building pre-train llm dataset for the indic languages: A case study on hindi

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.104136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.104136Z digest=sha256:64ae51bc40d07adf5a47c3766720286bc7ec49d1b1bd05034463094df96218e5

Observation 13ff38a0-299a-4cff-8429-f354bf0330b1 · outbound

This paper cites Hellaswag-th, 2023.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Hellaswag-th, 2023

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.162030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.162030Z digest=sha256:8ebcc1b95cfe33fdaebf567e2da52b670563d20866d6231f8cd3e363ae4a3b92

Observation 3636b3d9-ad11-4cf7-aa9b-70f4aad44a20 · outbound

This paper cites The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.208321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.208321Z digest=sha256:349151de4fa6f187b33a8b592f0cf4e75c46172a0898dcdc942cc44048c88f70

Observation 1421b3f6-cc9a-4d8e-9540-d4986d14c6f1 · outbound

This paper cites The fineweb datasets: Decanting the web for the finest text data at scale.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language The fineweb datasets: Decanting the web for the finest text data at scale

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.292670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.292670Z digest=sha256:ae30207411e01a212fd5b09b4e57d4fe49b2a10f0487ea5bd331cef610f5b5a6

Observation f6c30583-bdce-4a57-86ea-01fa20c397b2 · outbound

This paper cites Laonlp: Lao language natural language processing, July 2022.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Laonlp: Lao language natural language processing, July 2022

Reference 85

Resolution
verified exact
doi, observed 2026-08-06T22:46:43.524600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:46:41.328021Z digest=sha256:be4bd6589fdcf9ea0d57c9215fef72626523bd74b20f3b861fddd8d195f6feda

Observation d4910a25-953d-42a0-a5cc-a60dfbd80149 · outbound

This paper cites P y T hai NLP : T hai natural language processing in P ython, June 2024.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language P y T hai NLP : T hai natural language processing in P ython, June 2024

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.368981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.368981Z digest=sha256:4c06e40ffef44bb6542ff7b6f0d9d762398325500ca2d507b723daf92b24ba39

Observation 792e8fff-e19d-49a1-8874-bc9c2e996040 · outbound

This paper cites Typhoon: Thai Large Language Models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Typhoon: Thai Large Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.432866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.432866Z digest=sha256:d5f0301d2e1b7fd41b58ea6ad2c416f69cf525cb2d553af866010fd21fe3caf8

Observation 00ab1a9a-a5b7-47ee-a3ee-117d04220b57 · outbound

This paper cites Pllum: A family of polish large language models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Pllum: A family of polish large language models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.515692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.515692Z digest=sha256:5a7edc560590d8990915309d4b214a3ab5b569bf3add4f5edafa2598d4a1d80c

Observation 800b7eb2-19b6-4ae1-bbfa-70d87904a68d · outbound

This paper cites Chinesesquad.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Chinesesquad

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.584393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.584393Z digest=sha256:6f781fb4425b6c34dbe66e548a752f00cf1985c332791268c2d5deb5c14d4f0c

Observation 8ae140c9-8575-42bd-a403-2f666001c830 · outbound

This paper cites XCOPA: A Multilingual Dataset for Causal Commonsense Reasoning.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language XCOPA: A Multilingual Dataset for Causal Commonsense Reasoning

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.630958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.630958Z digest=sha256:b6e410db2edfd13046f479248de0bea3b5c859f4e94bcf7c4353d1fe06530558

Observation 213d288c-62ad-464c-a708-530c35a4b8d7 · outbound

This paper cites Stanza: A Python Natural Language Processing Toolkit for Many Human Languages.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Stanza: A Python Natural Language Processing Toolkit for Many Human Languages

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.693007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.693007Z digest=sha256:7682b94bf4cd38a544d82f512271810512c5df3bb13e4b42f179d5c58ff0e135

Observation 16b60db9-7c02-4cb9-90ab-2fdbfc3e98e6 · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.766456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.766456Z digest=sha256:6d3188d79bbbc15ccca114f3a7ba34d99b5965ef96273905589ac1d7a2d95a10

Observation b23d45d5-4bb6-4328-afa5-fa2979fd8043 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.821123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.821123Z digest=sha256:8d339e2b2f21e525d4695fe382a4ccf8979f902c832a94b870a61fda894cfc34

Observation 078640bb-dca3-4481-b332-812e91aec898 · outbound

This paper cites Impact of Pretraining Term Frequencies on Few-Shot Reasoning.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Impact of Pretraining Term Frequencies on Few-Shot Reasoning

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.895184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.895184Z digest=sha256:60c7c430018dce29170430302c02f13aa5529dc14bb79779078a6414f7e9508b

Observation f195de96-efca-4b7d-8ce1-43649ab4f262 · outbound

This paper cites How Much Knowledge Can You Pack Into the Parameters of a Language Model?.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language How Much Knowledge Can You Pack Into the Parameters of a Language Model?

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.932613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.932613Z digest=sha256:001f9d31cdc60894947d61c20351a96b7cf2f923c32ae396815f3d1852d3a381

Observation e9647ceb-cb4a-488a-a35e-0db91dae8727 · outbound

This paper cites How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:41.999731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:41.999731Z digest=sha256:06ccfa02eb3eb828dec098ef58abc908c95267c8a717be35157d68288027fe76

Observation 4b72f10a-e44f-434e-9e6c-32f17e657e5c · outbound

This paper cites Pyidaungsu: Python library for myanmar language, 2024.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Pyidaungsu: Python library for myanmar language, 2024

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:46:45.888138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:46:42.080031Z digest=sha256:4fcefa50200253f0c6c59f1a98925e91ddf6fe7dd2c0f8cf701fcb0941b8a3a1

Observation 63725564-5a04-489c-a39b-9f21105c3bc5 · outbound

This paper cites Compact language detector v3.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Compact language detector v3

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:46:45.879187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T22:46:42.103931Z digest=sha256:1d7f3843c8284c82e6f21e39322f0d37d6de39049fd12a73a775c0bcb806a747

Observation 8f76eddb-8eb6-4b46-a375-2dc8251a7cbd · outbound

This paper cites Mintaka: A Complex, Natural, and Multilingual Dataset for End-to-End Question Answering.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Mintaka: A Complex, Natural, and Multilingual Dataset for End-to-End Question Answering

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:42.192719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:42.192719Z digest=sha256:417e8ed7e7fe58bf2410c5d0ccad5d4cc648868330c1431a8f7469145825ebb2

Observation 6f7c332a-cbf0-48f2-8dd8-067bbe507e12 · outbound

This paper cites INDIC QA BENCHMARK: A Multilingual Benchmark to Evaluate Question Answering capability of LLMs for Indic Languages.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language INDIC QA BENCHMARK: A Multilingual Benchmark to Evaluate Question Answering capability of LLMs for Indic Languages

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:42.248951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:42.248951Z digest=sha256:ea23b955b402b9a9f3b8cf685d04cc1bf6756cc5e0171aaa7c5f02663b755250

Observation 982ab6e6-25d6-4030-8a74-e6e7278fcfaa · outbound

This paper cites Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:42.309351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:42.309351Z digest=sha256:d9838dc6ab1caba4847cccbd95bb1f354dca55b514a300c026778e3aeeb26856

Observation c20a06fc-aaf4-4213-a34f-b98e46657df4 · outbound

This paper cites Thquad: Turkish historic question answering dataset for reading comprehension.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Thquad: Turkish historic question answering dataset for reading comprehension

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:42.375506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:42.375506Z digest=sha256:ed4c006c54d604b18990621a3e46861c9e935f3fe299720a47769787855f6f9c

Pith citing papers

Observation 09e337de-4ea6-40da-aabe-c6305c9f5365 · inbound

From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation cites this paper.

From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:16:45.263479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:16:45.263479Z digest=sha256:b887879416d819d11bc9e05352f6f947ceb1e2ca4554673aac08adb0f1f82c4a

Observation 0c6bb196-3a1f-4b67-ae83-0f7f9f57eabe · inbound

Observation of momentum dependent charge density wave gap in EuTe4 cites this paper.

Observation of momentum dependent charge density wave gap in EuTe4 FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T22:45:15.379631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:45:15.379631Z digest=sha256:22dd66ff8d4be89a6a4134f6be081f05f2ba927e2fd054f840183c5dc4bc4db8

Observation 67cc670d-1c49-4e0f-9050-c5fdcc28d126 · inbound

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models cites this paper.

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:50:08.479194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T17:50:08.399160Z digest=sha256:7ddb18a80f228bbc02cd3ecbd30bee4bf832a9d825c53c952e18a79c2528ce73

Observation 7b73d5d6-df5b-429e-b414-072673213de2 · inbound

Entropy2Vec: Crosslingual Language Modeling Entropy as End-to-End Learnable Language Representations cites this paper.

Entropy2Vec: Crosslingual Language Modeling Entropy as End-to-End Learnable Language Representations FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T05:42:50.078953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:42:50.078953Z digest=sha256:137d28c155448733caf0f25bdf9ee47b97764b20dd118ecc87f9b44df876e8d9

Observation 6bc22003-f510-43ac-92f7-025f702bf74c · inbound

mmBERT: A Modern Multilingual Encoder with Annealed Language Learning cites this paper.

mmBERT: A Modern Multilingual Encoder with Annealed Language Learning FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T16:18:26.734265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:18:26.734265Z digest=sha256:18250201071856bd5dd2a40979691c6663b29e2f7498a148d948dc9f400cb2f7

Observation 45b04929-9776-4355-bbde-389fd154ecf0 · inbound

The Effect of Scripts and Formats on LLM Numeracy cites this paper.

The Effect of Scripts and Formats on LLM Numeracy FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T08:58:50.692729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:58:50.692729Z digest=sha256:f5babeb214bd1700efa69b543a9dd353e2919286ed680caa90e056a0b2174086

Observation e263f088-3f4e-4306-8aae-d48749d9a534 · inbound

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data cites this paper.

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 167

Resolution
unresolved
no resolver link, observed 2026-08-03T08:15:27.454474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:15:27.454474Z digest=sha256:d8f42bb58e3b0b0dca09366050dfb45c8c80c3db3085dd59c353053b17d745c4

Observation 79a8722b-9b1e-435f-9f17-7bc8c9414d93 · inbound

OpenLID-v3: Improving the Precision of Closely Related Language Identification -- An Experience Report cites this paper.

OpenLID-v3: Improving the Precision of Closely Related Language Identification -- An Experience Report FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T23:37:59.719719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T23:37:59.719719Z digest=sha256:46858e22eb2b4378d2f3fc50d41005e3dc04c1068ce11790746beabfde886057

Observation 763e30af-8044-4a58-8526-9f21ffcb16c1 · inbound

Scaling Laws for Mixture Pretraining Under Data Constraints cites this paper.

Scaling Laws for Mixture Pretraining Under Data Constraints FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:48:00.949304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-14T21:44:31.429223Z digest=sha256:b2324a20419f3bd33bb2525af76457df47ad8c7e1018324ed30c321e9ad06396

Observation 3a3d4225-0be4-4675-bc7f-70b881a3c0a0 · inbound

Scaling Laws for Mixture Pretraining Under Data Constraints cites this paper.

Scaling Laws for Mixture Pretraining Under Data Constraints FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:37:39.604606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-19T16:36:30.007014Z digest=sha256:6f02d4b5e9e580176f75de2e783f2e85699930166e9594ec4597c0b68239ea07

Observation 10a7a53c-b215-4103-8017-0c99cb34c20e · inbound

Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings cites this paper.

Mix, Don't Tune: Bilingual Pre-Training Outperforms Hyperparameter Search in Data-Constrained Settings FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:19:27.448259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T20:17:26.661595Z digest=sha256:d63f2efe97b22386969e564ccfdfc61c45d150940a07d690d45d30e524b8c6f0

Observation 009b4acd-a370-4dba-a981-20466c2442fe · inbound

Granite Embedding Multilingual R2 Models cites this paper.

Granite Embedding Multilingual R2 Models FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T18:17:34.932496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T18:16:49.148303Z digest=sha256:cd08493c8000498b0a541a06a89423882c8a789fe7d13baa96e5a06b00e2be31

Observation 9c65265e-e2be-427e-a18f-c02d71c8ba7e · inbound

A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAM$\Delta$ Integration into Upcycled MoE cites this paper.

A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAM$\Delta$ Integration into Upcycled MoE FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:13:13.616472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-20T11:09:22.027588Z digest=sha256:ccc49adbbbdd187f92215d3e484ca2e777c3682859af55ba75f3aef090940f02

Observation 55dab8fb-b24c-476f-a89e-a007b018028b · inbound

SomaliWeb v1: A Quality-Filtered Somali Web Corpus with a Matched Tokenizer and a Public Language-Identification Benchmark cites this paper.

SomaliWeb v1: A Quality-Filtered Somali Web Corpus with a Matched Tokenizer and a Public Language-Identification Benchmark FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:38:12.634240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:34:49.783942Z digest=sha256:656380542e6d776ceec945b40704649a04f074d77a6ca75e57e2fc0f2b3b3958

Observation ed65c45e-2831-430a-8d11-502aa052c297 · inbound

Multilingual Knowledge Transfer under Data Constraints via Lexical Interventions cites this paper.

Multilingual Knowledge Transfer under Data Constraints via Lexical Interventions FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:15:20.257197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-25T04:11:06.312503Z digest=sha256:3d7b8022ed32ed8b39f29a4bc681163f1338ea7d3f19ee5edc57a54d3eabadf2

Observation cf0129b7-a60b-40ca-9fb4-83adcd620143 · inbound

Mimir: Large-scale Multilingual Concept Modeling cites this paper.

Mimir: Large-scale Multilingual Concept Modeling FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T11:24:38.681356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T11:08:44.027943Z digest=sha256:a148d8c19087968f56f41fe9d5e018b71f90814657e6c39c9e8e5c16deca2eea

Observation 57f65719-6f44-406a-861c-fcf0faa9aabd · inbound

Ishigaki-IDS: An Open-Weight Verifier-Aware Model for Information Delivery Specification Drafting in Building Information Modeling cites this paper.

Ishigaki-IDS: An Open-Weight Verifier-Aware Model for Information Delivery Specification Drafting in Building Information Modeling FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 25

Resolution
malformed identifier
arxiv_id, observed 2026-07-02T23:17:29.057127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T18:24:53.592902Z digest=sha256:526303598a7b5f525e1dcbfe484e98798eeccbef74cec0bc4fa059f358bbc781

Observation df6a9db1-36aa-41d6-b236-56f1861d6a01 · inbound

Grouped Query Experts: Mixture-of-Experts on GQA Self-Attention cites this paper.

Grouped Query Experts: Mixture-of-Experts on GQA Self-Attention FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:49:29.893255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T17:40:06.036019Z digest=sha256:174cca4763b3631b82fe7b09bc0cef3ab79918fa53006a1b2f4a8fd7a6f5d4d6

Observation d2bd763f-f994-4834-9947-29e54e4eb04e · inbound

moBERTo: A Modern Encoder for Portuguese via Continued Pretraining of ModernBERT cites this paper.

moBERTo: A Modern Encoder for Portuguese via Continued Pretraining of ModernBERT FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:19:44.124861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T10:07:55.521194Z digest=sha256:8ae10dc292ca5c37fe43fe3db57f71c7b399206775c6916762c52206f8bcae91

Observation 1b5b1bdb-58fb-408b-9c29-2bf74907bb15 · inbound

LangMAP: A Language-Adaptive Approach to Tokenization cites this paper.

LangMAP: A Language-Adaptive Approach to Tokenization FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:49:46.653379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T08:26:33.185340Z digest=sha256:14f66af5b76cd86282a07200946073b3c44bdfb812296b79f1541281bd1d2749

Observation d3e45960-722e-4f56-bb55-437ffa4ab47a · inbound

SARA: Unlocking Multilingual Knowledge in Mixture-of-Experts via Semantically Anchored Routing Alignment cites this paper.

SARA: Unlocking Multilingual Knowledge in Mixture-of-Experts via Semantically Anchored Routing Alignment FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:00:09.022875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-25T20:48:04.003465Z digest=sha256:1982dde88d1fba086ecf833fcd3b37e165a36da671f7e131c6ebc3ac006dbd17

Observation 65447489-0aee-4602-976e-e5023ff38cb3 · inbound

MultiHashFormer: Hash-based Generative Language Models cites this paper.

MultiHashFormer: Hash-based Generative Language Models FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-29T04:23:05.508816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-29T04:13:05.082903Z digest=sha256:ac810f77ea735a60a2d0cc95ad05b43ec6fb5879969b0a45b8869348dcfd6051

Observation a12afaeb-34ce-473a-91f9-9ad1b79e3171 · inbound

In-Place Tokenizer Expansion for Pre-trained LLMs cites this paper.

In-Place Tokenizer Expansion for Pre-trained LLMs FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:16.706422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:16.706422Z digest=sha256:2ca962c099c4c1596b305e0ff7a874a63dad43081af2d19a13c6d5c84ead2da9

Observation ff36307a-718b-42a4-97fa-80e81435f081 · inbound

Instability of LLM Pre-Pretraining: It Doesn't Always Help. An Investigation on Multiple Languages cites this paper.

Instability of LLM Pre-Pretraining: It Doesn't Always Help. An Investigation on Multiple Languages FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:40.593917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T04:29:40.593917Z digest=sha256:7828d633e8157980f99a1e4493203cdabd63f4cfc1b08483dc4dfe6768c494f0