Pith. sign in

Paper Citation Record · LEDGER

CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2309.09400.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.09400 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:59:25.835410Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:39:51.301876Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ca0e8629-1ad9-4c67-808a-1fe0e87f714d · inbound

Yi: Open Foundation Models by 01.AI cites this paper.

Yi: Open Foundation Models by 01.AI CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:47:27.886303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T05:47:27.775529Z digest=sha256:7174c40bfe9adcc5e43b34db67414cd2f7450656a011d49933b9ba0c9113ee6d

Observation 1d903287-defd-46c8-b5b5-b8873a81dfec · inbound

Synthetic Document Question Answering in Hungarian cites this paper.

Synthetic Document Question Answering in Hungarian CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:25.835410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:25.835410Z digest=sha256:2607245bf7d85ddce7a74aaa5790ca2c8bb529ab99ebf21688a079f500c9b5ac

Observation 8503dc19-4009-46da-8a76-2c49283237d5 · inbound

BPE Stays on SCRIPT: Structured Encoding for Robust Multilingual Pretokenization cites this paper.

BPE Stays on SCRIPT: Structured Encoding for Robust Multilingual Pretokenization CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:31.714800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:31.714800Z digest=sha256:03c85e81164d387640e5c9d0bd94c162eea08a8ebd0f6af7fa281f335c83712c

Observation d26e90d6-4e49-4322-9b63-b76a565e415b · inbound

TokAlign: Efficient Vocabulary Adaptation via Token Alignment cites this paper.

TokAlign: Efficient Vocabulary Adaptation via Token Alignment CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:06:28.059185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:06:28.059185Z digest=sha256:93dab3a4457136ad90ea63feaa198dd3bb842f05e24aeda8b46f328d03b790e6

Observation 53bcebbb-a4be-48b0-a779-c31f70122194 · inbound

GeistBERT: Breathing Life into German NLP cites this paper.

GeistBERT: Breathing Life into German NLP CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T01:08:59.129461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:08:59.129461Z digest=sha256:27ff338e9d7115aa3c44eb4cee5dc44632ee43d6be362d1ca33dc37b76691116

Observation 8a241d9d-2f6e-4ed5-aec9-41ee2fe130fa · inbound

Mangosteen: An Open Thai Corpus for Language Model Pretraining cites this paper.

Mangosteen: An Open Thai Corpus for Language Model Pretraining CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:56:23.835272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:56:23.835272Z digest=sha256:75c208eef93d885432cacd1ad1fd9380e98047aaff1629374aba8eaa63ead090

Observation e2d4e365-3e12-4033-8bb8-f0b0d237ee85 · inbound

Preperiodic points, finiteness, and structures of semigroups of algebraic morphisms cites this paper.

Preperiodic points, finiteness, and structures of semigroups of algebraic morphisms CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T21:15:37.405672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:15:37.405672Z digest=sha256:890bcd2d0ba5ac9b84513beb545ac0dd08ca3138234d995b60055efe2556ff30

Observation ab296b74-5501-4841-819f-f5cf2a67d0c3 · inbound

Bridging Linguistic Gaps: Cross-Lingual Mapping in Pre-Training and Dataset for Enhanced Multilingual LLM Performance cites this paper.

Bridging Linguistic Gaps: Cross-Lingual Mapping in Pre-Training and Dataset for Enhanced Multilingual LLM Performance CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:16:09.169879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:02:54.374120Z digest=sha256:852b4c1676cb73a4e8ce62e21e315016e4daf3b419d606d52d8e83ab05a15a4b

Observation a8873b35-05f4-4f30-9077-b5adf23bd945 · inbound

TokAlign++: Advancing Vocabulary Adaptation via Better Token Alignment cites this paper.

TokAlign++: Advancing Vocabulary Adaptation via Better Token Alignment CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:37:52.257765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T19:35:10.026293Z digest=sha256:abb8f40371ceb7a995f1e732fd4d307653e1ca4bd2e9cefda93f1f6b28d25715

Observation 97dccc55-727c-44b0-ae83-fb2624bdc845 · inbound

TFGN: Task-Free, Replay-Free Continual Pre-Training Without Catastrophic Forgetting at LLM Scale cites this paper.

TFGN: Task-Free, Replay-Free Continual Pre-Training Without Catastrophic Forgetting at LLM Scale CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-19T17:07:41.547085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T17:06:12.280460Z digest=sha256:8ce53cdb10345ce30002773d14965010045cb3a673f50ceccdb43d85f6ccb959

Observation f76e6b07-e605-426c-8c2b-353fc618c1fe · inbound

A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAM$\Delta$ Integration into Upcycled MoE cites this paper.

A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAM$\Delta$ Integration into Upcycled MoE CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:13:13.613153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T11:09:22.027588Z digest=sha256:7ead2e080fb3957974f909be16a68c13756fcaab5e84942c984fb48d0c3fb267

Observation 255e4fe2-17bc-4f35-9f87-869f4ac4fa87 · inbound

The Tokenizer Tax Across 25 European Languages: Domain Invariance, Cross-Lingual Few-Shot Effects, and the Ukrainian Penalty cites this paper.

The Tokenizer Tax Across 25 European Languages: Domain Invariance, Cross-Lingual Few-Shot Effects, and the Ukrainian Penalty CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:14:40.882589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T13:08:33.172423Z digest=sha256:7c0f6ffde5d7b165040e4e0d1b0a3dd0952ded156faa470f419f1335d2dbc731

Observation 3d6325ca-f10c-4b6b-85a8-1563608b8c54 · inbound

Phonemes to the Rescue: Multilingual Tokenization Based on International Phonetic Alphabet cites this paper.

Phonemes to the Rescue: Multilingual Tokenization Based on International Phonetic Alphabet CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:39:34.631117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T16:53:25.556222Z digest=sha256:30d8910d47e6dfb2498650079fef9021da4b107cc4a00488e267720f41a3c3ff

Observation 945dcc70-a317-4826-b16e-5b7992a4ca69 · inbound

MinGram: A Minimalist Unigram Tokenizer with High Compression and Competitive Morphological Alignment cites this paper.

MinGram: A Minimalist Unigram Tokenizer with High Compression and Competitive Morphological Alignment CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:39:51.304281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T04:59:05.057552Z digest=sha256:d7605200e12326766b44dca3a5bcba53d3988f6f00f1f0b866060dd280f02f61