Pith. sign in

REVIEW 6 cited by

CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.09400 v1 pith:IV37MAHE submitted 2023-09-17 cs.CL cs.AI

classification cs.CLcs.AI
keywords llmsmultilingualtrainingculturaxdatasetdatasetslanguagecleaning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The driving factors behind the development of large language models (LLMs) with impressive learning capabilities are their colossal model sizes and extensive training datasets. Along with the progress in natural language processing, LLMs have been frequently made accessible to the public to foster deeper investigation and applications. However, when it comes to training datasets for these LLMs, especially the recent state-of-the-art models, they are often not fully disclosed. Creating training data for high-performing LLMs involves extensive cleaning and deduplication to ensure the necessary level of quality. The lack of transparency for training data has thus hampered research on attributing and addressing hallucination and bias issues in LLMs, hindering replication efforts and further advancements in the community. These challenges become even more pronounced in multilingual learning scenarios, where the available multilingual text datasets are often inadequately collected and cleaned. Consequently, there is a lack of open-source and readily usable dataset to effectively train LLMs in multiple languages. To overcome this issue, we present CulturaX, a substantial multilingual dataset with 6.3 trillion tokens in 167 languages, tailored for LLM development. Our dataset undergoes meticulous cleaning and deduplication through a rigorous pipeline of multiple stages to accomplish the best quality for model training, including language identification, URL-based filtering, metric-based cleaning, document refinement, and data deduplication. CulturaX is fully released to the public in HuggingFace to facilitate research and advancements in multilingual LLMs: https://huggingface.co/datasets/uonlp/CulturaX.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Preperiodic points, finiteness, and structures of semigroups of algebraic morphisms

    math.NT 2025-08 unverdicted novelty 6.0 of 10

    The paper proves finiteness and structural results for preperiodic points of algebraic morphisms, including Burnside-type and Northcott-type theorems.

  2. Mangosteen: An Open Thai Corpus for Language Model Pretraining

    cs.CL 2025-07 conditional novelty 6.0 of 10

    An open 47B-token Thai pre-training corpus and a Thai-adapted data cleaning pipeline, with ablations showing quality gains and an 8B model that improves on Thai benchmarks.

  3. TokAlign: Efficient Vocabulary Adaptation via Token Alignment

    cs.CL 2025-06 conditional novelty 6.0 of 10

    TokAlign aligns source and target BPE token vocabularies using GloVe co-occurrence embeddings and re-initializes LLM embeddings, recovering within 5k steps and enabling token-level distillation.

  4. BPE Stays on SCRIPT: Structured Encoding for Robust Multilingual Pretokenization

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A Unicode script and category based character encoding with constrained merging achieves compression competitive with byte-level BPE while removing the byte-premium penalty for non-Latin scripts.

  5. GeistBERT: Breathing Life into German NLP

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A 126M-parameter German BERT, pretrained further on 1.3TB of mixed German text, beats other base models on most tested German NLP benchmarks.

  6. Synthetic Document Question Answering in Hungarian

    cs.CV 2025-05 conditional novelty 4.0 of 10

    New Hungarian document VQA datasets (HuDocVQA, HuDocVQA-manual, HuCCPDF) reveal far lower accuracy for frontier VLMs than English DocVQA, with finetuning plus OCR data recovering part of the gap.

Pith tools