REVIEW 6 cited by
CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The driving factors behind the development of large language models (LLMs) with impressive learning capabilities are their colossal model sizes and extensive training datasets. Along with the progress in natural language processing, LLMs have been frequently made accessible to the public to foster deeper investigation and applications. However, when it comes to training datasets for these LLMs, especially the recent state-of-the-art models, they are often not fully disclosed. Creating training data for high-performing LLMs involves extensive cleaning and deduplication to ensure the necessary level of quality. The lack of transparency for training data has thus hampered research on attributing and addressing hallucination and bias issues in LLMs, hindering replication efforts and further advancements in the community. These challenges become even more pronounced in multilingual learning scenarios, where the available multilingual text datasets are often inadequately collected and cleaned. Consequently, there is a lack of open-source and readily usable dataset to effectively train LLMs in multiple languages. To overcome this issue, we present CulturaX, a substantial multilingual dataset with 6.3 trillion tokens in 167 languages, tailored for LLM development. Our dataset undergoes meticulous cleaning and deduplication through a rigorous pipeline of multiple stages to accomplish the best quality for model training, including language identification, URL-based filtering, metric-based cleaning, document refinement, and data deduplication. CulturaX is fully released to the public in HuggingFace to facilitate research and advancements in multilingual LLMs: https://huggingface.co/datasets/uonlp/CulturaX.
Forward citations
Cited by 6 Pith papers
-
Preperiodic points, finiteness, and structures of semigroups of algebraic morphisms
The paper proves finiteness and structural results for preperiodic points of algebraic morphisms, including Burnside-type and Northcott-type theorems.
-
Mangosteen: An Open Thai Corpus for Language Model Pretraining
An open 47B-token Thai pre-training corpus and a Thai-adapted data cleaning pipeline, with ablations showing quality gains and an 8B model that improves on Thai benchmarks.
-
TokAlign: Efficient Vocabulary Adaptation via Token Alignment
TokAlign aligns source and target BPE token vocabularies using GloVe co-occurrence embeddings and re-initializes LLM embeddings, recovering within 5k steps and enabling token-level distillation.
-
BPE Stays on SCRIPT: Structured Encoding for Robust Multilingual Pretokenization
A Unicode script and category based character encoding with constrained merging achieves compression competitive with byte-level BPE while removing the byte-premium penalty for non-Latin scripts.
-
GeistBERT: Breathing Life into German NLP
A 126M-parameter German BERT, pretrained further on 1.3TB of mixed German text, beats other base models on most tested German NLP benchmarks.
-
Synthetic Document Question Answering in Hungarian
New Hungarian document VQA datasets (HuDocVQA, HuDocVQA-manual, HuCCPDF) reveal far lower accuracy for frontier VLMs than English DocVQA, with finetuning plus OCR data recovering part of the gap.
Discussion (0). Continue with ORCID to comment.