Pith. sign in

REVIEW 3 major objections 7 minor 7 cited by

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language

T0 review · 3 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read One adaptive data-curation pipeline, whose filters and thresholds are set automatically from each language's own text statistics, produces pre-training corpora that outperform prior multilingual datasets on 11 of 14 languages tested.

desk verdict Strong engineering contribution with a solid internal ablation; the headline 11/14 claim needs an epoch-matched re-run before it can be taken at face value. read the letter →

arxiv 2506.20920 v1 pith:R3TB4SIM submitted 2025-06-26 cs.CL

classification cs.CL
keywords multilingualpre-trainingdatacurationpipelinelanguageidentificationdeduplicationrehydrationCommonCrawllow-resourcelanguagesevaluationtaskselection
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that the painstaking, per-language work of building pre-training data — language filters, confidence thresholds, deduplication settings, quality cutoffs — can be replaced with a single pipeline that measures each language's own text statistics and adapts automatically. The design is validated by more than two hundred ablation models trained on nine "canary" languages (Arabic, Chinese, French, Hindi, Russian, Swahili, Telugu, Thai, Turkish) chosen to span families, scripts, and resource levels, and evaluated on tasks selected by quantitative criteria so that early-training scores are meaningful. Models trained on the resulting corpora outperform those trained on prior multilingual datasets (CC-100, mC4, CulturaX, HPLT, HPLT2) on 11 of 14 evaluated languages, including five held-out languages (German, Indonesian, Italian, Japanese, Vietnamese) that played no role in design decisions; a final "rehydration" step upsamples documents by duplication count and quality, adding a further measured gain. If the claim is right, the released 20-terabyte, 5-billion-document FineWeb2 corpus — 1,868 language-script pairs drawn from 96 Common Crawl snapshots — gives more than a thousand languages the kind of carefully filtered pre-training data that English already had, without a hand-crafted pipeline per language.

What carries the argument

The carrying object is the FineWeb2 pipeline itself: a staged sequence in which every threshold is a function of statistics measured on the target language rather than a global constant. Its stages are GlotLID language identification with a per-language confidence cutoff; per-language MinHash deduplication using language-appropriate word tokenizers; heuristic filtering whose thresholds are adapted from FineWeb's English values to each language by distribution-matching (the 10Tail, Quantile, and MeanStd methods); a high-affinity wordlist filter for low-resource languages; and rehydration, which upsamples kept documents by duplication count with weights derived from filtering rates per cluster size as a quality proxy. What makes this a single mechanism for all languages is the claim that every adaptation is computed automatically from the language's own data, so no human tuning is needed for any of the 1,868 language-script pairs in the final dataset.

What would settle it

Two concrete checks would settle the claim. First, take a language outside the nine canary languages with enough data and a reliable benchmark (say, Korean, Ukrainian, or Persian), train identical small models on corpora kept at the formula-selected LID cutoff and at neighbouring cutoffs, and map the performance curve: each language where the formula's cutoff falls outside the highest-scoring range — as the paper's Table 15 already shows for Chinese and Hindi — weakens the claim of automatic adaptation to every language. Second, because the paper's evaluations rely on early-signal tasks that may change character with longer training, re-run a subset of the 14-language comparison at larger token budgets and check whether FineWeb2's 11-of-14 advantage persists.

Watch

Extended reading notes

Core claim

The paper's central claim is that a consistent, adaptive pipeline can replace hand-designed per-language processing: every stage derives its settings from the target language's own data instead of using fixed values for all languages. Language identification uses GlotLID, which covers 1,880 languages and labels script variants separately, and each language's confidence cutoff is set by the formula $\max\{0.3, \min\{0.9, \mathrm{Med}(X)-\sigma(X)\}\}$ on the distribution $X$ of that language's LID scores. Word tokenization, needed for filtering, deduplication, and evaluation, comes from native tokenizers where they exist and from proxy tokenizers propagated through the language-family tree otherwise. Documents are deduplicated globally per language with MinHash, and heuristic filters inherited from the English FineWeb pipeline have their thresholds re-derived by distribution-matching methods (10Tail, Quantile, MeanStd) so each language removes a comparable, quality-relevant fraction of data; low-resource languages additionally pass a high-affinity wordlist filter to strip misclassified high-resource content. The final stage, rehydration, uses filtering rates per MinHash cluster size as a quality proxy and upsamples documents with cluster-dependent weights, favouring neither unique nor massively duplicated documents. On this basis the paper concludes that models trained on FineWeb2 corpora are more performant than models trained on prior multilingual datasets on 11 of 14 evaluated languages, with each pipeline stage contributing a measured improvement, and that the same pipeline scales to 1,868 language-script pairs in the released 20-terabyte corpus.

Load-bearing premise

The load-bearing premise is that the automatic cutoff formula for language identification, which decides how much low-confidence text each language keeps, was validated on only nine languages — and on two of those nine (Chinese and Hindi) it already selects a cutoff outside the best-performing range (Section 4.2, Table 15) — yet it is applied without per-language checks to the other roughly 1,859 language-script pairs, so any gain or loss for those languages rests on the assumption that their confidence-score distributions resemble the seven languages where the formula worked.

Editorial extensions

If this is right

  • Teams training monolingual or multilingual LLMs can use the released pipeline and corpus directly, replacing hand-built per-language workflows with a single runnable recipe that matches or exceeds the early-signal performance of prior public multilingual web corpora.
  • The rehydration scheme turns an otherwise destructive deduplication step into a tunable one: because filter-based quality weights are cheap, any future deduplicated dataset that retains cluster-size metadata can apply the same upsampling trick.
  • The task-selection criteria (monotonicity, signal-to-noise ratio, non-randomness, ordering consistency) give a reusable, quantitative way to pick evaluation benchmarks for languages that lack validated ones, reducing reliance on noisy translated tasks for early-training comparisons.
  • For low-resource languages, the released corpora enable continued pre-training experiments even though most small language corpora are dominated by Bible or Wikipedia content; the paper's own audit finds that 70% of language-script pairs have over half their documents from those domains.
  • Combining FineWeb2 with FineWeb yields a single open, English-plus-everything-else pre-training mixture at 20TB scale, giving non-English languages the same kind of carefully filtered web data that English already had.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the adaptive thresholds generalize as claimed, the pipeline's adaptation logic could be rerun on each new Common Crawl snapshot release as a scheduled, fully automatic operation, making multilingual dataset refresh a maintenance task rather than a research project — the paper does not claim this, but it follows directly from the design.
  • The formula's known misses on Chinese and Hindi (Table 15 of the paper) open a measurable improvement path the authors did not take: adding a validation loop on more languages, or making the cutoff depend on distribution shape such as skewness, could recover the lost performance on those languages and reduce risk on the unvalidated ones.
  • The dominance of Bible and Wikipedia sources in the long tail implies that for very-low-resource languages the binding constraint is source diversity, not pipeline quality; a testable extension is to measure whether filtering and rehydration gains saturate for corpora under a few thousand documents, in which case augmenting with curated or synthetic data would matter more than further pipeline tun
  • The early-signal selection metrics could double as diagnostic audits of existing multilingual benchmarks: tasks that fail monotonicity or signal-to-noise thresholds on reference runs could be flagged and retired, which would change how multilingual model comparisons are reported.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. The paper introduces FineWeb2, a multilingual pre-training data curation pipeline that adapts FineWeb-style filtering, deduplication, and rehydration to over 1,000 languages. The pipeline is validated on nine canary languages with 207 small-model ablation runs and compared against CC-100, mC4, CulturaX, HPLT, HPLT2, raw Common Crawl, and several language-specific corpora on both the canary languages and five unseen languages. The authors report that FineWeb2 outperforms prior multilingual datasets on 11 of 14 languages, and release the 20TB dataset, the pipeline code, and training/evaluation codebases.

Significance. If the central claims hold, this is a substantial contribution to multilingual pre-training infrastructure: a released 20TB corpus, an automatically adaptable pipeline, and a rigorous ablation methodology with a novel, quantitatively motivated task-selection procedure. The unseen-language evaluation (de, id, it, ja, vi) provides genuinely independent evidence that the pipeline generalizes beyond the languages used for design decisions. The paper also contains useful transparency, including a limitation discussion and a release of the pre-filtering intermediate dataset. However, the headline 11/14 superiority claim is currently weakened by two load-bearing issues: epoch-count confounds in low-resource language comparisons and an unvalidated LID-threshold formula applied to nearly 1,900 language-script pairs.

major comments (3)
  1. [§3.2, §5, Figure 7, Table 45] The Telugu and Swahili comparisons are confounded by different numbers of epochs. Section 3.2 states that for Telugu and Swahili every baseline dataset except raw Common Crawl contained only a limited amount of data, while Table 45 shows FineWeb2 contains 14.4GB for Telugu and 3.1GB for Swahili. All models are trained for 29–30B tokens, so the CC-100/mC4/CulturaX/HPLT/HPLT2 models must pass over their unique tokens many more times than the FineWeb2 models. Since repeated exposure to a small corpus can change early-signal benchmark behavior, the FineWeb2 wins on these two languages cannot be cleanly attributed to corpus curation quality. Removing or controlling those comparisons reduces the reported 11/14 to 9/14, which is materially weaker. I request an epoch-matched re-run or an explicit analysis showing that the epoch-count difference does not explain the observed gains.
  2. [§4.2, Table 15] The LID threshold formula max{0.3, min{0.9, Med(X)−σ(X)}} is applied to all 1,868 language-script pairs, but Table 15 shows that on two of the nine canary languages (Chinese and Hindi) this formula selects a threshold outside the highest-performing range. The paper provides no validation of this formula on the remaining roughly 1,859 languages, so any quality gain or loss for those languages rests on an untested assumption that confidence-score distributions behave like the seven languages where the formula worked. Please either validate the formula on a held-out set of languages with downstream evaluations, provide a sensitivity analysis showing the choice is not load-bearing, or clearly qualify the pipeline's per-language claims.
  3. [§5, Figure 3, §A.10.1] The headline claim that FineWeb2 produces more performant models on 11 of 14 languages lacks uncertainty quantification. Each dataset/language comparison appears to be a single training run, and the approximate standard deviations reported for unseen-language tasks in Section A.10.1 are large enough that some of the aggregate differences in Figure 3 may be within noise. I request bootstrap confidence intervals across tasks or multiple training seeds for at least the key comparisons, or alternatively a softened statement that does not assert superiority on a per-language basis without such evidence.
minor comments (7)
  1. [§5] Typo: 'multilingal' should be 'multilingual' in the sentence 'FineWeb2 produces more performant models than prior multilingal datasets'.
  2. [§4.2] The threshold formula uses X without a formal definition; please state explicitly that X is the random variable of GlotLID confidence scores for documents initially classified as the target language.
  3. [Figure 7 caption] The caption says all models were trained for 30 billion tokens, while Section 5 and Table 4 say 29 billion tokens; please align these numbers.
  4. [§A.3] Typo: 'paramaters' should be 'parameters' in the sentence about embedding layer size.
  5. [Table 15] The notation '0.8 - 0.9' for the Russian range is ambiguous; please use an explicit interval such as [0.8, 0.9] or clarify that it means all tested thresholds in that span.
  6. [§3.3 and §A.5.3] The paper states that 84 benchmarks were selected out of 197 tested, but the per-language tables in A.5.3 appear to list a smaller number; please add a summary table or count verification.
  7. [§4.5] The rehydration weight assignment is described as based on filtering rates with interpolation between two endpoints, but the exact interpolation formula is not given; please provide the formula or a precise algorithmic description.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the pipeline is validated out-of-sample on unseen languages; canary tuning is transparently in-sample and the central generalization claim does not reduce to the tuning inputs.

full rationale

The paper's derivation chain is empirical rather than definitional: it selects pipeline components (GlotLID, per-language LID thresholds, filter adaptation methods, deduplication, rehydration weights) using ablations on nine canary languages, then evaluates the resulting corpora both on those canary languages and on five unseen languages (German, Indonesian, Italian, Japanese, Vietnamese) trained for 100B tokens against external baselines (CC-100, mC4, CulturaX, HPLT, HPLT2, raw Common Crawl). No equation in the paper defines the claimed output in terms of the input: the LID threshold formula max{0.3, min{0.9, Med(X)-sigma(X)}} is a heuristic fitted to canary score distributions and applied to 1,868 language-script pairs; its failure for Chinese and Hindi (Table 15) shows it is not forced by construction. The 11/14 headline includes eight canary-language wins, and those are in-sample because the canary languages were used for design decisions; however, the paper explicitly does not present canary results as the generalization test ('evaluating on unseen languages validates that the pipeline generalizes effectively'), and the unseen-language comparisons provide genuine out-of-sample support, with FineWeb2 winning on Italian, Japanese, and Vietnamese at 100B tokens. Self-citations to FineWeb (Penedo et al., 2024) provide the starting pipeline and MinHash hyperparameters, but the load-bearing filtering, deduplication, and rehydration steps are ablated in this paper, so the citations are not carrying the central claim. The epoch-count imbalance for Telugu and Swahili baselines is a confound for those two comparisons, but it is a correctness/robustness concern, not a circularity. Overall, no prediction reduces by construction to its inputs; score 1 reflects only the mild in-sample component of the composite headline.

Assumptions & free parameters 6 free parameters · 6 assumptions · 0 invented entities

The central claim rests on several fitted design choices (LID threshold formula, filter adaptation methods, rehydration weights, task selection thresholds, affinity gamma), all selected using the nine canary languages. No new physical or mathematical entities are introduced; the pipeline reuses established tools such as GlotLID, MinHash, and trafilatura.

free parameters (6)
  • LID threshold formula = Median - sigma, clipped to [0.3, 0.9]; e.g., Arabic 0.8812, Chinese 0.7415, Swahili 0.3
    Selected because it falls in the best threshold range for 7 of 9 canary languages (Table 15), then applied to all 1,868 language-script pairs.
  • Rehydration upsampling weights = Weight 10 at lowest removal-rate cluster size, weight 1 above global removal rate, linear interpolation between
    Derived from filtering-rate proxy for cluster quality (Section 4.5); validated only on canary languages.
  • Filter adaptation method per group = fwq: 10Tail on Wikipedia; goq: Quantile on Wikipedia/GlotLID-Corpus; gor: MeanStd on Common Crawl
    Chosen by average rank across canary languages (Table 25).
  • Affinity threshold gamma for precision wordlists = 0.85
    Hand-chosen threshold for high-affinity tokens used in low-resource contamination filtering (Section A.7.3).
  • Task selection thresholds = Monotonicity rho >= 0.5, SNR >= 20 (30 for generative), non-randomness >= 3
    Criteria chosen to filter 197 candidate tasks to 84; no external justification is given for these cutoffs (Section A.5.1).
  • Stopword frequency threshold = Adjusted so at least 8 stopwords remain after removing symbols and numbers
    Defined as words exceeding a frequency threshold in reference corpora; varies by language (Section A.7.1).
assumptions (6)
  • domain assumption Early-signal evaluation validity: quality at 29B or 100B tokens on 1.46B-parameter models predicts relative data quality at larger pre-training scales.
    Invoked in Section 3; the entire ablation methodology depends on this proxy.
  • domain assumption GlotLID V3 language labels and confidence scores are accurate enough across roughly 1,880 languages for the downstream pipeline.
    The pipeline's LID stage uses GlotLID, including its noise and undefined labels (Section 4.2).
  • ad hoc to paper The LID threshold formula max{0.3, min{0.9, Med(X)-sigma(X)}} generalizes beyond the nine canary languages.
    Applied to all 1,868 language-script pairs even though it misses the best threshold range for Chinese and Hindi (Table 15).
  • ad hoc to paper Filtering removal rate is a valid proxy for duplicate-cluster quality for rehydration weighting.
    Section 4.5 bases upsampling weights on this proxy; no independent validation outside canary languages is provided.
  • domain assumption Proxy word tokenizers assigned via Ethnologue language family propagate adequately for filtering, deduplication, and evaluation.
    Section A.1 assigns tokenizers from the closest language in the family tree to languages without native tokenizers.
  • domain assumption MinHash hyperparameters from FineWeb (14 buckets of size 8, 5-grams) transfer to all languages.
    Reused without re-tuning in Section 4.3; the transfer is not tested across scripts and tokenizers.

how reviews work

0 comments
Cite this review

Pith. "Pith review of FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language." pith.science (2026). https://pith.science/paper/R3TB4SIM

@misc{pith2026250620920,
  author       = {Pith},
  title        = {Pith review of: FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/R3TB4SIM}},
  note         = {Machine review of arXiv:2506.20920}
}
read the original abstract

Pre-training state-of-the-art large language models (LLMs) requires vast amounts of clean and diverse text data. While the open development of large high-quality English pre-training datasets has seen substantial recent progress, training performant multilingual LLMs remains a challenge, in large part due to the inherent difficulty of tailoring filtering and deduplication pipelines to a large number of languages. In this work, we introduce a new pre-training dataset curation pipeline based on FineWeb that can be automatically adapted to support any language. We extensively ablate our pipeline design choices on a set of nine diverse languages, guided by a set of meaningful and informative evaluation tasks that were chosen through a novel selection process based on measurable criteria. Ultimately, we show that our pipeline can be used to create non-English corpora that produce more performant models than prior datasets. We additionally introduce a straightforward and principled approach to rebalance datasets that takes into consideration both duplication count and quality, providing an additional performance uplift. Finally, we scale our pipeline to over 1000 languages using almost 100 Common Crawl snapshots to produce FineWeb2, a new 20 terabyte (5 billion document) multilingual dataset which we release along with our pipeline, training, and evaluation codebases.

Figures

Figures reproduced from arXiv: 2506.20920 by the authors.

Figure 1
Figure 1. The FineWeb2 pipeline: Evalua￾tion results of models trained on 350 billion to￾kens show that each pipeline step – Language Identification (LID), Deduplication (Dedup), Filtering, and Dedup-informed upsampling (Rehydration) – improves performance. One of the main drivers of the improving ca￾pabilities of large language models (LLMs) is increased scale, in terms of both model and pre-training dataset size. To satiate… view at source ↗
Figure 2
Figure 2. Filtering rates by MinHash cluster size for French documents. The global filtering rate represents the overall percentage of documents removed during the full filtering process. Individual filtering rates are shown for each cluster size, providing a proxy for cluster quality—higher removal rates may indicate lower-quality clusters. We assign upsampling weights to each cluster size based on the filtering rates. The f… view at source ↗
Figure 3
Figure 3. High-level performance comparison of FineWeb2 to other multilingual and [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Example tokenizer assignments based on language family data in Indo-European. [PITH_FULL_IMAGE:figures/full_fig_p026_4.png]
Figure 5
Figure 5. Figure 5: FT176 vs GlotLID without any threshold filtering applied to either classifier. While GlotLID seems to outperform in higher resource languages, FT176 performs slightly better on lower resource languages. However, GlotLID supports a considerably larger number of (lower-r…
Figure 6
Figure 6. Figure 6: Contamination scores for 1,900 languages via wordlist filtering. The plot indicates that the majority of the languages have their data in-language (non-contaminated). URL Matched Words http://www.supersport.com/football/nigeria-naija/news/121221/ Uefa don ban Malaga ni…
Figure 7
Figure 7. Figure 7: Per language comparison of FineWeb2 to other multilingual and language-specific datasets. All models were trained for 30 billion tokens. The plots have sliding window smoothing of size 3. 47 [PITH_FULL_IMAGE:figures/full_fig_p047_7.png]
Figure 8
Figure 8. Figure 8: Language composition of FineWeb2 Distribution of languages in the final FineWeb2 dataset. Percentages refer to total utf-8 bytes of each language or language family. 53 [PITH_FULL_IMAGE:figures/full_fig_p053_8.png]
Figure 9
Figure 9. Figure 9: Ratio of Wikipedia and Bible content per language Most languages have a small fraction of their content originating from Wikipedia (with some exceptions). Bible content, on the other hand, is a big part of the corpora of many lower-resource languages. 58 [PITH_FULL_IM…

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. In-Place Tokenizer Expansion for Pre-trained LLMs

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Continuing a model's own BPE merges and training only new embedding rows preserves quality while cutting token counts 2.4–4× for previously under-tokenized languages.

  2. The Effect of Scripts and Formats on LLM Numeracy

    cs.CL 2026-01 conditional novelty 6.0 of 10

    LLM arithmetic accuracy falls sharply when numerals leave the familiar Hindu–Arabic format, and few-shot prompting with examples narrows most of that gap.

  3. From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    KMMLU-Redux and KMMLU-Pro are new Korean benchmark datasets from national technical and professional licensure exams, with LLM evaluations reported against official pass thresholds.

  4. OpenLID-v3: Improving the Precision of Closely Related Language Identification -- An Experience Report

    cs.CL 2026-02 conditional novelty 5.0 of 10

    OpenLID-v3 matches or improves precision on closely related language identification by adding data, merging variants, and adding a noise class, while ensembling with GlotLID buys more precision at the cost of coverage.

  5. A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data

    cs.AI 2026-01 conditional novelty 5.0 of 10

    A metric-oriented survey that classifies intrinsic quality and trustworthiness metrics for LLM-generated data across six modalities and documents systematic evaluation gaps in the current literature.

  6. Entropy2Vec: Crosslingual Language Modeling Entropy as End-to-End Learnable Language Representations

    cs.CL 2025-09 conditional novelty 5.0 of 10

    Entropy2Vec turns the cross-lingual surprise of monolingual language models into dense language embeddings that resemble typological families and match curated vectors in downstream tasks.

  7. Observation of momentum dependent charge density wave gap in EuTe4

    cond-mat.mes-hall 2025-08 unverdicted novelty 4.0 of 10

    EuTe4 shows a momentum-dependent charge density wave gap at the Fermi level, largest along Gamma-Y and smallest along Gamma-X, plus a low-temperature magnetic phase diagram near TN = 6.9 K.

Reference graph

Works this paper leans on

118 extracted references · 16 canonical work pages · cited by 7 Pith papers

  1. [1]

    Yi: Open foundation models by 01.ai, 2025

    01.AI , Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Guoyin Wang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, Kaidong Yu, Peng Liu, Qiang Liu, Shawn Yue, Senbin Yang, Shiming Yang, Wen Xie, Wenhao Huang, Xiaohui Hu, Xiaoyi Ren, Xinyao Niu, Pengcheng Nie, Yanpeng Li, Yuchi Xu, Yudong Liu, Yue Wang, Yuxuan Cai, Zhenyu Gu, Zh...

  2. [2]

    Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, Alon Benhaim, Misha Bilenko, Johan Bjorck, Sébastien Bubeck, Martin Cai, Qin Cai, Vishrav Chaudhary, Dong Chen, Dongdong Chen, Weizhu Chen, Yen-Chun Chen, Yi-Ling Chen, Hao Cheng, Parul Chopra, Xiyang Dai, Matt...

  3. [3]

    Llama 3 model card

    AI@Meta. Llama 3 model card. 2024. URL https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md

  4. [4]

    A survey on data selection for language models

    Alon Albalak, Yanai Elazar, Sang Michael Xie, Shayne Longpre, Nathan Lambert, Xinyi Wang, Niklas Muennighoff, Bairu Hou, Liangming Pan, Haewon Jeong, et al. A survey on data selection for language models. Transactions on Machine Learning Research, 2024

  5. [5]

    Open llm turkish leaderboard v0.2

    Mohamad Alhajar. Open llm turkish leaderboard v0.2. https://huggingface.co/spaces/malhajar/OpenLLMTurkishLeaderboard, 2024

  6. [6]

    A l G hafa evaluation benchmark for A rabic language models

    Ebtesam Almazrouei, Ruxandra Cojocaru, Michele Baldo, Quentin Malartic, Hamza Alobeidli, Daniele Mazzotta, Guilherme Penedo, Giulia Campesan, Mugariya Farooq, Maitha Alhammadi, Julien Launay, and Badreddine Noune. A l G hafa evaluation benchmark for A rabic language models. In Hassan Sawaf, Samhaa El-Beltagy, Wajdi Zaghouani, Walid Magdy, Ahmed Abdelali, ...

  7. [7]

    101 billion arabic words dataset, 2024

    Manel Aloui, Hasna Chouikhi, Ghaith Chaabane, Haithem Kchaou, and Chehir Dhaouadi. 101 billion arabic words dataset, 2024

  8. [8]

    On the cross-lingual transferability of monolingual representations

    Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. On the cross-lingual transferability of monolingual representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, 2020 a . doi:10.18653/v1/2020.acl-main.421. URL http://dx.doi.org/10.18653/v1/2020.acl-main.421

Show all 118 references
  1. [9]

    A call for more rigor in unsupervised cross-lingual learning

    Mikel Artetxe, Sebastian Ruder, Dani Yogatama, Gorka Labaka, and Eneko Agirre. A call for more rigor in unsupervised cross-lingual learning. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp.\ 7375–7388. Association for Computationa...

  2. [10]

    Japanese massive multitask language understanding benchmark, 2023

    Kawahara Lab at Waseda University. Japanese massive multitask language understanding benchmark, 2023. URL https://huggingface.co/datasets/nlp-waseda/JMMLU

  3. [11]

    The belebele benchmark: a parallel reading comprehension dataset in 122 language variants

    Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, and Madian Khabsa. The belebele benchmark: a parallel reading comprehension dataset in 122 language variants. In Proceedings of ...

  4. [12]

    Building machine translation systems for the next thousand languages, 2022

    Ankur Bapna, Isaac Caswell, Julia Kreutzer, Orhan Firat, Daan van Esch, Aditya Siddhant, Mengmeng Niu, Pallavi Baljekar, Xavier Garcia, Wolfgang Macherey, Theresa Breiner, Vera Axelrod, Jason Riesa, Yuan Cao, Mia Xu Chen, Klaus Macherey, Maxim Krikun, Pidong Wang, Alexander Gu...

  5. [13]

    Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction

    Adrien Barbaresi. Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction . In Proceedings of the Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Nat...

  6. [14]

    On the resemblance and containment of documents

    Andrei Z Broder. On the resemblance and containment of documents. In Proceedings. Compression and Complexity of SEQUENCES 1997 (Cat. No. 97TB100171), pp.\ 21--29. IEEE, 1997

  7. [15]

    An open dataset and model for language identification

    Laurie Burchell, Alexandra Birch, Nikolay Bogoychev, and Kenneth Heafield. An open dataset and model for language identification. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguist...

  8. [16]

    An expanded massive multilingual dataset for high-performance language technologies, 2025

    Laurie Burchell, Ona de Gibert, Nikolay Arefyev, Mikko Aulamo, Marta Bañón, Pinzhen Chen, Mariia Fedorova, Liane Guillou, Barry Haddow, Jan Hajič, Jindřich Helcl, Erik Henriksson, Mateusz Klimaszewski, Ville Komulainen, Andrey Kutuzov, Joona Kytöniemi, Veronika Laippala, Pette...

  9. [17]

    PTT5 : Pretraining and validating the t5 model on brazilian portuguese data

    Diedre Carmo, Marcos Piau, Israel Campiotti, Rodrigo Nogueira, and Roberto Lotufo. PTT5 : Pretraining and validating the t5 model on brazilian portuguese data. arXiv preprint arXiv:2008.09144, 2020

  10. [18]

    Language ID in the wild: Unexpected challenges on the path to a thousand-language web text corpus

    Isaac Caswell, Theresa Breiner, Daan van Esch, and Ankur Bapna. Language ID in the wild: Unexpected challenges on the path to a thousand-language web text corpus. In Donia Scott, Nuria Bel, and Chengqing Zong (eds.), Proceedings of the 28th International Conference on Computat...

  11. [19]

    Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki

    Jonathan H. Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. TyDi QA : A benchmark for information-seeking question answering in typologically diverse languages. Transactions of the Association for Computational Ling...

  12. [20]

    Command r+

    Cohere. Command r+. Web, 2024. URL https://docs.cohere.com/docs/command-r-plus#model-details

  13. [21]

    Unsupervised cross-lingual representation learning at scale

    Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzm \'a n, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. Unsupervised cross-lingual representation learning at scale. In Dan Jurafsky, Joyce Chai, Natalie Schlu...

  14. [22]

    Neural learning for question answering in italian

    Danilo Croce, Alexandra Zelenanska, and Roberto Basili. Neural learning for question answering in italian. In Chiara Ghidini, Bernardo Magnini, Andrea Passerini, and Paolo Traverso (eds.), AI*IA 2018 -- Advances in Artificial Intelligence, pp.\ 389--402, Cham, 2018. Springer I...

  15. [23]

    Dataset for the first evaluation on chinese machine reading comprehension, 2018

    Yiming Cui, Ting Liu, Zhipeng Chen, Wentao Ma, Shijin Wang, and Guoping Hu. Dataset for the first evaluation on chinese machine reading comprehension, 2018. URL https://arxiv.org/abs/1709.08299

  16. [24]

    Daniels and William Bright (eds.)

    Peter T. Daniels and William Bright (eds.). The World's Writing Systems. Oxford University Press, New York, 1996

  17. [26]

    A new massive multilingual dataset for high-performance language technologies, 2024

    Ona de Gibert, Graeme Nail, Nikolay Arefyev, Marta Bañón, Jelmer van der Linde, Shaoxiong Ji, Jaume Zaragoza-Bernabeu, Mikko Aulamo, Gema Ramírez-Sánchez, Andrey Kutuzov, Sampo Pyysalo, Stephan Oepen, and Jörg Tiedemann. A new massive multilingual dataset for high-performance ...

  18. [27]

    BERTje : A dutch BERT model

    Wietse de Vries, Andreas van Cranenburgh, Arianna Bisazza, Tommaso Caselli, Gertjan van Noord, and Malvina Nissim. BERTje : A dutch BERT model. arXiv preprint arXiv:1912.09582, 2019

  19. [28]

    DeepSeek-AI, :, Xiao Bi, Deli Chen, Guanting Chen, Shanhuang Chen, Damai Dai, Chengqi Deng, Honghui Ding, Kai Dong, Qiushi Du, Zhe Fu, Huazuo Gao, Kaige Gao, Wenjun Gao, Ruiqi Ge, Kang Guan, Daya Guo, Jianzhong Guo, Guangbo Hao, Zhewen Hao, Ying He, Wenjie Hu, Panpan Huang, Er...

  20. [29]

    RobBERT : a dutch RoBERTa -based language model

    Pieter Delobelle, Thomas Winters, and Bettina Berendt. RobBERT : a dutch RoBERTa -based language model. arXiv preprint arXiv:2001.06286, 2020

  21. [30]

    Fquad: French question answering dataset, 2020

    Martin d'Hoffschmidt, Wacim Belblidia, Tom Brendlé, Quentin Heinrich, and Maxime Vidal. Fquad: French question answering dataset, 2020. URL https://arxiv.org/abs/2002.06071

  22. [31]

    Tran, Mike Zhang, Shiqi Chen, Tianyu Pang, Chao Du, Xinyi Wan, Wei Lu, and Min Lin

    Longxu Dou, Qian Liu, Fan Zhou, Changyu Chen, Zili Wang, Ziqi Jin, Zichen Liu, Tongyao Zhu, Cunxiao Du, Penghui Yang, Haonan Wang, Jiaheng Liu, Yongchi Zhao, Xiachong Feng, Xin Mao, Man Tsung Yeung, Kunat Pipatanakul, Fajri Koto, Min Si Thu, Hynek Kydlíček, Zeyi Liu, Qunshu Li...

  23. [32]

    Chinese tiny llm: Pretraining a chinese-centric large language model, 2024

    Xinrun Du, Zhouliang Yu, Songyang Gao, Ding Pan, Yuyang Cheng, Ziyang Ma, Ruibin Yuan, Xingwei Qu, Jiaheng Liu, Tianyu Zheng, Xinchen Luo, Guorui Zhou, Wenhu Chen, and Ge Zhang. Chinese tiny llm: Pretraining a chinese-centric large language model, 2024. URL https://arxiv.org/a...

  24. [33]

    Eberhard, Gary F

    David M. Eberhard, Gary F. Simons, and Charles D. Fenning. Ethnologue: Languages of the world, 2024. URL http://www.ethnologue.com

  25. [34]

    SberQuAD – Russian Reading Comprehension Dataset: Description and Analysis, pp.\ 3–15

    Pavel Efimov, Andrey Chertok, Leonid Boytsov, and Pavel Braslavski. SberQuAD – Russian Reading Comprehension Dataset: Description and Analysis, pp.\ 3–15. Springer International Publishing, 2020. ISBN 9783030582197. doi:10.1007/978-3-030-58219-7_1. URL http://dx.doi.org/10.100...

  26. [35]

    Arabicweb24: Creating a high quality arabic web-only pre-training dataset, 2024

    May Farhat, Said Taghadouini, Oskar Hallström, and Sonja Hajri-Gabouj. Arabicweb24: Creating a high quality arabic web-only pre-training dataset, 2024. URL www.lighton.ai/lighton-blogs/arabicweb24

  27. [36]

    Guerreiro, António Loison, Duarte M

    Manuel Faysse, Patrick Fernandes, Nuno M. Guerreiro, António Loison, Duarte M. Alves, Caio Corro, Nicolas Boizard, João Alves, Ricardo Rei, Pedro H. Martins, Antoni Bigata Casademunt, François Yvon, André F. T. Martins, Gautier Viaud, Céline Hudelot, and Pierre Colombo. Croiss...

  28. [37]

    Mera: A comprehensive llm evaluation in russian, 2024

    Alena Fenogenova, Artem Chervyakov, Nikita Martynov, Anastasia Kozlova, Maria Tikhonova, Albina Akhmetgareeva, Anton Emelyanov, Denis Shevelev, Pavel Lebedev, Leonid Sinev, Ulyana Isaeva, Katerina Kolomeytseva, Daniil Moskovskiy, Elizaveta Goncharova, Nikita Savushkin, Polina ...

  29. [38]

    Open llm leaderboard v2

    Clémentine Fourrier, Nathan Habib, Alina Lozovskaya, Konrad Szafer, and Thomas Wolf. Open llm leaderboard v2. https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard, 2024

  30. [39]

    Gemma Team , Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Léonard Hussenot, Pier Giuseppe Sessa, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Ale...

  31. [40]

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Art...

  32. [41]

    Studying large language model generalization with influence functions

    Roger Grosse, Juhan Bae, Cem Anil, Nelson Elhage, Alex Tamkin, Amirhossein Tajdini, Benoit Steiner, Dustin Li, Esin Durmus, Ethan Perez, et al. Studying large language model generalization with influence functions. arXiv preprint arXiv:2308.03296, 2023

  33. [42]

    Olmes: A standard for language model evaluations, 2025

    Yuling Gu, Oyvind Tafjord, Bailey Kuehl, Dany Haddad, Jesse Dodge, and Hannaneh Hajishirzi. Olmes: A standard for language model evaluations, 2025. URL https://arxiv.org/abs/2406.08446

  34. [43]

    Exams: A multi-subject high school examinations dataset for cross-lingual and multilingual question answering, 2020

    Momchil Hardalov, Todor Mihaylov, Dimitrina Zlatkova, Yoan Dinkov, Ivan Koychev, and Preslav Nakov. Exams: A multi-subject high school examinations dataset for cross-lingual and multilingual question answering, 2020. URL https://arxiv.org/abs/2011.03080

  35. [44]

    Measuring massive multitask language understanding

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. Proceedings of the International Conference on Learning Representations (ICLR), 2021

  36. [45]

    Khmer natural language processing tookit

    Phan Viet Hoang. Khmer natural language processing tookit. https://github.com/VietHoang1512/khmer-nltk, 2020

  37. [46]

    spaCy: Industrial-strength Natural Language Processing in Python

    Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd. spaCy: Industrial-strength Natural Language Processing in Python . 2020. doi:10.5281/zenodo.1212303

  38. [47]

    C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models, 2023

    Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Jiayi Lei, Yao Fu, Maosong Sun, and Junxian He. C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models, 2023. URL https://arxi...

  39. [48]

    Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas ...

  40. [49]

    Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne...

  41. [50]

    The state and fate of linguistic diversity and inclusion in the NLP world

    Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. The state and fate of linguistic diversity and inclusion in the NLP world. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (eds.), Proceedings of the 58th Annual Meeting of the ...

  42. [51]

    Fasttext.zip: Compressing text classification models, 2016

    Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hérve Jégou, and Tomas Mikolov. Fasttext.zip: Compressing text classification models, 2016. URL https://arxiv.org/abs/1612.03651

  43. [52]

    G lot LID : Language identification for low-resource languages

    Amir Hossein Kargaran, Ayyoob Imani, Fran c ois Yvon, and Hinrich Schuetze. G lot LID : Language identification for low-resource languages. In Houda Bouamor, Juan Pino, and Kalika Bali (eds.), Findings of the Association for Computational Linguistics: EMNLP 2023, pp.\ 6155--62...

  44. [53]

    Glot CC : An open broad-coverage commoncrawl corpus and pipeline for minority languages

    Amir Hossein Kargaran, Fran c ois Yvon, and Hinrich Schuetze. Glot CC : An open broad-coverage commoncrawl corpus and pipeline for minority languages. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2024. URL https://openr...

  45. [54]

    Mohammed Safi Ur Rahman Khan, Priyam Mehta, Ananth Sankar, Umashankar Kumaravelan, Sumanth Doddapaneni, Suriyaprasaad G, Varun Balan G, Sparsh Jain, Anoop Kunchukuttan, Pratyush Kumar, Raj Dabre, and Mitesh M. Khapra. Indicllmsuite: A blueprint for creating pre-training and fi...

  46. [55]

    Large language models only pass primary school exams in I ndonesia: A comprehensive test on I ndo MMLU

    Fajri Koto, Nurul Aisyah, Haonan Li, and Timothy Baldwin. Large language models only pass primary school exams in I ndonesia: A comprehensive test on I ndo MMLU . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), Singapore, 2023...

  47. [56]

    Arabicmmlu: Assessing massive multitask language understanding in arabic, 2024

    Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin. Arabicmmlu: Assessing massive multitask language understanding in ara...

  48. [57]

    The IndicNLP Library

    Anoop Kunchukuttan. The IndicNLP Library . https://github.com/anoopkunchukuttan/indic_nlp_library/blob/master/docs/indicnlp.pdf, 2020

  49. [58]

    JGLUE : J apanese general language understanding evaluation

    Kentaro Kurihara, Daisuke Kawahara, and Tomohide Shibata. JGLUE : J apanese general language understanding evaluation. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pp.\ 2957--2966, Marseille, France, June 2022. European Language Resources Asso...

  50. [59]

    Rossi, and Thien Huu Nguyen

    Viet Dac Lai, Chien Van Nguyen, Nghia Trung Ngo, Thuat Nguyen, Franck Dernoncourt, Ryan A. Rossi, and Thien Huu Nguyen. Okapi: Instruction-tuned large language models in multiple languages with reinforcement learning from human feedback, 2023. URL https://arxiv.org/abs/2307.16039

  51. [60]

    F lau BERT : Unsupervised language model pre-training for F rench

    Hang Le, Lo \" c Vial, Jibril Frej, Vincent Segonne, Maximin Coavoux, Benjamin Lecouteux, Alexandre Allauzen, Benoit Crabb \'e , Laurent Besacier, and Didier Schwab. F lau BERT : Unsupervised language model pre-training for F rench. In Proceedings of the 12th Language Resource...

  52. [61]

    Open-arabic-llm-leaderboard-v1

    Open Arabic LLM Leaderboard. Open-arabic-llm-leaderboard-v1. https://huggingface.co/spaces/OALL/Open-Arabic-LLM-Leaderboard-v1, 2024. Accessed: 2025-03-28

  53. [62]

    Deduplicating training data makes language models better, 2022

    Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. Deduplicating training data makes language models better, 2022. URL https://arxiv.org/abs/2107.06499

  54. [63]

    Kiwipiepy: Kiwi package for python, 2024

    Minchul Lee. Kiwipiepy: Kiwi package for python, 2024. URL https://github.com/bab2min/kiwipiepy

  55. [64]

    Mlqa: Evaluating cross-lingual extractive question answering, 2020

    Patrick Lewis, Barlas Oğuz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. Mlqa: Evaluating cross-lingual extractive question answering, 2020. URL https://arxiv.org/abs/1910.07475

  56. [65]

    Cmmlu: Measuring massive multitask language understanding in chinese, 2024 a

    Haonan Li, Yixuan Zhang, Fajri Koto, Yifei Yang, Hai Zhao, Yeyun Gong, Nan Duan, and Timothy Baldwin. Cmmlu: Measuring massive multitask language understanding in chinese, 2024 a . URL https://arxiv.org/abs/2306.09212

  57. [66]

    Dimakis, Yair Carmon, Achal Dave, Ludwig Schmidt, and Vaishaal Shankar

    Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Gadre, Hritik Bansal, Etash Guha, Sedrick Keh, Kushal Arora, Saurabh Garg, Rui Xin, Niklas Muennighoff, Reinhard Heckel, Jean Mercat, Mayee Chen, Suchin Gururangan, Mitchell Wortsman, Alon Albalak, Yonatan ...

  58. [67]

    Common sense beyond E nglish: Evaluating and improving multilingual language models for commonsense reasoning

    Bill Yuchen Lin, Seyeon Lee, Xiaoyang Qiao, and Xiang Ren. Common sense beyond E nglish: Evaluating and improving multilingual language models for commonsense reasoning. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (eds.), Proceedings of the 59th Annual Meeting o...

  59. [69]

    Few-shot learning with multilingual language models, 2022

    Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O'Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mon...

  60. [70]

    Fingpt: Large generative models for a small language

    Risto Luukkonen, Ville Komulainen, Jouni Luoma, Anni Eskelinen, Jenna Kanerva, Hanna-Mari Kupari, Filip Ginter, Veronika Laippala, Niklas Muennighoff, Aleksandra Piktus, et al. Fingpt: Large generative models for a small language. arXiv preprint arXiv:2311.05640, 2023

  61. [71]

    Singh, Rylan Schaeffer, Andrew Poulton, Sanmi Koyejo, Pontus Stenetorp, Sharan Narang, and Dieuwke Hupkes

    Lovish Madaan, Aaditya K. Singh, Rylan Schaeffer, Andrew Poulton, Sanmi Koyejo, Pontus Stenetorp, Sharan Narang, and Dieuwke Hupkes. Quantifying variance in evaluation benchmarks, 2024. URL https://arxiv.org/abs/2406.10229

  62. [72]

    C amem BERT : a tasty F rench language model

    Louis Martin, Benjamin Muller, Pedro Javier Ortiz Su \'a rez, Yoann Dupont, Laurent Romary, \'E ric de la Clergerie, Djam \'e Seddah, and Beno \^ t Sagot. C amem BERT : a tasty F rench language model. In Proceedings of the 58th Annual Meeting of the Association for Computation...

  63. [73]

    Between words and characters: A brief history of open-vocabulary modeling and tokenization in nlp

    Sabrina J Mielke, Zaid Alyafeai, Elizabeth Salesky, Colin Raffel, Manan Dey, Matthias Gall \'e , Arun Raja, Chenglei Si, Wilson Y Lee, Beno \^ t Sagot, et al. Between words and characters: A brief history of open-vocabulary modeling and tokenization in nlp. arXiv preprint arXi...

  64. [74]

    Mnbvc: Massive never-ending bt vast chinese corpus

    MOP-LIWU Community and MNBVC Team . Mnbvc: Massive never-ending bt vast chinese corpus. https://github.com/esbatmop/MNBVC, 2023

  65. [75]

    Neural A rabic question answering

    Hussein Mozannar, Elie Maamary, Karl El Hajal, and Hazem Hajj. Neural A rabic question answering. In Proceedings of the Fourth Arabic Natural Language Processing Workshop, pp.\ 108--118, Florence, Italy, August 2019. Association for Computational Linguistics. doi:10.18653/v1/W...

  66. [76]

    Crosslingual generalization through multitask finetuning, 2022

    Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward ...

  67. [77]

    Rossi, and Thien Huu Nguyen

    Thuat Nguyen, Chien Van Nguyen, Viet Dac Lai, Hieu Man, Nghia Trung Ngo, Franck Dernoncourt, Ryan A. Rossi, and Thien Huu Nguyen. C ultura X : A cleaned, enormous, and multilingual dataset for large language models in 167 languages. In Nicoletta Calzolari, Min-Yen Kan, Veroniq...

  68. [78]

    NLLB Team , Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Pran...

  69. [79]

    Omnia russica

    Omnia Russica Team . Omnia russica. https://omnia-russica.github.io/, 2024

  70. [80]

    Botok: State-of-the-art tokenizers for tibetan language, 2025

    OpenPecha . Botok: State-of-the-art tokenizers for tibetan language, 2025. URL https://github.com/OpenPecha/Botok. Support for various dialects, fully customizable word lists and adjustment rules

  71. [81]

    Building pre-train llm dataset for the indic languages: A case study on hindi

    Shantipriya Parida, Shakshi Panwar, Kusum Lata, Sanskruti Mishra, and Sambit Sekhar. Building pre-train llm dataset for the indic languages: A case study on hindi. https://huggingface.co/OdiaGenAI, 2024

  72. [82]

    Hellaswag-th, 2023

    Triamamornwooth Patteera. Hellaswag-th, 2023. URL https://huggingface.co/datasets/Patt/HellaSwag_TH. v1.0

  73. [83]

    The refinedweb dataset for falcon llm: Outperforming curated corpora with web data, and web data only, 2023

    Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. The refinedweb dataset for falcon llm: Outperforming curated corpora with web data, and web data only, 2023. UR...

  74. [84]

    The fineweb datasets: Decanting the web for the finest text data at scale

    Guilherme Penedo, Hynek Kydl\' c ek, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf. The fineweb datasets: Decanting the web for the finest text data at scale. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. ...

  75. [85]

    Laonlp: Lao language natural language processing, July 2022

    Wannaphong Phatthiyaphaibun. Laonlp: Lao language natural language processing, July 2022. URL https://doi.org/10.5281/zenodo.6833407

  76. [86]

    P y T hai NLP : T hai natural language processing in P ython, June 2024

    Wannaphong Phatthiyaphaibun, Korakot Chaovavanich, Charin Polpanumas, Arthit Suriyawongkul, Lalita Lowphansirikul, and Pattarawat Chormai. P y T hai NLP : T hai natural language processing in P ython, June 2024. URL https://github.com/PyThaiNLP/pythainlp/

  77. [87]

    Typhoon: Thai large language models

    Kunat Pipatanakul, Phatrasek Jirabovonvisut, Potsawee Manakul, Sittipong Sripaisarnmongkol, Ruangsak Patomwong, Pathomporn Chokchainant, and Kasima Tharnpipitchai. Typhoon: Thai large language models. arXiv preprint arXiv:2312.13951, 2023

  78. [88]

    Pllum: A family of polish large language models

    PLLuM Consortium . Pllum: A family of polish large language models. 2025

  79. [89]

    Chinesesquad

    Pluto-Junzeng. Chinesesquad. https://github.com/pluto-junzeng/ChineseSquad, 2019. Accessed: 2025-03-28

  80. [90]

    Xcopa: A multilingual dataset for causal commonsense reasoning, 2020

    Edoardo Maria Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulić, and Anna Korhonen. Xcopa: A multilingual dataset for causal commonsense reasoning, 2020. URL https://arxiv.org/abs/2005.00333

  81. [91]

    Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton, and Christopher D. Manning. Stanza: A python natural language processing toolkit for many human languages, 2020. URL https://arxiv.org/abs/2003.07082

  82. [92]

    Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks,...

  83. [93]

    Exploring the limits of transfer learning with a unified text-to-text transformer

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21 0 (140), 2020

  84. [94]

    Impact of pretraining term frequencies on few-shot reasoning

    Yasaman Razeghi, Robert L Logan IV, Matt Gardner, and Sameer Singh. Impact of pretraining term frequencies on few-shot reasoning. arXiv preprint arXiv:2202.07206, 2022

  85. [95]

    How much knowledge can you pack into the parameters of a language model? arXiv preprint arXiv:2002.08910, 2020

    Adam Roberts, Colin Raffel, and Noam Shazeer. How much knowledge can you pack into the parameters of a language model? arXiv preprint arXiv:2002.08910, 2020

  86. [96]

    How good is your tokenizer? on the monolingual performance of multilingual language models, 2021

    Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, and Iryna Gurevych. How good is your tokenizer? on the monolingual performance of multilingual language models, 2021. URL https://arxiv.org/abs/2012.15613

  87. [97]

    Pyidaungsu: Python library for myanmar language, 2024

    Oishi Sakana. Pyidaungsu: Python library for myanmar language, 2024. URL https://github.com/kaunghtetsan275/pyidaungsu

  88. [98]

    Compact language detector v3

    Alex Salcianu, Andy Golding, Anton Bakalov, Chris Alberti, Daniel Andor, David Weiss, Emily Pitler, Greg Coppola, Jason Riesa, Kuzman Ganchev, et al. Compact language detector v3. Technical report, 2018. URL https://chromium.googlesource.com/external/github.com/google/cld_3/. ...

  89. [99]

    Mintaka: A complex, natural, and multilingual dataset for end-to-end question answering, 2022

    Priyanka Sen, Alham Fikri Aji, and Amir Saffari. Mintaka: A complex, natural, and multilingual dataset for end-to-end question answering, 2022. URL https://arxiv.org/abs/2210.01613

  90. [100]

    Indic qa benchmark: A multilingual benchmark to evaluate question answering capability of llms for indic languages, 2025

    Abhishek Kumar Singh, Vishwajeet kumar, Rudra Murthy, Jaydeep Sen, Ashish Mittal, and Ganesh Ramakrishnan. Indic qa benchmark: A multilingual benchmark to evaluate question answering capability of llms for indic languages, 2025. URL https://arxiv.org/abs/2407.13522

  91. [101]

    Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh, Luke Zettlemoyer, Noah A

    Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Mu...

  92. [102]

    Thquad: Turkish historic question answering dataset for reading comprehension

    Fatih Soygazi, Okan Çiftçi, Uğurcan Kök, and Soner Cengiz. Thquad: Turkish historic question answering dataset for reading comprehension. In 2021 6th International Conference on Computer Science and Engineering (UBMK), pp.\ 215--220, 2021. doi:10.1109/UBMK52708.2021.9559013

  93. [103]

    Investigating prior knowledge for challenging C hinese machine reading comprehension

    Kai Sun, Dian Yu, Dong Yu, and Claire Cardie. Investigating prior knowledge for challenging C hinese machine reading comprehension. Transactions of the Association for Computational Linguistics, 8: 0 141--155, 2020. doi:10.1162/tacl_a_00305. URL https://aclanthology.org/2020.t...

  94. [104]

    Liping Tang, Nikhil Ranjan, Omkar Pangarkar, Xuezhi Liang, Zhen Wang, Li An, Bhaskar Rao, Linghao Jin, Huijuan Wang, Zhoujun Cheng, Suqi Sun, Cun Mu, Victor Miller, Xuezhe Ma, Yue Peng, Zhengzhong Liu, and Eric P. Xing. Txt360: A top-quality llm pre-training dataset requires t...

  95. [105]

    Tigerbot: A multi-language multi-task llm

    TigerResearch. Tigerbot: A multi-language multi-task llm. https://github.com/TigerResearch/TigerBot, 2023

  96. [106]

    Llama: Open and efficient foundation language models, 2023

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language...

  97. [107]

    Trakultaweekoon, S

    K. Trakultaweekoon, S. Thaiprayoon, P. Palingoon, and A. Rugchatjaroen. The first wikipedia questions and factoid answers corpus in the thai language. In 2019 14th International Joint Symposium on Artificial Intelligence and Natural Language Processing (iSAI-NLP), pp.\ 1--4. I...

  98. [108]

    Vbart: The turkish llm

    Meliksah Turker, Erdi Ari, and Aydin Han. Vbart: The turkish llm. arXiv preprint arXiv:2403.01308, 2024

  99. [109]

    Wanjawa, Lilian D

    Barack W. Wanjawa, Lilian D. A. Wanzare, Florence Indede, Owen Mconyango, Lawrence Muchemi, and Edward Ombui. Kenswquad—a question answering dataset for swahili low-resource language. ACM Transactions on Asian and Low-Resource Language Information Processing, 22 0 (4): 0 1–20,...

  100. [110]

    CCN et: Extracting high quality monolingual datasets from web crawl data

    Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzm \'a n, Armand Joulin, and Edouard Grave. CCN et: Extracting high quality monolingual datasets from web crawl data. In Nicoletta Calzolari, Fr \'e d \'e ric B \'e chet, Philippe Blache, Khal...

  101. [111]

    BigScience Workshop, :, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanch...

  102. [112]

    mt5: A massively multilingual pre-trained text-to-text transformer, 2021

    Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. mt5: A massively multilingual pre-trained text-to-text transformer, 2021. URL https://arxiv.org/abs/2010.11934

  103. [113]

    Qwen2.5 technical report

    An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Me...

  104. [114]

    Hellaswag: Can a machine really finish your sentence?, 2019

    Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence?, 2019. URL https://arxiv.org/abs/1905.07830

  105. [115]

    M3exam: A multilingual, multimodal, multilevel benchmark for examining large language models, 2023

    Wenxuan Zhang, Sharifah Mahani Aljunied, Chang Gao, Yew Ken Chia, and Lidong Bing. M3exam: A multilingual, multimodal, multilevel benchmark for examining large language models, 2023. URL https://arxiv.org/abs/2306.05179

  106. [117]

    Agieval: A human-centric benchmark for evaluating foundation models, 2023 b

    Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan. Agieval: A human-centric benchmark for evaluating foundation models, 2023 b . URL https://arxiv.org/abs/2304.06364

  107. [118]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

  108. [119]

    @esa (Ref

    \@ifxundefined[1] #1\@undefined \@firstoftwo \@secondoftwo \@ifnum[1] #1 \@firstoftwo \@secondoftwo \@ifx[1] #1 \@firstoftwo \@secondoftwo [2] @ #1 \@temptokena #2 #1 @ \@temptokena \@ifclassloaded agu2001 natbib The agu2001 class already includes natbib coding, so you should ...

  109. [120]

    \@lbibitem[] @bibitem@first@sw\@secondoftwo \@lbibitem[#1]#2 \@extra@b@citeb \@ifundefined br@#2\@extra@b@citeb \@namedef br@#2 \@nameuse br@#2\@extra@b@citeb \@ifundefined b@#2\@extra@b@citeb @num @parse #2 @tmp #1 NAT@b@open@#2 NAT@b@shut@#2 \@ifnum @merge>\@ne @bibitem@firs...

  110. [121]

    C MI:LLJ zfwj

    @open @close @open @close and [1] URL: #1 \@ifundefined chapter * \@mkboth \@ifxundefined @sectionbib * \@mkboth * \@mkboth\@gobbletwo \@ifclassloaded amsart * \@ifclassloaded amsbook * \@ifxundefined @heading @heading NAT@ctr thebibliography [1] @ \@biblabel @NAT@ctr \@bibset...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.