Pith. sign in

REVIEW 4 major objections 4 minor 3 cited by

KletterMix, a German corpus produced by translating a curated English pretraining mixture, yields measurable gains on German reasoning benchmarks under matched token budgets.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 12:25 UTC pith:VQUXHD4O

load-bearing objection KletterMix is a genuinely useful, well-documented dataset artifact, but the improvement claim rests on single-run comparisons with no contamination check against the translated benchmarks. the 4 major comments →

arxiv 2606.03773 v2 pith:VQUXHD4O submitted 2026-06-02 cs.CL

KletterMix: Climbing Toward High-Quality German Pretraining Data - The Full Report

classification cs.CL
keywords KletterMixGerman pretraining corpusmachine-translated training datatranslation quality estimationCOMETKiwi proxydata mixture transferGerman LLM pretrainingannealing
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper introduces KletterMix, a 725B-token German pretraining corpus created by translating ClimbMix, an English mixture, into German while keeping document boundaries, metadata, and source-cluster structure. The central claim is that carefully translated English data can transfer not only German surface form but also beneficial mixture structure: under matched 12B-token budgets with a 0.6B-parameter model, KletterMix-trained models score higher on a four-task German benchmark core average than models trained on FineWeb2-DE or GermanWeb. The gain concentrates on HellaSwag and ARC-C, tasks the authors read as needing coherent event continuation and compositional reasoning. Filtering translated documents by a proxy-predicted quality score further improves the core average point estimate, while annealing a FineWeb2-DE checkpoint on KletterMix outperforms annealing on GermanWeb. The paper treats this as evidence that translation-based data construction can meaningfully strengthen non-English pretraining data.

Core claim

KletterMix is a 725B-token German-language corpus built by machine-translating the ClimbMix English pretraining mixture with a document-preserving pipeline: length-aware routing, contextualized chunking, dynamic output budgeting, and shard-wise execution. The paper claims that models pretrained from scratch on matched 12B-token KletterMix subsets reach lower training and validation loss than on FineWeb2-DE or GermanWeb and achieve the strongest four-task core average across MMLU, PIQA, HellaSwag, and ARC-C among the compared runs, with the best point estimate (40.2) on the validation-selected proxy-filtered split. The authors interpret the task pattern — consistent gains on HellaSwag and ARC

What carries the argument

The central mechanism is the document-preserving translation pipeline paired with a target-only quality proxy. Documents are routed into length buckets, translated whole or as contextualized chunks, and scored by COMETKiwi on a stratified pilot sample; a gradient-boosted regressor trained on German-only text features (length, language-identification signals, character ratios, repetition) then predicts COMETKiwi-like scores for every document. This proxy is what allows the full 725B-token corpus to be filtered into controlled 12B-token training splits and permits the cluster-level and length-bucket quality diagnostics.

Load-bearing premise

The load-bearing premise is that the benchmark gains on translated German evaluations reflect general German-language competence; the paper does not check whether training documents overlap with the translated evaluation sets, so a memorization-based advantage cannot be ruled out.

What would settle it

Run a document-level overlap analysis between the KletterMix training subset and the German MMLU, PIQA, HellaSwag, and ARC-C evaluation sets using normalized n-gram or embedding similarity; if high-similarity documents exist in the training corpus and removing them erases the HellaSwag/ARC-C gains, the central claim of reasoning transfer would be falsified. A complementary check is to evaluate the same checkpoints on a native-German benchmark suite not derived from English translation.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the central claim holds, translated curated English mixtures provide a repeatable route to larger and more diverse German pretraining data without requiring additional native web crawling.
  • Under a fixed 12B-token budget, KletterMix improves validation loss and reasoning-style benchmark performance at the tested scale, and the annealing result suggests it can act as a late-stage steering corpus after training on native German web data.
  • Proxy-based filtering is a practical ranking signal for fixed-budget training mixtures: stricter thresholds improved the core average up to the validation-selected split, though not uniformly across every task.
  • The aligned English-German document identifiers make KletterMix a reusable testbed for studying how translation quality, source cluster, and document length affect downstream model behavior.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The benchmark gains are measured on German evaluations that are themselves translations of English benchmarks; the paper does not report a document-level overlap check between the training corpus and the evaluation sets, so part of the advantage could reflect memorized or near-duplicate translated content rather than transferable ability.
  • The reading that HellaSwag and ARC-C gains demonstrate 'reasoning transfer' is one interpretation; an equally testable one is that the translated corpus simply contains more structured exposition, which those benchmarks reward. A native-German reasoning suite and human naturalness judgments would separate the two.
  • Because the proxy is target-only, filtering cannot catch failures that survive translation while leaving the German text looking clean; sampling a subset and scoring it with a source-aware quality model would bound how much quality signal is being left unused.
  • The same pipeline could be applied to other languages, but whether the effect generalizes is unknown; if the gains shrink for a language typologically closer to English, mixture-structure transfer may be less important than translationese effects.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper introduces KletterMix, a 725B-token German pretraining corpus obtained by machine-translating the English ClimbMix mixture while preserving document boundaries, metadata, and source-cluster structure. It documents a scalable translation pipeline (length-aware routing, chunking, dynamic target budgeting, shard-wise execution), uses COMETKiwi scores to train a target-only quality proxy, and releases unfiltered plus threshold-filtered variants. The empirical core is a set of matched 12B-token pretraining and annealing ablations with Qwen3-0.6B, comparing KletterMix against FineWeb2-DE and GermanWeb on four German translated benchmarks (MMLU, PIQA, HellaSwag, ARC-C). The central claim is that models trained on KletterMix achieve measurable improvements on German-language downstream evaluations, with the strongest point estimates on HellaSwag and ARC-C.

Significance. If the central claim holds, the paper makes a useful contribution to non-English pretraining data: it provides a large, documented, reproducible translated corpus, a transparent quality-estimation pipeline, and controlled training ablations under matched token budgets. The dataset artifact itself, with preserved document identifiers and metadata, is valuable for future studies of translation-based data curation. The paper also deserves credit for reporting evaluation-set standard errors, releasing code and data links, documenting translation failure modes qualitatively, and validating the proxy on a disjoint split. However, the strength of the empirical claim is currently limited by the absence of a contamination analysis, by single-run point estimates whose aggregate gaps are within evaluation noise, and by validation-based threshold selection on the same benchmarks used for the headline comparison. These issues are fixable within the manuscript's scope but are load-bearing for the abstract's claim of 'measurable improvements.'

major comments (4)
  1. [Sec. 5, Table 1] The headline comparison is not statistically decisive under the reported uncertainties. KletterMix-Filt0.60 has Core Avg. 40.2±1.3 vs. FineWeb2-DE 38.3±1.3, a gap of 1.9 points, which is about one standard error of the difference; unfiltered KletterMix (38.7±1.4) is within 0.4 points of FineWeb2-DE. Since the 0.60 threshold was selected after inspecting these same validation benchmarks, the 'best filtered variant' comparison is a selected-maximum result and inflates the chance of a false positive. The paper should either report multiple seeds with seed-level variance, or present the threshold choice as a hypothesis on a separate validation benchmark, or explicitly weaken the 'measurable improvements' claim to a suggestive single-run result.
  2. [Sec. 5 / Sec. 3 / Table 1] No contamination analysis is reported between KletterMix training documents and the translated evaluation benchmarks. KletterMix is a translation of ClimbMix, an English web mixture, and the four German benchmarks are themselves translations of English benchmarks whose items often originate from web text. If source documents for benchmark items appear in ClimbMix, their German translations will appear in KletterMix, and the consistent HellaSwag/ARC-C gains could reflect near-duplicate or memorized content rather than transferable German-language competence. The paper should report document-level or n-gram overlap between the training subset and each evaluation set, remove or flag overlapping items, and rerun the key comparisons. This is the weakest load-bearing link for the 'reasoning transfer' interpretation.
  3. [Sec. 5, Tab. 11 and Fig. 6] The in-domain validation-loss comparisons do not support the claim that KletterMix is 'not merely easier to fit.' Each model is evaluated on its own corpus's held-out validation set, so the lower KletterMix perplexity (6.02 vs. 10.04 for FineWeb2-DE and 8.50 for GermanWeb) can reflect corpus difficulty or distributional differences, not better transferable modeling. The filtered rows are also compared on the KletterMix validation set without a common held-out corpus. A meaningful validation comparison would require a common held-out German benchmark or a carefully matched cross-corpus validation set; otherwise the optimization-dynamics discussion should be limited to training loss.
  4. [Sec. 5 / App. A.5] The aggregate Core Avg. is dominated by PIQA's small evaluation set (100 examples) and wide standard error. The claim that 'the recurring task-level pattern is the more stable signal' is reasonable, but the task-level pattern is itself based on single runs. For the HellaSwag and ARC-C gains to support the central claim, the paper needs either seed variance estimates or a more robust evaluation protocol. At minimum, the abstract and conclusion should not state 'measurable improvements' as an established fact while the only quantitative support is a selected single-run point estimate within noise.
minor comments (4)
  1. [Sec. 5, Dataset interpretation] 'The MMMLU result' appears to be a typo for 'MMLU result.'
  2. [Sec. 3, Proxy-filtered dataset variants] The choice of thresholds 0.50/0.55/0.60 and the dynamic budget parameters α=2.0, β=1024 are presented as ablations, which is fine, but the main text could state more explicitly that these are not derived from a principled criterion and that the 0.60 threshold is validation-selected.
  3. [Sec. 5, Setup] The description of the deterministic token-budgeted stratified sampler is clear, but it would help to report the exact size of each evaluation-set subset (e.g., the 100-example PIQA subset) in the main table caption or in App. A.5, since this materially affects how readers should interpret the standard errors.
  4. [Sec. 6, Limitations] The Limitations section explicitly acknowledges single-run comparisons and evaluation-set-only standard errors, which is commendable, but it should also acknowledge the absence of a benchmark-contamination analysis, given that both the training corpus and evaluation sets are translations of English web-derived data.

Circularity Check

0 steps flagged

No significant circularity: KletterMix is built from an external English mixture, filtered by external COMETKiwi-derived scores, and evaluated with controlled ablations; no result reduces to its inputs by construction.

full rationale

I walked the paper's derivation chain: source corpus (ClimbMix, external), translation pipeline, COMETKiwi-based quality labeling, target-only proxy filtering, and matched pretraining ablations. No step defines its output in terms of the downstream claim. KletterMix is not constructed from benchmark results; the proxy is supervised by reference-free COMETKiwi scores, not by downstream accuracy; and the filtering thresholds (0.50/0.55/0.60) are fixed proxy-score cutoffs rather than fitted benchmark parameters. The downstream benchmark numbers are empirical measurements, not identities. The closest concern is the 'validation-selected filtered split' in Sec. 5, where the 0.60-filtered variant reaches the best Core Avg.; this could reflect threshold selection on the same benchmark outcomes and is a statistical selection effect, but it is not an equation-level reduction and the paper does not claim to predict the benchmark from the filter. The absence of a contamination analysis between translated training data and translated benchmarks is a validity/correctness risk, not circularity: any overlap would be a contingent data artifact rather than a definitional consequence. Self-citations to the authors' earlier German benchmark work provide evaluation instruments, but those are external published artifacts and are not used to force the paper's conclusion. No circular step meets the evidence bar.

Axiom & Free-Parameter Ledger

2 free parameters · 5 axioms · 0 invented entities

The central claim rests on external quality assumptions (ClimbMix quality, COMETKiwi validity, proxy generalization, benchmark cleanliness) plus hand-chosen filter thresholds. No new physical or theoretical entities are introduced.

free parameters (2)
  • Proxy filter thresholds = 0.50, 0.55, 0.60; 0.60 reported as best (validation-selected)
    Hand-chosen thresholds define the filtering ablations; the 0.60 threshold is selected after observing validation performance, which inflates the reported best result.
  • Dynamic target budget coefficients (alpha, beta, min, Lmax) = alpha=2.0, beta=1024, min=2048, Lmax=32768
    Chosen by hand to allow moderate target-side expansion during translation; affects truncation risk but is not central to the benchmark comparison.
axioms (5)
  • domain assumption ClimbMix is a high-quality English pretraining mixture.
    KletterMix inherits its mixture design, source clusters, and metadata from ClimbMix [11]; if ClimbMix is not high-quality, KletterMix inherits its flaws. Invoked in Sec. 3 ('Source corpus and record structure').
  • domain assumption COMETKiwi reference-free scores are valid translation-quality labels.
    COMETKiwi scores supervise the proxy and are used as the quality signal; no human evaluation is performed on the full corpus. Invoked in Sec. 3 and App. A.4.
  • domain assumption The target-only proxy generalizes from the pilot subset to the full 725B-token corpus.
    Validated on a disjoint 18,275-document split (Tab. 6), but the full corpus is several orders of magnitude larger and includes long-tail domains. Invoked in Sec. 3 and App. A.4.
  • domain assumption The German evaluation benchmarks measure German LM quality without contamination from translated training data.
    No overlap/contamination check is reported between KletterMix and the translated benchmark sets (German MMLU, PIQA, HellaSwag, ARC-C). Invoked in Sec. 5 / Tab. 1.
  • domain assumption Single-run training results are representative of the data-mixture comparison.
    All pretraining and annealing comparisons are matched single runs without seed variance; the paper reports only evaluation-set standard errors. Invoked in Sec. 5 and Limitations.

pith-pipeline@v1.3.0-alltime-deepseek · 26079 in / 11846 out tokens · 119972 ms · 2026-08-02T12:25:13.136696+00:00 · methodology

0 comments
read the original abstract

High-quality pretraining data is a central ingredient in modern language models, but German-language resources remain far less developed than their English counterparts: they are often smaller, less carefully curated, weakly documented, and rarely validated through controlled training experiments. We introduce KletterMix, a high-quality German corpus for language model pretraining and annealing, designed as a reusable dataset artifact for the natural language processing and modeling community. KletterMix is built by translating a state-of-the-art English pretraining corpus into German while preserving document boundaries, metadata, source structure, and topical diversity. This construction yields a German corpus with the scale and diversity of a modern pretraining dataset, while enabling direct comparison to its English source. We document the dataset through a broad set of corpus-level analyses, including translation quality, document length distributions, topic coverage, source composition, and geographic metadata. Using COMETKiwi, we show that the translated documents achieve strong quality across diverse domains, suggesting that careful translation can preserve much of the semantic and stylistic richness of the original corpus. Beyond dataset construction, we evaluate KletterMix as training data. Through controlled pretraining and annealing ablations against established German corpora, we show that models trained on KletterMix achieve measurable improvements on German-language downstream evaluations. These results demonstrate that carefully curated translated data can substantially strengthen the German pretraining data ecosystem.

Figures

Figures reproduced from arXiv: 2606.03773 by Abbas Goher Khan, Kristian Kersting, Maurice Kraus, Mehdi Ali, Michael Fromm, Nicolas Flores-Herr, Ruben H\"arle, Sebastian Sztwiertnia.

Figure 1
Figure 1. Figure 1: Overview of the KletterMix pipeline. English source shards are routed into length-aware [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Proxy-score distribution and filtering thresholds used to construct the three 12B-token [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Corpus diagnostics for the full KletterMix release and the 12B-token subset. The length [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Training and annealing dynamics on matched 12B-token German subsets. Across both train [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: German token share by inherited source-cluster metadata. The plot shows how much [PITH_FULL_IMAGE:figures/full_fig_p021_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Extended results for the training ablations in Sec. 5. The main text reports the primary [PITH_FULL_IMAGE:figures/full_fig_p028_6.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. A Sovereign, Open-Source Foundation Model for German and English

    cs.CL 2026-07 conditional novelty 5.5

    Soofi S 30B-A3B, a fully documented German–English hybrid Mamba-MoE base model, matches dense 14–27B peers on bilingual aggregates while delivering 8–9× long-context decode throughput.

  2. A Sovereign, Open-Source Foundation Model for German and English

    cs.CL 2026-07 conditional novelty 5.5

    A fully documented German–English hybrid MoE base model matches dense 14–27B peers, leads open code scores, and sustains high long-context throughput at 3B active parameters.

  3. A Sovereign, Open-Source Foundation Model for German and English

    cs.CL 2026-07 conditional novelty 5.0

    Soofi S 30B-A3B, a hybrid Mamba-MoE model pretrained on ~27T tokens with deliberately up-weighted German, reports the highest English and German aggregate scores among fully open base models in its comparison while ma...

Reference graph

Works this paper leans on

43 extracted references · 9 linked inside Pith · cited by 1 Pith paper

  1. [1]

    Mortensen, Noah A

    Orevaoghene Ahia, Sachin Kumar, Hila Gonen, Jungo Kasai, David R. Mortensen, Noah A. Smith, and Yulia Tsvetkov. Do all languages cost the same? tokenization in the era of commer- cial language models. In Houda Bouamor, Juan Pino, and Kalika Bali, editors,Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Si...

  2. [2]

    Teuken-7b-base & teuken-7b-instruct: Towards european llms

    Mehdi Ali, Michael Fromm, Klaudia Thellmann, Jan Ebert, Alexander Arno Weber, Richard Rutmann, Charvi Jain, Max Lübbering, Daniel Steinigen, Johannes Leveling, Katrin Klug, Jasper Schulze Buschhoff, Lena Jurkschat, Hammam Abdelwahab, Benny Jörg Stein, Karl- Heinz Sylla, Pavel Denisov, Nicolo’ Brandizzi, Qasid Saleem, Anirban Bhowmick, Lennard Helmer, Chel...

  3. [3]

    Occiglot at WMT24: european open-source large language models evaluated on translation

    Eleftherios Avramidis, Annika Grützner-Zahn, Manuel Brack, Patrick Schramowski, Pedro Ortiz Suarez, Malte Ostendorff, Fabio Barth, Shushen Manakhimova, Vivien Macketanz, Georg Rehm, and Kristian Kersting. Occiglot at WMT24: european open-source large language models evaluated on translation. In Barry Haddow, Tom Kocmi, Philipp Koehn, and Christof Monz, ed...

  4. [4]

    Bender and Batya Friedman

    Emily M. Bender and Batya Friedman. Data statements for natural language processing: Toward mitigating system bias and enabling better science.Trans. Assoc. Comput. Linguistics, 6:587–604, 2018

  5. [5]

    Burns, Letitia Parcalabescu, Stephan Wäldchen, Michael Barlow, Gregor Ziegltrum, V olker Stampa, Bastian Harren, and Björn Deiseroth

    Thomas F. Burns, Letitia Parcalabescu, Stephan Wäldchen, Michael Barlow, Gregor Ziegltrum, V olker Stampa, Bastian Harren, and Björn Deiseroth. Aleph-alpha-germanweb: Improving german-language LLM pre-training with model-based data curation and synthetic data gen- eration. In Vera Demberg, Kentaro Inui, and Lluís Marquez, editors,Proceedings of the 19th C...

  6. [6]

    Tyler A. Chang, Catherine Arnett, Abdelrahman Eldesokey, Abdelrahman Sadallah, Abeer Kashar, Abolade Daud, Abosede Grace Olanihun, Adamu Labaran Mohammed, Adeyemi Praise, Adhikarinayum Meerajita Sharma, Aditi Gupta, Afitab Iyigun, Afonso Simplício, Ahmed Essouaied, Aicha Chorana, Akhil Eppa, Akintunde Oladipo, Akshay Ramesh, Aleksei Dorkin, 11 Alfred Male...

  7. [7]

    Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018

    Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018. arXiv:1803.05457

  8. [8]

    Smith, Ahmad Idrissi-Yaghir, Constantin Seibold, Jianning Li, Lars Heiliger, Christoph M

    Amin Dada, Aokun Chen, Cheng Peng, Kaleb E. Smith, Ahmad Idrissi-Yaghir, Constantin Seibold, Jianning Li, Lars Heiliger, Christoph M. Friedrich, Daniel Truhn, Jan Egger, Jiang Bian, Jens Kleesiek, and Yonghui Wu. On the impact of cross-domain data on german language models. In Houda Bouamor, Juan Pino, and Kalika Bali, editors,Findings of the Association ...

  9. [9]

    A new massive multilingual dataset for high- performance language technologies

    Ona de Gibert, Graeme Nail, Nikolay Arefyev, Marta Bañón, Jelmer van der Linde, Shaoxiong Ji, Jaume Zaragoza-Bernabeu, Mikko Aulamo, Gema Ramírez-Sánchez, Andrey Kutuzov, Sampo Pyysalo, Stephan Oepen, and Jörg Tiedemann. A new massive multilingual dataset for high- performance language technologies. In Nicoletta Calzolari, Min-Yen Kan, Véronique Hoste, Al...

  10. [10]

    WMT24++: Expanding the language coverage of WMT24 to 55 languages & dialects

    Daniel Deutsch, Eleftheria Briakou, Isaac Rayburn Caswell, Mara Finkelstein, Rebecca Galor, Juraj Juraska, Geza Kovacs, Alison Lui, Ricardo Rei, Jason Riesa, Shruti Rijhwani, Parker Riley, Elizabeth Salesky, Firas Trabelsi, Stephanie Winkler, Biao Zhang, and Markus Freitag. WMT24++: Expanding the language coverage of WMT24 to 55 languages & dialects. In F...

  11. [11]

    Nemotron-climb: Clustering-based iterative data mixture bootstrapping for language model pre-training, 2025

    Shizhe Diao, Yu Yang, Yonggan Fu, Xin Dong, Dan Su, Markus Kliegl, Zijia Chen, Peter Belcak, Yoshi Suhara, Hongxu Yin, Mostofa Patwary, Yingyan, Lin, Jan Kautz, and Pavlo Molchanov. Nemotron-climb: Clustering-based iterative data mixture bootstrapping for language model pre-training, 2025. arXiv:2504.13161

  12. [12]

    Documenting large webtext corpora: A case study on the colossal clean crawled corpus

    Jesse Dodge, Maarten Sap, Ana Marasovic, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. Documenting large webtext corpora: A case study on the colossal clean crawled corpus. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih, editors,Proceedings of the 2021 Conference on Empirical Methods in...

  13. [13]

    Pretraining language models using translationese

    Meet Doshi, Raj Dabre, and Pushpak Bhattacharyya. Pretraining language models using translationese. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors,Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, pages 5843–5862. Association for Computational Linguistic...

  14. [14]

    The pile: An 800gb dataset of diverse text for language modeling, 2020

    Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. The pile: An 800gb dataset of diverse text for language modeling, 2020. arXiv:2101.00027

  15. [15]

    The language model evaluation harness, 07 2024

    Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. The languag...

  16. [16]

    Wallach, Hal Daumé III, and Kate Crawford

    Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna M. Wallach, Hal Daumé III, and Kate Crawford. Datasheets for datasets.Commun. ACM, 64(12): 86–92, 2021

  17. [17]

    The german commons - 154 billion tokens of openly licensed text for german language models, 2025

    Lukas Gienapp, Christopher Schröder, Stefan Schweter, Christopher Akiki, Ferdinand Schlatt, Arden Zimmermann, Phillipe Genêt, and Martin Potthast. The german commons - 154 billion tokens of openly licensed text for german language models, 2025. arXiv:2510.13996

  18. [18]

    Measuring massive multitask language understanding

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021

  19. [19]

    Rae, and Laurent Sifre

    Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katherine Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Oriol Vinyals, Jack W. Rae, and Laurent S...

  20. [20]

    Glotlid: Language identification for low-resource languages

    Amir Hossein Kargaran, Ayyoob Imani, François Yvon, and Hinrich Schütze. Glotlid: Language identification for low-resource languages. In Houda Bouamor, Juan Pino, and Kalika Bali, editors,Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, Findings of ACL, pages 6155–6218. Association for Computational Li...

  21. [21]

    Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii- Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Javier Ortiz Suárez, Iroro Orife, Kelechi Ogueji, An- d...

  22. [22]

    Yamshchikov

    Pierre-Carl Langlais, Pavel Chizhov, Catherine Arnett, Carlos Rosas Hinostroza, Mattia Nee, Eliot Krzysztof Jones, Irène Girard, David Mach, Anastasia Stasenko, and Ivan P. Yamshchikov. Common corpus: The largest collection of ethical data for LLM pre-training. InThe Fourteenth International Conference on Learning Representations, 2026

  23. [23]

    The bigscience ROOTS corpus: A 1.6tb composite multilingual dataset

    Hugo Laurençon, Lucile Saulnier, Thomas Wang, Christopher Akiki, Albert Villanova del Moral, Teven Le Scao, Leandro von Werra, Chenghao Mou, Eduardo González Ponferrada, Huu Nguyen, Jörg Frohberg, Mario Sasko, Quentin Lhoest, Angelina McMillan-Major, Gérard Dupont, Stella Biderman, Anna Rogers, Loubna Ben Allal, Francesco De Toni, Giada Pistilli, Olivier ...

  24. [24]

    Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Yitzhak Gadre, Hritik Bansal, Etash Kumar Guha, Sedrick Scott Keh, Kushal Arora, Saurabh Garg, Rui Xin, Niklas Muennighoff, Reinhard Heckel, Jean Mercat, Mayee F. Chen, Suchin Gururangan, Mitchell Wortsman, Alon Albalak, Yonatan Bitton, Marianna Nezhurina, Amro Abbas, Cheng-Yu Hsieh, D...

  25. [25]

    Guerreiro, Ricardo Rei, Duarte M

    Pedro Henrique Martins, Patrick Fernandes, João Alves, Nuno M. Guerreiro, Ricardo Rei, Duarte M. Alves, José Pombal, Amin Farajian, Manuel Faysse, Mateusz Klimaszewski, Pierre Colombo, Barry Haddow, José G. C. de Souza, Alexandra Birch, and André F. T. Martins. EuroLLM: Multilingual language models for europe, 2024. arXiv:2409.16235

  26. [26]

    Rossi, and Thien Huu Nguyen

    Thuat Nguyen, Chien Van Nguyen, Viet Dac Lai, Hieu Man, Nghia Trung Ngo, Franck Dernoncourt, Ryan A. Rossi, and Thien Huu Nguyen. CulturaX: A cleaned, enormous, and multilingual dataset for large language models in 167 languages. In Nicoletta Calzolari, Min-Yen Kan, Véronique Hoste, Alessandro Lenci, Sakriani Sakti, and Nianwen Xue, editors, Proceedings o...

  27. [27]

    Hplt 3.0: Very large- scale multilingual resources for llm and mt

    Stephan Oepen, Nikolay Arefev, Mikko Aulamo, Marta Bañón, Maja Buljan, Laurie Burchell, Lucas Charpentier, Pinzhen Chen, Mariya Fedorova, Ona de Gibert, et al. Hplt 3.0: Very large- scale multilingual resources for llm and mt. mono-and bi-lingual data, multilingual evaluation, and pre-trained models.arXiv preprint arXiv:2511.01066, 2025

  28. [28]

    FineWeb2: One pipeline to scale them all — adapting pre-training data processing to every language

    Guilherme Penedo, Hynek Kydlí ˇcek, Vinko Sabol ˇcec, Bettina Messmer, Negar Foroutan, Amir Hossein Kargaran, Colin Raffel, Martin Jaggi, Leandro V on Werra, and Thomas Wolf. FineWeb2: One pipeline to scale them all — adapting pre-training data processing to every language. InSecond Conference on Language Modeling, 2025

  29. [29]

    Llämmlein: Transparent, compact and compet- itive german-only language models from scratch

    Jan Pfister, Julia Wunderle, and Andreas Hotho. Llämmlein: Transparent, compact and compet- itive german-only language models from scratch. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors,Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna...

  30. [30]

    Leolm: Igniting german-language llm research

    Björn Plüster. Leolm: Igniting german-language llm research. LAION Blog, September 2023. URLhttps://laion.ai/blog/leo-lm/. Accessed: May 5, 2026

  31. [31]

    Treviso, Nuno Miguel Guerreiro, Chrysoula Zerva, Ana C

    Ricardo Rei, Marcos V . Treviso, Nuno Miguel Guerreiro, Chrysoula Zerva, Ana C. Farinha, Christine Maroti, José G. C. de Souza, Taisiya Glushkova, Duarte M. Alves, Luísa Coheur, Alon Lavie, and André F. T. Martins. CometKiwi: Ist-unbabel 2022 submission for the quality 14 estimation shared task. In Philipp Koehn, Loïc Barrault, Ondrej Bojar, Fethi Bougare...

  32. [32]

    How good is your tokenizer? on the monolingual performance of multilingual language models

    Phillip Rust, Jonas Pfeiffer, Ivan Vulic, Sebastian Ruder, and Iryna Gurevych. How good is your tokenizer? on the monolingual performance of multilingual language models. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli, editors,Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joi...

  33. [33]

    Gottbert: a pure german language model

    Raphael Scheible, Johann Frei, Fabian Thomczyk, Henry He, Patric Tippmann, Jochen Knaus, Victor Jaravine, Frank Kramer, and Martin Boeker. Gottbert: a pure german language model. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors,Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, ...

  34. [34]

    Megatron-LM: Training multi-billion parameter language models using model parallelism, 2020

    Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron-LM: Training multi-billion parameter language models using model parallelism, 2020. arXiv:1909.08053

  35. [35]

    Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, Raymond Ng, Shayne Longpre, Sebastian Ruder, Wei-Yin Ko, Antoine Bosselut, Alice Oh, André F. T. Martins, Leshem Choshen, Daphne Ippolito, Enzo Ferrante, Marzieh Fadaee,...

  36. [36]

    Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh, Luke Zettlemoyer, Noah A

    Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Raghavi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Rav...

  37. [37]

    A monolingual approach to contextualized word embeddings for mid-resource languages

    Pedro Javier Ortiz Suárez, Laurent Romary, and Benoît Sagot. A monolingual approach to contextualized word embeddings for mid-resource languages. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel R. Tetreault, editors,Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 170...

  38. [38]

    Towards multilingual llm evaluation for european languages, 2024

    Klaudia Thellmann, Bernhard Stadler, Michael Fromm, Jasper Schulze Buschhoff, Alex Jude, Fabio Barth, Johannes Leveling, Nicolas Flores-Herr, Joachim Köhler, René Jäkel, and Mehdi Ali. Towards multilingual llm evaluation for european languages, 2024. URL https://arxiv. org/abs/2410.08928

  39. [39]

    A shocking amount of the web is machine translated: Insights from multi-way parallelism

    Brian Thompson, Mehak Preet Dhaliwal, Peter Frisch, Tobias Domhan, and Marcello Federico. A shocking amount of the web is machine translated: Insights from multi-way parallelism. 15 In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors,Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11...

  40. [40]

    Multilingual language model pretraining using machine- translated data

    Jiayi Wang, Yao Lu, Maurice Weber, Max Ryabinin, David Ifeoluwa Adelani, Yihong Chen, Raphael Tang, and Pontus Stenetorp. Multilingual language model pretraining using machine- translated data. In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng, editors,Proceedings of the 2025 Conference on Empirical Methods in Natural Langu...

  41. [41]

    mT5: A massively multilingual pre-trained text-to-text transformer

    Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. mT5: A massively multilingual pre-trained text-to-text transformer. In Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hakkani-Tür, Iz Beltagy, Steven Bethard, Ryan Cotterell, Tanmoy Chakraborty, and Yichao Zhou, editors, Procee...

  42. [42]

    Qwen3 technical report, 2025

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jing Zhou, Jingren Zhou, Junyang Lin, Kai Dang, Keqin Bao, Kexin Yang, ...

  43. [43]

    Mixed:" and describe the dominant themes. - Use evidence from multiple samples, not a single outlier. Return valid JSON only, with this schema: {

    Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. HellaSwag: can a machine really finish your sentence? In Anna Korhonen, David R. Traum, and Lluís Màrquez, editors,Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 4791...