REVIEW 4 major objections 4 minor 3 cited by
KletterMix, a German corpus produced by translating a curated English pretraining mixture, yields measurable gains on German reasoning benchmarks under matched token budgets.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 12:25 UTC pith:VQUXHD4O
load-bearing objection KletterMix is a genuinely useful, well-documented dataset artifact, but the improvement claim rests on single-run comparisons with no contamination check against the translated benchmarks. the 4 major comments →
KletterMix: Climbing Toward High-Quality German Pretraining Data - The Full Report
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
KletterMix is a 725B-token German-language corpus built by machine-translating the ClimbMix English pretraining mixture with a document-preserving pipeline: length-aware routing, contextualized chunking, dynamic output budgeting, and shard-wise execution. The paper claims that models pretrained from scratch on matched 12B-token KletterMix subsets reach lower training and validation loss than on FineWeb2-DE or GermanWeb and achieve the strongest four-task core average across MMLU, PIQA, HellaSwag, and ARC-C among the compared runs, with the best point estimate (40.2) on the validation-selected proxy-filtered split. The authors interpret the task pattern — consistent gains on HellaSwag and ARC
What carries the argument
The central mechanism is the document-preserving translation pipeline paired with a target-only quality proxy. Documents are routed into length buckets, translated whole or as contextualized chunks, and scored by COMETKiwi on a stratified pilot sample; a gradient-boosted regressor trained on German-only text features (length, language-identification signals, character ratios, repetition) then predicts COMETKiwi-like scores for every document. This proxy is what allows the full 725B-token corpus to be filtered into controlled 12B-token training splits and permits the cluster-level and length-bucket quality diagnostics.
Load-bearing premise
The load-bearing premise is that the benchmark gains on translated German evaluations reflect general German-language competence; the paper does not check whether training documents overlap with the translated evaluation sets, so a memorization-based advantage cannot be ruled out.
What would settle it
Run a document-level overlap analysis between the KletterMix training subset and the German MMLU, PIQA, HellaSwag, and ARC-C evaluation sets using normalized n-gram or embedding similarity; if high-similarity documents exist in the training corpus and removing them erases the HellaSwag/ARC-C gains, the central claim of reasoning transfer would be falsified. A complementary check is to evaluate the same checkpoints on a native-German benchmark suite not derived from English translation.
If this is right
- If the central claim holds, translated curated English mixtures provide a repeatable route to larger and more diverse German pretraining data without requiring additional native web crawling.
- Under a fixed 12B-token budget, KletterMix improves validation loss and reasoning-style benchmark performance at the tested scale, and the annealing result suggests it can act as a late-stage steering corpus after training on native German web data.
- Proxy-based filtering is a practical ranking signal for fixed-budget training mixtures: stricter thresholds improved the core average up to the validation-selected split, though not uniformly across every task.
- The aligned English-German document identifiers make KletterMix a reusable testbed for studying how translation quality, source cluster, and document length affect downstream model behavior.
Where Pith is reading between the lines
- The benchmark gains are measured on German evaluations that are themselves translations of English benchmarks; the paper does not report a document-level overlap check between the training corpus and the evaluation sets, so part of the advantage could reflect memorized or near-duplicate translated content rather than transferable ability.
- The reading that HellaSwag and ARC-C gains demonstrate 'reasoning transfer' is one interpretation; an equally testable one is that the translated corpus simply contains more structured exposition, which those benchmarks reward. A native-German reasoning suite and human naturalness judgments would separate the two.
- Because the proxy is target-only, filtering cannot catch failures that survive translation while leaving the German text looking clean; sampling a subset and scoring it with a source-aware quality model would bound how much quality signal is being left unused.
- The same pipeline could be applied to other languages, but whether the effect generalizes is unknown; if the gains shrink for a language typologically closer to English, mixture-structure transfer may be less important than translationese effects.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces KletterMix, a 725B-token German pretraining corpus obtained by machine-translating the English ClimbMix mixture while preserving document boundaries, metadata, and source-cluster structure. It documents a scalable translation pipeline (length-aware routing, chunking, dynamic target budgeting, shard-wise execution), uses COMETKiwi scores to train a target-only quality proxy, and releases unfiltered plus threshold-filtered variants. The empirical core is a set of matched 12B-token pretraining and annealing ablations with Qwen3-0.6B, comparing KletterMix against FineWeb2-DE and GermanWeb on four German translated benchmarks (MMLU, PIQA, HellaSwag, ARC-C). The central claim is that models trained on KletterMix achieve measurable improvements on German-language downstream evaluations, with the strongest point estimates on HellaSwag and ARC-C.
Significance. If the central claim holds, the paper makes a useful contribution to non-English pretraining data: it provides a large, documented, reproducible translated corpus, a transparent quality-estimation pipeline, and controlled training ablations under matched token budgets. The dataset artifact itself, with preserved document identifiers and metadata, is valuable for future studies of translation-based data curation. The paper also deserves credit for reporting evaluation-set standard errors, releasing code and data links, documenting translation failure modes qualitatively, and validating the proxy on a disjoint split. However, the strength of the empirical claim is currently limited by the absence of a contamination analysis, by single-run point estimates whose aggregate gaps are within evaluation noise, and by validation-based threshold selection on the same benchmarks used for the headline comparison. These issues are fixable within the manuscript's scope but are load-bearing for the abstract's claim of 'measurable improvements.'
major comments (4)
- [Sec. 5, Table 1] The headline comparison is not statistically decisive under the reported uncertainties. KletterMix-Filt0.60 has Core Avg. 40.2±1.3 vs. FineWeb2-DE 38.3±1.3, a gap of 1.9 points, which is about one standard error of the difference; unfiltered KletterMix (38.7±1.4) is within 0.4 points of FineWeb2-DE. Since the 0.60 threshold was selected after inspecting these same validation benchmarks, the 'best filtered variant' comparison is a selected-maximum result and inflates the chance of a false positive. The paper should either report multiple seeds with seed-level variance, or present the threshold choice as a hypothesis on a separate validation benchmark, or explicitly weaken the 'measurable improvements' claim to a suggestive single-run result.
- [Sec. 5 / Sec. 3 / Table 1] No contamination analysis is reported between KletterMix training documents and the translated evaluation benchmarks. KletterMix is a translation of ClimbMix, an English web mixture, and the four German benchmarks are themselves translations of English benchmarks whose items often originate from web text. If source documents for benchmark items appear in ClimbMix, their German translations will appear in KletterMix, and the consistent HellaSwag/ARC-C gains could reflect near-duplicate or memorized content rather than transferable German-language competence. The paper should report document-level or n-gram overlap between the training subset and each evaluation set, remove or flag overlapping items, and rerun the key comparisons. This is the weakest load-bearing link for the 'reasoning transfer' interpretation.
- [Sec. 5, Tab. 11 and Fig. 6] The in-domain validation-loss comparisons do not support the claim that KletterMix is 'not merely easier to fit.' Each model is evaluated on its own corpus's held-out validation set, so the lower KletterMix perplexity (6.02 vs. 10.04 for FineWeb2-DE and 8.50 for GermanWeb) can reflect corpus difficulty or distributional differences, not better transferable modeling. The filtered rows are also compared on the KletterMix validation set without a common held-out corpus. A meaningful validation comparison would require a common held-out German benchmark or a carefully matched cross-corpus validation set; otherwise the optimization-dynamics discussion should be limited to training loss.
- [Sec. 5 / App. A.5] The aggregate Core Avg. is dominated by PIQA's small evaluation set (100 examples) and wide standard error. The claim that 'the recurring task-level pattern is the more stable signal' is reasonable, but the task-level pattern is itself based on single runs. For the HellaSwag and ARC-C gains to support the central claim, the paper needs either seed variance estimates or a more robust evaluation protocol. At minimum, the abstract and conclusion should not state 'measurable improvements' as an established fact while the only quantitative support is a selected single-run point estimate within noise.
minor comments (4)
- [Sec. 5, Dataset interpretation] 'The MMMLU result' appears to be a typo for 'MMLU result.'
- [Sec. 3, Proxy-filtered dataset variants] The choice of thresholds 0.50/0.55/0.60 and the dynamic budget parameters α=2.0, β=1024 are presented as ablations, which is fine, but the main text could state more explicitly that these are not derived from a principled criterion and that the 0.60 threshold is validation-selected.
- [Sec. 5, Setup] The description of the deterministic token-budgeted stratified sampler is clear, but it would help to report the exact size of each evaluation-set subset (e.g., the 100-example PIQA subset) in the main table caption or in App. A.5, since this materially affects how readers should interpret the standard errors.
- [Sec. 6, Limitations] The Limitations section explicitly acknowledges single-run comparisons and evaluation-set-only standard errors, which is commendable, but it should also acknowledge the absence of a benchmark-contamination analysis, given that both the training corpus and evaluation sets are translations of English web-derived data.
Circularity Check
No significant circularity: KletterMix is built from an external English mixture, filtered by external COMETKiwi-derived scores, and evaluated with controlled ablations; no result reduces to its inputs by construction.
full rationale
I walked the paper's derivation chain: source corpus (ClimbMix, external), translation pipeline, COMETKiwi-based quality labeling, target-only proxy filtering, and matched pretraining ablations. No step defines its output in terms of the downstream claim. KletterMix is not constructed from benchmark results; the proxy is supervised by reference-free COMETKiwi scores, not by downstream accuracy; and the filtering thresholds (0.50/0.55/0.60) are fixed proxy-score cutoffs rather than fitted benchmark parameters. The downstream benchmark numbers are empirical measurements, not identities. The closest concern is the 'validation-selected filtered split' in Sec. 5, where the 0.60-filtered variant reaches the best Core Avg.; this could reflect threshold selection on the same benchmark outcomes and is a statistical selection effect, but it is not an equation-level reduction and the paper does not claim to predict the benchmark from the filter. The absence of a contamination analysis between translated training data and translated benchmarks is a validity/correctness risk, not circularity: any overlap would be a contingent data artifact rather than a definitional consequence. Self-citations to the authors' earlier German benchmark work provide evaluation instruments, but those are external published artifacts and are not used to force the paper's conclusion. No circular step meets the evidence bar.
Axiom & Free-Parameter Ledger
free parameters (2)
- Proxy filter thresholds =
0.50, 0.55, 0.60; 0.60 reported as best (validation-selected)
- Dynamic target budget coefficients (alpha, beta, min, Lmax) =
alpha=2.0, beta=1024, min=2048, Lmax=32768
axioms (5)
- domain assumption ClimbMix is a high-quality English pretraining mixture.
- domain assumption COMETKiwi reference-free scores are valid translation-quality labels.
- domain assumption The target-only proxy generalizes from the pilot subset to the full 725B-token corpus.
- domain assumption The German evaluation benchmarks measure German LM quality without contamination from translated training data.
- domain assumption Single-run training results are representative of the data-mixture comparison.
read the original abstract
High-quality pretraining data is a central ingredient in modern language models, but German-language resources remain far less developed than their English counterparts: they are often smaller, less carefully curated, weakly documented, and rarely validated through controlled training experiments. We introduce KletterMix, a high-quality German corpus for language model pretraining and annealing, designed as a reusable dataset artifact for the natural language processing and modeling community. KletterMix is built by translating a state-of-the-art English pretraining corpus into German while preserving document boundaries, metadata, source structure, and topical diversity. This construction yields a German corpus with the scale and diversity of a modern pretraining dataset, while enabling direct comparison to its English source. We document the dataset through a broad set of corpus-level analyses, including translation quality, document length distributions, topic coverage, source composition, and geographic metadata. Using COMETKiwi, we show that the translated documents achieve strong quality across diverse domains, suggesting that careful translation can preserve much of the semantic and stylistic richness of the original corpus. Beyond dataset construction, we evaluate KletterMix as training data. Through controlled pretraining and annealing ablations against established German corpora, we show that models trained on KletterMix achieve measurable improvements on German-language downstream evaluations. These results demonstrate that carefully curated translated data can substantially strengthen the German pretraining data ecosystem.
Figures
Forward citations
Cited by 3 Pith papers
-
A Sovereign, Open-Source Foundation Model for German and English
Soofi S 30B-A3B, a fully documented German–English hybrid Mamba-MoE base model, matches dense 14–27B peers on bilingual aggregates while delivering 8–9× long-context decode throughput.
-
A Sovereign, Open-Source Foundation Model for German and English
A fully documented German–English hybrid MoE base model matches dense 14–27B peers, leads open code scores, and sustains high long-context throughput at 3B active parameters.
-
A Sovereign, Open-Source Foundation Model for German and English
Soofi S 30B-A3B, a hybrid Mamba-MoE model pretrained on ~27T tokens with deliberately up-weighted German, reports the highest English and German aggregate scores among fully open base models in its comparison while ma...
Reference graph
Works this paper leans on
-
[1]
Mortensen, Noah A
Orevaoghene Ahia, Sachin Kumar, Hila Gonen, Jungo Kasai, David R. Mortensen, Noah A. Smith, and Yulia Tsvetkov. Do all languages cost the same? tokenization in the era of commer- cial language models. In Houda Bouamor, Juan Pino, and Kalika Bali, editors,Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Si...
2023
-
[2]
Teuken-7b-base & teuken-7b-instruct: Towards european llms
Mehdi Ali, Michael Fromm, Klaudia Thellmann, Jan Ebert, Alexander Arno Weber, Richard Rutmann, Charvi Jain, Max Lübbering, Daniel Steinigen, Johannes Leveling, Katrin Klug, Jasper Schulze Buschhoff, Lena Jurkschat, Hammam Abdelwahab, Benny Jörg Stein, Karl- Heinz Sylla, Pavel Denisov, Nicolo’ Brandizzi, Qasid Saleem, Anirban Bhowmick, Lennard Helmer, Chel...
2025
-
[3]
Occiglot at WMT24: european open-source large language models evaluated on translation
Eleftherios Avramidis, Annika Grützner-Zahn, Manuel Brack, Patrick Schramowski, Pedro Ortiz Suarez, Malte Ostendorff, Fabio Barth, Shushen Manakhimova, Vivien Macketanz, Georg Rehm, and Kristian Kersting. Occiglot at WMT24: european open-source large language models evaluated on translation. In Barry Haddow, Tom Kocmi, Philipp Koehn, and Christof Monz, ed...
2024
-
[4]
Bender and Batya Friedman
Emily M. Bender and Batya Friedman. Data statements for natural language processing: Toward mitigating system bias and enabling better science.Trans. Assoc. Comput. Linguistics, 6:587–604, 2018
2018
-
[5]
Burns, Letitia Parcalabescu, Stephan Wäldchen, Michael Barlow, Gregor Ziegltrum, V olker Stampa, Bastian Harren, and Björn Deiseroth
Thomas F. Burns, Letitia Parcalabescu, Stephan Wäldchen, Michael Barlow, Gregor Ziegltrum, V olker Stampa, Bastian Harren, and Björn Deiseroth. Aleph-alpha-germanweb: Improving german-language LLM pre-training with model-based data curation and synthetic data gen- eration. In Vera Demberg, Kentaro Inui, and Lluís Marquez, editors,Proceedings of the 19th C...
2026
-
[6]
Tyler A. Chang, Catherine Arnett, Abdelrahman Eldesokey, Abdelrahman Sadallah, Abeer Kashar, Abolade Daud, Abosede Grace Olanihun, Adamu Labaran Mohammed, Adeyemi Praise, Adhikarinayum Meerajita Sharma, Aditi Gupta, Afitab Iyigun, Afonso Simplício, Ahmed Essouaied, Aicha Chorana, Akhil Eppa, Akintunde Oladipo, Akshay Ramesh, Aleksei Dorkin, 11 Alfred Male...
Pith/arXiv arXiv 2025
-
[7]
Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018. arXiv:1803.05457
Pith/arXiv arXiv 2018
-
[8]
Smith, Ahmad Idrissi-Yaghir, Constantin Seibold, Jianning Li, Lars Heiliger, Christoph M
Amin Dada, Aokun Chen, Cheng Peng, Kaleb E. Smith, Ahmad Idrissi-Yaghir, Constantin Seibold, Jianning Li, Lars Heiliger, Christoph M. Friedrich, Daniel Truhn, Jan Egger, Jiang Bian, Jens Kleesiek, and Yonghui Wu. On the impact of cross-domain data on german language models. In Houda Bouamor, Juan Pino, and Kalika Bali, editors,Findings of the Association ...
2023
-
[9]
A new massive multilingual dataset for high- performance language technologies
Ona de Gibert, Graeme Nail, Nikolay Arefyev, Marta Bañón, Jelmer van der Linde, Shaoxiong Ji, Jaume Zaragoza-Bernabeu, Mikko Aulamo, Gema Ramírez-Sánchez, Andrey Kutuzov, Sampo Pyysalo, Stephan Oepen, and Jörg Tiedemann. A new massive multilingual dataset for high- performance language technologies. In Nicoletta Calzolari, Min-Yen Kan, Véronique Hoste, Al...
2024
-
[10]
WMT24++: Expanding the language coverage of WMT24 to 55 languages & dialects
Daniel Deutsch, Eleftheria Briakou, Isaac Rayburn Caswell, Mara Finkelstein, Rebecca Galor, Juraj Juraska, Geza Kovacs, Alison Lui, Ricardo Rei, Jason Riesa, Shruti Rijhwani, Parker Riley, Elizabeth Salesky, Firas Trabelsi, Stephanie Winkler, Biao Zhang, and Markus Freitag. WMT24++: Expanding the language coverage of WMT24 to 55 languages & dialects. In F...
2025
-
[11]
Shizhe Diao, Yu Yang, Yonggan Fu, Xin Dong, Dan Su, Markus Kliegl, Zijia Chen, Peter Belcak, Yoshi Suhara, Hongxu Yin, Mostofa Patwary, Yingyan, Lin, Jan Kautz, and Pavlo Molchanov. Nemotron-climb: Clustering-based iterative data mixture bootstrapping for language model pre-training, 2025. arXiv:2504.13161
Pith/arXiv arXiv 2025
-
[12]
Documenting large webtext corpora: A case study on the colossal clean crawled corpus
Jesse Dodge, Maarten Sap, Ana Marasovic, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. Documenting large webtext corpora: A case study on the colossal clean crawled corpus. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih, editors,Proceedings of the 2021 Conference on Empirical Methods in...
2021
-
[13]
Pretraining language models using translationese
Meet Doshi, Raj Dabre, and Pushpak Bhattacharyya. Pretraining language models using translationese. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors,Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, pages 5843–5862. Association for Computational Linguistic...
2024
-
[14]
The pile: An 800gb dataset of diverse text for language modeling, 2020
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. The pile: An 800gb dataset of diverse text for language modeling, 2020. arXiv:2101.00027
Pith/arXiv arXiv 2020
-
[15]
The language model evaluation harness, 07 2024
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. The languag...
2024
-
[16]
Wallach, Hal Daumé III, and Kate Crawford
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna M. Wallach, Hal Daumé III, and Kate Crawford. Datasheets for datasets.Commun. ACM, 64(12): 86–92, 2021
2021
-
[17]
The german commons - 154 billion tokens of openly licensed text for german language models, 2025
Lukas Gienapp, Christopher Schröder, Stefan Schweter, Christopher Akiki, Ferdinand Schlatt, Arden Zimmermann, Phillipe Genêt, and Martin Potthast. The german commons - 154 billion tokens of openly licensed text for german language models, 2025. arXiv:2510.13996
arXiv 2025
-
[18]
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. In9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net, 2021
2021
-
[19]
Rae, and Laurent Sifre
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katherine Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Oriol Vinyals, Jack W. Rae, and Laurent S...
2022
-
[20]
Glotlid: Language identification for low-resource languages
Amir Hossein Kargaran, Ayyoob Imani, François Yvon, and Hinrich Schütze. Glotlid: Language identification for low-resource languages. In Houda Bouamor, Juan Pino, and Kalika Bali, editors,Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, Findings of ACL, pages 6155–6218. Association for Computational Li...
2023
-
[21]
Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii- Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitriou, Salomey Osei, Pedro Javier Ortiz Suárez, Iroro Orife, Kelechi Ogueji, An- d...
2022
-
[22]
Yamshchikov
Pierre-Carl Langlais, Pavel Chizhov, Catherine Arnett, Carlos Rosas Hinostroza, Mattia Nee, Eliot Krzysztof Jones, Irène Girard, David Mach, Anastasia Stasenko, and Ivan P. Yamshchikov. Common corpus: The largest collection of ethical data for LLM pre-training. InThe Fourteenth International Conference on Learning Representations, 2026
2026
-
[23]
The bigscience ROOTS corpus: A 1.6tb composite multilingual dataset
Hugo Laurençon, Lucile Saulnier, Thomas Wang, Christopher Akiki, Albert Villanova del Moral, Teven Le Scao, Leandro von Werra, Chenghao Mou, Eduardo González Ponferrada, Huu Nguyen, Jörg Frohberg, Mario Sasko, Quentin Lhoest, Angelina McMillan-Major, Gérard Dupont, Stella Biderman, Anna Rogers, Loubna Ben Allal, Francesco De Toni, Giada Pistilli, Olivier ...
2022
-
[24]
Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Yitzhak Gadre, Hritik Bansal, Etash Kumar Guha, Sedrick Scott Keh, Kushal Arora, Saurabh Garg, Rui Xin, Niklas Muennighoff, Reinhard Heckel, Jean Mercat, Mayee F. Chen, Suchin Gururangan, Mitchell Wortsman, Alon Albalak, Yonatan Bitton, Marianna Nezhurina, Amro Abbas, Cheng-Yu Hsieh, D...
2024
-
[25]
Guerreiro, Ricardo Rei, Duarte M
Pedro Henrique Martins, Patrick Fernandes, João Alves, Nuno M. Guerreiro, Ricardo Rei, Duarte M. Alves, José Pombal, Amin Farajian, Manuel Faysse, Mateusz Klimaszewski, Pierre Colombo, Barry Haddow, José G. C. de Souza, Alexandra Birch, and André F. T. Martins. EuroLLM: Multilingual language models for europe, 2024. arXiv:2409.16235
Pith/arXiv arXiv 2024
-
[26]
Rossi, and Thien Huu Nguyen
Thuat Nguyen, Chien Van Nguyen, Viet Dac Lai, Hieu Man, Nghia Trung Ngo, Franck Dernoncourt, Ryan A. Rossi, and Thien Huu Nguyen. CulturaX: A cleaned, enormous, and multilingual dataset for large language models in 167 languages. In Nicoletta Calzolari, Min-Yen Kan, Véronique Hoste, Alessandro Lenci, Sakriani Sakti, and Nianwen Xue, editors, Proceedings o...
2024
-
[27]
Hplt 3.0: Very large- scale multilingual resources for llm and mt
Stephan Oepen, Nikolay Arefev, Mikko Aulamo, Marta Bañón, Maja Buljan, Laurie Burchell, Lucas Charpentier, Pinzhen Chen, Mariya Fedorova, Ona de Gibert, et al. Hplt 3.0: Very large- scale multilingual resources for llm and mt. mono-and bi-lingual data, multilingual evaluation, and pre-trained models.arXiv preprint arXiv:2511.01066, 2025
Pith/arXiv arXiv 2025
-
[28]
FineWeb2: One pipeline to scale them all — adapting pre-training data processing to every language
Guilherme Penedo, Hynek Kydlí ˇcek, Vinko Sabol ˇcec, Bettina Messmer, Negar Foroutan, Amir Hossein Kargaran, Colin Raffel, Martin Jaggi, Leandro V on Werra, and Thomas Wolf. FineWeb2: One pipeline to scale them all — adapting pre-training data processing to every language. InSecond Conference on Language Modeling, 2025
2025
-
[29]
Llämmlein: Transparent, compact and compet- itive german-only language models from scratch
Jan Pfister, Julia Wunderle, and Andreas Hotho. Llämmlein: Transparent, compact and compet- itive german-only language models from scratch. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors,Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna...
2025
-
[30]
Leolm: Igniting german-language llm research
Björn Plüster. Leolm: Igniting german-language llm research. LAION Blog, September 2023. URLhttps://laion.ai/blog/leo-lm/. Accessed: May 5, 2026
2023
-
[31]
Treviso, Nuno Miguel Guerreiro, Chrysoula Zerva, Ana C
Ricardo Rei, Marcos V . Treviso, Nuno Miguel Guerreiro, Chrysoula Zerva, Ana C. Farinha, Christine Maroti, José G. C. de Souza, Taisiya Glushkova, Duarte M. Alves, Luísa Coheur, Alon Lavie, and André F. T. Martins. CometKiwi: Ist-unbabel 2022 submission for the quality 14 estimation shared task. In Philipp Koehn, Loïc Barrault, Ondrej Bojar, Fethi Bougare...
2022
-
[32]
How good is your tokenizer? on the monolingual performance of multilingual language models
Phillip Rust, Jonas Pfeiffer, Ivan Vulic, Sebastian Ruder, and Iryna Gurevych. How good is your tokenizer? on the monolingual performance of multilingual language models. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli, editors,Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joi...
2021
-
[33]
Gottbert: a pure german language model
Raphael Scheible, Johann Frei, Fabian Thomczyk, Henry He, Patric Tippmann, Jochen Knaus, Victor Jaravine, Frank Kramer, and Martin Boeker. Gottbert: a pure german language model. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors,Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, ...
2024
-
[34]
Megatron-LM: Training multi-billion parameter language models using model parallelism, 2020
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. Megatron-LM: Training multi-billion parameter language models using model parallelism, 2020. arXiv:1909.08053
Pith/arXiv arXiv 2020
-
[35]
Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, Raymond Ng, Shayne Longpre, Sebastian Ruder, Wei-Yin Ko, Antoine Bosselut, Alice Oh, André F. T. Martins, Leshem Choshen, Daphne Ippolito, Enzo Ferrante, Marzieh Fadaee,...
2025
-
[36]
Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh, Luke Zettlemoyer, Noah A
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Raghavi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Rav...
2024
-
[37]
A monolingual approach to contextualized word embeddings for mid-resource languages
Pedro Javier Ortiz Suárez, Laurent Romary, and Benoît Sagot. A monolingual approach to contextualized word embeddings for mid-resource languages. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel R. Tetreault, editors,Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 170...
2020
-
[38]
Towards multilingual llm evaluation for european languages, 2024
Klaudia Thellmann, Bernhard Stadler, Michael Fromm, Jasper Schulze Buschhoff, Alex Jude, Fabio Barth, Johannes Leveling, Nicolas Flores-Herr, Joachim Köhler, René Jäkel, and Mehdi Ali. Towards multilingual llm evaluation for european languages, 2024. URL https://arxiv. org/abs/2410.08928
Pith/arXiv arXiv 2024
-
[39]
A shocking amount of the web is machine translated: Insights from multi-way parallelism
Brian Thompson, Mehak Preet Dhaliwal, Peter Frisch, Tobias Domhan, and Marcello Federico. A shocking amount of the web is machine translated: Insights from multi-way parallelism. 15 In Lun-Wei Ku, Andre Martins, and Vivek Srikumar, editors,Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11...
2024
-
[40]
Multilingual language model pretraining using machine- translated data
Jiayi Wang, Yao Lu, Maurice Weber, Max Ryabinin, David Ifeoluwa Adelani, Yihong Chen, Raphael Tang, and Pontus Stenetorp. Multilingual language model pretraining using machine- translated data. In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng, editors,Proceedings of the 2025 Conference on Empirical Methods in Natural Langu...
2025
-
[41]
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. mT5: A massively multilingual pre-trained text-to-text transformer. In Kristina Toutanova, Anna Rumshisky, Luke Zettlemoyer, Dilek Hakkani-Tür, Iz Beltagy, Steven Bethard, Ryan Cotterell, Tanmoy Chakraborty, and Yichao Zhou, editors, Procee...
2021
-
[42]
An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jing Zhou, Jingren Zhou, Junyang Lin, Kai Dang, Keqin Bao, Kexin Yang, ...
Pith/arXiv arXiv 2025
-
[43]
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. HellaSwag: can a machine really finish your sentence? In Anna Korhonen, David R. Traum, and Lluís Màrquez, editors,Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 4791...
arXiv 2019
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.