REVIEW 3 major objections 7 minor 7 cited by
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
T0 review · 3 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read One adaptive data-curation pipeline, whose filters and thresholds are set automatically from each language's own text statistics, produces pre-training corpora that outperform prior multilingual datasets on 11 of 14 languages tested.
desk verdict Strong engineering contribution with a solid internal ablation; the headline 11/14 claim needs an epoch-matched re-run before it can be taken at face value. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying object is the FineWeb2 pipeline itself: a staged sequence in which every threshold is a function of statistics measured on the target language rather than a global constant. Its stages are GlotLID language identification with a per-language confidence cutoff; per-language MinHash deduplication using language-appropriate word tokenizers; heuristic filtering whose thresholds are adapted from FineWeb's English values to each language by distribution-matching (the 10Tail, Quantile, and MeanStd methods); a high-affinity wordlist filter for low-resource languages; and rehydration, which upsamples kept documents by duplication count with weights derived from filtering rates per cluster size as a quality proxy. What makes this a single mechanism for all languages is the claim that every adaptation is computed automatically from the language's own data, so no human tuning is needed for any of the 1,868 language-script pairs in the final dataset.
What would settle it
Two concrete checks would settle the claim. First, take a language outside the nine canary languages with enough data and a reliable benchmark (say, Korean, Ukrainian, or Persian), train identical small models on corpora kept at the formula-selected LID cutoff and at neighbouring cutoffs, and map the performance curve: each language where the formula's cutoff falls outside the highest-scoring range — as the paper's Table 15 already shows for Chinese and Hindi — weakens the claim of automatic adaptation to every language. Second, because the paper's evaluations rely on early-signal tasks that may change character with longer training, re-run a subset of the 14-language comparison at larger token budgets and check whether FineWeb2's 11-of-14 advantage persists.
Extended reading notes
Core claim
The paper's central claim is that a consistent, adaptive pipeline can replace hand-designed per-language processing: every stage derives its settings from the target language's own data instead of using fixed values for all languages. Language identification uses GlotLID, which covers 1,880 languages and labels script variants separately, and each language's confidence cutoff is set by the formula $\max\{0.3, \min\{0.9, \mathrm{Med}(X)-\sigma(X)\}\}$ on the distribution $X$ of that language's LID scores. Word tokenization, needed for filtering, deduplication, and evaluation, comes from native tokenizers where they exist and from proxy tokenizers propagated through the language-family tree otherwise. Documents are deduplicated globally per language with MinHash, and heuristic filters inherited from the English FineWeb pipeline have their thresholds re-derived by distribution-matching methods (10Tail, Quantile, MeanStd) so each language removes a comparable, quality-relevant fraction of data; low-resource languages additionally pass a high-affinity wordlist filter to strip misclassified high-resource content. The final stage, rehydration, uses filtering rates per MinHash cluster size as a quality proxy and upsamples documents with cluster-dependent weights, favouring neither unique nor massively duplicated documents. On this basis the paper concludes that models trained on FineWeb2 corpora are more performant than models trained on prior multilingual datasets on 11 of 14 evaluated languages, with each pipeline stage contributing a measured improvement, and that the same pipeline scales to 1,868 language-script pairs in the released 20-terabyte corpus.
Load-bearing premise
The load-bearing premise is that the automatic cutoff formula for language identification, which decides how much low-confidence text each language keeps, was validated on only nine languages — and on two of those nine (Chinese and Hindi) it already selects a cutoff outside the best-performing range (Section 4.2, Table 15) — yet it is applied without per-language checks to the other roughly 1,859 language-script pairs, so any gain or loss for those languages rests on the assumption that their confidence-score distributions resemble the seven languages where the formula worked.
Editorial extensions
If this is right
- Teams training monolingual or multilingual LLMs can use the released pipeline and corpus directly, replacing hand-built per-language workflows with a single runnable recipe that matches or exceeds the early-signal performance of prior public multilingual web corpora.
- The rehydration scheme turns an otherwise destructive deduplication step into a tunable one: because filter-based quality weights are cheap, any future deduplicated dataset that retains cluster-size metadata can apply the same upsampling trick.
- The task-selection criteria (monotonicity, signal-to-noise ratio, non-randomness, ordering consistency) give a reusable, quantitative way to pick evaluation benchmarks for languages that lack validated ones, reducing reliance on noisy translated tasks for early-training comparisons.
- For low-resource languages, the released corpora enable continued pre-training experiments even though most small language corpora are dominated by Bible or Wikipedia content; the paper's own audit finds that 70% of language-script pairs have over half their documents from those domains.
- Combining FineWeb2 with FineWeb yields a single open, English-plus-everything-else pre-training mixture at 20TB scale, giving non-English languages the same kind of carefully filtered web data that English already had.
Reading between the lines
- If the adaptive thresholds generalize as claimed, the pipeline's adaptation logic could be rerun on each new Common Crawl snapshot release as a scheduled, fully automatic operation, making multilingual dataset refresh a maintenance task rather than a research project — the paper does not claim this, but it follows directly from the design.
- The formula's known misses on Chinese and Hindi (Table 15 of the paper) open a measurable improvement path the authors did not take: adding a validation loop on more languages, or making the cutoff depend on distribution shape such as skewness, could recover the lost performance on those languages and reduce risk on the unvalidated ones.
- The dominance of Bible and Wikipedia sources in the long tail implies that for very-low-resource languages the binding constraint is source diversity, not pipeline quality; a testable extension is to measure whether filtering and rehydration gains saturate for corpora under a few thousand documents, in which case augmenting with curated or synthetic data would matter more than further pipeline tun
- The early-signal selection metrics could double as diagnostic audits of existing multilingual benchmarks: tasks that fail monotonicity or signal-to-noise thresholds on reference runs could be flagged and retired, which would change how multilingual model comparisons are reported.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces FineWeb2, a multilingual pre-training data curation pipeline that adapts FineWeb-style filtering, deduplication, and rehydration to over 1,000 languages. The pipeline is validated on nine canary languages with 207 small-model ablation runs and compared against CC-100, mC4, CulturaX, HPLT, HPLT2, raw Common Crawl, and several language-specific corpora on both the canary languages and five unseen languages. The authors report that FineWeb2 outperforms prior multilingual datasets on 11 of 14 languages, and release the 20TB dataset, the pipeline code, and training/evaluation codebases.
Significance. If the central claims hold, this is a substantial contribution to multilingual pre-training infrastructure: a released 20TB corpus, an automatically adaptable pipeline, and a rigorous ablation methodology with a novel, quantitatively motivated task-selection procedure. The unseen-language evaluation (de, id, it, ja, vi) provides genuinely independent evidence that the pipeline generalizes beyond the languages used for design decisions. The paper also contains useful transparency, including a limitation discussion and a release of the pre-filtering intermediate dataset. However, the headline 11/14 superiority claim is currently weakened by two load-bearing issues: epoch-count confounds in low-resource language comparisons and an unvalidated LID-threshold formula applied to nearly 1,900 language-script pairs.
major comments (3)
- [§3.2, §5, Figure 7, Table 45] The Telugu and Swahili comparisons are confounded by different numbers of epochs. Section 3.2 states that for Telugu and Swahili every baseline dataset except raw Common Crawl contained only a limited amount of data, while Table 45 shows FineWeb2 contains 14.4GB for Telugu and 3.1GB for Swahili. All models are trained for 29–30B tokens, so the CC-100/mC4/CulturaX/HPLT/HPLT2 models must pass over their unique tokens many more times than the FineWeb2 models. Since repeated exposure to a small corpus can change early-signal benchmark behavior, the FineWeb2 wins on these two languages cannot be cleanly attributed to corpus curation quality. Removing or controlling those comparisons reduces the reported 11/14 to 9/14, which is materially weaker. I request an epoch-matched re-run or an explicit analysis showing that the epoch-count difference does not explain the observed gains.
- [§4.2, Table 15] The LID threshold formula max{0.3, min{0.9, Med(X)−σ(X)}} is applied to all 1,868 language-script pairs, but Table 15 shows that on two of the nine canary languages (Chinese and Hindi) this formula selects a threshold outside the highest-performing range. The paper provides no validation of this formula on the remaining roughly 1,859 languages, so any quality gain or loss for those languages rests on an untested assumption that confidence-score distributions behave like the seven languages where the formula worked. Please either validate the formula on a held-out set of languages with downstream evaluations, provide a sensitivity analysis showing the choice is not load-bearing, or clearly qualify the pipeline's per-language claims.
- [§5, Figure 3, §A.10.1] The headline claim that FineWeb2 produces more performant models on 11 of 14 languages lacks uncertainty quantification. Each dataset/language comparison appears to be a single training run, and the approximate standard deviations reported for unseen-language tasks in Section A.10.1 are large enough that some of the aggregate differences in Figure 3 may be within noise. I request bootstrap confidence intervals across tasks or multiple training seeds for at least the key comparisons, or alternatively a softened statement that does not assert superiority on a per-language basis without such evidence.
minor comments (7)
- [§5] Typo: 'multilingal' should be 'multilingual' in the sentence 'FineWeb2 produces more performant models than prior multilingal datasets'.
- [§4.2] The threshold formula uses X without a formal definition; please state explicitly that X is the random variable of GlotLID confidence scores for documents initially classified as the target language.
- [Figure 7 caption] The caption says all models were trained for 30 billion tokens, while Section 5 and Table 4 say 29 billion tokens; please align these numbers.
- [§A.3] Typo: 'paramaters' should be 'parameters' in the sentence about embedding layer size.
- [Table 15] The notation '0.8 - 0.9' for the Russian range is ambiguous; please use an explicit interval such as [0.8, 0.9] or clarify that it means all tested thresholds in that span.
- [§3.3 and §A.5.3] The paper states that 84 benchmarks were selected out of 197 tested, but the per-language tables in A.5.3 appear to list a smaller number; please add a summary table or count verification.
- [§4.5] The rehydration weight assignment is described as based on filtering rates with interpolation between two endpoints, but the exact interpolation formula is not given; please provide the formula or a precise algorithmic description.
Circularity Check
No significant circularity: the pipeline is validated out-of-sample on unseen languages; canary tuning is transparently in-sample and the central generalization claim does not reduce to the tuning inputs.
full rationale
The paper's derivation chain is empirical rather than definitional: it selects pipeline components (GlotLID, per-language LID thresholds, filter adaptation methods, deduplication, rehydration weights) using ablations on nine canary languages, then evaluates the resulting corpora both on those canary languages and on five unseen languages (German, Indonesian, Italian, Japanese, Vietnamese) trained for 100B tokens against external baselines (CC-100, mC4, CulturaX, HPLT, HPLT2, raw Common Crawl). No equation in the paper defines the claimed output in terms of the input: the LID threshold formula max{0.3, min{0.9, Med(X)-sigma(X)}} is a heuristic fitted to canary score distributions and applied to 1,868 language-script pairs; its failure for Chinese and Hindi (Table 15) shows it is not forced by construction. The 11/14 headline includes eight canary-language wins, and those are in-sample because the canary languages were used for design decisions; however, the paper explicitly does not present canary results as the generalization test ('evaluating on unseen languages validates that the pipeline generalizes effectively'), and the unseen-language comparisons provide genuine out-of-sample support, with FineWeb2 winning on Italian, Japanese, and Vietnamese at 100B tokens. Self-citations to FineWeb (Penedo et al., 2024) provide the starting pipeline and MinHash hyperparameters, but the load-bearing filtering, deduplication, and rehydration steps are ablated in this paper, so the citations are not carrying the central claim. The epoch-count imbalance for Telugu and Swahili baselines is a confound for those two comparisons, but it is a correctness/robustness concern, not a circularity. Overall, no prediction reduces by construction to its inputs; score 1 reflects only the mild in-sample component of the composite headline.
Assumptions & free parameters
free parameters (6)
- LID threshold formula =
Median - sigma, clipped to [0.3, 0.9]; e.g., Arabic 0.8812, Chinese 0.7415, Swahili 0.3
- Rehydration upsampling weights =
Weight 10 at lowest removal-rate cluster size, weight 1 above global removal rate, linear interpolation between
- Filter adaptation method per group =
fwq: 10Tail on Wikipedia; goq: Quantile on Wikipedia/GlotLID-Corpus; gor: MeanStd on Common Crawl
- Affinity threshold gamma for precision wordlists =
0.85
- Task selection thresholds =
Monotonicity rho >= 0.5, SNR >= 20 (30 for generative), non-randomness >= 3
- Stopword frequency threshold =
Adjusted so at least 8 stopwords remain after removing symbols and numbers
assumptions (6)
- domain assumption Early-signal evaluation validity: quality at 29B or 100B tokens on 1.46B-parameter models predicts relative data quality at larger pre-training scales.
- domain assumption GlotLID V3 language labels and confidence scores are accurate enough across roughly 1,880 languages for the downstream pipeline.
- ad hoc to paper The LID threshold formula max{0.3, min{0.9, Med(X)-sigma(X)}} generalizes beyond the nine canary languages.
- ad hoc to paper Filtering removal rate is a valid proxy for duplicate-cluster quality for rehydration weighting.
- domain assumption Proxy word tokenizers assigned via Ethnologue language family propagate adequately for filtering, deduplication, and evaluation.
- domain assumption MinHash hyperparameters from FineWeb (14 buckets of size 8, 5-grams) transfer to all languages.
Cite this review
Pith. "Pith review of FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language." pith.science (2026). https://pith.science/paper/R3TB4SIM
@misc{pith2026250620920,
author = {Pith},
title = {Pith review of: FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language},
year = {2026},
howpublished = {\url{https://pith.science/paper/R3TB4SIM}},
note = {Machine review of arXiv:2506.20920}
}
read the original abstract
Pre-training state-of-the-art large language models (LLMs) requires vast amounts of clean and diverse text data. While the open development of large high-quality English pre-training datasets has seen substantial recent progress, training performant multilingual LLMs remains a challenge, in large part due to the inherent difficulty of tailoring filtering and deduplication pipelines to a large number of languages. In this work, we introduce a new pre-training dataset curation pipeline based on FineWeb that can be automatically adapted to support any language. We extensively ablate our pipeline design choices on a set of nine diverse languages, guided by a set of meaningful and informative evaluation tasks that were chosen through a novel selection process based on measurable criteria. Ultimately, we show that our pipeline can be used to create non-English corpora that produce more performant models than prior datasets. We additionally introduce a straightforward and principled approach to rebalance datasets that takes into consideration both duplication count and quality, providing an additional performance uplift. Finally, we scale our pipeline to over 1000 languages using almost 100 Common Crawl snapshots to produce FineWeb2, a new 20 terabyte (5 billion document) multilingual dataset which we release along with our pipeline, training, and evaluation codebases.
Figures
Figures from the paper (6 more)
Forward citations
Cited by 7 Pith papers
-
In-Place Tokenizer Expansion for Pre-trained LLMs
Continuing a model's own BPE merges and training only new embedding rows preserves quality while cutting token counts 2.4–4× for previously under-tokenized languages.
-
The Effect of Scripts and Formats on LLM Numeracy
LLM arithmetic accuracy falls sharply when numerals leave the familiar Hindu–Arabic format, and few-shot prompting with examples narrows most of that gap.
-
From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation
KMMLU-Redux and KMMLU-Pro are new Korean benchmark datasets from national technical and professional licensure exams, with LLM evaluations reported against official pass thresholds.
-
OpenLID-v3: Improving the Precision of Closely Related Language Identification -- An Experience Report
OpenLID-v3 matches or improves precision on closely related language identification by adding data, merging variants, and adding a noise class, while ensembling with GlotLID buys more precision at the cost of coverage.
-
A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data
A metric-oriented survey that classifies intrinsic quality and trustworthiness metrics for LLM-generated data across six modalities and documents systematic evaluation gaps in the current literature.
-
Entropy2Vec: Crosslingual Language Modeling Entropy as End-to-End Learnable Language Representations
Entropy2Vec turns the cross-lingual surprise of monolingual language models into dense language embeddings that resemble typological families and match curated vectors in downstream tasks.
-
Observation of momentum dependent charge density wave gap in EuTe4
EuTe4 shows a momentum-dependent charge density wave gap at the Fermi level, largest along Gamma-Y and smallest along Gamma-X, plus a low-temperature magnetic phase diagram near TN = 6.9 K.
Reference graph
Works this paper leans on
-
[1]
Yi: Open foundation models by 01.ai, 2025
01.AI , Alex Young, Bei Chen, Chao Li, Chengen Huang, Ge Zhang, Guanwei Zhang, Guoyin Wang, Heng Li, Jiangcheng Zhu, Jianqun Chen, Jing Chang, Kaidong Yu, Peng Liu, Qiang Liu, Shawn Yue, Senbin Yang, Shiming Yang, Wen Xie, Wenhao Huang, Xiaohui Hu, Xiaoyi Ren, Xinyao Niu, Pengcheng Nie, Yanpeng Li, Yuchi Xu, Yudong Liu, Yue Wang, Yuxuan Cai, Zhenyu Gu, Zh...
arXiv 2025
-
[2]
Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, Alon Benhaim, Misha Bilenko, Johan Bjorck, Sébastien Bubeck, Martin Cai, Qin Cai, Vishrav Chaudhary, Dong Chen, Dongdong Chen, Weizhu Chen, Yen-Chun Chen, Yi-Ling Chen, Hao Cheng, Parul Chopra, Xiyang Dai, Matt...
arXiv 2024
-
[3]
Llama 3 model card
AI@Meta. Llama 3 model card. 2024. URL https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md
2024
-
[4]
A survey on data selection for language models
Alon Albalak, Yanai Elazar, Sang Michael Xie, Shayne Longpre, Nathan Lambert, Xinyi Wang, Niklas Muennighoff, Bairu Hou, Liangming Pan, Haewon Jeong, et al. A survey on data selection for language models. Transactions on Machine Learning Research, 2024
2024
-
[5]
Open llm turkish leaderboard v0.2
Mohamad Alhajar. Open llm turkish leaderboard v0.2. https://huggingface.co/spaces/malhajar/OpenLLMTurkishLeaderboard, 2024
2024
-
[6]
A l G hafa evaluation benchmark for A rabic language models
Ebtesam Almazrouei, Ruxandra Cojocaru, Michele Baldo, Quentin Malartic, Hamza Alobeidli, Daniele Mazzotta, Guilherme Penedo, Giulia Campesan, Mugariya Farooq, Maitha Alhammadi, Julien Launay, and Badreddine Noune. A l G hafa evaluation benchmark for A rabic language models. In Hassan Sawaf, Samhaa El-Beltagy, Wajdi Zaghouani, Walid Magdy, Ahmed Abdelali, ...
2023
-
[7]
101 billion arabic words dataset, 2024
Manel Aloui, Hasna Chouikhi, Ghaith Chaabane, Haithem Kchaou, and Chehir Dhaouadi. 101 billion arabic words dataset, 2024
2024
-
[8]
On the cross-lingual transferability of monolingual representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. On the cross-lingual transferability of monolingual representations. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, 2020 a . doi:10.18653/v1/2020.acl-main.421. URL http://dx.doi.org/10.18653/v1/2020.acl-main.421
Show all 118 references
-
[9]
A call for more rigor in unsupervised cross-lingual learning
Mikel Artetxe, Sebastian Ruder, Dani Yogatama, Gorka Labaka, and Eneko Agirre. A call for more rigor in unsupervised cross-lingual learning. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp.\ 7375–7388. Association for Computationa...
2020 doi
-
[10]
Japanese massive multitask language understanding benchmark, 2023
Kawahara Lab at Waseda University. Japanese massive multitask language understanding benchmark, 2023. URL https://huggingface.co/datasets/nlp-waseda/JMMLU
2023
-
[11]
The belebele benchmark: a parallel reading comprehension dataset in 122 language variants
Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, and Madian Khabsa. The belebele benchmark: a parallel reading comprehension dataset in 122 language variants. In Proceedings of ...
2024
-
[12]
Building machine translation systems for the next thousand languages, 2022
Ankur Bapna, Isaac Caswell, Julia Kreutzer, Orhan Firat, Daan van Esch, Aditya Siddhant, Mengmeng Niu, Pallavi Baljekar, Xavier Garcia, Wolfgang Macherey, Theresa Breiner, Vera Axelrod, Jason Riesa, Yuan Cao, Mia Xu Chen, Klaus Macherey, Maxim Krikun, Pidong Wang, Alexander Gu...
2022 arXiv
-
[13]
Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction
Adrien Barbaresi. Trafilatura: A Web Scraping Library and Command-Line Tool for Text Discovery and Extraction . In Proceedings of the Joint Conference of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Nat...
2021
-
[14]
On the resemblance and containment of documents
Andrei Z Broder. On the resemblance and containment of documents. In Proceedings. Compression and Complexity of SEQUENCES 1997 (Cat. No. 97TB100171), pp.\ 21--29. IEEE, 1997
1997
-
[15]
An open dataset and model for language identification
Laurie Burchell, Alexandra Birch, Nikolay Bogoychev, and Kenneth Heafield. An open dataset and model for language identification. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguist...
2023 doi
-
[16]
An expanded massive multilingual dataset for high-performance language technologies, 2025
Laurie Burchell, Ona de Gibert, Nikolay Arefyev, Mikko Aulamo, Marta Bañón, Pinzhen Chen, Mariia Fedorova, Liane Guillou, Barry Haddow, Jan Hajič, Jindřich Helcl, Erik Henriksson, Mateusz Klimaszewski, Ville Komulainen, Andrey Kutuzov, Joona Kytöniemi, Veronika Laippala, Pette...
2025 arXiv
-
[17]
PTT5 : Pretraining and validating the t5 model on brazilian portuguese data
Diedre Carmo, Marcos Piau, Israel Campiotti, Rodrigo Nogueira, and Roberto Lotufo. PTT5 : Pretraining and validating the t5 model on brazilian portuguese data. arXiv preprint arXiv:2008.09144, 2020
2008 arXiv
-
[18]
Language ID in the wild: Unexpected challenges on the path to a thousand-language web text corpus
Isaac Caswell, Theresa Breiner, Daan van Esch, and Ankur Bapna. Language ID in the wild: Unexpected challenges on the path to a thousand-language web text corpus. In Donia Scott, Nuria Bel, and Chengqing Zong (eds.), Proceedings of the 28th International Conference on Computat...
2020 doi
-
[19]
Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki
Jonathan H. Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki. TyDi QA : A benchmark for information-seeking question answering in typologically diverse languages. Transactions of the Association for Computational Ling...
2020 doi
-
[20]
Command r+
Cohere. Command r+. Web, 2024. URL https://docs.cohere.com/docs/command-r-plus#model-details
2024
-
[21]
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzm \'a n, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. Unsupervised cross-lingual representation learning at scale. In Dan Jurafsky, Joyce Chai, Natalie Schlu...
2020 doi
-
[22]
Neural learning for question answering in italian
Danilo Croce, Alexandra Zelenanska, and Roberto Basili. Neural learning for question answering in italian. In Chiara Ghidini, Bernardo Magnini, Andrea Passerini, and Paolo Traverso (eds.), AI*IA 2018 -- Advances in Artificial Intelligence, pp.\ 389--402, Cham, 2018. Springer I...
2018
-
[23]
Dataset for the first evaluation on chinese machine reading comprehension, 2018
Yiming Cui, Ting Liu, Zhipeng Chen, Wentao Ma, Shijin Wang, and Guoping Hu. Dataset for the first evaluation on chinese machine reading comprehension, 2018. URL https://arxiv.org/abs/1709.08299
2018 arXiv
-
[24]
Daniels and William Bright (eds.)
Peter T. Daniels and William Bright (eds.). The World's Writing Systems. Oxford University Press, New York, 1996
1996
-
[26]
A new massive multilingual dataset for high-performance language technologies, 2024
Ona de Gibert, Graeme Nail, Nikolay Arefyev, Marta Bañón, Jelmer van der Linde, Shaoxiong Ji, Jaume Zaragoza-Bernabeu, Mikko Aulamo, Gema Ramírez-Sánchez, Andrey Kutuzov, Sampo Pyysalo, Stephan Oepen, and Jörg Tiedemann. A new massive multilingual dataset for high-performance ...
2024 arXiv
-
[27]
BERTje : A dutch BERT model
Wietse de Vries, Andreas van Cranenburgh, Arianna Bisazza, Tommaso Caselli, Gertjan van Noord, and Malvina Nissim. BERTje : A dutch BERT model. arXiv preprint arXiv:1912.09582, 2019
1912 arXiv
-
[28]
DeepSeek-AI, :, Xiao Bi, Deli Chen, Guanting Chen, Shanhuang Chen, Damai Dai, Chengqi Deng, Honghui Ding, Kai Dong, Qiushi Du, Zhe Fu, Huazuo Gao, Kaige Gao, Wenjun Gao, Ruiqi Ge, Kang Guan, Daya Guo, Jianzhong Guo, Guangbo Hao, Zhewen Hao, Ying He, Wenjie Hu, Panpan Huang, Er...
2024 arXiv
-
[29]
RobBERT : a dutch RoBERTa -based language model
Pieter Delobelle, Thomas Winters, and Bettina Berendt. RobBERT : a dutch RoBERTa -based language model. arXiv preprint arXiv:2001.06286, 2020
2001 arXiv
-
[30]
Fquad: French question answering dataset, 2020
Martin d'Hoffschmidt, Wacim Belblidia, Tom Brendlé, Quentin Heinrich, and Maxime Vidal. Fquad: French question answering dataset, 2020. URL https://arxiv.org/abs/2002.06071
2020 arXiv
-
[31]
Tran, Mike Zhang, Shiqi Chen, Tianyu Pang, Chao Du, Xinyi Wan, Wei Lu, and Min Lin
Longxu Dou, Qian Liu, Fan Zhou, Changyu Chen, Zili Wang, Ziqi Jin, Zichen Liu, Tongyao Zhu, Cunxiao Du, Penghui Yang, Haonan Wang, Jiaheng Liu, Yongchi Zhao, Xiachong Feng, Xin Mao, Man Tsung Yeung, Kunat Pipatanakul, Fajri Koto, Min Si Thu, Hynek Kydlíček, Zeyi Liu, Qunshu Li...
2025 arXiv
-
[32]
Chinese tiny llm: Pretraining a chinese-centric large language model, 2024
Xinrun Du, Zhouliang Yu, Songyang Gao, Ding Pan, Yuyang Cheng, Ziyang Ma, Ruibin Yuan, Xingwei Qu, Jiaheng Liu, Tianyu Zheng, Xinchen Luo, Guorui Zhou, Wenhu Chen, and Ge Zhang. Chinese tiny llm: Pretraining a chinese-centric large language model, 2024. URL https://arxiv.org/a...
2024 arXiv
-
[33]
Eberhard, Gary F
David M. Eberhard, Gary F. Simons, and Charles D. Fenning. Ethnologue: Languages of the world, 2024. URL http://www.ethnologue.com
2024
-
[34]
SberQuAD – Russian Reading Comprehension Dataset: Description and Analysis, pp.\ 3–15
Pavel Efimov, Andrey Chertok, Leonid Boytsov, and Pavel Braslavski. SberQuAD – Russian Reading Comprehension Dataset: Description and Analysis, pp.\ 3–15. Springer International Publishing, 2020. ISBN 9783030582197. doi:10.1007/978-3-030-58219-7_1. URL http://dx.doi.org/10.100...
2020 doi
-
[35]
Arabicweb24: Creating a high quality arabic web-only pre-training dataset, 2024
May Farhat, Said Taghadouini, Oskar Hallström, and Sonja Hajri-Gabouj. Arabicweb24: Creating a high quality arabic web-only pre-training dataset, 2024. URL www.lighton.ai/lighton-blogs/arabicweb24
2024
-
[36]
Guerreiro, António Loison, Duarte M
Manuel Faysse, Patrick Fernandes, Nuno M. Guerreiro, António Loison, Duarte M. Alves, Caio Corro, Nicolas Boizard, João Alves, Ricardo Rei, Pedro H. Martins, Antoni Bigata Casademunt, François Yvon, André F. T. Martins, Gautier Viaud, Céline Hudelot, and Pierre Colombo. Croiss...
2024
-
[37]
Mera: A comprehensive llm evaluation in russian, 2024
Alena Fenogenova, Artem Chervyakov, Nikita Martynov, Anastasia Kozlova, Maria Tikhonova, Albina Akhmetgareeva, Anton Emelyanov, Denis Shevelev, Pavel Lebedev, Leonid Sinev, Ulyana Isaeva, Katerina Kolomeytseva, Daniil Moskovskiy, Elizaveta Goncharova, Nikita Savushkin, Polina ...
2024 arXiv
-
[38]
Open llm leaderboard v2
Clémentine Fourrier, Nathan Habib, Alina Lozovskaya, Konrad Szafer, and Thomas Wolf. Open llm leaderboard v2. https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard, 2024
2024
-
[39]
Gemma Team , Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Léonard Hussenot, Pier Giuseppe Sessa, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Ale...
2024 arXiv
-
[40]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Art...
2024 arXiv
-
[41]
Studying large language model generalization with influence functions
Roger Grosse, Juhan Bae, Cem Anil, Nelson Elhage, Alex Tamkin, Amirhossein Tajdini, Benoit Steiner, Dustin Li, Esin Durmus, Ethan Perez, et al. Studying large language model generalization with influence functions. arXiv preprint arXiv:2308.03296, 2023
2023 arXiv
-
[42]
Olmes: A standard for language model evaluations, 2025
Yuling Gu, Oyvind Tafjord, Bailey Kuehl, Dany Haddad, Jesse Dodge, and Hannaneh Hajishirzi. Olmes: A standard for language model evaluations, 2025. URL https://arxiv.org/abs/2406.08446
2025 arXiv
-
[43]
Exams: A multi-subject high school examinations dataset for cross-lingual and multilingual question answering, 2020
Momchil Hardalov, Todor Mihaylov, Dimitrina Zlatkova, Yoan Dinkov, Ivan Koychev, and Preslav Nakov. Exams: A multi-subject high school examinations dataset for cross-lingual and multilingual question answering, 2020. URL https://arxiv.org/abs/2011.03080
2020 arXiv
-
[44]
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. Proceedings of the International Conference on Learning Representations (ICLR), 2021
2021
-
[45]
Khmer natural language processing tookit
Phan Viet Hoang. Khmer natural language processing tookit. https://github.com/VietHoang1512/khmer-nltk, 2020
2020
-
[46]
spaCy: Industrial-strength Natural Language Processing in Python
Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd. spaCy: Industrial-strength Natural Language Processing in Python . 2020. doi:10.5281/zenodo.1212303
2020 doi
-
[47]
C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models, 2023
Yuzhen Huang, Yuzhuo Bai, Zhihao Zhu, Junlei Zhang, Jinghan Zhang, Tangjun Su, Junteng Liu, Chuancheng Lv, Yikai Zhang, Jiayi Lei, Yao Fu, Maosong Sun, and Junxian He. C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models, 2023. URL https://arxi...
2023 arXiv
-
[48]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas ...
2023 arXiv
-
[49]
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne...
2024 arXiv
-
[50]
The state and fate of linguistic diversity and inclusion in the NLP world
Pratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali, and Monojit Choudhury. The state and fate of linguistic diversity and inclusion in the NLP world. In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (eds.), Proceedings of the 58th Annual Meeting of the ...
2020 doi
-
[51]
Fasttext.zip: Compressing text classification models, 2016
Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hérve Jégou, and Tomas Mikolov. Fasttext.zip: Compressing text classification models, 2016. URL https://arxiv.org/abs/1612.03651
2016 arXiv
-
[52]
G lot LID : Language identification for low-resource languages
Amir Hossein Kargaran, Ayyoob Imani, Fran c ois Yvon, and Hinrich Schuetze. G lot LID : Language identification for low-resource languages. In Houda Bouamor, Juan Pino, and Kalika Bali (eds.), Findings of the Association for Computational Linguistics: EMNLP 2023, pp.\ 6155--62...
2023 doi
-
[53]
Glot CC : An open broad-coverage commoncrawl corpus and pipeline for minority languages
Amir Hossein Kargaran, Fran c ois Yvon, and Hinrich Schuetze. Glot CC : An open broad-coverage commoncrawl corpus and pipeline for minority languages. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2024. URL https://openr...
2024
-
[54]
Mohammed Safi Ur Rahman Khan, Priyam Mehta, Ananth Sankar, Umashankar Kumaravelan, Sumanth Doddapaneni, Suriyaprasaad G, Varun Balan G, Sparsh Jain, Anoop Kunchukuttan, Pratyush Kumar, Raj Dabre, and Mitesh M. Khapra. Indicllmsuite: A blueprint for creating pre-training and fi...
2024 arXiv
-
[55]
Large language models only pass primary school exams in I ndonesia: A comprehensive test on I ndo MMLU
Fajri Koto, Nurul Aisyah, Haonan Li, and Timothy Baldwin. Large language models only pass primary school exams in I ndonesia: A comprehensive test on I ndo MMLU . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), Singapore, 2023...
2023
-
[56]
Arabicmmlu: Assessing massive multitask language understanding in arabic, 2024
Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin. Arabicmmlu: Assessing massive multitask language understanding in ara...
2024 arXiv
-
[57]
The IndicNLP Library
Anoop Kunchukuttan. The IndicNLP Library . https://github.com/anoopkunchukuttan/indic_nlp_library/blob/master/docs/indicnlp.pdf, 2020
2020
-
[58]
JGLUE : J apanese general language understanding evaluation
Kentaro Kurihara, Daisuke Kawahara, and Tomohide Shibata. JGLUE : J apanese general language understanding evaluation. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pp.\ 2957--2966, Marseille, France, June 2022. European Language Resources Asso...
2022
-
[59]
Rossi, and Thien Huu Nguyen
Viet Dac Lai, Chien Van Nguyen, Nghia Trung Ngo, Thuat Nguyen, Franck Dernoncourt, Ryan A. Rossi, and Thien Huu Nguyen. Okapi: Instruction-tuned large language models in multiple languages with reinforcement learning from human feedback, 2023. URL https://arxiv.org/abs/2307.16039
2023 arXiv
-
[60]
F lau BERT : Unsupervised language model pre-training for F rench
Hang Le, Lo \" c Vial, Jibril Frej, Vincent Segonne, Maximin Coavoux, Benjamin Lecouteux, Alexandre Allauzen, Benoit Crabb \'e , Laurent Besacier, and Didier Schwab. F lau BERT : Unsupervised language model pre-training for F rench. In Proceedings of the 12th Language Resource...
2020
-
[61]
Open-arabic-llm-leaderboard-v1
Open Arabic LLM Leaderboard. Open-arabic-llm-leaderboard-v1. https://huggingface.co/spaces/OALL/Open-Arabic-LLM-Leaderboard-v1, 2024. Accessed: 2025-03-28
2024
-
[62]
Deduplicating training data makes language models better, 2022
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. Deduplicating training data makes language models better, 2022. URL https://arxiv.org/abs/2107.06499
2022 arXiv
-
[63]
Kiwipiepy: Kiwi package for python, 2024
Minchul Lee. Kiwipiepy: Kiwi package for python, 2024. URL https://github.com/bab2min/kiwipiepy
2024
-
[64]
Mlqa: Evaluating cross-lingual extractive question answering, 2020
Patrick Lewis, Barlas Oğuz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. Mlqa: Evaluating cross-lingual extractive question answering, 2020. URL https://arxiv.org/abs/1910.07475
2020 arXiv
-
[65]
Cmmlu: Measuring massive multitask language understanding in chinese, 2024 a
Haonan Li, Yixuan Zhang, Fajri Koto, Yifei Yang, Hai Zhao, Yeyun Gong, Nan Duan, and Timothy Baldwin. Cmmlu: Measuring massive multitask language understanding in chinese, 2024 a . URL https://arxiv.org/abs/2306.09212
2024 arXiv
-
[66]
Dimakis, Yair Carmon, Achal Dave, Ludwig Schmidt, and Vaishaal Shankar
Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Gadre, Hritik Bansal, Etash Guha, Sedrick Keh, Kushal Arora, Saurabh Garg, Rui Xin, Niklas Muennighoff, Reinhard Heckel, Jean Mercat, Mayee Chen, Suchin Gururangan, Mitchell Wortsman, Alon Albalak, Yonatan ...
2024 arXiv
-
[67]
Common sense beyond E nglish: Evaluating and improving multilingual language models for commonsense reasoning
Bill Yuchen Lin, Seyeon Lee, Xiaoyang Qiao, and Xiang Ren. Common sense beyond E nglish: Evaluating and improving multilingual language models for commonsense reasoning. In Chengqing Zong, Fei Xia, Wenjie Li, and Roberto Navigli (eds.), Proceedings of the 59th Annual Meeting o...
2021
-
[69]
Few-shot learning with multilingual language models, 2022
Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O'Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mon...
2022 arXiv
-
[70]
Fingpt: Large generative models for a small language
Risto Luukkonen, Ville Komulainen, Jouni Luoma, Anni Eskelinen, Jenna Kanerva, Hanna-Mari Kupari, Filip Ginter, Veronika Laippala, Niklas Muennighoff, Aleksandra Piktus, et al. Fingpt: Large generative models for a small language. arXiv preprint arXiv:2311.05640, 2023
2023 arXiv
-
[71]
Singh, Rylan Schaeffer, Andrew Poulton, Sanmi Koyejo, Pontus Stenetorp, Sharan Narang, and Dieuwke Hupkes
Lovish Madaan, Aaditya K. Singh, Rylan Schaeffer, Andrew Poulton, Sanmi Koyejo, Pontus Stenetorp, Sharan Narang, and Dieuwke Hupkes. Quantifying variance in evaluation benchmarks, 2024. URL https://arxiv.org/abs/2406.10229
2024 arXiv
-
[72]
C amem BERT : a tasty F rench language model
Louis Martin, Benjamin Muller, Pedro Javier Ortiz Su \'a rez, Yoann Dupont, Laurent Romary, \'E ric de la Clergerie, Djam \'e Seddah, and Beno \^ t Sagot. C amem BERT : a tasty F rench language model. In Proceedings of the 58th Annual Meeting of the Association for Computation...
2020 doi
-
[73]
Between words and characters: A brief history of open-vocabulary modeling and tokenization in nlp
Sabrina J Mielke, Zaid Alyafeai, Elizabeth Salesky, Colin Raffel, Manan Dey, Matthias Gall \'e , Arun Raja, Chenglei Si, Wilson Y Lee, Beno \^ t Sagot, et al. Between words and characters: A brief history of open-vocabulary modeling and tokenization in nlp. arXiv preprint arXi...
2021 arXiv
-
[74]
Mnbvc: Massive never-ending bt vast chinese corpus
MOP-LIWU Community and MNBVC Team . Mnbvc: Massive never-ending bt vast chinese corpus. https://github.com/esbatmop/MNBVC, 2023
2023
-
[75]
Neural A rabic question answering
Hussein Mozannar, Elie Maamary, Karl El Hajal, and Hazem Hajj. Neural A rabic question answering. In Proceedings of the Fourth Arabic Natural Language Processing Workshop, pp.\ 108--118, Florence, Italy, August 2019. Association for Computational Linguistics. doi:10.18653/v1/W...
2019 doi
-
[76]
Crosslingual generalization through multitask finetuning, 2022
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward ...
2022
-
[77]
Rossi, and Thien Huu Nguyen
Thuat Nguyen, Chien Van Nguyen, Viet Dac Lai, Hieu Man, Nghia Trung Ngo, Franck Dernoncourt, Ryan A. Rossi, and Thien Huu Nguyen. C ultura X : A cleaned, enormous, and multilingual dataset for large language models in 167 languages. In Nicoletta Calzolari, Min-Yen Kan, Veroniq...
2024
-
[78]
NLLB Team , Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Pran...
2022 arXiv
-
[79]
Omnia russica
Omnia Russica Team . Omnia russica. https://omnia-russica.github.io/, 2024
2024
-
[80]
Botok: State-of-the-art tokenizers for tibetan language, 2025
OpenPecha . Botok: State-of-the-art tokenizers for tibetan language, 2025. URL https://github.com/OpenPecha/Botok. Support for various dialects, fully customizable word lists and adjustment rules
2025
-
[81]
Building pre-train llm dataset for the indic languages: A case study on hindi
Shantipriya Parida, Shakshi Panwar, Kusum Lata, Sanskruti Mishra, and Sambit Sekhar. Building pre-train llm dataset for the indic languages: A case study on hindi. https://huggingface.co/OdiaGenAI, 2024
2024
-
[82]
Hellaswag-th, 2023
Triamamornwooth Patteera. Hellaswag-th, 2023. URL https://huggingface.co/datasets/Patt/HellaSwag_TH. v1.0
2023
-
[83]
The refinedweb dataset for falcon llm: Outperforming curated corpora with web data, and web data only, 2023
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. The refinedweb dataset for falcon llm: Outperforming curated corpora with web data, and web data only, 2023. UR...
2023 arXiv
-
[84]
The fineweb datasets: Decanting the web for the finest text data at scale
Guilherme Penedo, Hynek Kydl\' c ek, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf. The fineweb datasets: Decanting the web for the finest text data at scale. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. ...
2024
-
[85]
Laonlp: Lao language natural language processing, July 2022
Wannaphong Phatthiyaphaibun. Laonlp: Lao language natural language processing, July 2022. URL https://doi.org/10.5281/zenodo.6833407
2022 doi
-
[86]
P y T hai NLP : T hai natural language processing in P ython, June 2024
Wannaphong Phatthiyaphaibun, Korakot Chaovavanich, Charin Polpanumas, Arthit Suriyawongkul, Lalita Lowphansirikul, and Pattarawat Chormai. P y T hai NLP : T hai natural language processing in P ython, June 2024. URL https://github.com/PyThaiNLP/pythainlp/
2024
-
[87]
Typhoon: Thai large language models
Kunat Pipatanakul, Phatrasek Jirabovonvisut, Potsawee Manakul, Sittipong Sripaisarnmongkol, Ruangsak Patomwong, Pathomporn Chokchainant, and Kasima Tharnpipitchai. Typhoon: Thai large language models. arXiv preprint arXiv:2312.13951, 2023
2023 arXiv
-
[88]
Pllum: A family of polish large language models
PLLuM Consortium . Pllum: A family of polish large language models. 2025
2025
-
[89]
Chinesesquad
Pluto-Junzeng. Chinesesquad. https://github.com/pluto-junzeng/ChineseSquad, 2019. Accessed: 2025-03-28
2019
-
[90]
Xcopa: A multilingual dataset for causal commonsense reasoning, 2020
Edoardo Maria Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulić, and Anna Korhonen. Xcopa: A multilingual dataset for causal commonsense reasoning, 2020. URL https://arxiv.org/abs/2005.00333
2020 arXiv
-
[91]
Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton, and Christopher D. Manning. Stanza: A python natural language processing toolkit for many human languages, 2020. URL https://arxiv.org/abs/2003.07082
2020 arXiv
-
[92]
Jack W. Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, Eliza Rutherford, Tom Hennigan, Jacob Menick, Albin Cassirer, Richard Powell, George van den Driessche, Lisa Anne Hendricks,...
2022 arXiv
-
[93]
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21 0 (140), 2020
2020
-
[94]
Impact of pretraining term frequencies on few-shot reasoning
Yasaman Razeghi, Robert L Logan IV, Matt Gardner, and Sameer Singh. Impact of pretraining term frequencies on few-shot reasoning. arXiv preprint arXiv:2202.07206, 2022
2022 arXiv
-
[95]
How much knowledge can you pack into the parameters of a language model? arXiv preprint arXiv:2002.08910, 2020
Adam Roberts, Colin Raffel, and Noam Shazeer. How much knowledge can you pack into the parameters of a language model? arXiv preprint arXiv:2002.08910, 2020
2002 arXiv
-
[96]
How good is your tokenizer? on the monolingual performance of multilingual language models, 2021
Phillip Rust, Jonas Pfeiffer, Ivan Vulić, Sebastian Ruder, and Iryna Gurevych. How good is your tokenizer? on the monolingual performance of multilingual language models, 2021. URL https://arxiv.org/abs/2012.15613
2021 arXiv
-
[97]
Pyidaungsu: Python library for myanmar language, 2024
Oishi Sakana. Pyidaungsu: Python library for myanmar language, 2024. URL https://github.com/kaunghtetsan275/pyidaungsu
2024
-
[98]
Compact language detector v3
Alex Salcianu, Andy Golding, Anton Bakalov, Chris Alberti, Daniel Andor, David Weiss, Emily Pitler, Greg Coppola, Jason Riesa, Kuzman Ganchev, et al. Compact language detector v3. Technical report, 2018. URL https://chromium.googlesource.com/external/github.com/google/cld_3/. ...
2018
-
[99]
Mintaka: A complex, natural, and multilingual dataset for end-to-end question answering, 2022
Priyanka Sen, Alham Fikri Aji, and Amir Saffari. Mintaka: A complex, natural, and multilingual dataset for end-to-end question answering, 2022. URL https://arxiv.org/abs/2210.01613
2022 arXiv
-
[100]
Indic qa benchmark: A multilingual benchmark to evaluate question answering capability of llms for indic languages, 2025
Abhishek Kumar Singh, Vishwajeet kumar, Rudra Murthy, Jaydeep Sen, Ashish Mittal, and Ganesh Ramakrishnan. Indic qa benchmark: A multilingual benchmark to evaluate question answering capability of llms for indic languages, 2025. URL https://arxiv.org/abs/2407.13522
2025 arXiv
-
[101]
Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh, Luke Zettlemoyer, Noah A
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Mu...
2024 arXiv
-
[102]
Thquad: Turkish historic question answering dataset for reading comprehension
Fatih Soygazi, Okan Çiftçi, Uğurcan Kök, and Soner Cengiz. Thquad: Turkish historic question answering dataset for reading comprehension. In 2021 6th International Conference on Computer Science and Engineering (UBMK), pp.\ 215--220, 2021. doi:10.1109/UBMK52708.2021.9559013
2021
-
[103]
Investigating prior knowledge for challenging C hinese machine reading comprehension
Kai Sun, Dian Yu, Dong Yu, and Claire Cardie. Investigating prior knowledge for challenging C hinese machine reading comprehension. Transactions of the Association for Computational Linguistics, 8: 0 141--155, 2020. doi:10.1162/tacl_a_00305. URL https://aclanthology.org/2020.t...
2020 doi
-
[104]
Liping Tang, Nikhil Ranjan, Omkar Pangarkar, Xuezhi Liang, Zhen Wang, Li An, Bhaskar Rao, Linghao Jin, Huijuan Wang, Zhoujun Cheng, Suqi Sun, Cun Mu, Victor Miller, Xuezhe Ma, Yue Peng, Zhengzhong Liu, and Eric P. Xing. Txt360: A top-quality llm pre-training dataset requires t...
2024
-
[105]
Tigerbot: A multi-language multi-task llm
TigerResearch. Tigerbot: A multi-language multi-task llm. https://github.com/TigerResearch/TigerBot, 2023
2023
-
[106]
Llama: Open and efficient foundation language models, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. Llama: Open and efficient foundation language...
2023 arXiv
-
[107]
Trakultaweekoon, S
K. Trakultaweekoon, S. Thaiprayoon, P. Palingoon, and A. Rugchatjaroen. The first wikipedia questions and factoid answers corpus in the thai language. In 2019 14th International Joint Symposium on Artificial Intelligence and Natural Language Processing (iSAI-NLP), pp.\ 1--4. I...
2019
-
[108]
Vbart: The turkish llm
Meliksah Turker, Erdi Ari, and Aydin Han. Vbart: The turkish llm. arXiv preprint arXiv:2403.01308, 2024
2024 arXiv
-
[109]
Wanjawa, Lilian D
Barack W. Wanjawa, Lilian D. A. Wanzare, Florence Indede, Owen Mconyango, Lawrence Muchemi, and Edward Ombui. Kenswquad—a question answering dataset for swahili low-resource language. ACM Transactions on Asian and Low-Resource Language Information Processing, 22 0 (4): 0 1–20,...
2023 doi
-
[110]
CCN et: Extracting high quality monolingual datasets from web crawl data
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzm \'a n, Armand Joulin, and Edouard Grave. CCN et: Extracting high quality monolingual datasets from web crawl data. In Nicoletta Calzolari, Fr \'e d \'e ric B \'e chet, Philippe Blache, Khal...
2020
-
[111]
BigScience Workshop, :, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanch...
2023 arXiv
-
[112]
mt5: A massively multilingual pre-trained text-to-text transformer, 2021
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. mt5: A massively multilingual pre-trained text-to-text transformer, 2021. URL https://arxiv.org/abs/2010.11934
2021 arXiv
-
[113]
Qwen2.5 technical report
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Me...
2024 arXiv
-
[114]
Hellaswag: Can a machine really finish your sentence?, 2019
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence?, 2019. URL https://arxiv.org/abs/1905.07830
2019 arXiv
-
[115]
M3exam: A multilingual, multimodal, multilevel benchmark for examining large language models, 2023
Wenxuan Zhang, Sharifah Mahani Aljunied, Chang Gao, Yew Ken Chia, and Lidong Bing. M3exam: A multilingual, multimodal, multilevel benchmark for examining large language models, 2023. URL https://arxiv.org/abs/2306.05179
2023 arXiv
-
[117]
Agieval: A human-centric benchmark for evaluating foundation models, 2023 b
Wanjun Zhong, Ruixiang Cui, Yiduo Guo, Yaobo Liang, Shuai Lu, Yanlin Wang, Amin Saied, Weizhu Chen, and Nan Duan. Agieval: A human-centric benchmark for evaluating foundation models, 2023 b . URL https://arxiv.org/abs/2304.06364
2023 arXiv
-
[118]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
-
[119]
@esa (Ref
\@ifxundefined[1] #1\@undefined \@firstoftwo \@secondoftwo \@ifnum[1] #1 \@firstoftwo \@secondoftwo \@ifx[1] #1 \@firstoftwo \@secondoftwo [2] @ #1 \@temptokena #2 #1 @ \@temptokena \@ifclassloaded agu2001 natbib The agu2001 class already includes natbib coding, so you should ...
-
[120]
\@lbibitem[] @bibitem@first@sw\@secondoftwo \@lbibitem[#1]#2 \@extra@b@citeb \@ifundefined br@#2\@extra@b@citeb \@namedef br@#2 \@nameuse br@#2\@extra@b@citeb \@ifundefined b@#2\@extra@b@citeb @num @parse #2 @tmp #1 NAT@b@open@#2 NAT@b@shut@#2 \@ifnum @merge>\@ne @bibitem@firs...
-
[121]
C MI:LLJ zfwj
@open @close @open @close and [1] URL: #1 \@ifundefined chapter * \@mkboth \@ifxundefined @sectionbib * \@mkboth * \@mkboth\@gobbletwo \@ifclassloaded amsart * \@ifclassloaded amsbook * \@ifxundefined @heading @heading NAT@ctr thebibliography [1] @ \@biblabel @NAT@ctr \@bibset...
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.