REVIEW 4 major objections 4 minor 53 references
MELABenchv1: Benchmarking Large Language Models against Smaller Fine-Tuned Models for Low-Resource Maltese NLP
T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read For Maltese, prompting large public LLMs consistently loses to small task-specific fine-tuned models, and prior Maltese training exposure explains more of the gap than model size or language coverage.
desk verdict A useful first Maltese benchmark and a broad, honest evaluation; the fine-tuned-vs-prompted finding holds up, but the 'Maltese exposure is the dominant factor' claim is left vulnerable to benchmark contamination. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is MELABenchv1, a benchmark of 11 existing Maltese datasets split into discriminative tasks (classification and reading comprehension) and generative tasks (translation, data-to-text, and summarisation). The analytic device that carries the argument is a Maltese-exposure taxonomy: every one of the 55 models is labelled according to whether Maltese appeared during pre-training, during instruction-tuning, both, or neither, and the paper attributes performance differences to those labels. The comparators are small fine-tuned models, one per task, built from BERTu, mBERT, and mT5-Small, which supply the baseline that prompted LLMs are measured against.
What would settle it
Run the same benchmark on a family of models whose training data is fully controlled, varying only the number of Maltese tokens while holding architecture, size, and instruction-tuning constant; if performance does not rise with Maltese token count, the exposure claim fails. A complementary check is a memorization probe that searches for near-duplicate benchmark instances in training corpora, which would show whether the best scores come from contamination.
Extended reading notes
Core claim
The paper claims that on Maltese, prompted LLMs consistently lag behind smaller fine-tuned models, especially in generative tasks where their performance is significantly lower, and that the level of exposure to Maltese in a model's training has the largest bearing on performance, particularly exposure during instruction-tuning. This is demonstrated by aggregating zero-shot and one-shot results across 55 public models on 11 discriminative and generative tasks, and comparing them with per-task fine-tuned baselines. The paper also claims that fine-tuning carries a higher initial computational cost but yields better scores and much lower per-instance inference cost, so it becomes the more efficient choice as the number of inference samples grows.
Load-bearing premise
The conclusion that Maltese training exposure is the most important factor depends on the paper's labels of which models actually saw Maltese data being correct, and on those models not having silently memorized the benchmark's public test sets.
Editorial extensions
If this is right
- For Maltese, prompting even large open LLMs is not a substitute for task-specific fine-tuning, and the gap is widest when the model must generate new Maltese text.
- A model's prior Maltese exposure, especially in instruction-tuning, predicts performance better than parameter count or breadth of language coverage.
- One-shot prompting narrows the gap but does not close it, and it helps generative tasks consistently while sometimes hurting discriminative ones.
- Fine-tuning small models beats prompting on both score and per-instance cost once enough inference samples are needed, despite its higher upfront compute.
- Switching the prompt language from English to Maltese generally does not help and often hurts, so even successful models remain awkward for native Maltese speakers to use.
Reading between the lines
- The paper treats Maltese exposure as a categorical yes-or-no label; a natural extension would be a dose-response test that varies the actual number of Maltese training tokens while holding architecture and size fixed.
- If exposure is really the dominant factor, then adding even modest Maltese text to pre-training or instruction-tuning data may buy more for low-resource languages than scaling model size or adding more languages.
- Because the benchmark reuses public datasets, the same harness could double as a contamination probe: exact-match or near-duplicate retrieval of test instances would show whether top scores come from memorization rather than ability.
- The efficiency calculation implies a practical decision rule for a research group with limited compute: choose prompting only for very small inference volumes, and fine-tune a small model for anything larger.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces MELABenchv1, a Maltese benchmark of 11 discriminative and generative tasks, and uses it to evaluate 55 open-weight LLMs under zero-shot and one-shot prompting, with both English and Maltese instructions, comparing them against fine-tuned BERTu, mBERT, and mT5-Small baselines. The authors analyze how model training type (pre-trained vs instruction-tuned), Maltese exposure stage, model size, and multilingual coverage relate to performance, and they add a FLOP-based efficiency comparison of prompting versus fine-tuning. They conclude that prompted LLMs consistently lag behind smaller fine-tuned models, particularly on generative tasks, and that prior exposure to Maltese during pre-training and instruction-tuning is the most important performance factor. The paper releases evaluation code and fine-tuned models.
Significance. The benchmark and the breadth of models tested are a useful contribution to low-resource language evaluation: the per-task results for 11 tasks, the documented prompt templates in Appendix A, and the fine-tuning details in Appendix B provide a reusable resource for Maltese NLP. The efficiency analysis also gives a concrete cost comparison between fine-tuning and prompting. If the exposure finding were robust, it would be an important data point for how training-data decisions affect low-resource performance. However, the headline claim about Maltese exposure is currently vulnerable to benchmark contamination, and the aggregate scoring mixes incompatible metrics; the paper's strongest scientific assertion is therefore not yet fully supported.
major comments (4)
- [Section 8; Sections 4.1 and 7] The paper's central claim that Maltese exposure is the most important factor is not protected against dataset contamination. The top-scoring IT/PT and IT/IT models (mT0-XXL and Aya-101) were instruction-tuned on xP3 and the Aya Dataset, which aggregate public task collections, and several benchmark tasks (FLORES-200, SIB-200, MultiEURLEX, WebNLG) predate these models. The Limitations section states that 'it is likely that increased performance in some models is due to data contamination during training' and that instruction-tuned models 'may have used certain datasets deliberately,' but no analysis quantifies how much of the observed IT advantage remains after removing potentially contaminated instances. Since Section 7 elevates exposure to the 'largest bearing' factor, an overlap analysis (for example, n-gram overlap between training corpora and benchmark test splits, or a sensitivity analysis restricted to tasks unlikely to be in training data) is needed before this claim can be accepted.
- [Section 3, Figure 1, Appendix E] The aggregate scores used in Figure 1 and in the regression analyses of Appendix E average raw task scores that come from incompatible metrics (macro-F1, accuracy, ChrF, Rouge-L) without normalization or weighting. The footnote in Section 3 says the metrics are 'normalised within the same range,' but equal ranges do not imply equal scales or equal difficulty; the near-zero floor on generative tasks means that the generative aggregate is essentially a scaled indicator of whether a model produces Maltese text at all. Using these aggregates as the dependent variable in the model-size and multilinguality regressions can therefore produce correlations that reflect task mix rather than the factor of interest. I request task-z-scored or rank-based aggregates, or separate analyses per metric family.
- [Section 3.2, footnote 5, Figure 5] The one-shot analysis excludes mT0 and BLOOMZ models 'due to their zero-shot instruction-tuning' negative effects. These model families include some of the strongest zero-shot performers on generative tasks (for example, mT0-XXL in Table 6), so their removal from Figure 5 and the statement that 'consistent performance improvements with one-shot across all generative tasks' may not hold for the full model set. I ask that the one-shot analysis be reported both with and without these models, or that the exclusion be justified by a criterion that is not outcome-dependent.
- [Table 2; Sections 4.1, 4.2, 7] The categorization of Aya-101 as IT/IT should be checked. If Aya-101 is initialized from mT5, which the paper's own Table 2 lists as PT/PT with Maltese pre-training, then Aya-101 should be classified as IT/PT rather than IT/IT. The IT/IT group is the strongest-performing category, and Section 7 concludes that exposure during instruction-tuning is especially important, so a misclassification here would change the interpretation. Please state the base model used for Aya-101 and make the PT/IT labels consistent with it.
minor comments (4)
- [Section 2.2 and Table 2] The paper states that 55 models are evaluated, but the models enumerated in Table 2 (and the labels in Figure 1) sum to 53; please reconcile the count.
- [Table 2 and Figures 1, 9] MaLA-500's parameter size is given as 8.6B in Table 2 but displayed as 10.0B in Figures 1 and 9; the labels should be made consistent.
- [Section 3.2, Tables 7-8, Figures 3 and 10] The one-shot experiments appear to omit Gemma 2 9B, which is included in the zero-shot experiments, without an explanation; please state the reason for this omission.
- [Section 4.2] The observation that IT/IT models degrade most with Maltese prompts despite being the most exposed to Maltese is reported but not analyzed; adding a brief explanation (for example, instruction-following in English or tokenizer effects) would help readers interpret Figure 7.
Circularity Check
No significant circularity: MELABenchv1 is an empirical measurement study whose claims rest on direct model evaluations, not on fitted parameters, self-citations, or definitional reductions.
full rationale
The paper's central claims—prompted LLMs lag smaller fine-tuned models on Maltese, and prior Maltese exposure is the largest performance factor—are inferred from direct benchmark measurements (Sections 3, 4, and Appendix E). No equation in the paper is fitted to a subset of the reported scores and then reused as a prediction; no parameter is renamed as a result; and no author-derived uniqueness theorem is invoked to force a conclusion. The self-citations (BERTu, the Maltese news datasets, and the language report) supply models, data, and background that are evaluated or used as resources in this paper, but they do not encode the paper's conclusions: BERTu's superiority is measured here on held-out test splits, and the news datasets are external benchmark tasks. The limitations section explicitly flags dataset contamination and the categorical treatment of Maltese exposure as threats to the causal interpretation (Section 8), and Appendix E checks robustness by removing Maltese-trained models; these are validity caveats, not circular dependencies. The comparison and ranking claims are therefore self-contained empirical findings rather than reductions to their own inputs.
Assumptions & free parameters
free parameters (1)
- Task aggregation weights =
equal (1/11 across tasks)
assumptions (4)
- domain assumption Publicly released benchmark datasets (SIB-200, Belebele, Flores-200, etc.) are correctly constructed and their labels are accurate.
- domain assumption Model training data metadata reported in model cards and papers is accurate for categorizing Maltese exposure (PT/IT/NO/NK).
- domain assumption The LM Evaluation Harness computes metrics correctly.
- domain assumption The FLOPs estimation method of Liu et al. (2022) is a valid measure for the efficiency comparison.
Cite this review
Pith. "Pith review of MELABenchv1: Benchmarking Large Language Models against Smaller Fine-Tuned Models for Low-Resource Maltese NLP." pith.science (2026). https://pith.science/paper/5NHQC7HQ
@misc{pith2026250604385,
author = {Pith},
title = {Pith review of: MELABenchv1: Benchmarking Large Language Models against Smaller Fine-Tuned Models for Low-Resource Maltese NLP},
year = {2026},
howpublished = {\url{https://pith.science/paper/5NHQC7HQ}},
note = {Machine review of arXiv:2506.04385}
}
read the original abstract
Large Language Models (LLMs) have demonstrated remarkable performance across various Natural Language Processing (NLP) tasks, largely due to their generalisability and ability to perform tasks without additional training. However, their effectiveness for low-resource languages remains limited. In this study, we evaluate the performance of 55 publicly available LLMs on Maltese, a low-resource language, using a newly introduced benchmark covering 11 discriminative and generative tasks. Our experiments highlight that many models perform poorly, particularly on generative tasks, and that smaller fine-tuned models often perform better across all tasks. From our multidimensional analysis, we investigate various factors impacting performance. We conclude that prior exposure to Maltese during pre-training and instruction-tuning emerges as the most important factor. We also examine the trade-offs between fine-tuning and prompting, highlighting that while fine-tuning requires a higher initial cost, it yields better performance and lower inference costs. Through this work, we aim to highlight the need for more inclusive language technologies and recommend that researchers working with low-resource languages consider more "traditional" language modelling approaches.
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[1]
Kurt Abela, Kurt Micallef, Marc Tanti, and Claudia Borg. 2024. https://doi.org/10.18653/v1/2024.loresmt-1.11 Tokenisation in machine translation does matter: The impact of different tokenisation approaches for M altese . In Proceedings of the Seventh Workshop on Technologies for Machine Translation of Low-Resource Languages (LoResMT 2024), pages 109--120,...
-
[2]
Alabi, Yanke Mao, Haonan Gao, and En-Shiun Annie Lee
David Ifeoluwa Adelani, Hannah Liu, Xiaoyu Shen, Nikita Vassilyev, Jesujoba O. Alabi, Yanke Mao, Haonan Gao, and En-Shiun Annie Lee. 2024. https://aclanthology.org/2024.eacl-long.14/ SIB -200: A simple, inclusive, and big evaluation dataset for topic classification in 200+ languages and dialects . In Proceedings of the 18th Conference of the European Chap...
2024
-
[3]
Kabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng, Krithika Ramesh, Prachi Jain, Akshay Nambi, Tanuja Ganu, Sameer Segal, Mohamed Ahmed, et al. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.258 MEGA : Multilingual evaluation of generative AI . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 4...
-
[4]
Mehdi Ali, Michael Fromm, Klaudia Thellmann, Jan Ebert, Alexander Arno Weber, Richard Rutmann, Charvi Jain, Max Lübbering, Daniel Steinigen, Johannes Leveling, et al. 2024. https://arxiv.org/abs/2410.03730 Teuken-7B-Base & Teuken-7B-Instruct : Towards E uropean LLMs . Preprint, arXiv:2410.03730
arXiv 2024
-
[5]
Viraat Aryabumi, John Dang, Dwarak Talupuru, Saurabh Dash, David Cairuz, Hangyu Lin, Bharat Venkitesh, Madeline Smith, Jon Ander Campos, Yi Chern Tan, et al. 2024. https://arxiv.org/abs/2405.15032 Aya 23 : Open weight releases to further multilingual progress . Preprint, arXiv:2405.15032
arXiv 2024
-
[6]
Akari Asai, Sneha Kudugunta, Xinyan Yu, Terra Blevins, Hila Gonen, Machel Reid, Yulia Tsvetkov, Sebastian Ruder, and Hannaneh Hajishirzi. 2024. https://doi.org/10.18653/v1/2024.naacl-long.100 BUFFET : Benchmarking large language models for few-shot cross-lingual transfer . In Proceedings of the 2024 Conference of the North American Chapter of the Associat...
-
[7]
Dennis Aumiller, Ashish Chouhan, and Michael Gertz. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.519 EUR -lex-sum: A multi- and cross-lingual dataset for long-form summarization in the legal domain . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 7626--7639, Abu Dhabi, United Arab Emirates. Associatio...
-
[8]
Simone Balloccu, Patr \'i cia Schmidtov \'a , Mateusz Lango, and Ondrej Dusek. 2024. https://aclanthology.org/2024.eacl-long.5/ Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source LLM s . In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), ...
2024
Show all 53 references
-
[9]
Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, and Madian Khabsa. 2024. https://doi.org/10.18653/v1/2024.acl-long.44 The belebele benchmark: a parallel reading comprehension d...
2024 doi
-
[10]
BigScience Workshop , Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, et al. 2023. https://arxiv.org/abs/2211.05100 BLOOM : A 176 B -parameter open-access multilingual language m...
2023 arXiv
-
[11]
Terra Blevins and Luke Zettlemoyer. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.233 Language contamination helps explains the cross-lingual capabilities of E nglish pretrained models . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Process...
2022 doi
-
[12]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. https://proceedings.neurips.cc/paper_files/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf Language...
2020
-
[13]
Ilias Chalkidis, Manos Fergadiotis, and Ion Androutsopoulos. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.559 M ulti EURLEX - a multi-lingual and multi-label legal document classification dataset for zero-shot cross-lingual transfer . In Proceedings of the 2021 Conference...
2021 doi
-
[14]
Chau, Lucy H
Ethan C. Chau, Lucy H. Lin, and Noah A. Smith. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.118 Parsing with multilingual BERT , a small corpus, and a small treebank . In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1324--1334, Online. ...
2020 doi
-
[15]
Amit Kumar Chaudhary, Kurt Micallef, and Claudia Borg. 2024. https://aclanthology.org/2024.lrec-main.1414/ Topic classification and headline generation for M altese using a public news corpus . In Proceedings of the 2024 Joint International Conference on Computational Linguist...
2024
-
[16]
Liam Cripwell, Anya Belz, Claire Gardent, Albert Gatt, Claudia Borg, Marthese Borg, John Judge, Michela Lorandi, Anna Nikiforovskaya, and William Soto Martinez. 2023. https://aclanthology.org/2023.mmnlg-1.6/ The 2023 W eb NLG shared task on low resource languages. overview and...
2023
-
[17]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423 BERT : Pre-training of deep bidirectional transformers for language understanding . In Proceedings of the 2019 Conference of the North A merican Chapter of the Associat...
2019 doi
-
[18]
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac'h, et al. 2024. https://doi.org/10.5281/zenodo.12608602 A framework for few-shot language model evaluation
2024 doi
-
[19]
Gemma Team , Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al. 2024. https://arxiv.org/abs/2408.00118 Gemma 2 : Improving open language models at a practical size . ...
2024 arXiv
-
[20]
Aitor Gonzalez-Agirre, Marc Pàmies, Joan Llop, Irene Baucells, Severino Da Dalt, Daniel Tamayo, José Javier Saiz, Ferran Espuña, Jaume Prats, Javier Aula-Blasco, et al. 2025. https://arxiv.org/abs/2502.08489 Salamandra technical report . Preprint, arXiv:2502.08489
2025 arXiv
-
[21]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. 2024. https://arxiv.org/abs/2407.21783 The Llama 3 herd of models . Preprint, arXiv:2407.21783
2024 arXiv
-
[22]
Xu, Jun Araki, and Graham Neubig
Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig. 2020. https://doi.org/10.1162/tacl_a_00324 How can we know what language models know? Transactions of the Association for Computational Linguistics, 8:423--438
2020 doi
-
[23]
Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. https://arxiv.org/abs/2001.08361 Scaling laws for neural language models . Preprint, arXiv:2001.08361
2020 arXiv
-
[24]
Viet Dac Lai, Nghia Ngo, Amir Pouran Ben Veyseh, Hieu Man, Franck Dernoncourt, Trung Bui, and Thien Huu Nguyen. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.878 C hat GPT beyond E nglish: Towards a comprehensive evaluation of large language models in multilingual lear...
2023 doi
-
[25]
Haonan Li, Fajri Koto, Minghao Wu, Alham Fikri Aji, and Timothy Baldwin. 2023. https://arxiv.org/abs/2305.15011 Bactrian-X : Multilingual replicable instruction-following models with low-rank adaptation . Preprint, arXiv:2305.15011
2023 arXiv
-
[26]
Peiqin Lin, Shaoxiong Ji, Jörg Tiedemann, André F. T. Martins, and Hinrich Schütze. 2024. https://arxiv.org/abs/2401.13303 MaLA-500 : Massive language adaptation of large language models . Preprint, arXiv:2401.13303
2024 arXiv
-
[27]
Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, et al. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.616 Few-shot learning with multilingual generative language models . In Proceedi...
2022 doi
-
[28]
Haokun Liu, Derek Tam, Muqeeth Mohammed, Jay Mohta, Tenghao Huang, Mohit Bansal, and Colin Raffel. 2022. https://openreview.net/forum?id=rBCvMG-JsPd Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning . In Advances in Neural Information Proc...
2022
-
[29]
Chunlan Ma, Ayyoob ImaniGooghari, Haotian Ye, Renhao Pei, Ehsaneddin Asgari, and Hinrich Schütze. 2024. https://arxiv.org/abs/2305.08487 Taxi1500 : A multilingual dataset for text classification in 1500 languages . Preprint, arXiv:2305.08487
2024 arXiv
-
[30]
Antonio Mart \'i nez-Garc \'i a, Toni Badia, and Jeremy Barnes. 2021. https://doi.org/10.18653/v1/2021.acl-long.244 Evaluating morphological typology in zero-shot cross-lingual transfer . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistic...
2021 doi
-
[31]
Guerreiro, Ricardo Rei, Duarte M
Pedro Henrique Martins, Patrick Fernandes, João Alves, Nuno M. Guerreiro, Ricardo Rei, Duarte M. Alves, José Pombal, Amin Farajian, Manuel Faysse, Mateusz Klimaszewski, et al. 2025. https://doi.org/10.1016/j.procs.2025.02.260 Eurollm: Multilingual language models for europe . ...
2025 doi
-
[32]
Kurt Micallef, Albert Gatt, Marc Tanti, Lonneke van der Plas, and Claudia Borg. 2022. https://doi.org/10.18653/v1/2022.deeplo-1.10 Pre-training data quality and quantity for a low-resource language: New corpus and BERT models for M altese . In Proceedings of the Third Workshop...
2022 doi
-
[33]
Mistral AI Team . 2024. Un M inistral, des M inistraux. https://mistral.ai/news/ministraux/. Accessed: 2024-12-20
2024
-
[34]
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, et al. 2023. https://doi.org/10.18653/v1/2023.acl-long.891 Crosslingual generalization through multitask finetuning . ...
2023 doi
-
[35]
Benjamin Muller, Antonios Anastasopoulos, Beno \^i t Sagot, and Djam \'e Seddah. 2021. https://doi.org/10.18653/v1/2021.naacl-main.38 When being unseen from m BERT is just the beginning: Handling new languages with multilingual language models . In Proceedings of the 2021 Conf...
2021 doi
-
[36]
Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, et al
NLLB Team , Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, et al. 2022. https://arxiv.org/abs/2207.04672 No language left behind: Scaling human-centered machine translation . Preprint, ...
2022 arXiv
-
[37]
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, et al. 2024. https://arxiv.org/abs/2303.08774 GPT-4 technical report . Preprint, arXiv:2303.08774
2024 arXiv
-
[38]
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023. https://openreview.net/forum?id=HPuSIXJaa9 Direct preference optimization: Your language model is secretly a reward model . In Thirty-seventh Conference on Neural Infor...
2023
-
[39]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. http://jmlr.org/papers/v21/20-074.html Exploring the limits of transfer learning with a unified text-to-text transformer . Journal of Machine Lea...
2020
-
[40]
Michael Rosner and Claudia Borg. 2023. https://doi.org/10.1007/978-3-031-28819-7_27 Language Report Maltese . In Georg Rehm and Andy Way, editors, European Language Equality: A Strategic Agenda for Digital Language Equality, pages 183--186. Springer International Publishing, Cham
2023 doi
-
[41]
Logan IV, Eric Wallace, and Sameer Singh
Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.346 A uto P rompt: E liciting K nowledge from L anguage M odels with A utomatically G enerated P rompts . In Proceedings of the 2020 Conference o...
2020 doi
-
[42]
Oleh Shliazhko, Alena Fenogenova, Maria Tikhonova, Anastasia Kozlova, Vladislav Mikhailov, and Tatiana Shavrina. 2024. https://doi.org/10.1162/tacl_a_00633 m GPT : Few-shot learners go multilingual . Transactions of the Association for Computational Linguistics, 12:58--79
2024 doi
-
[43]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. https://arxiv.org/abs/2307.09288 Llama 2 : Open foundation and fine-tuned chat models . Preprint, arXiv:2307.09288
2023 arXiv
-
[44]
Ahmet \"U st \"u n, Viraat Aryabumi, Zheng Yong, Wei-Yin Ko, Daniel D ' souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, et al. 2024. https://doi.org/10.18653/v1/2024.acl-long.845 Aya model: An instruction finetuned open-access multilingual lan...
2024 doi
-
[45]
Xiangpeng Wei, Haoran Wei, Huan Lin, Tianhao Li, Pei Zhang, Xingzhang Ren, Mei Li, Yu Wan, Zhiwei Cao, Binbin Xie, et al. 2023. https://arxiv.org/abs/2307.06018 PolyLM : An open source polyglot large language model . Preprint, arXiv:2307.06018
2023 arXiv
-
[46]
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, et al. 2020. https://doi.org/10.18653/v1/2020.emnlp-demos.6 Transformers: State-of-the-art natural language processing . In Proceed...
2020 doi
-
[47]
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. https://doi.org/10.18653/v1/2021.naacl-main.41 m T 5: A massively multilingual pre-trained text-to-text transformer . In Proceedings of the 2021 Conferenc...
2021 doi
-
[48]
Miaoran Zhang, Vagrant Gautam, Mingyang Wang, Jesujoba Alabi, Xiaoyu Shen, Dietrich Klakow, and Marius Mosbach. 2024. https://doi.org/10.18653/v1/2024.findings-acl.438 The impact of demonstrations on multilingual in-context learning: A multidimensional analysis . In Findings o...
2024 doi
-
[49]
Ruochen Zhang, Samuel Cahyawijaya, Jan Christian Blaise Cruz, Genta Winata, and Alham Fikri Aji. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.774 Multilingual large language models are not (yet) code-switchers . In Proceedings of the 2023 Conference on Empirical Methods i...
2023 doi
-
[50]
Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. https://proceedings.mlr.press/v139/zhao21c.html Calibrate before use: Improving few-shot performance of language models . In Proceedings of the 38th International Conference on Machine Learning, volume 139 ...
2021
-
[51]
Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B
Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul Christiano, and Geoffrey Irving. 2020. https://arxiv.org/abs/1909.08593 Fine-tuning language models from human preferences . Preprint, arXiv:1909.08593
2020 arXiv
-
[52]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[53]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.