REVIEW 1 major objections 5 minor 66 references
Overview of the First Workshop on Language Models for Low-Resource Languages (LoResLM 2025)
T0 review · 1 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The first Workshop on Language Models for Low-Resource Languages reports 35 accepted papers from 52 submissions, spanning eight language families and 13 NLP research areas.
desk verdict A competent, honest workshop report whose coverage statistics are not externally reproducible because 'low-resource' is context-dependent; useful as an index, not as a research contribution. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument rests on two classification tables. Table 1 assigns each accepted paper to a language family, with the Indo-European family split into branches, yielding a count of 28 low-resource languages plus a 'Multiple' category for papers working on more than five languages. Table 2 assigns each paper to one or more of 13 NLP research areas drawn from the call-for-papers topics of leading conferences in 2024. These tables are the machinery because the paper's claims about coverage and about where the community's effort concentrates are simply the sums and intersections of these mappings.
What would settle it
Inspect the papers cited in Table 1 and re-classify their languages against the stated context-dependent criterion; a different count of low-resource languages (28) or families (8) would refute the report's central claim.
Extended reading notes
Core claim
The central discovery is the workshop's own uptake and coverage: 35 accepted papers from 52 submissions, 28 low-resource languages grouped into eight families, and 13 research areas, with Language Modelling (11 papers) and Machine Translation and Translation Aids (6 papers) as the dominant topics. The paper also documents that some languages usually treated as high-resource, such as Arabic, German, and Portuguese, were counted as low-resource in particular domains, dialects, or research areas. The authors present this distribution as evidence that the workshop succeeded in gathering a broad range of work and as a baseline for steering future editions toward underrepresented families such as Uralic, Dravidian, and Indigenous languages.
Load-bearing premise
The paper assumes that the organizers' context-dependent classification of a language as low-resource is consistent enough that the reported counts of families, languages, and areas are meaningful and reproducible.
Editorial extensions
If this is right
- If the reported uptake is accurate, there is enough active research to sustain a recurring workshop on low-resource language models.
- The concentration of papers in Language Modelling and Machine Translation suggests those subfields are the entry points for low-resource work, while speech, information extraction, and dialogue remain open niches.
- The organizers' stated plans imply future editions will actively recruit work on Uralic, Dravidian, and Indigenous languages of the Americas.
- The inclusion of context-dependent low-resource languages means the workshop's scope is broader than a list of conventionally resource-poor languages.
Reading between the lines
- Because 'low-resource' is applied contextually rather than by a fixed threshold, the headline counts are not directly comparable with other workshops that use a stricter definition of low-resource languages.
- The over-representation of Indo-European languages within the low-resource set suggests that the workshop's coverage, while broad, still tilts toward languages familiar to a Western research community; a strict resource-based definition might shift the balance.
- Papers classified as 'Multiple' (more than five languages across several families) may be demonstrating task-level methods that transfer across languages, which could be a separate line of work from language-specific resource creation.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper is the organizers' overview of the first Workshop on Language Models for Low-Resource Languages (LoResLM 2025), held at COLING 2025 in Abu Dhabi. It reports that the workshop received 52 submissions and accepted 35 papers (28 long and 7 short), that the accepted papers cover 28 low-resource languages from eight language families, and that they span 13 NLP research areas. Section 2.1 gives a language-level breakdown in Table 1, Section 2.2 gives a research-area breakdown in Table 2, and the reference list contains the 35 accepted workshop papers. The paper frames the workshop as a step toward reducing the English-centric bias of NLP research.
Significance. The paper is a descriptive workshop report rather than a research contribution, so its value is primarily archival and documentary. Its strengths are the internal consistency of the headline statistics: the reference list contains exactly the 35 LoResLM papers, the 28-language figure is reproducible by counting the language rows in Table 1 (excluding 'Multiple'), the eight family labels match the table, and the 13 area headers match Table 2. The authors are also transparent about excluding high-resource comparison languages and about the 'Multiple' category. The main substantive weakness is that the 'low-resource' classification is context-dependent and not operationalized, which makes the language and family counts non-reproducible; there are also small completeness and presentation issues in Table 2. These concerns are fixable and do not undermine the archival value of the overview.
major comments (1)
- [Section 2.1 (Table 1) and abstract] The abstract's headline figure of '28 low-resource languages' (repeated in Section 2.1 and the conclusions) is not externally reproducible because Section 2.1 provides no operational definition of 'low-resource'. The text concedes that typically high-resource languages such as Arabic and German were counted when resources were limited in a specific domain, dialect, or research area, but it does not specify which of the 28 languages were included on this basis or which resource criterion was applied. As a result, the count is sensitive to the organizers' judgment: Table 1 includes, for example, German, Italian, Portuguese, and Korean without any stated low-resource context, and excluding any of these would change the reported total (and potentially the number of families, e.g., if Koreanic were removed). Since this count is the main quantitative summary of the workshop, please add a per-language note giving the relevant context or a defined resource threshold.
minor comments (5)
- [Section 2.2 (Table 2)] The sentence 'Table 2 shows the distribution of the accepted papers' is not fully accurate because Veitsman and Hartmann (2025) is excluded from the table; the note after the table acknowledges this, but the caption and surrounding text should be amended to reflect the exception.
- [Section 2.1 (Table 1) and Section 2.2 (Table 2)] Turumtaev (2025) appears in Table 2 but has no language row in Table 1 and is not listed in the 'Multiple' row; please clarify its language coverage or add it to the appropriate row so the two tables can be reconciled.
- [Section 2.1] The category 'Isolate' is described as a language family, but language isolates are by definition not members of a family; consider renaming the category 'Language isolates' or adding an explanatory note.
- [Section 3 (Conclusions)] The final sentence says 'empower linguistic diversity for millions of low-resource languages'; since the paper itself notes there are approximately 7,000 spoken languages, this should be 'millions of speakers of low-resource languages' or similar.
- [Section 2.2 (Table 2)] In the linearized text, the check marks in Table 2 are not visually aligned with the column headings, so the per-area counts (e.g., eleven papers in 'Language Modelling') cannot be verified from the text alone; please ensure the published table has clear column alignment.
Circularity Check
No circularity: the workshop statistics are self-reported event data, not the output of any derivation, fit, or self-citation chain.
full rationale
This paper is a workshop overview. Its central claims are counts of submissions, accepted papers, language families, languages, and research areas. These are self-reported summaries of the organizers' own records and of the accepted papers' contents, not predictions derived from a model or from cited work. The reference list contains exactly the 35 LoResLM 2025 papers discussed, so the acceptance statistics are internally checkable against the proceedings themselves. The authors' prior publications appear only as background citations about data scarcity and Sinhala resources; they do not carry any load-bearing argument. The one genuine weakness, flagged in Section 2.1, is that languages such as Arabic, German, Italian, and Portuguese are counted as low-resource when resources are limited in a particular domain, dialect, or research area, without a stated threshold or external resource measure. That is a reproducibility and operationalization concern about how the coverage statistics are defined, not a circularity: the classification is an input assumption, and the counts are summaries of that assumption rather than results forced to equal an input by construction. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors' prior work, and no ansatz is smuggled in via citation. Accordingly, the circularity score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Accepted papers' self-reported language and task categories are accurate.
- domain assumption Standard genealogical language family taxonomy with isolates treated as a family.
- ad hoc to paper Context-dependent low-resource status can include typically high-resource languages.
Cite this review
Pith. "Pith review of Overview of the First Workshop on Language Models for Low-Resource Languages (LoResLM 2025)." pith.science (2026). https://pith.science/paper/GXCQWA4L
@misc{pith2026241216365,
author = {Pith},
title = {Pith review of: Overview of the First Workshop on Language Models for Low-Resource Languages (LoResLM 2025)},
year = {2026},
howpublished = {\url{https://pith.science/paper/GXCQWA4L}},
note = {Machine review of arXiv:2412.16365}
}
read the original abstract
The first Workshop on Language Models for Low-Resource Languages (LoResLM 2025) was held in conjunction with the 31st International Conference on Computational Linguistics (COLING 2025) in Abu Dhabi, United Arab Emirates. This workshop mainly aimed to provide a forum for researchers to share and discuss their ongoing work on language models (LMs) focusing on low-resource languages, following the recent advancements in neural language models and their linguistic biases towards high-resource languages. LoResLM 2025 attracted notable interest from the natural language processing (NLP) community, resulting in 35 accepted papers from 52 submissions. These contributions cover a broad range of low-resource languages from eight language families and 13 diverse research areas, paving the way for future possibilities and promoting linguistic inclusivity in NLP.
Figures
Reference graph
Works this paper leans on
-
[1]
Kabir Ahuja, Shanu Kumar, Sandipan Dandapat, and Monojit Choudhury. 2022. https://doi.org/10.18653/v1/2022.acl-long.374 Multi Task Learning For Zero Shot Performance Prediction of Multilingual Models . In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5454--5467, Dublin, Ireland. Asso...
-
[2]
Sadia Alam, Md Farhan Ishmam, Navid Hasin Alvee, Md Shahnewaz Siddique, Md Azam Hossain, and Abu Raihan Mostofa Kamal. 2025. BnSentMix: A Diverse Bengali-English Code-Mixed Dataset for Sentiment Analysis . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), Abu Dhabi, UAE. International Committee on Computationa...
work page 2025
-
[3]
Muhammad Saad Amin, Luca Anselma, and Alessandro Mazzei. 2025. Exploiting Task Reversibility of DRS Parsing and Generation: Challenges and Insights from a Multi-lingual Perspective . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), Abu Dhabi, UAE. International Committee on Computational Linguistics
work page 2025
-
[4]
Sina Bagheri Nezhad, Ameeta Agrawal, and Rhitabrat Pokharel. 2025. Beyond Data Quantity: Key Factors Driving Performance in Multilingual Language Models . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), Abu Dhabi, UAE. International Committee on Computational Linguistics
work page 2025
-
[5]
Emily M. Bender. 2011. https://doi.org/10.33011/lilt.v6i.1239 On Achieving and Evaluating Language-Independence in NLP . Linguistic Issues in Language Technology, 6
-
[6]
Damian Blasi, Antonios Anastasopoulos, and Graham Neubig. 2022. https://doi.org/10.18653/v1/2022.acl-long.376 Systematic Inequalities in Language Technology Performance across the World ' s Languages . In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5486--5505, Dublin, Ireland. Asso...
-
[7]
Latofat Bobojonova, Arofat Akhundjanova, Phil Sidney Ostheimer, and Sophie Fellenz. 2025. BBPOS: BERT-based Part-of-Speech Tagging for Uzbek . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), Abu Dhabi, UAE. International Committee on Computational Linguistics
work page 2025
-
[8]
Bharathi Raja Chakravarthi, B Bharathi, John P McCrae, Manel Zarrouk, Kalika Bali, and Paul Buitelaar, editors. 2022. https://aclanthology.org/2022.ltedi-1.0 Proceedings of the Second Workshop on Language Technology for Equality, Diversity and Inclusion . Association for Computational Linguistics, Dublin, Ireland
work page 2022
Show all 66 references
-
[9]
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzm \'a n, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. https://doi.org/10.18653/v1/2020.acl-main.747 Unsupervised Cross-lingual Representation Learning ...
2020 doi
-
[10]
Jan Christian Blaise Cruz. 2025. Extracting General-use Transformers for Low-resource Languages via Knowledge Distillation . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), Abu Dhabi, UAE. International Committee on Computational...
2025
-
[11]
Yuqian Dai, Chun Fai Chan, Ying Ki Wong, and Tsz Ho Pun. 2025. Next-Level Cantonese-to-Mandarin Translation: Fine-Tuning and Post-Processing with LLMs . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), Abu Dhabi, UAE. Internationa...
2025
-
[12]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423 BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding . In Proceedings of the 2019 Conference of the North A merican Chapter of the Associat...
2019 doi
-
[13]
Vikrant Dewangan, Bharath Raj S, Garvit Suri, and Raghav Sonavane. 2025. When Every Token Counts: Optimal Segmentation for Low-Resource Language Models . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), Abu Dhabi, UAE. Internation...
2025
-
[14]
Vinura Dhananjaya, Piyumal Demotte, Surangika Ranathunga, and Sanath Jayasena. 2022. https://aclanthology.org/2022.lrec-1.803 BERT ifying S inhala - A Comprehensive Analysis of Pre-trained Language Models for S inhala Text Classification . In Proceedings of the Thirteenth Lang...
2022
-
[15]
Alphaeus Dmonte, Shrey Satapara, Rehab Alsudais, Tharindu Ranasinghe, and Marcos Zampieri. 2025. Does Machine Translation Impact Offensive Language Identification? The Case of Indo-Aryan Languages . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languag...
2025
-
[16]
Patel, Joon Young Doh, Eid Rodan, Kevin Zhu, and Sean O'Brien
Sundesh Donthi, Maximilian Spencer, Om B. Patel, Joon Young Doh, Eid Rodan, Kevin Zhu, and Sean O'Brien. 2025. Improving LLM Abilities in Idiomatic Translation . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), Abu Dhabi, UAE. Int...
2025
-
[17]
Lance Calvin Lim Gamboa and Mark Lee. 2025. Filipino Benchmarks for Measuring Sexist and Homophobic Bias in Multilingual Language Models from Southeast Asia . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), Abu Dhabi, UAE. Intern...
2025
-
[18]
Zahra Habibzadeh and Masoud Asadpour. 2025. Using Language Models for Assessment of Users' Satisfaction with Their Partner in Persian . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), Abu Dhabi, UAE. International Committee on Co...
2025
-
[19]
Anika Harju and Rob van der Goot. 2025. How to age BERT Well: Continuous Training for Historical Language Adaptation . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), Abu Dhabi, UAE. International Committee on Computational Linguistics
2025
-
[20]
Hansi Hettiarachchi, Mariam Adedoyin-Olowe, Jagdev Bhogal, and Mohamed Medhat Gaber. 2023. TTL : transformer-based two-phase transfer learning for cross-lingual news event detection . International Journal of Machine Learning and Cybernetics, 14(8):2739--2760
2023
-
[21]
Hansi Hettiarachchi, Damith Premasiri, Lasitha Randunu Chandrakantha Uyangodage, and Tharindu Ranasinghe. 2024. https://aclanthology.org/2024.lrec-main.1076 NS ina: A news corpus for S inhala . In Proceedings of the 2024 Joint International Conference on Computational Linguist...
2024
-
[22]
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7B . arXiv preprint arXiv:2310.06825
2023 arXiv
-
[23]
Daria Kryvosheieva and Roger Levy. 2025. Controlled Evaluation of Syntactic Knowledge in Multilingual Language Models . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), Abu Dhabi, UAE. International Committee on Computational Linguistics
2025
-
[24]
Uriel Anderson Lasheras and Vladia Pinheiro. 2025. CaLQuest.PT: Towards the Collection and Evaluation of Natural Causal Ladder Questions in Portuguese for AI Agents . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), Abu Dhabi, UAE...
2025
-
[25]
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. https://doi.org/10.18653/v1/2020.acl-main.703 BART : Denoising sequence-to-sequence pre-training for natural language generation, translatio...
2020 doi
-
[26]
Constantine Lignos, Nolan Holley, Chester Palen-Michel, and Jonne S \"a lev \"a . 2022. https://doi.org/10.18653/v1/2022.findings-acl.44 Toward More Meaningful Resources for Lower-resourced Languages . In Findings of the Association for Computational Linguistics: ACL 2022, pag...
2022 doi
-
[27]
Y Liu, M Ott, N Goyal, J Du, M Joshi, D Chen, O Levy, M Lewis, L Zettlemoyer, and V Stoyanov. 2019. Roberta: A robustly optimized bert pretraining approach. arXiv preprint arXiv:1907.11692, 364
2019 arXiv
-
[28]
Yifan Liu, Gelila Tilahun, Xinxiang Gao, Qianfeng Wen, and Michael Gervers. 2025. A Comparative Study of Static and Contextual Embeddings for Analyzing Semantic Changes in Medieval Latin Charters . In Proceedings of the 1st Workshop on Language Models for Low-Resource Language...
2025
-
[29]
Alexandre Magueresse, Vincent Carles, and Evan Heetderks. 2020. Low-resource languages: A review of past work and future challenges . arXiv preprint arXiv:2006.07264
2020 arXiv
-
[30]
Maria Keet, Imaan Sayed, and Alexander Van Der Leek
Zola Mahlaza, C. Maria Keet, Imaan Sayed, and Alexander Van Der Leek. 2025. IsiZulu Noun Classification Based on Replicating the Ensemble Approach for Runyankore . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), Abu Dhabi, UAE. I...
2025
-
[31]
Alexis Matzopoulos, Charl Hendriks, Hishaam Mahomed, and Francois Meyer. 2025. BabyLMs for isiXhosa: Data-Efficient Language Modelling in a Low-Resource Context . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), Abu Dhabi, UAE. In...
2025
-
[32]
Maite Melero, Sakriani Sakti, and Claudia Soria, editors. 2024. https://aclanthology.org/2024.sigul-1.0 Proceedings of the 3rd Annual Meeting of the Special Interest Group on Under-resourced Languages @ LREC-COLING 2024 . ELRA and ICCL, Torino, Italia
2024
-
[33]
Ibrahim Merad, Amos Wolf, Ziad Mazzawi, and Yannick Léo. 2025. Language verY Rare for All . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), Abu Dhabi, UAE. International Committee on Computational Linguistics
2025
-
[34]
Shervin Minaee, Tomas Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Amatriain, and Jianfeng Gao. 2024. Large language models: A survey . arXiv preprint arXiv:2402.06196
2024 arXiv
-
[35]
Hojjat Mokhtarabadi, Ziba Zamani, Abbas Maazallahi, and Mohammad Hossein Manshaei. 2025. Empowering Persian LLMs for Instruction Following: A Novel Dataset and Training Approach . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), A...
2025
-
[36]
Atharva Mutsaddi and Aditya Prashant Choudhary. 2025. Enhancing Plagiarism Detection in Marathi with a Weighted Ensemble of TF-IDF and BERT Embeddings for Low-Resource Language Processing . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoRes...
2025
-
[37]
Al-Batati, Arwa Alsehibani, Nour Qandos, Omar Elshehy, Mohamed Abdelkader, and Anis Koubaa
Omer Nacar, Serry Taiseer Sibaee, Samar Ahmed, Safa Ben Atitallah, Adel Ammar, Yasser Alhabashi, Abdulrahman S. Al-Batati, Arwa Alsehibani, Nour Qandos, Omar Elshehy, Mohamed Abdelkader, and Anis Koubaa. 2025. Towards Inclusive Arabic LLMs: A Culturally Aligned Benchmark in Ar...
2025
-
[38]
Dat Quoc Nguyen and Anh Tuan Nguyen. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.92 P ho BERT : Pre-trained language models for V ietnamese . In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1037--1042, Online. Association for Computati...
2020 doi
-
[39]
Ojha, Chao-hong Liu, Ekaterina Vylomova, Flammie Pirinen, Jade Abbott, Jonathan Washington, Nathaniel Oco, Valentin Malykh, Varvara Logacheva, and Xiaobing Zhao, editors
Atul Kr. Ojha, Chao-hong Liu, Ekaterina Vylomova, Flammie Pirinen, Jade Abbott, Jonathan Washington, Nathaniel Oco, Valentin Malykh, Varvara Logacheva, and Xiaobing Zhao, editors. 2023. https://aclanthology.org/2023.loresmt-1.0 Proceedings of the Sixth Workshop on Technologies...
2023
-
[40]
Bissyandé
Maimouna Ouattara, Abdoul Kader Kaboré, Jacques Klein, and Tegawendé F. Bissyandé. 2025. Bridging Literacy Gaps in African Informal Business Management with Low-Resource Conversational Agents . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (L...
2025
-
[41]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res., 21(1)
2020
-
[42]
Tharindu Ranasinghe, Isuri Anuradha, Damith Premasiri, Kanishka Silva, Hansi Hettiarachchi, Lasitha Uyangodage, and Marcos Zampieri. 2024. Sold: Sinhala offensive language dataset . Language Resources and Evaluation, pages 1--41
2024
-
[43]
Surangika Ranathunga, En-Shiun Annie Lee, Marjana Prifti Skenduli, Ravi Shekhar, Mehreen Alam, and Rishemjit Kaur. 2023. https://doi.org/10.1145/3567592 Neural Machine Translation for Low-resource Languages: A Survey . ACM Comput. Surv., 55(11)
2023 doi
-
[44]
Maciej Rapacz and Aleksander Smywiński-Pohl. 2025. Low-Resource Interlinear Translation: Morphology-Enhanced Neural Models for Ancient Greek . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), Abu Dhabi, UAE. International Committe...
2025
-
[45]
Sebastian Ruder, Ivan Vuli \'c , and Anders S gaard. 2022. https://doi.org/10.18653/v1/2022.findings-acl.184 Square One Bias in NLP : Towards a Multi-Dimensional Exploration of the Research Manifold . In Findings of the Association for Computational Linguistics: ACL 2022, page...
2022 doi
-
[46]
Jayanta Sadhu, Maneesha Rani Saha, and Rifat Shahriyar. 2025. Social Bias in Large Language Models for Bangla: An Empirical Study on Gender and Religious Bias . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), Abu Dhabi, UAE. Inte...
2025
-
[47]
Sani Abdullahi Sani, Shamsuddeen Hassan Muhammad, and Devon Jarvis. 2025. Investigating the Impact of Language-Adaptive Fine-Tuning on Sentiment Analysis in Hausa Language Using AfriBERTa . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoRes...
2025
-
[48]
Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ili \'c , Daniel Hesslow, Roman Castagn \'e , Alexandra Sasha Luccioni, Fran c ois Yvon, et al. 2022. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100
2022 arXiv
-
[49]
Guokan Shang, Hadi Abdine, Yousef Khoubrane, Amr Mohamed, Yassine ABBAHADDOU, Sofiane Ennadir, Imane Momayiz, Xuguang Ren, Eric Moulines, Preslav Nakov, Michalis Vazirgiannis, and Eric Xing. 2025. Atlas-Chat: Adapting Large Language Models for Low-Resource Moroccan Arabic Dial...
2025
-
[50]
Claude E Shannon. 1951. Prediction and entropy of printed English . Bell system technical journal, 30(1):50--64
1951
-
[51]
Archchana Sindhujan, Diptesh Kanojia, Constantin Orasan, and Shenbin Qian. 2025. When LLMs Struggle: Reference-less Translation Evaluation for Low-resource Languages . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), Abu Dhabi, UA...
2025
-
[52]
Tsegaye Misikir Tashu and Andreea Ioana Tudor. 2025. Mapping Cross-Lingual Sentence Representations for Low-Resource Language Pairs Using Pre-trained Language Models . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), Abu Dhabi, UA...
2025
-
[53]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models . arXiv preprint arXiv:2307.09288
2023 arXiv
-
[54]
Van-Hien Tran, Raj Dabre, Hour Kaing, Haiyue Song, Hideki Tanaka, and Masao Utiyama. 2025. Exploiting Word Sense Disambiguation in Large Language Models for Machine Translation . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), Ab...
2025
-
[55]
Galim Turumtaev. 2025. Stop Jostling: Adaptive Negative Sampling Reduces the Marginalization of Low-Resource Language Tokens by Cross-Entropy Loss . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), Abu Dhabi, UAE. International Co...
2025
-
[56]
Daan van Esch, Tamar Lucassen, Sebastian Ruder, Isaac Caswell, and Clara Rivera. 2022. https://aclanthology.org/2022.lrec-1.538 Writing System and Speaker Metadata for 2,800+ Language Varieties . In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pa...
2022
-
[57]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. 2017. https://proceedings.neurips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf Attention is All you Need . In Advances in Neural Informa...
2017
-
[58]
Yana Veitsman and Mareike Hartmann. 2025. Recent Advancements and Challenges of Turkic Central Asian Language Processing . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), Abu Dhabi, UAE. International Committee on Computational L...
2025
-
[59]
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. https://doi.org/10.18653/v1/2021.naacl-main.41 m T 5: A massively multilingual pre-trained text-to-text transformer . In Proceedings of the 2021 Conferenc...
2021 doi
-
[60]
Kamyar Zeinalipour, Neda Jamshidi, Fahimeh Akbari, Marco Maggini, Monica Bianchini, and Marco Gori. 2025 a . PersianMCQ-Instruct: A Comprehensive Resource for Generating Multiple-Choice Questions in Persian . In Proceedings of the 1st Workshop on Language Models for Low-Resour...
2025
-
[61]
Kamyar Zeinalipour, Moahmmad Saad, Marco Maggini, and Marco Gori. 2025 b . From Arabic Text to Puzzles: LLM-Driven Development of Arabic Educational Crosswords . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), Abu Dhabi, UAE. Int...
2025
-
[62]
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023. A survey of large language models . arXiv preprint arXiv:2303.18223
2023 arXiv
-
[63]
Hongpu Zhu, Yuqi Liang, Wenjing Xu, and Hongzhi Xu. 2025. Evaluating Large Language Models for In-Context Learning of Linguistic Patterns In Unseen Low Resource Languages . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), Abu Dhab...
2025
-
[64]
Matt, and Bela Gipp
Anastasia Zhukova, Christian E. Matt, and Bela Gipp. 2025. Automated Collection of Evaluation Dataset for Semantic Search in Low-Resource Domain Language . In Proceedings of the 1st Workshop on Language Models for Low-Resource Languages (LoResLM2025), Abu Dhabi, UAE. Internati...
2025
-
[65]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[66]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.