Pith. sign in

REVIEW 1 cited by

Prompt, Translate, Fine-Tune, Re-Initialize, or Instruction-Tune? Adapting LLMs for In-Context Learning in Low-Resource Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.19187 v1 pith:A5YBG6AR submitted 2025-06-23 cs.CL

classification cs.CL
keywords languagesin-contextlearningllmsadaptationlow-resourcepromptingtrained
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

LLMs are typically trained in high-resource languages, and tasks in lower-resourced languages tend to underperform the higher-resource language counterparts for in-context learning. Despite the large body of work on prompting settings, it is still unclear how LLMs should be adapted cross-lingually specifically for in-context learning in the low-resource target languages. We perform a comprehensive study spanning five diverse target languages, three base LLMs, and seven downstream tasks spanning over 4,100 GPU training hours (9,900+ TFLOPs) across various adaptation techniques: few-shot prompting, translate-test, fine-tuning, embedding re-initialization, and instruction fine-tuning. Our results show that the few-shot prompting and translate-test settings tend to heavily outperform the gradient-based adaptation methods. To better understand this discrepancy, we design a novel metric, Valid Output Recall (VOR), and analyze model outputs to empirically attribute the degradation of these trained models to catastrophic forgetting. To the extent of our knowledge, this is the largest study done on in-context learning for low-resource languages with respect to train compute and number of adaptation techniques considered. We make all our datasets and trained models available for public use.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TigerCoder: A Novel Suite of LLMs for Code Generation in Bangla

    cs.CL 2025-09 conditional novelty 6.0 of 10

    TigerCoder is a fine-tuned Bangla code-generation LLM family that posts 0.82 Pass@1 on the new MBPP-Bangla benchmark, but its gains partly come from model selection on the test set.

Reference graph

Works this paper leans on

74 extracted references · 5 canonical work pages · cited by 1 Pith paper

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    David Ifeoluwa Adelani, Jade Abbott, Graham Neubig, Daniel D'souza, Julia Kreutzer, Constantine Lignos, Chester Palen-Michel, Happy Buzaaba, Shruti Rijhwani, Sebastian Ruder, Stephen Mayhew, Israel Abebe Azime, Shamsuddeen Muhammad, Chris Chinenye Emezue, Joyce Nakatumba-Nabende, Perez Ogayo, Anuoluwapo Aremu, Catherine Gitau, Derguene Mbaye, Jesujoba Ala...

  4. [4]

    Seza Doğruöz, André Coneglian, and Atul Kr

    David Ifeoluwa Adelani, A. Seza Doğruöz, André Coneglian, and Atul Kr. Ojha. 2024 a . https://arxiv.org/abs/2404.18286 Comparing llm prompting with cross-lingual transfer performance on indigenous and low-resource brazilian languages . Preprint, arXiv:2404.18286

  5. [5]

    David Ifeoluwa Adelani, Jessica Ojo, Israel Abebe Azime, Jian Yun Zhuang, Jesujoba O. Alabi, Xuanli He, Millicent Ochieng, Sara Hooker, Andiswa Bukula, En-Shiun Annie Lee, Chiamaka Chukwuneke, Happy Buzaaba, Blessing Sibanda, Godson Kalipe, Jonathan Mukiibi, Salomon Kabongo, Foutse Yuehgoh, Mmasibidi Setaka, Lolwethu Ndolela, Nkiruka Odu, Rooweither Mabuy...

  6. [6]

    Kabir Ahuja, Harshita Diddee, Rishav Hada, Millicent Ochieng, Krithika Ramesh, Prachi Jain, Akshay Nambi, Tanuja Ganu, Sameer Segal, Maxamed Axmed, Kalika Bali, and Sunayana Sitaram. 2023. https://arxiv.org/abs/2303.12528 Mega: Multilingual evaluation of generative ai . Preprint, arXiv:2303.12528

  7. [7]

    Sanchit Ahuja, Divyanshu Aggarwal, Varun Gumma, Ishaan Watts, Ashutosh Sathe, Millicent Ochieng, Rishav Hada, Prachi Jain, Maxamed Axmed, Kalika Bali, and Sunayana Sitaram. 2024. https://arxiv.org/abs/2311.07463 Megaverse: Benchmarking large language models across languages, modalities, models and tasks . Preprint, arXiv:2311.07463

  8. [8]

    Alabi, David Ifeoluwa Adelani, Marius Mosbach, and Dietrich Klakow

    Jesujoba O. Alabi, David Ifeoluwa Adelani, Marius Mosbach, and Dietrich Klakow. 2022. https://aclanthology.org/2022.coling-1.382 Adapting pre-trained language models to A frican languages via multilingual adaptive fine-tuning . In Proceedings of the 29th International Conference on Computational Linguistics, pages 4336--4349, Gyeongju, Republic of Korea. ...

Show all 74 references
  1. [9]

    Tyers, and Gregor Weber

    Rosana Ardila, Megan Branson, Kelly Davis, Michael Henretty, Michael Kohler, Josh Meyer, Reuben Morais, Lindsay Saunders, Francis M. Tyers, and Gregor Weber. 2020. https://arxiv.org/abs/1912.06670 Common voice: A massively-multilingual speech corpus . Preprint, arXiv:1912.06670

  2. [10]

    Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020. https://doi.org/10.18653/v1/2020.acl-main.421 On the cross-lingual transferability of monolingual representations . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4623--...

  3. [11]

    Akari Asai, Sneha Kudugunta, Xinyan Velocity Yu, Terra Blevins, Hila Gonen, Machel Reid, Yulia Tsvetkov, Sebastian Ruder, and Hannaneh Hajishirzi. 2023. https://arxiv.org/abs/2305.14857 Buffet: Benchmarking large language models for few-shot cross-lingual transfer . Preprint, ...

  4. [12]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert - Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jef...

  5. [13]

    José Cañete, Gabriel Chaperon, Rodrigo Fuentes, Jou-Hui Ho, Hojin Kang, and Jorge Pérez. 2023. https://arxiv.org/abs/2308.02976 Spanish pre-trained bert model and evaluation data . Preprint, arXiv:2308.02976

  6. [14]

    Yang Chen, Vedaant Shah, and Alan Ritter. 2024. https://arxiv.org/abs/2305.13582 Translation and fusion improves zero-shot cross-lingual information extraction . Preprint, arXiv:2305.13582

  7. [15]

    Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzm \' a n, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019. https://arxiv.org/abs/1911.02116 Unsupervised cross-lingual representation learning at scale . C...

  8. [16]

    Bowman, Holger Schwenk, and Veselin Stoyanov

    Alexis Conneau, Guillaume Lample, Ruty Rinott, Adina Williams, Samuel R. Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. https://arxiv.org/abs/1809.05053 Xnli: Evaluating cross-lingual sentence representations . Preprint, arXiv:1809.05053

  9. [17]

    Yiming Cui, Ziqing Yang, and Xin Yao. 2024. https://arxiv.org/abs/2304.08177 Efficient and effective text encoding for chinese llama and alpaca . Preprint, arXiv:2304.08177

  10. [18]

    Severino Da Dalt, Joan Llop, Irene Baucells, Marc Pamies, Yishi Xu, Aitor Gonzalez-Agirre, and Marta Villegas. 2024. https://aclanthology.org/2024.lrec-main.650/ FLOR : On the effectiveness of language adaptation . In Proceedings of the 2024 Joint International Conference on C...

  11. [19]

    Wietse de Vries and Malvina Nissim. 2021. https://doi.org/10.18653/v1/2021.findings-acl.74 As good as new. how to successfully recycle english GPT -2 to make models for other languages . In Findings of the Association for Computational Linguistics: ACL - IJCNLP 2021 . Associat...

  12. [20]

    DeepSpeed. 2021. D eep S peed Z e R O -3 O ffload --- deepspeed.ai. https://www.deepspeed.ai/2021/03/07/zero3-offload.html

  13. [21]

    Tim Dettmers. 2022. https://github.com/TimDettmers/bitsandbytes bitsandbytes . GitHub repository

  14. [22]

    Jacob Devlin. 2019. https://github.com/google-research/bert/blob/master/multilingual.md Bert/multilingual.md at master · google-research/bert

  15. [23]

    Konstantin Dobler and Gerard de Melo. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.829 FOCUS : Effective embedding initialization for monolingual specialization of multilingual models . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Process...

  16. [24]

    Meet Doshi, Raj Dabre, and Pushpak Bhattacharyya. 2024. https://arxiv.org/abs/2403.13638 Do not worry if you do not have data: Building pretrained language models using translationese . Preprint, arXiv:2403.13638

  17. [25]

    Samuel Gehman, Suchin Gururangan, Maarten Sap, Yejin Choi, and Noah A. Smith. 2020. https://arxiv.org/abs/2009.11462 Realtoxicityprompts: Evaluating neural toxic degeneration in language models . Preprint, arXiv:2009.11462

  18. [26]

    Gurpreet Gosal, Yishi Xu, Gokul Ramakrishnan, Rituraj Joshi, Avraham Sheinin, Zhiming, Chen, Biswajit Mishra, Natalia Vassilieva, Joel Hestness, Neha Sengupta, Sunil Kumar Sahu, Bokang Jia, Onkar Pandit, Satheesh Katipomu, Samta Kamboj, Samujjwal Ghosh, Rahul Pal, Parvez Mulla...

  19. [27]

    Saiful Islam, Kazi Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M

    Tahmid Hasan, Abhik Bhattacharjee, Md. Saiful Islam, Kazi Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar. 2021. https://doi.org/10.18653/v1/2021.findings-acl.413 XL -sum: Large-scale multilingual abstractive summarization for 44 languages . In Fin...

  20. [28]

    Aung Htet and Mark Dras. 2024. https://doi.org/10.21203/rs.3.rs-4329843/v1 Myanmar xnli: Building a dataset and exploring low-resource approaches to natural language inference with myanmar

  21. [29]

    Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020. https://arxiv.org/abs/2003.11080 Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalization . Preprint, arXiv:2003.11080

  22. [30]

    Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2025. https://doi.org/10.1145/3703155 A survey on hallucination in large language models: Principles, taxonomy, challenges, and o...

  23. [31]

    Mojan Javaheripi and Sébastien Bubec. 2023. P hi-2: T he surprising power of small language models --- microsoft.com. https://www.microsoft.com/en-us/research/blog/phi-2-the-surprising-power-of-small-language-models/ . []

  24. [32]

    Raviraj Joshi, Kanishk Singla, Anusha Kamath, Raunak Kalani, Rakesh Paul, Utkarsh Vaidya, Sanjay Singh Chauhan, Niranjan Wartikar, and Eileen Long. 2024. https://arxiv.org/abs/2410.14815 Adapting multilingual llms to low-resource languages using continued pre-training and synt...

  25. [33]

    Khapra, and Pratyush Kumar

    Divyanshu Kakwani, Anoop Kunchukuttan, Satish Golla, Gokul N.C., Avik Bhattacharyya, Mitesh M. Khapra, and Pratyush Kumar. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.445 I ndic NLPS uite: Monolingual corpora, evaluation benchmarks and pre-trained multilingual langua...

  26. [34]

    Fajri Koto, Afshin Rahimi, Jey Han Lau, and Timothy Baldwin. 2020. https://doi.org/10.18653/v1/2020.coling-main.66 I ndo LEM and I ndo BERT : A benchmark dataset and pre-trained language model for I ndonesian NLP . In Proceedings of the 28th International Conference on Computa...

  27. [35]

    Hele-Andra Kuulmets, Taido Purason, Agnes Luhtaru, and Mark Fishel. 2024. https://arxiv.org/abs/2404.04042 Teaching llama a new language through cross-lingual knowledge transfer . Preprint, arXiv:2404.04042

  28. [36]

    Viet Dac Lai, Nghia Trung Ngo, Amir Pouran Ben Veyseh, Hieu Man, Franck Dernoncourt, Trung Bui, and Thien Huu Nguyen. 2023. https://arxiv.org/abs/2304.05613 Chatgpt beyond english: Towards a comprehensive evaluation of large language models in multilingual learning . Preprint,...

  29. [37]

    Guillaume Lample and Alexis Conneau. 2019. https://arxiv.org/abs/1901.07291 Cross-lingual language model pretraining . Preprint, arXiv:1901.07291

  30. [38]

    Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O'Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mon...

  31. [39]

    Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. https://arxiv.org/abs/2001.08210 Multilingual denoising pre-training for neural machine translation . Preprint, arXiv:2001.08210

  32. [40]

    Ilya Loshchilov and Frank Hutter. 2019. https://arxiv.org/abs/1711.05101 Decoupled weight decay regularization . Preprint, arXiv:1711.05101

  33. [41]

    Yun Luo, Zhen Yang, Fandong Meng, Yafu Li, Jie Zhou, and Yue Zhang. 2025. https://arxiv.org/abs/2308.08747 An empirical study of catastrophic forgetting in large language models during continual fine-tuning . Preprint, arXiv:2308.08747

  34. [42]

    Alexandre Magueresse, Vincent Carles, and Evan Heetderks. 2020. https://arxiv.org/abs/2006.07264 Low-resource languages: A review of past work and future challenges . Preprint, arXiv:2006.07264

  35. [43]

    Louis Martin, Benjamin Muller, Pedro Javier Ortiz Su \'a rez, Yoann Dupont, Laurent Romary, \'E ric de la Clergerie, Djam \'e Seddah, and Beno \^ t Sagot. 2020. https://doi.org/10.18653/v1/2020.acl-main.645 C amem BERT : a tasty F rench language model . In Proceedings of the 5...

  36. [44]

    Michael McCloskey and Neal J. Cohen. 1989. https://doi.org/10.1016/S0079-7421(08)60536-8 Catastrophic interference in connectionist networks: The sequential learning problem . volume 24 of Psychology of Learning and Motivation, pages 109--165. Academic Press

  37. [45]

    Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, and Hao Wu. 2018. https://arxiv.org/abs/1710.03740 Mixed precision training . Preprint, arXiv:1710.03740

  38. [46]

    MosaicAI. 2023. I ntroducing M P T -7 B : A N ew S tandard for O pen- S ource, C ommercially U sable L L M s --- databricks.com. https://www.databricks.com/blog/mpt-7b

  39. [47]

    Nandini Mundra, Aditya Nanda Kishore, Raj Dabre, Ratish Puduppully, Anoop Kunchukuttan, and Mitesh M. Khapra. 2024. https://arxiv.org/abs/2407.05841 An empirical comparison of vocabulary expansion and initialization approaches for language models . Preprint, arXiv:2407.05841

  40. [48]

    Kelechi Ogueji, Yuxin Zhu, and Jimmy Lin. 2021. https://doi.org/10.18653/v1/2021.mrl-1.11 Small data? no problem! exploring the viability of pretrained multilingual language models for low-resourced languages . In Proceedings of the 1st Workshop on Multilingual Representation ...

  41. [49]

    Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji. 2017. https://doi.org/10.18653/v1/P17-1178 Cross-lingual name tagging and linking for 282 languages . In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (...

  42. [50]

    Le, and Luu Anh Tuan

    Trinh Pham, Khoi M. Le, and Luu Anh Tuan. 2024. https://arxiv.org/abs/2406.09717 Unibridge: A unified approach to cross-lingual transfer learning for low-resource languages . Preprint, arXiv:2406.09717

  43. [51]

    Marco Polignano, Valerio Basile, Pierpaolo Basile, Marco de Gemmis, and Giovanni Semeraro. 2019. https://doi.org/10.4000/ijcol.472 AlBERTo : Modeling italian social media language with BERT . Italian Journal of Computational Linguistics, 5(2):11--31

  44. [52]

    Edoardo Maria Ponti, Goran Glava s , Olga Majewska, Qianchu Liu, Ivan Vuli \'c , and Anna Korhonen. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.185 XCOPA : A multilingual dataset for causal commonsense reasoning . In Proceedings of the 2020 Conference on Empirical Method...

  45. [53]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. https://arxiv.org/abs/1910.10683 Exploring the limits of transfer learning with a unified text-to-text transformer . Preprint, arXiv:1910.10683

  46. [54]

    Leonardo Ranaldi and Giulia Pucci. 2023. https://doi.org/10.18653/v1/2023.mrl-1.14 Does the E nglish matter? elicit cross-lingual abilities of large language models . In Proceedings of the 3rd Workshop on Multi-lingual Representation Learning (MRL), pages 173--183, Singapore. ...

  47. [55]

    Leonardo Ranaldi, Giulia Pucci, and Andre Freitas. 2023. https://arxiv.org/abs/2308.14186 Empowering cross-lingual abilities of instruction-tuned large language models by translation-following demonstrations . Preprint, arXiv:2308.14186

  48. [56]

    Evgeniia Razumovskaia, Ivan Vulić, and Anna Korhonen. 2024. https://arxiv.org/abs/2403.01929 Analyzing and adapting large language models for few-shot multilingual nlu: Are we there yet? Preprint, arXiv:2403.01929

  49. [57]

    François Remy, Pieter Delobelle, Bettina Berendt, Kris Demuynck, and Thomas Demeester. 2023. https://arxiv.org/abs/2310.03477 Tik-to-tok: Translating language models one token at a time: An embedding initialization strategy for efficient language adaptation . Preprint, arXiv:2...

  50. [58]

    Samin Mahdizadeh Sani, Pouya Sadeghi, Thuy-Trang Vu, Yadollah Yaghoobzadeh, and Gholamreza Haffari. 2025. https://arxiv.org/abs/2412.13375 Extending llms to new languages: A case study of llama and persian adaptation . Preprint, arXiv:2412.13375

  51. [59]

    Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, Dipanjan Das, and Jason Wei. 2022. https://arxiv.org/abs/2210.03057 Language models are multilingual chain-of-thought reasoners . Prepri...

  52. [60]

    Hashimoto

    Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Stanford alpaca: An instruction-following llama model. https://github.com/tatsu-lab/stanford_alpaca

  53. [61]

    NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prang...

  54. [62]

    Atula Tejaswi, Nilesh Gupta, and Eunsol Choi. 2024. https://arxiv.org/abs/2406.14670 Exploring design choices for building language-specific llms . Preprint, arXiv:2406.14670

  55. [63]

    Prajwal Thapa, Jinu Nyachhyon, Mridul Sharma, and Bal Krishna Bal. 2024. https://arxiv.org/abs/2411.15734 Development of pre-trained transformer-based models for the nepali language . Preprint, arXiv:2411.15734

  56. [64]

    Cagri Toraman. 2024. https://doi.org/10.18653/v1/2024.mrl-1.3 Adapting open-source generative large language models for low-resource languages: A case study for T urkish . In Proceedings of the Fourth Workshop on Multilingual Representation Learning (MRL 2024), pages 30--44, M...

  57. [65]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 a . https://arxiv.org/abs/2302.13971 Lla...

  58. [66]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, W...

  59. [67]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. https://arxiv.org/abs/1706.03762 Attention is all you need . CoRR, abs/1706.03762

  60. [68]

    Domen Vreš, Martin Božič, Aljaž Potočnik, Tomaž Martinčič, and Marko Robnik-Šikonja. 2024. https://arxiv.org/abs/2410.06898 Generative model for less-resourced language with 1 billion parameters . Preprint, arXiv:2410.06898

  61. [69]

    Shumin Wang, Yuexiang Xie, Bolin Ding, Jinyang Gao, and Yanyong Zhang. 2025. https://aclanthology.org/2025.coling-main.480/ Language adaptation of large language models: An empirical study on LL a MA 2 . In Proceedings of the 31st International Conference on Computational Ling...

  62. [70]

    Bryan Wilie, Karissa Vincentio, Genta Indra Winata, Samuel Cahyawijaya, Xiaohong Li, Zhi Yuan Lim, Sidik Soleman, Rahmad Mahendra, Pascale Fung, Syafri Bahar, and Ayu Purwarianti. 2020. https://aclanthology.org/2020.aacl-main.85 I ndo NLU : Benchmark and resources for evaluati...

  63. [71]

    Shijie Wu and Mark Dredze. 2020. https://doi.org/10.18653/v1/2020.repl4nlp-1.16 Are all languages created equal in multilingual BERT ? In Proceedings of the 5th Workshop on Representation Learning for NLP, pages 120--130, Online. Association for Computational Linguistics

  64. [72]

    Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. https://arxiv.org/abs/2010.11934 mt5: A massively multilingual pre-trained text-to-text transformer . Preprint, arXiv:2010.11934

  65. [73]

    Atsuki Yamaguchi, Aline Villavicencio, and Nikolaos Aletras. 2024 a . https://arxiv.org/abs/2402.10712 An empirical study on cross-lingual vocabulary adaptation for efficient language model inference . Preprint, arXiv:2402.10712

  66. [74]

    Atsuki Yamaguchi, Aline Villavicencio, and Nikolaos Aletras. 2024 b . https://arxiv.org/abs/2406.11477 How can we effectively expand the vocabulary of llms with 0.01gb of target language text? Preprint, arXiv:2406.11477

Pith tools