Pith. sign in

REVIEW 3 cited by

Efficiently Adapting Pretrained Language Models To New Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.05741 v2 pith:OL2LC2U2 submitted 2023-11-09 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords languagelanguagesmodelsadaptingdataenglishpretrainedefficiency
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent large language models (LLM) exhibit sub-optimal performance on low-resource languages, as the training data of these models is usually dominated by English and other high-resource languages. Furthermore, it is challenging to train models for low-resource languages, especially from scratch, due to a lack of high quality training data. Adapting pretrained LLMs reduces the need for data in the new language while also providing cross lingual transfer capabilities. However, naively adapting to new languages leads to catastrophic forgetting and poor tokenizer efficiency. In this work, we study how to efficiently adapt any existing pretrained LLM to a new language without running into these issues. In particular, we improve the encoding efficiency of the tokenizer by adding new tokens from the target language and study the data mixing recipe to mitigate forgetting. Our experiments on adapting an English LLM to Hungarian and Thai show that our recipe can reach better performance than open source models on the target language, with minimal regressions on English.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Typhoon 2: A Family of Open Text and Multimodal Thai Large Language Models

    cs.CL 2024-12 conditional novelty 5.0 of 10

    Typhoon 2 improves Thai LLM performance through continual pre-training on curated Thai data and post-training, releasing text, vision, audio, and safety models.

  2. Extending LLMs to New Languages: A Case Study of Llama and Persian Adaptation

    cs.CL 2024-12 conditional novelty 4.0 of 10

    Bilingual embedding alignment plus instruction tuning improves Persian classification in Llama-2, while English-to-Persian transfer is marginal and task-dependent.

  3. BgGPT 1.0: Extending English-centric LLMs to other languages

    cs.CL 2024-12 conditional novelty 4.0 of 10

    Continually pretraining Gemma-2 on a curated Bulgarian corpus and merging with instruction-tuned models yields open Bulgarian-English models that beat larger open models on Bulgarian benchmarks.

Pith tools