REVIEW 20 cited by
Crosslingual Generalization through Multitask Finetuning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Multitask prompted finetuning (MTF) has been shown to help large language models generalize to new tasks in a zero-shot setting, but so far explorations of MTF have focused on English data and models. We apply MTF to the pretrained multilingual BLOOM and mT5 model families to produce finetuned variants called BLOOMZ and mT0. We find finetuning large multilingual language models on English tasks with English prompts allows for task generalization to non-English languages that appear only in the pretraining corpus. Finetuning on multilingual tasks with English prompts further improves performance on English and non-English tasks leading to various state-of-the-art zero-shot results. We also investigate finetuning on multilingual tasks with prompts that have been machine-translated from English to match the language of each dataset. We find training on these machine-translated prompts leads to better performance on human-written prompts in the respective languages. Surprisingly, we find models are capable of zero-shot generalization to tasks in languages they have never intentionally seen. We conjecture that the models are learning higher-level capabilities that are both task- and language-agnostic. In addition, we introduce xP3, a composite of supervised datasets in 46 languages with English and machine-translated prompts. Our code, datasets and models are freely available at https://github.com/bigscience-workshop/xmtf.
Forward citations
Cited by 20 Pith papers
-
MameLoshnLM: Yiddish Language Model and Evaluation Benchmark
MameLoshnLM, trained by continuing pretraining Llama 3.1 8B on a curated Yiddish corpus, outperforms similar-scale open models on a new multi-task Yiddish benchmark and better retains Yiddish-specific loshn-koydesh vo...
-
ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language Models
ChiKhaPo is an 8-subtask benchmark that measures word-level comprehension and generation in 2,700+ languages and shows state-of-the-art models perform poorly on low-resource languages.
-
UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu
A new minimal-pair benchmark for Urdu grammar shows that multilingual models vary widely across syntactic phenomena, with LLaMA-3-70B best at 94.7% accuracy.
-
Evaluating LLMs on Chinese Idiom Translation
Across 900 annotated translation pairs from nine MT systems, the best system still mistranslates Chinese idioms in 28% of cases, and standard metrics miss these errors (Pearson correlation below 0.48).
-
Preperiodic points, finiteness, and structures of semigroups of algebraic morphisms
The paper proves finiteness and structural results for preperiodic points of algebraic morphisms, including Burnside-type and Northcott-type theorems.
-
Revisiting LLM Value Probing Strategies: Are They Robust and Expressive?
Value representations from token logits, sequence perplexity, and text generation are all sensitive to prompt and option changes, and their correlation with model behavior in value scenarios is weak.
-
The Emergence of Abstract Thought in Large Language Models Beyond Any Language
Across 20 open LLMs, shared multilingual neurons grow in number and per-neuron importance over release generations, which the authors interpret as evidence of language-agnostic abstract thought and use to guide neuron...
-
Can LLMs Ground when they (Don't) Know: A Study on Direct and Loaded Political Questions
In German political questions, three LLMs frequently accommodated false presuppositions, and correct answers to direct questions did not guarantee rejection of the false assumption.
-
MaXIFE: Multilingual and Cross-lingual Instruction Following Evaluation
MaXIFE is a 23-language, 1,667-task benchmark for multilingual and cross-lingual instruction-following evaluation, with baseline scores for five commercial LLMs.
-
CC-Tuning: A Cross-Lingual Connection Mechanism for Improving Joint Multilingual Supervised Fine-Tuning
CC-Tuning fuses English feed-forward activations into non-English inputs during multilingual supervised fine-tuning, using a trainable Decision Maker and a least-squares Transform Matrix to simulate the connection at ...
-
Disentangling Language and Culture for Evaluating Multilingual Large Language Models
A new dual-axis evaluation framework shows multilingual LLMs answer culture-specific questions best when the question language matches the cultural context, with partial neuron-level evidence for the effect.
-
How and Where to Translate? The Impact of Translation Strategies in Cross-lingual LLM Prompting
For multilingual RAG intent classification, the best translation strategy depends on the model and language; translating instructions into the user's language helps some models, while making the model answer in low-re...
-
Seed-X: Building Strong Multilingual Translation LLM with 7B Parameters
A 7B open-weight translation model matches or outperforms far larger commercial systems across 28 languages in automatic and human evaluations.
-
Multilingual Test-Time Scaling via Initial Thought Transfer
MITT, a prefix-tuning method for multilingual test-time scaling, is evaluated on questions whose English reasoning was used for training, confounding the reported gains.
-
Beyond English: The Impact of Prompt Translation Strategies across Languages and Tasks in Multilingual LLMs
Selective pre-translation, translating only some prompt components into English, generally outperforms both full prompt translation and direct inference across tasks and languages, with the largest gains for low-resou...
-
QUST_NLP at SemEval-2025 Task 7: A Three-Stage Retrieval Framework for Monolingual and Crosslingual Fact-Checked Claim Retrieval
A three-stage ensemble of retrieval models, rerankers, and weighted voting achieves strong multilingual fact-checked claim retrieval results at SemEval-2025 Task 7.
-
SingaKids: A Multilingual Multimodal Dialogic Tutor for Language Learning
The paper presents SingaKids, a four-language dialogic tutoring system, and reports component-level improvements plus a 35-student pilot study of its scaffolding behavior.
-
Customized Generative AI Agent for Transportation Engineering Practice: A Development and Continued Pre-training Guideline
LoRA continued pretraining on a small U.S. transportation corpus lifts BLEU-4 and ROUGE for Qwen2.5-7B and LLaMA-3.1-8B far above the other four models tested.
-
Learning Text Styles: A Study on Transfer, Attribution, and Verification
A thesis compiles published work claiming that lightweight adapters, contrastive disentanglement, and instruction tuning improve text style transfer, authorship attribution, and authorship verification.
-
Large Language models for Time Series Analysis: Techniques, Applications, and Challenges
A review of LLM-based time series analysis that proposes several taxonomies, but is undermined by citation errors and a lack of systematic methodology.
Discussion (0). Continue with ORCID to comment.