Pith. sign in

REVIEW 3 cited by

Is It Good Data for Multilingual Instruction Tuning or Just Bad Multilingual Evaluation for Large Language Models?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.12822 v3 pith:J4SLE43O submitted 2024-06-18 cs.CL cs.AI

classification cs.CLcs.AI
keywords datainstructiontranslatedmodelsmultilingualnativecannotevaluation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multilingual large language models are designed, claimed, and expected to cater to speakers of varied languages. We hypothesise that the current practices of fine-tuning and evaluating these models may not perfectly align with this objective owing to a heavy reliance on translation, which cannot cover language-specific knowledge but can introduce translation defects. It remains unknown whether the nature of the instruction data has an impact on the model output; conversely, it is questionable whether translated test sets can capture such nuances. Due to the often coupled practices of using translated data in both stages, such imperfections could have been overlooked. This work investigates these issues using controlled native or translated data during the instruction tuning and evaluation stages. We show that native or generation benchmarks reveal a notable difference between native and translated instruction data especially when model performance is high, whereas other types of test sets cannot. The comparison between round-trip and single-pass translations reflects the importance of knowledge from language-native resources. Finally, we demonstrate that regularization is beneficial to bridging this gap on structured but not generative tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation

    cs.CL 2024-12 conditional novelty 6.0 of 10

    This paper quantifies Western-centric bias in MMLU, releases Global-MMLU across 42 languages with human-verified translations, and shows model rankings shift on culturally sensitive versus agnostic subsets.

  2. Salamandra Technical Report

    cs.CL 2025-02 conditional novelty 5.0 of 10

    Salamandra is an open, from-scratch multilingual LLM family with 2B, 7B, and 40B checkpoints, instruction-tuned variants, a vision proof-of-concept, and detailed evaluations across Iberian and European languages.

  3. Train More Parameters But Mind Their Placement: Insights into Language Adaptation with PEFT

    cs.CL 2024-12 conditional novelty 5.0 of 10

    Feed-forward LoRA with more parameters adapts Llama-3.2-1B to Icelandic better than attention LoRA, bottleneck adapters, IA3, or prefix tuning, and final-layer placement mitigates context-length degradation.

Pith tools