Pith. sign in

REVIEW 1 cited by

On Evaluating Multilingual Compositional Generalization with Translated Datasets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.11420 v1 pith:NAI6TEEM submitted 2023-06-20 cs.CL

classification cs.CL
keywords compositionalgeneralizationcross-lingualdatasetdatasetsenglishevaluatingmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Compositional generalization allows efficient learning and human-like inductive biases. Since most research investigating compositional generalization in NLP is done on English, important questions remain underexplored. Do the necessary compositional generalization abilities differ across languages? Can models compositionally generalize cross-lingually? As a first step to answering these questions, recent work used neural machine translation to translate datasets for evaluating compositional generalization in semantic parsing. However, we show that this entails critical semantic distortion. To address this limitation, we craft a faithful rule-based translation of the MCWQ dataset from English to Chinese and Japanese. Even with the resulting robust benchmark, which we call MCWQ-R, we show that the distribution of compositions still suffers due to linguistic divergences, and that multilingual models still struggle with cross-lingual compositional generalization. Our dataset and methodology will be useful resources for the study of cross-lingual compositional generalization in other tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Synthetic Data Can Mislead Evaluations: Membership Inference as Machine Text Detection

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Membership inference attacks on LLMs score synthetic text as more 'member-like' than real training data, so using synthetic data as non-members produces misleading memorization conclusions.

Pith tools