Pith. sign in

REVIEW 3 cited by

Don't Trust ChatGPT when Your Question is not in English: A Study of Multilingual Abilities and Types of LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.16339 v2 pith:DDWEWKAL submitted 2023-05-24 cs.CL cs.AI

classification cs.CLcs.AI
keywords llmslanguageabilitiesmulti-lingualmultilingualperformancedemonstratedenglish
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) have demonstrated exceptional natural language understanding abilities and have excelled in a variety of natural language processing (NLP)tasks in recent years. Despite the fact that most LLMs are trained predominantly in English, multiple studies have demonstrated their comparative performance in many other languages. However, fundamental questions persist regarding how LLMs acquire their multi-lingual abilities and how performance varies across different languages. These inquiries are crucial for the study of LLMs since users and researchers often come from diverse language backgrounds, potentially influencing their utilization and interpretation of LLMs' results. In this work, we propose a systematic way of qualifying the performance disparities of LLMs under multilingual settings. We investigate the phenomenon of across-language generalizations in LLMs, wherein insufficient multi-lingual training data leads to advanced multi-lingual capabilities. To accomplish this, we employ a novel back-translation-based prompting method. The results show that GPT exhibits highly translating-like behaviour in multilingual settings.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating Chinese Large Language Models: The Influence of Persona Assignment on Stereotypes and Safeguards

    cs.CY 2025-06 conditional novelty 6.0 of 10

    Assigning personas to Chinese LLMs amplifies toxic output relative to default behavior, while refusal rates shift systematically with persona gender and target social group.

  2. CausalAbstain: Enhancing Multilingual LLMs with Causal Reasoning for Trustworthy Abstention

    cs.CL 2025-05 conditional novelty 6.0 of 10

    CausalAbstain filters multilingual self-feedback by comparing how much it changes the model's abstention decision, improving abstention accuracy over baselines on two benchmarks.

  3. A Qualitative Investigation into LLM-Generated Multilingual Code Comments and Automatic Evaluation Metrics

    cs.SE 2025-05 conditional novelty 6.0 of 10

    Neural metrics for evaluating code comments are unreliable for multilingual output, often scoring random noise as high as real generated comments.

Pith tools