Pith. sign in

REVIEW 5 cited by

Are Multilingual LLMs Culturally-Diverse Reasoners? An Investigation into Multicultural Proverbs and Sayings

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.08591 v2 pith:HHLXQH33 submitted 2023-09-15 cs.CL

classification cs.CL
keywords proverbssayingscontextllmsmllmsconversationallanguagesreasoning
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Large language models (LLMs) are highly adept at question answering and reasoning tasks, but when reasoning in a situational context, human expectations vary depending on the relevant cultural common ground. As languages are associated with diverse cultures, LLMs should also be culturally-diverse reasoners. In this paper, we study the ability of a wide range of state-of-the-art multilingual LLMs (mLLMs) to reason with proverbs and sayings in a conversational context. Our experiments reveal that: (1) mLLMs "know" limited proverbs and memorizing proverbs does not mean understanding them within a conversational context; (2) mLLMs struggle to reason with figurative proverbs and sayings, and when asked to select the wrong answer (instead of asking it to select the correct answer); and (3) there is a "culture gap" in mLLMs when reasoning about proverbs and sayings translated from other languages. We construct and release our evaluation dataset MAPS (MulticultrAl Proverbs and Sayings) for proverb understanding with conversational context for six different languages.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Every Token Counts: Exact Likert-Scale Distributions for Measuring LLM Attitudes and Biases

    cs.CL 2026-08 conditional novelty 6.0 of 10

    A fully crossed factorial experiment on exact token probability distributions, combined with a distributional ANOVA, measures LLM ethnocentrism without text-sampling noise.

  2. CultureVLM: Characterizing and Improving Cultural Understanding of Vision-Language Models for over 100 Countries

    cs.AI 2025-01 conditional novelty 6.0 of 10

    CultureVerse is a 188-country, 19k-concept visual QA benchmark, and fine-tuning open VLMs on it improves cultural accuracy, but the main evaluation shares concepts between training and test sets.

  3. Do Large Language Models Understand Morality Across Cultures?

    cs.CL 2025-07 reject novelty 4.0 of 10

    Small language models compress cross-cultural moral differences, producing more uniformly permissive and less varied judgments than international survey data.

  4. LLMs as mirrors of societal moral standards: reflection of cultural divergence and agreement across ethical topics

    cs.AI 2024-12 conditional novelty 4.0 of 10

    Across five LLMs and two global surveys, model-generated moral judgments poorly matched cross-cultural patterns of agreement and disagreement.

  5. A Survey on Human-Centric LLMs

    cs.CL 2024-11 conditional novelty 1.0 of 10

    A review that sorts existing evidence on how well large language models imitate individual human skills and collective social dynamics into one taxonomy.

Pith tools