REVIEW 5 cited by
Are Multilingual LLMs Culturally-Diverse Reasoners? An Investigation into Multicultural Proverbs and Sayings
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Large language models (LLMs) are highly adept at question answering and reasoning tasks, but when reasoning in a situational context, human expectations vary depending on the relevant cultural common ground. As languages are associated with diverse cultures, LLMs should also be culturally-diverse reasoners. In this paper, we study the ability of a wide range of state-of-the-art multilingual LLMs (mLLMs) to reason with proverbs and sayings in a conversational context. Our experiments reveal that: (1) mLLMs "know" limited proverbs and memorizing proverbs does not mean understanding them within a conversational context; (2) mLLMs struggle to reason with figurative proverbs and sayings, and when asked to select the wrong answer (instead of asking it to select the correct answer); and (3) there is a "culture gap" in mLLMs when reasoning about proverbs and sayings translated from other languages. We construct and release our evaluation dataset MAPS (MulticultrAl Proverbs and Sayings) for proverb understanding with conversational context for six different languages.
Forward citations
Cited by 5 Pith papers
-
Every Token Counts: Exact Likert-Scale Distributions for Measuring LLM Attitudes and Biases
A fully crossed factorial experiment on exact token probability distributions, combined with a distributional ANOVA, measures LLM ethnocentrism without text-sampling noise.
-
CultureVLM: Characterizing and Improving Cultural Understanding of Vision-Language Models for over 100 Countries
CultureVerse is a 188-country, 19k-concept visual QA benchmark, and fine-tuning open VLMs on it improves cultural accuracy, but the main evaluation shares concepts between training and test sets.
-
Do Large Language Models Understand Morality Across Cultures?
Small language models compress cross-cultural moral differences, producing more uniformly permissive and less varied judgments than international survey data.
-
LLMs as mirrors of societal moral standards: reflection of cultural divergence and agreement across ethical topics
Across five LLMs and two global surveys, model-generated moral judgments poorly matched cross-cultural patterns of agreement and disagreement.
-
A Survey on Human-Centric LLMs
A review that sorts existing evidence on how well large language models imitate individual human skills and collective social dynamics into one taxonomy.
Discussion (0). Continue with ORCID to comment.