REVIEW 9 cited by
NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
To be effectively and safely deployed to global user populations, large language models (LLMs) may need to adapt outputs to user values and cultures, not just know about them. We introduce NormAd, an evaluation framework to assess LLMs' cultural adaptability, specifically measuring their ability to judge social acceptability across varying levels of cultural norm specificity, from abstract values to explicit social norms. As an instantiation of our framework, we create NormAd-Eti, a benchmark of 2.6k situational descriptions representing social-etiquette related cultural norms from 75 countries. Through comprehensive experiments on NormAd-Eti, we find that LLMs struggle to accurately judge social acceptability across these varying degrees of cultural contexts and show stronger adaptability to English-centric cultures over those from the Global South. Even in the simplest setting where the relevant social norms are provided, the best LLMs' performance (< 82\%) lags behind humans (> 95\%). In settings with abstract values and country information, model performance drops substantially (< 60\%), while human accuracy remains high (> 90\%). Furthermore, we find that models are better at recognizing socially acceptable versus unacceptable situations. Our findings showcase the current pitfalls in socio-cultural reasoning of LLMs which hinder their adaptability for global audiences.
Forward citations
Cited by 9 Pith papers
-
Trusting sovereign language models as scientific instruments: evidence from Portugal's AMALIA
Portugal's 9B-parameter national language model AMALIA agrees with human annotators on coding moral authority but fails a construct-validity test showing it reaches correct codes via surface correlates rather than the...
-
PLURAL: A Global Dataset for Value Alignment
Synthetic preference data generated from the Integrated Values Survey preserves cross-country value differences and enables DPO fine-tuning that improves LLM cultural alignment across five countries.
-
XCR-Bench: Benchmarking Cross-Cultural Reasoning in LLMs via Culture-Specific Items and Hall's Triad
XCR-Bench provides 4,100+ parallel sentences with 1,098 culture-specific items mapped to Hall's Triad, and shows LLMs struggle most with deeper, semi-visible cultural norms.
-
Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages
Across nine Asian languages, multilingual LLMs favor Western cultural entities in 30-40% of culturally grounded contexts, with model-specific sentiment biases and extraction accuracy gaps.
-
CultureSynth: A Hierarchical Taxonomy-Guided and Retrieval-Augmented Framework for Cultural Question-Answer Synthesis
A taxonomy-guided retrieval-augmented framework generates CultureSynth-7, a multilingual cultural QA benchmark, and its evaluation of 14 LLMs suggests cultural competence emerges around 3B parameters.
-
EtiCor++: Towards Understanding Etiquettical Bias in LLMs
A new English etiquette corpus and bias metrics show that LLMs over-prefer Western norms and under-predict etiquettes from low-resource regions.
-
CulFiT: A Fine-grained Cultural-aware LLM Training Paradigm via Multilingual Critique Data Synthesis
A multilingual critique-data training paradigm with a knowledge-unit reward improves LLM cultural alignment on several benchmarks, but its headline benchmark is evaluated with the same LLM-judged metric used to select...
-
PerCul: A Story-Driven Cultural Evaluation of LLMs in Persian
PerCul is a Persian cultural story-comprehension benchmark on which the best LLMs lag human performance by 11.3 to 21.3 percentage points.
-
Do Large Language Models Know Folktales? A Case Study of Yokai in Japanese Folktales
A benchmark of 809 yokai questions shows Japanese-centric LLMs, particularly Llama-3-based continual pretraining models, outperform English-centric models on Japanese folktale knowledge.
Discussion (0). Continue with ORCID to comment.