Pith. sign in

REVIEW 10 cited by

NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.12464 v10 pith:PTQS4CSE submitted 2024-04-18 cs.CL

classification cs.CL
keywords culturalllmsadaptabilitysocialframeworkglobalmodelsnorms
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

To be effectively and safely deployed to global user populations, large language models (LLMs) may need to adapt outputs to user values and cultures, not just know about them. We introduce NormAd, an evaluation framework to assess LLMs' cultural adaptability, specifically measuring their ability to judge social acceptability across varying levels of cultural norm specificity, from abstract values to explicit social norms. As an instantiation of our framework, we create NormAd-Eti, a benchmark of 2.6k situational descriptions representing social-etiquette related cultural norms from 75 countries. Through comprehensive experiments on NormAd-Eti, we find that LLMs struggle to accurately judge social acceptability across these varying degrees of cultural contexts and show stronger adaptability to English-centric cultures over those from the Global South. Even in the simplest setting where the relevant social norms are provided, the best LLMs' performance (< 82\%) lags behind humans (> 95\%). In settings with abstract values and country information, model performance drops substantially (< 60\%), while human accuracy remains high (> 90\%). Furthermore, we find that models are better at recognizing socially acceptable versus unacceptable situations. Our findings showcase the current pitfalls in socio-cultural reasoning of LLMs which hinder their adaptability for global audiences.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Trusting sovereign language models as scientific instruments: evidence from Portugal's AMALIA

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Portugal's 9B-parameter national language model AMALIA agrees with human annotators on coding moral authority but fails a construct-validity test showing it reaches correct codes via surface correlates rather than the...

  2. PLURAL: A Global Dataset for Value Alignment

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Synthetic preference data generated from the Integrated Values Survey preserves cross-country value differences and enables DPO fine-tuning that improves LLM cultural alignment across five countries.

  3. XCR-Bench: Benchmarking Cross-Cultural Reasoning in LLMs via Culture-Specific Items and Hall's Triad

    cs.CL 2026-01 conditional novelty 6.0 of 10

    XCR-Bench provides 4,100+ parallel sentences with 1,098 culture-specific items mapped to Hall's Triad, and shows LLMs struggle most with deeper, semi-visible cultural norms.

  4. Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages

    cs.CL 2025-10 conditional novelty 6.0 of 10

    Across nine Asian languages, multilingual LLMs favor Western cultural entities in 30-40% of culturally grounded contexts, with model-specific sentiment biases and extraction accuracy gaps.

  5. CultureSynth: A Hierarchical Taxonomy-Guided and Retrieval-Augmented Framework for Cultural Question-Answer Synthesis

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A taxonomy-guided retrieval-augmented framework generates CultureSynth-7, a multilingual cultural QA benchmark, and its evaluation of 14 LLMs suggests cultural competence emerges around 3B parameters.

  6. EtiCor++: Towards Understanding Etiquettical Bias in LLMs

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new English etiquette corpus and bias metrics show that LLMs over-prefer Western norms and under-predict etiquettes from low-resource regions.

  7. CulFiT: A Fine-grained Cultural-aware LLM Training Paradigm via Multilingual Critique Data Synthesis

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A multilingual critique-data training paradigm with a knowledge-unit reward improves LLM cultural alignment on several benchmarks, but its headline benchmark is evaluated with the same LLM-judged metric used to select...

  8. PerCul: A Story-Driven Cultural Evaluation of LLMs in Persian

    cs.CL 2025-02 conditional novelty 6.0 of 10

    PerCul is a Persian cultural story-comprehension benchmark on which the best LLMs lag human performance by 11.3 to 21.3 percentage points.

  9. Meta-Cultural Competence: Climbing the Right Hill of Cultural Awareness

    cs.CY 2025-02 conditional novelty 6.0 of 10

    The paper argues that LLMs should be evaluated and built for meta-cultural competence rather than static knowledge of specific cultures, and gives a first, illustrative measurement of one component.

  10. Do Large Language Models Know Folktales? A Case Study of Yokai in Japanese Folktales

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A benchmark of 809 yokai questions shows Japanese-centric LLMs, particularly Llama-3-based continual pretraining models, outperform English-centric models on Japanese folktale knowledge.

Pith tools