REVIEW 7 cited by
Massively Multi-Cultural Knowledge Acquisition & LM Benchmarking
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Pretrained large language models have revolutionized many applications but still face challenges related to cultural bias and a lack of cultural commonsense knowledge crucial for guiding cross-culture communication and interactions. Recognizing the shortcomings of existing methods in capturing the diverse and rich cultures across the world, this paper introduces a novel approach for massively multicultural knowledge acquisition. Specifically, our method strategically navigates from densely informative Wikipedia documents on cultural topics to an extensive network of linked pages. Leveraging this valuable source of data collection, we construct the CultureAtlas dataset, which covers a wide range of sub-country level geographical regions and ethnolinguistic groups, with data cleaning and preprocessing to ensure textual assertion sentence self-containment, as well as fine-grained cultural profile information extraction. Our dataset not only facilitates the evaluation of language model performance in culturally diverse contexts but also serves as a foundational tool for the development of culturally sensitive and aware language models. Our work marks an important step towards deeper understanding and bridging the gaps of cultural disparities in AI, to promote a more inclusive and balanced representation of global cultures in the digital domain.
Forward citations
Cited by 7 Pith papers
-
Cross-Lingual Transfer of Cultural Knowledge: An Asymmetric Phenomenon
Cross-lingual transfer of cultural knowledge is bidirectional for high-resource languages and asymmetric for low-resource ones, with corpus frequency correlating with transfer success.
-
Tears or Cheers? Benchmarking LLMs via Culturally Elicited Distinct Affective Responses
CEDAR is a 7-language, 2-modality benchmark of 10,962 culturally divergent emotion scenarios; 17 LLMs perform poorly, and prompt-language matching does not fix cultural misalignment.
-
The Alignment Veto: How Safety Training Suppresses Cultural Knowledge in LLMs
The full text builds the MENA Values benchmark (864 questions, 7 models) and reports that LLM cultural answers shift with language, decline with reasoning prompts, and hide strong internal preferences behind refusals—...
-
Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages
Across nine Asian languages, multilingual LLMs favor Western cultural entities in 30-40% of culturally grounded contexts, with model-specific sentiment biases and extraction accuracy gaps.
-
CulFiT: A Fine-grained Cultural-aware LLM Training Paradigm via Multilingual Critique Data Synthesis
A multilingual critique-data training paradigm with a knowledge-unit reward improves LLM cultural alignment on several benchmarks, but its headline benchmark is evaluated with the same LLM-judged metric used to select...
-
PerCul: A Story-Driven Cultural Evaluation of LLMs in Persian
PerCul is a Persian cultural story-comprehension benchmark on which the best LLMs lag human performance by 11.3 to 21.3 percentage points.
-
Evaluating Vision-Language Models for Emotion Recognition
Vision-language models are weak and prompt-sensitive at evoked emotion recognition, and many fine-grained errors are best explained by noisy dataset labels.
Discussion (0). Continue with ORCID to comment.