Pith. sign in

REVIEW 10 cited by

Survey of Cultural Awareness in Language Models: Text and Beyond

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.00860 v1 pith:KGTE3HEP submitted 2024-10-30 cs.CL cs.CV

classification cs.CLcs.CV
keywords culturalllmsawarenessanthropologybeenpsychologyresearchalignment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large-scale deployment of large language models (LLMs) in various applications, such as chatbots and virtual assistants, requires LLMs to be culturally sensitive to the user to ensure inclusivity. Culture has been widely studied in psychology and anthropology, and there has been a recent surge in research on making LLMs more culturally inclusive in LLMs that goes beyond multilinguality and builds on findings from psychology and anthropology. In this paper, we survey efforts towards incorporating cultural awareness into text-based and multimodal LLMs. We start by defining cultural awareness in LLMs, taking the definitions of culture from anthropology and psychology as a point of departure. We then examine methodologies adopted for creating cross-cultural datasets, strategies for cultural inclusion in downstream tasks, and methodologies that have been used for benchmarking cultural awareness in LLMs. Further, we discuss the ethical implications of cultural alignment, the role of Human-Computer Interaction in driving cultural inclusion in LLMs, and the role of cultural alignment in driving social science research. We finally provide pointers to future research based on our findings about gaps in the literature.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Cross-Lingual Transfer of Cultural Knowledge: An Asymmetric Phenomenon

    cs.CL 2025-06 conditional novelty 7.0 of 10

    Cross-lingual transfer of cultural knowledge is bidirectional for high-resource languages and asymmetric for low-resource ones, with corpus frequency correlating with transfer success.

  2. Show Me the Work: Fact-Checkers' Requirements for Explainable Automated Fact-Checking

    cs.HC 2025-02 conditional novelty 7.0 of 10

    Fact-checkers want automated fact-checking explanations that trace the reasoning path, cite checkable evidence, and clearly flag uncertainty and information gaps, not just confidence scores.

  3. CultureSynth: A Hierarchical Taxonomy-Guided and Retrieval-Augmented Framework for Cultural Question-Answer Synthesis

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A taxonomy-guided retrieval-augmented framework generates CultureSynth-7, a multilingual cultural QA benchmark, and its evaluation of 14 LLMs suggests cultural competence emerges around 3B parameters.

  4. Evaluating Chinese Large Language Models: The Influence of Persona Assignment on Stereotypes and Safeguards

    cs.CY 2025-06 conditional novelty 6.0 of 10

    Assigning personas to Chinese LLMs amplifies toxic output relative to default behavior, while refusal rates shift systematically with persona gender and target social group.

  5. CulFiT: A Fine-grained Cultural-aware LLM Training Paradigm via Multilingual Critique Data Synthesis

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A multilingual critique-data training paradigm with a knowledge-unit reward improves LLM cultural alignment on several benchmarks, but its headline benchmark is evaluated with the same LLM-judged metric used to select...

  6. Prompt Programming for Cultural Bias and Alignment of Large Language Models

    cs.AI 2026-03 conditional novelty 5.0 of 10

    Automatically optimized prompts (DSPy) reduce survey-measured cultural distance for open-weight LLMs more often than manual cultural prompting, with MIPROv2 and a large proposer model giving the most consistent gains.

  7. Exploring LLM-generated Culture-specific Affective Human-Robot Tactile Interaction

    cs.HC 2025-07 conditional novelty 5.0 of 10

    LLM-generated touch descriptions conveyed six of twelve emotions above chance in matched cultures, while cultural mismatch reduced decoding accuracy and perceived appropriateness.

  8. BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation

    cs.LG 2025-05 conditional novelty 5.0 of 10

    BenchHub is an automatically categorized, customizable LLM benchmark suite covering 303K questions across 38 benchmarks in English and Korean.

  9. Igniting Creative Writing in Small Language Models: LLM-as-a-Judge versus Multi-Agent Refined Rewards

    cs.CL 2025-08 conditional novelty 4.0 of 10

    An adversarially tuned LLM-as-a-Judge reward signal outperforms a multi-agent-refined reward model for fine-tuning a 7B SLM on Chinese greeting generation, though the comparison is weakened by circular evaluation and ...

  10. An evaluation of LLMs for generating movie reviews: GPT-4o, Gemini-2.0 and DeepSeek-V3

    cs.CL 2025-05 conditional novelty 4.0 of 10

    LLMs can produce fluent movie reviews that readers often mistake for human-written ones, but the models differ in emotional balance and depth.

Pith tools