Pith. sign in

REVIEW 10 cited by

CulturalTeaming: AI-Assisted Interactive Red-Teaming for Challenging LLMs' (Lack of) Multicultural Knowledge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.06664 v1 pith:2L4C4TAI submitted 2024-04-10 cs.CL cs.AIcs.HC

classification cs.CLcs.AIcs.HC
keywords llmsmulticulturalculturalknowledgeannotatorsassistanceculturalteamingdataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Frontier large language models (LLMs) are developed by researchers and practitioners with skewed cultural backgrounds and on datasets with skewed sources. However, LLMs' (lack of) multicultural knowledge cannot be effectively assessed with current methods for developing benchmarks. Existing multicultural evaluations primarily rely on expensive and restricted human annotations or potentially outdated internet resources. Thus, they struggle to capture the intricacy, dynamics, and diversity of cultural norms. LLM-generated benchmarks are promising, yet risk propagating the same biases they are meant to measure. To synergize the creativity and expert cultural knowledge of human annotators and the scalability and standardizability of LLM-based automation, we introduce CulturalTeaming, an interactive red-teaming system that leverages human-AI collaboration to build truly challenging evaluation dataset for assessing the multicultural knowledge of LLMs, while improving annotators' capabilities and experiences. Our study reveals that CulturalTeaming's various modes of AI assistance support annotators in creating cultural questions, that modern LLMs fail at, in a gamified manner. Importantly, the increased level of AI assistance (e.g., LLM-generated revision hints) empowers users to create more difficult questions with enhanced perceived creativity of themselves, shedding light on the promises of involving heavier AI assistance in modern evaluation dataset creation procedures. Through a series of 1-hour workshop sessions, we gather CULTURALBENCH-V0.1, a compact yet high-quality evaluation dataset with users' red-teaming attempts, that different families of modern LLMs perform with accuracy ranging from 37.7% to 72.2%, revealing a notable gap in LLMs' multicultural proficiency.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. We Politely Insist: Your LLM Must Learn the Persian Art of Taarof

    cs.CL 2025-09 conditional novelty 7.0 of 10

    A new 450-scenario benchmark shows that LLMs lag native Persian speakers by 40 to 48 points on taarof-expected interactions, and that fine-tuning on the benchmark narrows the gap.

  2. Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages

    cs.CL 2025-10 conditional novelty 6.0 of 10

    Across nine Asian languages, multilingual LLMs favor Western cultural entities in 30-40% of culturally grounded contexts, with model-specific sentiment biases and extraction accuracy gaps.

  3. CultureSynth: A Hierarchical Taxonomy-Guided and Retrieval-Augmented Framework for Cultural Question-Answer Synthesis

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A taxonomy-guided retrieval-augmented framework generates CultureSynth-7, a multilingual cultural QA benchmark, and its evaluation of 14 LLMs suggests cultural competence emerges around 3B parameters.

  4. Nunchi-Bench: Benchmarking Language Models on Cultural Reasoning with a Focus on Korean Superstition

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A new benchmark, Nunchi-Bench, shows that LLMs know Korean superstition facts but frequently fail to apply them in practical cultural contexts, and that explicit cultural framing beats prompt language alone.

  5. CulFiT: A Fine-grained Cultural-aware LLM Training Paradigm via Multilingual Critique Data Synthesis

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A multilingual critique-data training paradigm with a knowledge-unit reward improves LLM cultural alignment on several benchmarks, but its headline benchmark is evaluated with the same LLM-judged metric used to select...

  6. On The Origin of Cultural Biases in Language Models: From Pre-training Data to Linguistic Phenomena

    cs.CL 2025-01 conditional novelty 6.0 of 10

    Arab cultural entities that double as everyday Arabic words are harder for language models to recognize, especially when tokenized as single tokens.

  7. WeAudit: Scaffolding User Auditors and AI Practitioners in Auditing Generative AI

    cs.HC 2025-01 conditional novelty 6.0 of 10

    WeAudit scaffolds everyday users to audit generative AI through comparison, examples, discussion, verification, and structured reports, and practitioners found the resulting audit reports actionable.

  8. SafeWorld: Geo-Diverse Safety Alignment

    cs.CL 2024-12 conditional novelty 6.0 of 10

    This paper introduces a geo-diverse cultural and legal safety benchmark and shows that a DPO-trained 7B model can outperform GPT-4o on it, with caveats about the GPT-4-based evaluation loop.

  9. INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge

    cs.CL 2024-11 conditional novelty 6.0 of 10

    INCLUDE is a multilingual benchmark of 197,243 exam questions from local sources that evaluates how well LLMs handle regional and cultural knowledge.

  10. A Survey on Human-Centric LLMs

    cs.CL 2024-11 conditional novelty 1.0 of 10

    A review that sorts existing evidence on how well large language models imitate individual human skills and collective social dynamics into one taxonomy.

Pith tools