Pith. sign in

REVIEW 9 cited by

CulturalBench: A Robust, Diverse, and Challenging Cultural Benchmark by Human-AI CulturalTeaming

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.02677 v2 pith:F45YC7WV submitted 2024-10-03 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords culturalbenchdiversequestionsacrosschallengingculturalaccuracychinese
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Robust, diverse, and challenging cultural knowledge benchmarks are essential for measuring our progress towards making LMs that are helpful across diverse cultures. We introduce CulturalBench: a set of 1,696 human-written and human-verified questions to assess LMs' cultural knowledge, covering 45 global regions including underrepresented ones like Bangladesh, Zimbabwe, and Peru. Questions are each verified by five independent annotators and span 17 diverse topics ranging from food preferences to greeting etiquette. We construct CulturalBench using methods inspired by Human-AI Red-Teaming. Compared to human performance (92.4% accuracy), the hard version of CulturalBench is challenging even for the best-performing frontier LMs, ranging from 28.7% to 61.5% in accuracy. We find that LMs often struggle with tricky questions that have multiple correct answers (e.g., What utensils do the Chinese usually use?), revealing a tendency to overfit to a single answer. Our results indicate that GPT-4o substantially outperform other models across cultures, besting local providers (e.g., Mistral on European culture and DeepSeek on Chinese culture). Across the board, models under-perform on questions related to North Africa, South America and Middle East.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CultureSynth: A Hierarchical Taxonomy-Guided and Retrieval-Augmented Framework for Cultural Question-Answer Synthesis

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A taxonomy-guided retrieval-augmented framework generates CultureSynth-7, a multilingual cultural QA benchmark, and its evaluation of 14 LLMs suggests cultural competence emerges around 3B parameters.

  2. MyCulture: Exploring Malaysia's Diverse Culture under Low-Resource Language Constraints

    cs.CL 2025-08 reject novelty 6.0 of 10

    MyCulture, a new Malay-language cultural benchmark, shows LLM accuracy drops by at least 17% when multiple-choice questions are converted to an open-ended format.

  3. Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language Models

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Marco-Bench-MIF localizes IFEval into 30 languages with cultural adaptation and finds large resource gaps and scale effects in multilingual instruction following.

  4. Disentangling Language and Culture for Evaluating Multilingual Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new dual-axis evaluation framework shows multilingual LLMs answer culture-specific questions best when the question language matches the cultural context, with partial neuron-level evidence for the effect.

  5. CulFiT: A Fine-grained Cultural-aware LLM Training Paradigm via Multilingual Critique Data Synthesis

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A multilingual critique-data training paradigm with a knowledge-unit reward improves LLM cultural alignment on several benchmarks, but its headline benchmark is evaluated with the same LLM-judged metric used to select...

  6. SEADialogues: A Multilingual Culturally Grounded Multi-turn Dialogue Dataset on Southeast Asian Languages

    cs.CL 2025-08 unverdicted novelty 5.0 of 10

    SEADialogues is a culturally grounded multi-turn dialogue dataset covering eight Southeast Asian languages.

  7. BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation

    cs.LG 2025-05 conditional novelty 5.0 of 10

    BenchHub is an automatically categorized, customizable LLM benchmark suite covering 303K questions across 38 benchmarks in English and Korean.

  8. Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models?

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Large reasoning models default to English or Chinese as internal 'reasoning hubs', and forcing non-hub reasoning lowers math accuracy, especially for low-resource languages, while sometimes improving safety and cultur...

  9. Prompt Robustness Is Task-Dependent: Comparing Objective and Belief-Style Questions in LLM Evaluation

    cs.CL 2026-07 conditional novelty 4.5 of 10

    Prompt robustness in LLMs is systematically lower for subjective survey items than for objective questions, with the largest gap under option-order perturbations and strong model–dataset–prompt interactions.

Pith tools