Pith. sign in

REVIEW 4 cited by

BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.09948 v2 pith:PPOPZE4J submitted 2024-06-14 cs.CL

classification cs.CL
keywords languagesllmsculturesblenddiverseknowledgebenchmarkbetter
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) often lack culture-specific knowledge of daily life, especially across diverse regions and non-English languages. Existing benchmarks for evaluating LLMs' cultural sensitivities are limited to a single language or collected from online sources such as Wikipedia, which do not reflect the mundane everyday lifestyles of diverse regions. That is, information about the food people eat for their birthday celebrations, spices they typically use, musical instruments youngsters play, or the sports they practice in school is common cultural knowledge but uncommon in easily collected online sources, especially for underrepresented cultures. To address this issue, we introduce BLEnD, a hand-crafted benchmark designed to evaluate LLMs' everyday knowledge across diverse cultures and languages. BLEnD comprises 52.6k question-answer pairs from 16 countries/regions, in 13 different languages, including low-resource ones such as Amharic, Assamese, Azerbaijani, Hausa, and Sundanese. We construct the benchmark to include two formats of questions: short-answer and multiple-choice. We show that LLMs perform better for cultures that are highly represented online, with a maximum 57.34% difference in GPT-4, the best-performing model, in the short-answer format. For cultures represented by mid-to-high-resource languages, LLMs perform better in their local languages, but for cultures represented by low-resource languages, LLMs perform better in English than the local languages. We make our dataset publicly available at: https://github.com/nlee0212/BLEnD.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 5 citations worldwide. Full citation record

  1. CultureSynth: A Hierarchical Taxonomy-Guided and Retrieval-Augmented Framework for Cultural Question-Answer Synthesis

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A taxonomy-guided retrieval-augmented framework generates CultureSynth-7, a multilingual cultural QA benchmark, and its evaluation of 14 LLMs suggests cultural competence emerges around 3B parameters.

  2. Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language Models

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Marco-Bench-MIF localizes IFEval into 30 languages with cultural adaptation and finds large resource gaps and scale effects in multilingual instruction following.

  3. Nunchi-Bench: Benchmarking Language Models on Cultural Reasoning with a Focus on Korean Superstition

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A new benchmark, Nunchi-Bench, shows that LLMs know Korean superstition facts but frequently fail to apply them in practical cultural contexts, and that explicit cultural framing beats prompt language alone.

  4. CulFiT: A Fine-grained Cultural-aware LLM Training Paradigm via Multilingual Critique Data Synthesis

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A multilingual critique-data training paradigm with a knowledge-unit reward improves LLM cultural alignment on several benchmarks, but its headline benchmark is evaluated with the same LLM-judged metric used to select...

Pith tools