Pith. sign in

REVIEW 3 major objections 4 minor 18 references

Musical ethnocentrism in Large Language Models

T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read Large language models show a strong, consistent preference for Western music cultures in both top-artist lists and country-level rating tasks.

desk verdict A modest but honest first measurement of geocultural bias in LLM music output; the Top-100 evidence is solid, the rating experiment needs validation before its results carry weight. read the letter →

arxiv 2501.13720 v2 pith:HPPPJWLK submitted 2025-01-23 cs.CL cs.AIcs.SDeess.AS

classification cs.CLcs.AIcs.SDeess.AS
keywords largelanguagemodelsculturalbiasgeoculturalmusicalethnocentrismChatGPTMixtralpromptevaluationmusicculture
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper asks whether large language models, which are already known to mirror geographic biases, also carry a musical version of that bias. The author prompts two models, ChatGPT-4 and Mixtral-8x7B, to produce lists of the "Top 100" performers in several categories and to rate the music of every country on six attributes, with prompts in English, Spanish, Chinese, and French. The reported result is a strong preference for Western music cultures, above all the U.S., in both tasks and across nearly all prompt versions. A sympathetic reader would care because the same models are used to summarize, recommend, and write about music, so an unnoticed Western skew would quietly shape which music cultures users encounter.

What carries the argument

The machinery is a pair of prompt-based probes plus a postprocessing pipeline. The "Top 100" probe asks the model to enumerate performers in five categories and then attaches countries of origin; the rating probe adapts the prompt protocol of Manvi et al. (2024), asking the model to rate a sampled country relative to all populated locations on Earth on six named musical attributes. Mention frequencies and normalized, run-averaged ratings are then mapped to world maps. The paired probes are meant to catch two different things: open-ended generation reveals which cultures come to mind, while numeric ratings are meant to reveal implicit judgments about them.

What would settle it

Run the same two prompts on a model trained predominantly on non-Western web text, for example a Chinese LLM, with the country list translated into Chinese. If the "Top 100" lists and country ratings still place U.S. and European music on top and Asia and Africa at the bottom, then the bias is not explained by the regional composition of training data; if the model favors its home region, the paper's training-data explanation gains support.

Watch

Extended reading notes

Core claim

The paper's central claim is that current LLMs display musical ethnocentrism in a specific, measurable way. In the first experiment, asking for top bands, singers, solo artists, instrumentalists, and composers produces country-of-origin distributions concentrated in Western countries, especially the U.S., with Asia and Africa almost absent and South America in between. In the second, asking for numeric ratings of agreeableness, successfulness, musical creativity, global influence, musical tradition, and musical complexity yields the same Western-leaning ordering, with India as the main exception under "Tradition." The author interprets the consistency across the two models and four languages as evidence that the bias comes from the composition and value judgments of shared training data rather than from any single prompt.

Load-bearing premise

The rating experiment assumes that asking an LLM to put a number on a country's musical agreeableness, success, creativity, influence, tradition, or complexity extracts a learned cultural judgment rather than a plausible-sounding number with no stable semantics; the paper itself admits the task is not well-posed.

Editorial extensions

If this is right

  • LLM-generated music rankings, recommendations, and cultural overviews will systematically underrepresent Asian and African music while overrepresenting U.S. and European music.
  • Changing the prompt language to Spanish, Chinese, or French does not fix the skew, so mitigation cannot rely on localization of prompts alone.
  • The skew is not specific to one provider: two models trained on different continents still land on the same Western-heavy distribution, suggesting a shared training-data cause.
  • Because the bias is implicit, it can leak into downstream tasks such as writing assistance and recommendation pipelines, where it is harder to notice than an explicitly biased filter.
  • For tasks like "Top 100," users may expect the biased answer, which poses a design choice about whether models should mirror human majority taste or offer broader diversity.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An implication the paper leaves implicit is that the "Top 100" probe likely conflates commercial success or streaming volume with musical importance; a testable extension would compare the model's country distribution against global recorded-music market shares to separate learned consumption statistics from cultural valuation.
  • The rating probe could be validated by asking models to justify a few ratings in free text before giving numbers, or by rating fictional countries; if scores do not track the justification or track a random-seeming baseline, the numeric scale is measuring prompt compliance rather than a stable belief about music culture.
  • The same two-prompt design could be transferred to other cultural domains such as cuisine, literature, or cinema; if the Western skew reproduces there, the paper's result would generalize from music to a broader cultural ethnocentrism in LLMs.
  • Because the author tested only Western-trained models, the decisive comparison would come from a large model trained primarily on Chinese, Arabic, or Hindi web text; the paper's training-data explanation predicts that model would favor its own cultural region, while a "global elite" variant of the bias would look different.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper investigates whether LLMs exhibit a Western (especially U.S.) bias in music-related judgments. In Experiment 1, the authors prompt ChatGPT-4 and Mixtral-8x7B, through their online interfaces, to generate 'Top 100' lists of musical contributors (bands, singers, instrumentalists, composers, solo artists) in several languages, then map the countries of origin. In Experiment 2, the same models are asked to rate countries on six musical attributes (agreeableness, successfulness, creativity, global influence, tradition, complexity) using a prompt adapted from Manvi et al. (2024). The results are presented as choropleth maps and informal comparisons. The paper concludes that LLMs show a strong preference for Western music cultures in both experiments, with the U.S. dominating, and that this bias persists across models and languages.

Significance. If the central claim is established, this is a useful contribution to the emerging literature on geocultural bias in LLMs, extending prior geographic-bias findings (Manvi et al., 2024) to a previously understudied domain, music. The paper also ships a public repository with data and analysis notebooks, which supports reproducibility and follow-up work. However, the current evidence is descriptive: no statistical tests, confidence intervals, or variance measures are reported, and the rating experiment rests on an assumption that the authors themselves partially disavow in Section 6. The strength of the claim ('strong preference in both experiments') is therefore not yet commensurate with the analysis.

major comments (3)
  1. [Section 4.2, Listing 1, Section 6] The above is a single complete comment.
  2. [Section 4] The above is a single complete comment.
  3. [Section 3.1, Section 6] The above is a single complete comment.
minor comments (4)
  1. [Section 3.2] The above is a single complete comment.
  2. [Section 3.3] The above is a single complete comment.
  3. [Section 4.2] The above is a single complete comment.
  4. [Appendix Figures 3 and 4] The above is a single complete comment.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the study directly measures LLM outputs and draws no derivation that presupposes its conclusion.

full rationale

The paper's claim is an empirical measurement: the models are prompted to list top musical contributors or rate musical culture aspects, and the resulting country distributions are reported. There is no fitted parameter, no equation linking inputs to conclusions, and no self-citation that carries a load-bearing argument. The authors explicitly acknowledge in Section 6 that the rating task is 'not well-posed,' but that is a validity caveat, not a circularity: the experiment still measures what the model outputs under the prompt. The two cited external works (Manvi et al., 2024; Tao et al., 2024) motivate the methodology and prior findings but do not supply the present paper's result. The only self-citation (Kruspe, 2024) appears in a future-work suggestion about transparency and is not used to justify the empirical claim. Because the analysis is directly observational, the central result is self-contained and not reducible to its inputs by construction.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

This is an empirical bias audit, not a derivation. The load-bearing assumptions are about the validity of the prompting/normalization pipeline and the interpretation of model outputs as cultural judgments.

assumptions (4)
  • domain assumption The rating prompt methodology from Manvi et al. (2024) transfers to music attributes and elicits genuine learned judgments rather than random or prompted-output noise.
    The second experiment is a direct adaptation of Manvi et al.; the paper does not validate that subjective musical attributes behave like the geographic attributes in that study.
  • domain assumption Three repeated runs per condition are sufficient to approximate the model's typical output distribution for a prompt.
    The paper averages over three runs but reports no variance or convergence check.
  • domain assumption The model-generated country-of-origin labels in the Top 100 experiment are accurate enough to compute meaningful country frequencies.
    The paper keeps only the first country named per performer and does not verify the labels against external knowledge.
  • domain assumption Normalized high/low ratings on attributes like 'complexity' or 'agreeableness' are interpretable as cultural preference or bias.
    The central conclusion interprets the rating pattern as bias, which presumes the attribute ratings are value judgments rather than mere geographic correlations in the model's knowledge.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Musical ethnocentrism in Large Language Models." pith.science (2026). https://pith.science/paper/HPPPJWLK

@misc{pith2026250113720,
  author       = {Pith},
  title        = {Pith review of: Musical ethnocentrism in Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HPPPJWLK}},
  note         = {Machine review of arXiv:2501.13720}
}
read the original abstract

Large Language Models (LLMs) reflect the biases in their training data and, by extension, those of the people who created this training data. Detecting, analyzing, and mitigating such biases is becoming a focus of research. One type of bias that has been understudied so far are geocultural biases. Those can be caused by an imbalance in the representation of different geographic regions and cultures in the training data, but also by value judgments contained therein. In this paper, we make a first step towards analyzing musical biases in LLMs, particularly ChatGPT and Mixtral. We conduct two experiments. In the first, we prompt LLMs to provide lists of the "Top 100" musical contributors of various categories and analyze their countries of origin. In the second experiment, we ask the LLMs to numerically rate various aspects of the musical cultures of different countries. Our results indicate a strong preference of the LLMs for Western music cultures in both experiments.

Figures

Figures reproduced from arXiv: 2501.13720 by the authors.

Figure 1
Figure 1. Example results of the “Top 100” experiments for singers, prompted on GPT in English and Chinese, and [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Example results of the rating experiments for musical complexity, prompted on GPT in English and [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. “Top 100” result graphs. Gray means None, and darker colors indicate higher numbers. [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Rating result graphs. Scale runs from dark red (low rating) to bright yellow (high rating). [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

18 extracted references · 5 canonical work pages

  1. [1]

    Muhammad Farid Adilazuarda, Sagnik Mukherjee, Pradhyumna Lavania, Siddhant Singh, Alham Fikri Aji, Jacki O'Neill, Ashutosh Modi, and Monojit Choudhury. 2024. https://arxiv.org/abs/2403.15412 Towards Measuring and Modeling "Culture" in LLMs: A Survey . Preprint, arXiv:2403.15412

  2. [2]

    Pablo Biedma, Xiaoyuan Yi, Linus Huang, Maosong Sun, and Xing Xie. 2024. Beyond Human Norms: Unveiling Unique Values of Large Language Models through Interdisciplinary Approaches . arXiv, 2404.12744v1

  3. [3]

    Tolga Bolukbasi, Kai-Wei Chang, James Zou, Venkatesh Saligrama, and Adam Kalai. 2016. Man is to computer programmer as woman is to homemaker? Debiasing word embeddings . In Proceedings of the 30th International Conference on Neural Information Processing Systems, NIPS'16

  4. [4]

    Hadas Kotek, Rikker Dockum, and David Sun. 2023. https://doi.org/10.1145/3582269.3615599 Gender bias and stereotypes in Large Language Models . In Proceedings of The ACM Collective Intelligence Conference, CI '23, page 12–24, New York, NY, USA. Association for Computing Machinery

  5. [5]

    Anna Kruspe. 2024. Towards detecting unanticipated bias in language models. arXiv preprint arXiv:2404.02650

  6. [7]

    Cheng Li, Damien Teney, Linyi Yang, Qingsong Wen, Xing Xie, and Jindong Wang. 2024 b . CulturePark: Boosting Cross-cultural Understanding in Large Language Models . arXiv, 2405.15145

  7. [8]

    Jiajia Li, Lu Yang, Mingni Tang, Cong Chen, Zuchao Li, Ping Wang, and Hai Zhao. 2024 c . https://arxiv.org/abs/2406.15885 The Music Maestro or The Musically Challenged, A Massive Music Evaluation Benchmark for Large Language Models . Preprint, arXiv:2406.15885

  8. [9]

    Rohin Manvi, Samar Khanna, Marshall Burke, David Lobell, and Stefano Ermon. 2024. https://arxiv.org/abs/2402.02680 Large Language Models are Geographically Biased . arXiv preprint arXiv:2402.02680

Show all 18 references
  1. [10]

    Tarek Naous, Michael Joseph Ryan, and Wei Xu. 2023. https://api.semanticscholar.org/CorpusID:258865272 Having Beer after Prayer? Measuring Cultural Bias in Large Language Models . In Annual Meeting of the Association for Computational Linguistics

  2. [11]

    Omiye, Jenna C

    Jesutofunmi A. Omiye, Jenna C. Lester, Simon Spichak, Veronica Rotemberg, and Roxana Daneshjou. 2023. https://doi.org/10.1038/s41746-023-00939-z Large Language Models propagate race-based medicine . npj Digital Medicine, 6(195)

  3. [12]

    Abhinav Rao, Akhila Yerukola, Vishwa Shah, Katharina Reinecke, and Maarten Sap. 2024. NormAd: A Benchmark for Measuring the Cultural Adaptability of Large Language Models . arXiv, 2404.12464

  4. [13]

    Huaman Sun, Jiaxin Pei, Minje Choi, and David Jurgens. 2023. https://arxiv.org/abs/2311.09730 Aligning with Whom? Large Language Models Have Gender and Racial Biases in Subjective NLP Tasks . Preprint, arXiv:2311.09730

  5. [14]

    Kizilcec

    Yan Tao, Olga Viberg, Ryan S Baker, and René F. Kizilcec. 2024. https://doi.org/10.1093/pnasnexus/pgae346 Cultural bias and cultural alignment of Large Language Models . PNAS Nexus, 3(9):pgae346

  6. [15]

    Yuhang Wang, Yanxu Zhu, Chao Kong, Shuyu Wei, Xiaoyuan Yi, Xing Xie, and Jitao Sang. 2023. CDEval: A Benchmark for Measuring the Cultural Dimensions of Large Language Models . arXiv preprint arXiv:2402.10946

  7. [16]

    Melissa Warr, Nicole Jakubczyk Oster, and Roger Isaac. 2023. Implicit Bias in Large Language Models: Experimental Proof and Implications for Education . SSRN Electronic Journal

  8. [17]

    Lingxuan Zhu, Weiming Mou, Yancheng Lai, Junda Lin, and Peng Luo. 2024. Language and cultural bias in AI: Comparing the performance of Large Language Models developed in different countries on Traditional Chinese Medicine highlights the need for localized models . Journal of T...

  9. [18]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...

  10. [19]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.