REVIEW 3 major objections 4 minor 18 references
Musical ethnocentrism in Large Language Models
T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Large language models show a strong, consistent preference for Western music cultures in both top-artist lists and country-level rating tasks.
desk verdict A modest but honest first measurement of geocultural bias in LLM music output; the Top-100 evidence is solid, the rating experiment needs validation before its results carry weight. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a pair of prompt-based probes plus a postprocessing pipeline. The "Top 100" probe asks the model to enumerate performers in five categories and then attaches countries of origin; the rating probe adapts the prompt protocol of Manvi et al. (2024), asking the model to rate a sampled country relative to all populated locations on Earth on six named musical attributes. Mention frequencies and normalized, run-averaged ratings are then mapped to world maps. The paired probes are meant to catch two different things: open-ended generation reveals which cultures come to mind, while numeric ratings are meant to reveal implicit judgments about them.
What would settle it
Run the same two prompts on a model trained predominantly on non-Western web text, for example a Chinese LLM, with the country list translated into Chinese. If the "Top 100" lists and country ratings still place U.S. and European music on top and Asia and Africa at the bottom, then the bias is not explained by the regional composition of training data; if the model favors its home region, the paper's training-data explanation gains support.
Extended reading notes
Core claim
The paper's central claim is that current LLMs display musical ethnocentrism in a specific, measurable way. In the first experiment, asking for top bands, singers, solo artists, instrumentalists, and composers produces country-of-origin distributions concentrated in Western countries, especially the U.S., with Asia and Africa almost absent and South America in between. In the second, asking for numeric ratings of agreeableness, successfulness, musical creativity, global influence, musical tradition, and musical complexity yields the same Western-leaning ordering, with India as the main exception under "Tradition." The author interprets the consistency across the two models and four languages as evidence that the bias comes from the composition and value judgments of shared training data rather than from any single prompt.
Load-bearing premise
The rating experiment assumes that asking an LLM to put a number on a country's musical agreeableness, success, creativity, influence, tradition, or complexity extracts a learned cultural judgment rather than a plausible-sounding number with no stable semantics; the paper itself admits the task is not well-posed.
Editorial extensions
If this is right
- LLM-generated music rankings, recommendations, and cultural overviews will systematically underrepresent Asian and African music while overrepresenting U.S. and European music.
- Changing the prompt language to Spanish, Chinese, or French does not fix the skew, so mitigation cannot rely on localization of prompts alone.
- The skew is not specific to one provider: two models trained on different continents still land on the same Western-heavy distribution, suggesting a shared training-data cause.
- Because the bias is implicit, it can leak into downstream tasks such as writing assistance and recommendation pipelines, where it is harder to notice than an explicitly biased filter.
- For tasks like "Top 100," users may expect the biased answer, which poses a design choice about whether models should mirror human majority taste or offer broader diversity.
Reading between the lines
- An implication the paper leaves implicit is that the "Top 100" probe likely conflates commercial success or streaming volume with musical importance; a testable extension would compare the model's country distribution against global recorded-music market shares to separate learned consumption statistics from cultural valuation.
- The rating probe could be validated by asking models to justify a few ratings in free text before giving numbers, or by rating fictional countries; if scores do not track the justification or track a random-seeming baseline, the numeric scale is measuring prompt compliance rather than a stable belief about music culture.
- The same two-prompt design could be transferred to other cultural domains such as cuisine, literature, or cinema; if the Western skew reproduces there, the paper's result would generalize from music to a broader cultural ethnocentrism in LLMs.
- Because the author tested only Western-trained models, the decisive comparison would come from a large model trained primarily on Chinese, Arabic, or Hindi web text; the paper's training-data explanation predicts that model would favor its own cultural region, while a "global elite" variant of the bias would look different.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates whether LLMs exhibit a Western (especially U.S.) bias in music-related judgments. In Experiment 1, the authors prompt ChatGPT-4 and Mixtral-8x7B, through their online interfaces, to generate 'Top 100' lists of musical contributors (bands, singers, instrumentalists, composers, solo artists) in several languages, then map the countries of origin. In Experiment 2, the same models are asked to rate countries on six musical attributes (agreeableness, successfulness, creativity, global influence, tradition, complexity) using a prompt adapted from Manvi et al. (2024). The results are presented as choropleth maps and informal comparisons. The paper concludes that LLMs show a strong preference for Western music cultures in both experiments, with the U.S. dominating, and that this bias persists across models and languages.
Significance. If the central claim is established, this is a useful contribution to the emerging literature on geocultural bias in LLMs, extending prior geographic-bias findings (Manvi et al., 2024) to a previously understudied domain, music. The paper also ships a public repository with data and analysis notebooks, which supports reproducibility and follow-up work. However, the current evidence is descriptive: no statistical tests, confidence intervals, or variance measures are reported, and the rating experiment rests on an assumption that the authors themselves partially disavow in Section 6. The strength of the claim ('strong preference in both experiments') is therefore not yet commensurate with the analysis.
major comments (3)
- [Section 4.2, Listing 1, Section 6] The above is a single complete comment.
- [Section 4] The above is a single complete comment.
- [Section 3.1, Section 6] The above is a single complete comment.
minor comments (4)
- [Section 3.2] The above is a single complete comment.
- [Section 3.3] The above is a single complete comment.
- [Section 4.2] The above is a single complete comment.
- [Appendix Figures 3 and 4] The above is a single complete comment.
Circularity Check
No circularity: the study directly measures LLM outputs and draws no derivation that presupposes its conclusion.
full rationale
The paper's claim is an empirical measurement: the models are prompted to list top musical contributors or rate musical culture aspects, and the resulting country distributions are reported. There is no fitted parameter, no equation linking inputs to conclusions, and no self-citation that carries a load-bearing argument. The authors explicitly acknowledge in Section 6 that the rating task is 'not well-posed,' but that is a validity caveat, not a circularity: the experiment still measures what the model outputs under the prompt. The two cited external works (Manvi et al., 2024; Tao et al., 2024) motivate the methodology and prior findings but do not supply the present paper's result. The only self-citation (Kruspe, 2024) appears in a future-work suggestion about transparency and is not used to justify the empirical claim. Because the analysis is directly observational, the central result is self-contained and not reducible to its inputs by construction.
Assumptions & free parameters
assumptions (4)
- domain assumption The rating prompt methodology from Manvi et al. (2024) transfers to music attributes and elicits genuine learned judgments rather than random or prompted-output noise.
- domain assumption Three repeated runs per condition are sufficient to approximate the model's typical output distribution for a prompt.
- domain assumption The model-generated country-of-origin labels in the Top 100 experiment are accurate enough to compute meaningful country frequencies.
- domain assumption Normalized high/low ratings on attributes like 'complexity' or 'agreeableness' are interpretable as cultural preference or bias.
Cite this review
Pith. "Pith review of Musical ethnocentrism in Large Language Models." pith.science (2026). https://pith.science/paper/HPPPJWLK
@misc{pith2026250113720,
author = {Pith},
title = {Pith review of: Musical ethnocentrism in Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/HPPPJWLK}},
note = {Machine review of arXiv:2501.13720}
}
read the original abstract
Large Language Models (LLMs) reflect the biases in their training data and, by extension, those of the people who created this training data. Detecting, analyzing, and mitigating such biases is becoming a focus of research. One type of bias that has been understudied so far are geocultural biases. Those can be caused by an imbalance in the representation of different geographic regions and cultures in the training data, but also by value judgments contained therein. In this paper, we make a first step towards analyzing musical biases in LLMs, particularly ChatGPT and Mixtral. We conduct two experiments. In the first, we prompt LLMs to provide lists of the "Top 100" musical contributors of various categories and analyze their countries of origin. In the second experiment, we ask the LLMs to numerically rate various aspects of the musical cultures of different countries. Our results indicate a strong preference of the LLMs for Western music cultures in both experiments.
Figures
Reference graph
Works this paper leans on
-
[1]
Muhammad Farid Adilazuarda, Sagnik Mukherjee, Pradhyumna Lavania, Siddhant Singh, Alham Fikri Aji, Jacki O'Neill, Ashutosh Modi, and Monojit Choudhury. 2024. https://arxiv.org/abs/2403.15412 Towards Measuring and Modeling "Culture" in LLMs: A Survey . Preprint, arXiv:2403.15412
arXiv 2024
-
[2]
Pablo Biedma, Xiaoyuan Yi, Linus Huang, Maosong Sun, and Xing Xie. 2024. Beyond Human Norms: Unveiling Unique Values of Large Language Models through Interdisciplinary Approaches . arXiv, 2404.12744v1
arXiv 2024
-
[3]
Tolga Bolukbasi, Kai-Wei Chang, James Zou, Venkatesh Saligrama, and Adam Kalai. 2016. Man is to computer programmer as woman is to homemaker? Debiasing word embeddings . In Proceedings of the 30th International Conference on Neural Information Processing Systems, NIPS'16
work page 2016
-
[4]
Hadas Kotek, Rikker Dockum, and David Sun. 2023. https://doi.org/10.1145/3582269.3615599 Gender bias and stereotypes in Large Language Models . In Proceedings of The ACM Collective Intelligence Conference, CI '23, page 12–24, New York, NY, USA. Association for Computing Machinery
arXiv 2023
-
[5]
Anna Kruspe. 2024. Towards detecting unanticipated bias in language models. arXiv preprint arXiv:2404.02650
work page Pith review arXiv 2024
-
[7]
Cheng Li, Damien Teney, Linyi Yang, Qingsong Wen, Xing Xie, and Jindong Wang. 2024 b . CulturePark: Boosting Cross-cultural Understanding in Large Language Models . arXiv, 2405.15145
arXiv 2024
-
[8]
Jiajia Li, Lu Yang, Mingni Tang, Cong Chen, Zuchao Li, Ping Wang, and Hai Zhao. 2024 c . https://arxiv.org/abs/2406.15885 The Music Maestro or The Musically Challenged, A Massive Music Evaluation Benchmark for Large Language Models . Preprint, arXiv:2406.15885
arXiv 2024
-
[9]
Rohin Manvi, Samar Khanna, Marshall Burke, David Lobell, and Stefano Ermon. 2024. https://arxiv.org/abs/2402.02680 Large Language Models are Geographically Biased . arXiv preprint arXiv:2402.02680
arXiv 2024
Show all 18 references
-
[10]
Tarek Naous, Michael Joseph Ryan, and Wei Xu. 2023. https://api.semanticscholar.org/CorpusID:258865272 Having Beer after Prayer? Measuring Cultural Bias in Large Language Models . In Annual Meeting of the Association for Computational Linguistics
2023
-
[11]
Omiye, Jenna C
Jesutofunmi A. Omiye, Jenna C. Lester, Simon Spichak, Veronica Rotemberg, and Roxana Daneshjou. 2023. https://doi.org/10.1038/s41746-023-00939-z Large Language Models propagate race-based medicine . npj Digital Medicine, 6(195)
2023 doi
-
[12]
Abhinav Rao, Akhila Yerukola, Vishwa Shah, Katharina Reinecke, and Maarten Sap. 2024. NormAd: A Benchmark for Measuring the Cultural Adaptability of Large Language Models . arXiv, 2404.12464
2024 arXiv
-
[13]
Huaman Sun, Jiaxin Pei, Minje Choi, and David Jurgens. 2023. https://arxiv.org/abs/2311.09730 Aligning with Whom? Large Language Models Have Gender and Racial Biases in Subjective NLP Tasks . Preprint, arXiv:2311.09730
2023 arXiv
-
[14]
Kizilcec
Yan Tao, Olga Viberg, Ryan S Baker, and René F. Kizilcec. 2024. https://doi.org/10.1093/pnasnexus/pgae346 Cultural bias and cultural alignment of Large Language Models . PNAS Nexus, 3(9):pgae346
2024 doi
-
[15]
Yuhang Wang, Yanxu Zhu, Chao Kong, Shuyu Wei, Xiaoyuan Yi, Xing Xie, and Jitao Sang. 2023. CDEval: A Benchmark for Measuring the Cultural Dimensions of Large Language Models . arXiv preprint arXiv:2402.10946
2023 arXiv
-
[16]
Melissa Warr, Nicole Jakubczyk Oster, and Roger Isaac. 2023. Implicit Bias in Large Language Models: Experimental Proof and Implications for Education . SSRN Electronic Journal
2023
-
[17]
Lingxuan Zhu, Weiming Mou, Yancheng Lai, Junda Lin, and Peng Luo. 2024. Language and cultural bias in AI: Comparing the performance of Large Language Models developed in different countries on Traditional Chinese Medicine highlights the need for localized models . Journal of T...
2024
-
[18]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[19]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.