Pith. sign in

REVIEW 3 major objections 6 minor 196 references

Music Recommendation with Large Language Models: Challenges, Opportunities, and Evaluation

T0 review · 3 major / 6 minor · reviewed 2026-08-03 · deepseek-v4-flash

Pith's one-line read LLM-based music recommenders cannot be judged by accuracy alone; evaluation must move to groundedness, discovery, and risk diagnostics.

desk verdict A credible, well-organized position paper that gives the MRS community a useful vocabulary for evaluating LLM-based recommenders, but its proposed new metrics are mostly sketches awaiting validation. read the letter →

arxiv 2511.16478 v2 pith:PTKXT6YT submitted 2025-11-20 cs.IR cs.CL

classification cs.IRcs.CL
keywords MusicrecommendersystemsLargelanguagemodelsEvaluationmetricsGroundednessHallucinationPersonalizationNaturalrecommendationBeyond-accuracy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This review argues that LLM-driven music recommender systems break the field's dominant evaluation paradigm, which equates a good recommendation with accurately predicting what users consume. Because LLMs generate text instead of ranking items, standard accuracy metrics miss hallucinations, non-determinism, and personalization, and train/test separation is undermined by opaque training data. The paper proposes a structured alternative: six success dimensions (groundedness, discovery, personalization gain, profile fidelity, cultural/linguistic coverage, classical relevance) plus a complementary set of risk diagnostics, adapted from NLP evaluation practices. If adopted, studies of LLM recommenders would report whether recommendations are real, grounded, diverse, and controllable, not just whether they match logged consumption.

What carries the argument

The central object is the paper's proposed evaluation framework: a decision tree for selecting NLP-derived metrics, plus a catalog of six success dimensions (G1–G6) and eight risk diagnostics. For the natural-language recommendation scenario, the task is defined as producing a ranked list of k resolvable catalog items from a prompt that may include a user request, a natural-language profile, in-context examples, retrieved documents, or a reasoning scaffold. The framework's work is to separate faithful grounding from fluent text, and to make hallucination, bias, and instability measurable rather than hidden inside a single accuracy number.

What would settle it

A large-scale user study would settle it: collect natural-language music requests, have an LLM produce recommendation lists, then compare human judgments of quality against both standard accuracy metrics and the paper's groundedness and discovery metrics; if accuracy correlates as strongly with human preference as the proposed metrics do, the claim that accuracy metrics are inadequate collapses.

Watch

Extended reading notes

Core claim

The paper's central claim is that evaluating an LLM-based music recommender as if it were a retrieval system—scoring precision and recall against consumption logs—no longer measures whether recommendations are good. LLMs are generative, non-deterministic, and prone to hallucination, so the field must rethink evaluation from the ground up. The paper establishes that success should be assessed along six dimensions: query adherence and groundedness, discovery quality, personalization gain, profile fidelity and controllability, cultural and linguistic coverage, and classical relevance. It then adds a complementary set of risk diagnostics covering hallucination, popularity/temporal/language bias,

Load-bearing premise

The framework's prescriptions depend on the assumption that the proposed evaluation measures can be reliably computed and validated—especially LLM-based judges, whose reliability for music content is itself an open question the paper flags.

Editorial extensions

If this is right

  • If the paper is right, accuracy-only benchmarks become insufficient for LLM recommenders; studies would need to report groundedness, discovery, personalization gain, and coverage alongside classical relevance.
  • Entity resolution against a catalog or knowledge base should become a standard evaluation step, so hallucinated tracks and misattributed metadata are penalized explicitly.
  • LLM-as-judge methods should not be adopted in music without first being validated against human judgments, because position, verbosity, and self-enhancement biases can distort scores.
  • Each prompting setting needs specialized tests: shot calibration and prompt sensitivity for in-context learning, evidence grounding and document-swap tests for retrieval-augmented generation, and reasoning faithfulness plus self-consistency for chain-of-thought prompting.
  • Because black-box LLMs may have been exposed to public benchmark datasets during training, offline accuracy results are hard to interpret, and evaluation should be reported with contamination awareness.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper, this evaluation framework likely generalizes to other generative recommender domains such as movies, books, and e-commerce, since the core problem—ranking-based metrics failing on generative outputs—is not music-specific.
  • The risk diagnostics could be operationalized into a standardized 'hallucination rate' and 'bias report card' for deployed music assistants, enabling direct comparison across systems in a way that accuracy metrics do not.
  • The emphasis on profile fidelity suggests a testable interactive protocol: let users edit their natural-language profiles, then measure how quickly and consistently the recommender's outputs adapt; such a protocol is not yet standard in offline evaluation.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper argues that LLM-driven music recommender systems (MRS) require a fundamental rethinking of evaluation, because LLMs are generative, non-deterministic, prone to hallucination, and trained on opaque data, which breaks the assumptions behind traditional accuracy/IR metrics. It reviews LLM applications to user modeling, item modeling, and natural-language recommendation in music, surveys NLP evaluation methods (reference-based, reference-free, human/LLM-as-judge) and their risks, and proposes a structured set of six success dimensions (G1–G6) and eight risk dimensions, together with setting-specific bundles for in-context learning, retrieval-augmented generation, and chain-/tree-of-thought prompting. The paper is explicitly a position/perspective piece: it provides no empirical validation of the proposed metrics, but it is transparent about open questions, especially the reliability of LLM-as-judge evaluation.

Significance. If the central claim is accepted, the paper is a timely and useful synthesis. It compiles a wide range of MRS-specific challenges and connects them to evaluation practices from NLP, going beyond a simple accuracy-oriented framing. The structured catalog of success and risk dimensions provides a concrete reference point for future work, and the separation of catalog, resolver/knowledge base, and prompting components is a helpful modeling choice. Strengths include the explicit treatment of non-determinism and decoding variability, the recognition that LLM-as-judge must be validated before adoption, and the inclusion of risk dimensions such as popularity/temporal/language bias and sampling-bias amplification. As a position paper, the lack of empirical validation is not fatal, but the informal status of several proposed metrics should be clarified so that readers do not mistake a research agenda for a validated framework.

major comments (3)
  1. [3.2.1] The proposed success metrics for S-RAG and S-COT — EvidenceGrounding@k, Δdoc, TTFS, Day-0 Lift, ΔU_CoT — are introduced with informal hedges ('one might define', 'one might introduce') and no formal definitions, aggregation rules, or validation protocol. Since the paper's constructive contribution is precisely this catalog, the prescriptive force of G1–G6 depends on these metrics being operationalizable. Please either (a) provide precise definitions (pseudo-code or formulas) with a validation plan against human judgment, or (b) explicitly label Section 3.2 as a research agenda and add a paragraph on the validation steps needed. As written, the reader cannot tell which parts of the framework are intended as standards and which as conjectures.
  2. [3.1.3, 3.2.2(4)] The paper correctly states in Section 3.1.3 that LLM-as-judge must be rigorously validated against human judgments before adoption, and Section 3.2.2(4) lists evaluator bias as a risk dimension. However, several of the proposed success/risk measures (groundedness, faithfulness, reasoning faithfulness) are described as LLM-based assessments. This creates a circular dependency if the framework itself relies on the very evaluation approach whose reliability it flags as an open problem. Please state explicitly how the circularity is broken: which metrics should be computed only with human annotation, which can use LLM-as-judge after calibration, and what acceptance criteria (e.g., agreement thresholds) are proposed.
  3. [3.2.1–3.2.2] The framework provides separate success (G1–G6) and risk dimensions but no guidance on joint reporting or trade-offs. For example, a system may achieve high G6 accuracy while exhibiting high popularity bias (risk 2), or high G2 discovery at the cost of latency/token budget in S-RAG. A short subsection on how to report these dimensions together (e.g., a required disclosure format, or a minimum-quality threshold on risk dimensions) would make the framework actionable. Without this, the catalog remains a list of dimensions rather than a coherent evaluation protocol.
minor comments (6)
  1. [Figure 4 (Section 3.1)] The diagram appears to contain a stray browser/export header '20/11/2025, 16:37 nlp-eval Page 1 of 2 https://app.diagrams.net/' and the text is rendered as an image rather than typeset text. This will make the figure hard to read and contains an unintended URL/timestamp artifact. Please replace with a clean, typeset figure.
  2. [References] Several references are incomplete: [20] and [161] use 'et al.' with no author names, and [174] contains the placeholder 'arXiv:2502.xxxxx'. Also, [165] and [166] duplicate the same self-consistency work, and [55] lists 'Haofen Wang' twice. Please clean up the reference list.
  3. [Table in 3.2.1] The table's setting 'S-Prompting — In-Context Learning & Chain-of-Thought' merges ICL and CoT, while the text uses separate S-ICL and S-COT subsections. Align the nomenclature to avoid confusion.
  4. [3.2.3] The phrase 'most straightaway methods of evaluating' should be 'most straightforward methods'.
  5. [G4] The citation list in G4 includes the same reference twice as [151,151]; check and deduplicate.
  6. [Introduction/Abstract] The paper sometimes says 'accuracy-only metrics are questionable' but later retains G6 classical relevance as a core dimension. Make explicit that the argument is for supplementation rather than abandonment of accuracy metrics, to avoid misreading.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a position/survey proposing evaluation dimensions, with no derivation that reduces to its inputs.

full rationale

This is a survey/position paper, not a derivation. Its central claim—that LLM-driven MRS requires rethinking evaluation—is argued from external observations about generative models, known NLP evaluation practice, and cited empirical studies. The proposed success and risk dimensions (G1–G6, S-ICL, S-RAG, S-CoT) are explicitly framed as a catalog, with hedged definitions such as 'one might define EvidenceGrounding@k' and 'one might introduce freshness... metrics such as TTFS and Day-0 Lift.' These are proposed metrics, not fitted parameters presented as predictions, and the paper does not claim to have validated them. The self-citations (e.g., Deldjoo et al. [30] on ChatGPT bias, Sguerra et al. [144] on LLM-generated taste profiles) are used as empirical evidence from peer-reviewed venues, not as a self-referential uniqueness theorem or ansatz; they support the motivation for risk dimensions but are not the sole basis for any derived result. The paper itself flags the key open problem that LLM-as-judge must be validated before adoption ('LLMs should be rigorously validated against human judgments before being adopted as automatic evaluators'), which is an acknowledged limitation rather than a circular step. No equation or definition equates an output to an input by construction, and no fitted value is renamed as a prediction. Therefore, no significant circularity is present.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

This is a review/position paper: it contains no fitted parameters and postulates no new entities. Its central claim rests on domain assumptions about the recommendation task, the availability of knowledge bases, the applicability of NLP evaluation, and the availability of ground-truth references.

assumptions (4)
  • domain assumption The primary evaluation scenario is a ranked list of k resolvable catalog items produced in response to an NL request (possibly conditioned on profile, ICL, RAG, or CoT).
    Section 2.3 defines the task components (catalog, resolver/KB, NL request, profile, in-context examples, RAG docs, reasoning scaffold) and the whole framework in Section 3 is scoped to this scenario.
  • domain assumption External knowledge bases (e.g., MusicBrainz) are available and reliable for entity resolution and grounding.
    Section 2.3 says 'we explicitly separate the catalog... from knowledge bases' and the evaluation relies on KB entity resolution and EntitiesResolvedShare.
  • domain assumption NLP evaluation metrics (reference-based, reference-free, human/LLM-as-judge) are applicable to MRS outputs.
    Section 3.1 adapts NLP evaluation to MRS; the paper itself notes the open question of whether LLM-based evaluators reliably assess music-related text (Section 3.1.3).
  • domain assumption Ground-truth references or human judgments are available or obtainable for evaluation.
    Reference-based evaluation in Section 3.1.1 relies on ground-truth outputs y*; the paper acknowledges human evaluation is costly and often unavailable.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Music Recommendation with Large Language Models: Challenges, Opportunities, and Evaluation." pith.science (2026). https://pith.science/paper/PTKXT6YT

@misc{pith2026251116478,
  author       = {Pith},
  title        = {Pith review of: Music Recommendation with Large Language Models: Challenges, Opportunities, and Evaluation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PTKXT6YT}},
  note         = {Machine review of arXiv:2511.16478}
}
read the original abstract

Music Recommender Systems (MRSs) have long relied on an information retrieval framing, where progress is measured mainly through accuracy on retrieval-oriented subtasks. While effective, this reductionist paradigm struggles to address the deeper question of what makes a good recommendation. Attempts to broaden evaluation, through user studies or fairness analyses, have had limited impact. The emergence of Large Language Models (LLMs) disrupts this framework: LLMs are generative rather than ranking-based, making standard accuracy metrics questionable. They also introduce challenges such as hallucinations, knowledge cutoffs, non-determinism, and opaque training data, rendering traditional train or test protocols difficult to interpret. At the same time, LLMs create new opportunities, enabling natural language (NL) interaction and even allowing models to act as evaluators. This work argues that the shift toward LLM-driven MRSs requires rethinking evaluation. We first review how LLMs reshape user modeling, item modeling, and NL-based recommendation in music. We then examine evaluation practices from NLP, highlighting methodologies and open challenges relevant to MRSs. Finally, we synthesize insights, focusing on how LLM prompting applies to MRSs, to outline a structured set of success and risk dimensions. Our goal is to provide the MRSs community with an updated, pedagogical, and cross-disciplinary perspective on evaluation.

Figures

Figures reproduced from arXiv: 2511.16478 by the authors.

Figure 1
Figure 1. Paper’s overview as a generic diagram presenting music recommendation with LLMs [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Representation of the method to derive NL user preference profiles from consumption data. [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Prompt example with the different task components highlighted. [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: NLP-driven evaluation framework: decision process and metrics overview. [PITH_FULL_IMAGE:figures/full_fig_p019_4.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

196 extracted references · 17 canonical work pages

  1. [67]

    Hui Huang, Xingyuan Bu, Hongli Zhou, Yingqi Qu, Jing Liu, Muyun Yang, Bing Xu, and Tiejun Zhao. 2025. An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4. arXiv:2403.02839 [cs.CL] https://arxiv.org/abs/2403.02839

  2. [1]

    Marah I Abdin, Suriya Gunasekar, Varun Chandrasekaran, Jerry Li, Mert Yuksekgonul, Rahee Ghosh Peshawaria, Ranjita Naik, and Besmira Nushi. 2024. KITAB: Evaluating LLMs on Constraint Satisfaction for Information Retrieval. InThe Twelfth International Conference on Learning Representations (ICLR). https://openreview.net/forum?id=b3kDP3IytM

  3. [2]

    Himan Abdollahpouri, Masoud Mansoury, Robin Burke, and Bamshad Mobasher. 2019. Managing Popularity Bias in Recommender Systems with Personalized Re-ranking. InProceedings of the 29th International Florida Artificial Intelligence Research Society Conference (FLAIRS). 413–418

  4. [3]

    Khetam Al Sharou, Zhenhao Li, and Lucia Specia. 2021. Towards a Better Understanding of Noise in Natural Language Processing. InProceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2021), Ruslan Mitkov and Galia Angelova (Eds.). INCOMA Ltd., Held Online, 53–62. https://aclanthology.org/2021.ranlp-1.7/

  5. [4]

    Geetha Sai Aluri, Siddharth Sharma, Tarun Sharma, and Joaquin Delgado. 2024. Playlist Search Reinvented: LLMs Behind the Curtain. InProceedings of the 18th ACM Conference on Recommender Systems(Bari, Italy)(RecSys ’24). Association for Computing Machinery, New York, NY, USA, 813–815. doi:10.1145/3640457.3688047

  6. [5]

    Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2024. Self-rag: Learning to retrieve, generate, and critique through self-reflection. (2024)

  7. [6]

    Yuyang Bai, Shangbin Feng, Vidhisha Balachandran, Zhaoxuan Tan, Shiqi Lou, Tianxing He, and Yulia Tsvetkov. 2024. KGQuiz: Evaluating the Generalization of Encoded Knowledge in Large Language Models. InProceedings of the ACM Web Conference 2024(Singapore, Singapore)(WWW ’24). Association for Computing Machinery, New York, NY, USA, 2226–2237. doi:10.1145/35...

  8. [8]

    Keqin Bao, Ming Yan, Yang Zhang, Jizhi Zhang, Wenjie Wang, Fuli Feng, and Xiangnan He. 2024. Real-Time Personalization for LLM-based Recommendation with Customized In-Context Learning.arXiv preprint arXiv:2410.23136(2024)

Show all 196 references
  1. [9]

    Christine Bauer and Markus Schedl. 2019. Global and country-specific mainstreaminess measures: Definitions, analysis, and usage for improving personalized music recommendation systems.PLOS ONE14, 6 (June 2019), e0217389. doi:10.1371/journal.pone.0217389 Publisher: Public Libra...

  2. [10]

    Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fernández, Albert Gatt, Esam Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, Andre Martins, Philipp Mondorf, Vera Neplenbroek, Sandro Pezzelle, Barbara Plank, David Schlangen, Alessand...

  3. [11]

    Giovanni Maria Biancofiore, Tommaso Di Noia, Eugenio Di Sciascio, Fedelucio Narducci, and Paolo Pastore. 2022. Aspect based sentiment analysis in music: a case study with spotify. InProceedings of the 37th ACM/SIGAPP Symposium on Applied Computing(Virtual Event)(SAC ’22). Asso...

  4. [12]

    Tony Brooke. 2014. descriptive metadata in the music industry: Why it is broken and how to fix it—part one.Journal of Digital Media Management 2, 3 (2014), 263–282

  5. [13]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...

  6. [14]

    Pedro G Campos, Fernando Díez, and Iván Cantador. 2014. Time-aware recommender systems: a comprehensive survey and analysis of existing evaluation protocols.User Modeling and User-Adapted Interaction24, 1 (2014), 67–119

  7. [15]

    Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al. 2021. Extracting Training Data from Large Language Models. In30th USENIX Security Symposium (USENIX Security 21). ...

  8. [16]

    Brandon James Carone and Pablo Ripollés. 2024. SoundSignature: What Type of Music do you Like?. In2024 IEEE 5th International Symposium on the Internet of Sounds (IS2). IEEE, 1–10

  9. [17]

    Elif Celen, Pol van Rijn, Harin Lee, and Nori Jacoby. 2025. Are Expressions for Music Emotions the Same Across Cultures?. InProceedings of the Annual Meeting of the Cognitive Science Society (CogSci)

  10. [18]

    2009.Music recommendation and discovery in the long tail

    Òscar Celma Herrada et al. 2009.Music recommendation and discovery in the long tail. Universitat Pompeu Fabra

  11. [19]

    Ching-Wei Chen, Paul Lamere, Markus Schedl, and Hamed Zamani. 2018. Recsys challenge 2018: Automatic music playlist continuation. In Proceedings of the 12th ACM Conference on Recommender Systems. 527–528

  12. [20]

    et al. Chen. 2025. MLA-Trust: Benchmarking Trustworthiness of Multimodal Large-scale Agents.arXiv preprint arXiv:2506.01616(2025)

  13. [21]

    Jiawei Chen, Hande Dong, Xiang Wang, Fuli Feng, Meng Wang, and Xiangnan He. 2023. Bias and debias in recommender system: A survey and future directions.ACM Transactions on Information Systems41, 3 (2023), 1–39

  14. [22]

    Yanran Chen and Steffen Eger. 2023. MENLI: Robust Evaluation Metrics from Natural Language Inference.Transactions of the Association for Computational Linguistics11 (2023), 804–825. doi:10.1162/tacl_a_00576

  15. [23]

    Yi Chen, Rui Wang, Haiyun Jiang, Shuming Shi, and Ruifeng Xu. 2023. Exploring the Use of Large Language Models for Reference-Free Text Quality Evaluation: An Empirical Study. InFindings of the Association for Computational Linguistics: IJCNLP-AACL 2023 (Findings), Jong C. Park...

  16. [24]

    Cheng-Han Chiang and Hung-yi Lee. 2023. Can Large Language Models Be an Alternative to Human Evaluations?. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Ed...

  17. [25]

    Keunwoo Choi, Seungheon Doh, and Juhan Nam. 2025. TalkPlayData 2: An Agentic Synthetic Data Pipeline for Multimodal Conversational Music Recommendation.arXiv preprint arXiv:2509.09685(2025)

  18. [26]

    Pierre Colombo, Guillaume Staerman, Chloé Clavel, and Pablo Piantanida. 2021. Automatic Text Evaluation through the Lens of Wasserstein Barycenters. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Marie-Francine Moens, Xuanjing Huang, ...

  19. [27]

    Microsoft Corporation. 2025. GroundednessEvaluator — Azure AI Evaluation SDK for Python. https://learn.microsoft.com/en-us/python/api/azure- ai-evaluation/azure.ai.evaluation.groundnessevaluator?view=azure-python. Accessed: 2025-10-25

  20. [28]

    Paolo Cremonesi and Dietmar Jannach. 2021. Progress in recommender systems research: Crisis? What crisis?AI Magazine42, 3 (2021), 43–54

  21. [29]

    Mathieu Delcluze, Antoine Khoury, Clémence Vast, Valerio Arnaudo, Léa Briand, Walid Bendada, and Thomas Bouabça. 2025. Text2Playlist: Generating Personalized Playlists from Text on Deezer. InThe 47th European Conference on Information Retrieval (ECIR 2025). Manuscript submitte...

  22. [30]

    Yashar Deldjoo. 2024. Understanding biases in chatgpt-based recommender systems: Provider fairness, temporal stability, and recency.ACM Transactions on Recommender Systems(2024)

  23. [31]

    Yashar Deldjoo, Zhankui He, Julian McAuley, Anton Korikov, Scott Sanner, Arnau Ramisa, Rene Vidal, Maheswaran Sathiamoorthy, Atoosa Kasrizadeh, Silvia Milano, et al. 2024. Recommendation with Generative Models.arXiv preprint arXiv:2409.15173(2024)

  24. [32]

    Yashar Deldjoo, Nikhil Mehta, Maheswaran Sathiamoorthy, Shuai Zhang, Pablo Castells, and Julian McAuley. 2025. Toward Holistic Evaluation of Recommender Systems Powered by Generative Models. InProceedings of the 48th International ACM SIGIR Conference on Research and Developme...

  25. [33]

    Yashar Deldjoo, Markus Schedl, and Peter Knees. 2024. Content-driven music recommendation: Evolution, state of the art, and challenges.Computer Science Review51 (2024), 100618. doi:10.1016/j.cosrev.2024.100618

  26. [34]

    Zihao Deng, Yinghao Ma, Yudong Liu, Rongchen Guo, Ge Zhang, Wenhu Chen, Wenhao Huang, and Emmanouil Benetos. 2024. MusiLingo: Bridging Music and Text with Pre-trained Language Models for Music Captioning and Query Response. InFindings of the Association for Computational Lingu...

  27. [35]

    Daniel Deutsch, Rotem Dror, and Dan Roth. 2022. On the Limitations of Reference-Free Evaluations of Generated Text. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for...

  28. [36]

    Karlijn Dinnissen and Christine Bauer. 2022. Fairness in music recommender systems: A stakeholder-centered mini review.Frontiers in big Data5 (2022), 913608

  29. [37]

    Seungheon Doh, Keunwoo Choi, Jongpil Lee, and Juhan Nam. 2023. LP-MusicCaps: LLM-Based Pseudo Music Captioning. InIsmir 2023 Hybrid Conference

  30. [38]

    Seungheon Doh, Junwon Lee, and Juhan Nam. 2021. Music Playlist Title Generation: A Machine-Translation Approach. InProceedings of the 2nd Workshop on NLP for Music and Spoken Audio (NLP4MusA), Sergio Oramas, Elena Epure, Luis Espinosa-Anke, Rosie Jones, Massimo Quadrana, Moham...

  31. [39]

    SeungHeon Doh, Minhee Lee, Dasaem Jeong, and Juhan Nam. 2024. Enriching music descriptions with a finetuned-llm and metadata for text-to-music retrieval. InICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 826–830

  32. [40]

    Elena Epure and Romain Hennequin. 2023. A human subject study of named entity recognition in conversational music recommendation queries. InProceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 1281–1296

  33. [41]

    Epure, Anis Khlif, and Romain Hennequin

    Elena V. Epure, Anis Khlif, and Romain Hennequin. 2019. Leveraging knowledge bases and parallel annotations for music genre translation. In International Society for Music Information Retrieval Conference

  34. [42]

    Elena V Epure, Guillaume Salha, and Romain Hennequin. 2020. Multilingual music genre embeddings for effective cross-lingual music item annotation.arXiv preprint arXiv:2009.07755(2020)

  35. [43]

    Elena V Epure, Guillaume Salha, Manuel Moussallam, and Romain Hennequin. 2020. Modeling the Music Genre Perception across Language-Bound Cultures. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 4765–4779

  36. [44]

    Shahul Es, Jithin James, Luis Espinosa-Anke, and Steven Schockaert. 2023. RAGAS: Automated Evaluation of Retrieval Augmented Generation. arXiv preprint arXiv:2309.15217(2023)

  37. [45]

    Shahul Es, Jithin James, Luis Espinosa Anke, and Steven Schockaert. 2024. RAGAs: Automated Evaluation of Retrieval Augmented Generation. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, Nikol...

  38. [46]

    Andres Ferraro. 2019. Music cold-start and long-tail recommendation: bias in deep representations. InProceedings of the 13th ACM conference on recommender systems. 586–590

  39. [47]

    Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, George Foster, Alon Lavie, and Ondřej Bojar. 2021. Results of the WMT21 Metrics Shared Task: Evaluating Metrics with Expert-based Human Evaluations on TED and News Domain. InProceedings of the Sixth Confere...

  40. [48]

    Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2024. GPTScore: Evaluate as You Desire. InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Kevin Duh...

  41. [49]

    Yumeng Fu, Junjie Wu, Zhongjie Wang, Meishan Zhang, Lili Shan, Yulin Wu, and Bingquan Liu. 2025. LaERC-S: Improving LLM-based Emotion Recognition in Conversation with Speaker Characteristics. InProceedings of the 31st International Conference on Computational Linguistics, Owen...

  42. [50]

    Christian Fuentes, Johan Hagberg, and Hans Kjellberg. 2019. Soundtracking: music listening practices in the digital age.European Journal of Marketing53, 3 (2019), 483–503

  43. [51]

    Escrig, and M

    Nieves Fuentes-Sánchez, Raúl Pastor, Tuomas Eerola, Miguel A. Escrig, and M. Carmen Pastor. 2022. Musical preference but not familiarity influences subjective ratings and psychophysiological correlates of music-induced emotions.Personality and Individual Differences198 (Nov. 2...

  44. [52]

    Giovanni Gabbolini and Derek Bridge. 2021. Generating Interesting Song-to-Song Segues With Dave. InProceedings of the 29th ACM Conference on User Modeling, Adaptation and Personalization(Utrecht, Netherlands)(UMAP ’21). Association for Computing Machinery, New York, NY, USA, 9...

  45. [53]

    Giovanni Gabbolini, Romain Hennequin, and Elena Epure. 2022. Data-Efficient Playlist Captioning With Musical and Linguistic Knowledge. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Ed...

  46. [54]

    Qingqing Gao, Jiuxin Cao, Biwei Cao, Xin Guan, and Bo Liu. 2024. CEPT: A Contrast-Enhanced Prompt-Tuning Framework for Emotion Recognition in Conversation. InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation ...

  47. [55]

    Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yixin Dai, Jiawei Sun, Haofen Wang, and Haofen Wang. 2023. Retrieval- augmented generation for large language models: A survey.arXiv preprint arXiv:2312.109972, 1 (2023)

  48. [56]

    Zhaolin Gao, Joyce Zhou, Yijia Dai, and Thorsten Joachims. 2024. End-to-end Training for Recommendation with Language-based User Profiles. arXiv preprint arXiv:2410.18870(2024)

  49. [57]

    Iker García-Ferrero, Begoña Altuna, Javier Alvez, Itziar Gonzalez-Dios, and German Rigau. 2023. This is not a Dataset: A Large Negation Benchmark to Challenge Large Language Models. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Houda...

  50. [58]

    Josh Gardner, Simon Durand, Daniel Stoller, and Rachel Bittner. 2024. LLARK: a multimodal instruction-following language model for music. In Proceedings of the 41st International Conference on Machine Learning(Vienna, Austria)(ICML’24). JMLR.org, Article 603, 46 pages

  51. [59]

    Tao Ge, Xin Chan, Xiaoyang Wang, Dian Yu, Haitao Mi, and Dong Yu. 2025. Scaling Synthetic Data Creation with 1,000,000,000 Personas. doi:10.48550/arXiv.2406.20094 arXiv:2406.20094 [cs]

  52. [60]

    Yue Guo, Tal August, Gondy Leroy, Trevor Cohen, and Lucy Lu Wang. 2024. APPLS: Evaluating Evaluation Metrics for Plain Language Summarization. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung...

  53. [61]

    Simon Hachmeier and Robert Jäschke. 2025. A Benchmark and Robustness Study of In-Context-Learning with Large Language Models in Music Entity Detection. InProceedings of the 31st International Conference on Computational Linguistics. 9845–9859

  54. [62]

    Anna Hausberger, Hannah Strauss, and Markus Schedl. 2025. ExIM: Exploring Intent of Music Listening for Retrieving User-generated Playlists. In Proceedings of the 2025 ACM SIGIR Conference on Human Information Interaction and Retrieval (CHIIR ’25). Association for Computing Ma...

  55. [63]

    Xiangnan He, Lizi Liao, Hanwang Zhang, Liqiang Nie, Xia Hu, and Tat-Seng Chua. 2017. Neural collaborative filtering. InProceedings of the 26th international conference on world wide web. 173–182

  56. [64]

    Zhankui He, Zhouhang Xie, Rahul Jha, Harald Steck, Dawen Liang, Yesu Feng, Bodhisattwa Prasad Majumder, Nathan Kallus, and Julian McAuley

  57. [65]

    Yupeng Hou, Junjie Zhang, Zihan Lin, Hongyu Lu, Ruobing Xie, Julian McAuley, and Wayne Xin Zhao. 2024. Large language models are zero-shot rankers for recommender systems. InEuropean Conference on Information Retrieval. Springer, 364–381

  58. [66]

    Yifan Hu, Yehuda Koren, and Chris Volinsky. 2008. Collaborative filtering for implicit feedback datasets. In2008 Eighth IEEE international conference on data mining. Ieee, 263–272

  59. [68]

    Jin Huang, Harrie Oosterhuis, Masoud Mansoury, Herke Van Hoof, and Maarten de Rijke. 2024. Going beyond popularity and positivity bias: Correcting for multifactorial bias in recommender systems. InProceedings of the 47th International ACM SIGIR Conference on Research and Devel...

  60. [69]

    Wenyu Huang, Guancheng Zhou, Mirella Lapata, Pavlos Vougiouklis, Sebastien Montella, and Jeff Z. Pan. 2025. Prompting large language models with knowledge graphs for question answering involving long-tail facts.Knowledge-Based Systems324 (Aug. 2025), 113648. doi:10.1016/j.knos...

  61. [70]

    Karim M Ibrahim, Elena V Epure, Geoffroy Peeters, and Gael Richard. 2020. SHOULD WE CONSIDER THE USERS IN CONTEXTUAL MUSIC AUTO-TAGGING MODELS?. In21st International Society for Music Information Retrieval Conference

  62. [71]

    Ari Jacovi and Yoav Goldberg. 2020. Aligning Faithful Explanations: Towards Explanation as Verification.Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP 2020)(2020), 514–525. https://aclanthology.org/2020.emnlp-main.40

  63. [72]

    Aryan Jadon and Avinash Patil. 2024. A comprehensive survey of evaluation techniques for recommendation systems. InInternational Conference on Computation of Artificial Intelligence & Machine Learning. Springer, 281–304

  64. [73]

    Dietmar Jannach, Ahtsham Manzoor, Wanling Cai, and Li Chen. 2021. A survey on conversational recommender systems.ACM Computing Surveys (CSUR)54, 5 (2021), 1–36. Manuscript submitted to ACM Music Recommendation with Large Language Models 31

  65. [74]

    Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation.ACM computing surveys55, 12 (2023), 1–38

  66. [75]

    Marius Kaminskas and Derek Bridge. 2016. Diversity, serendipity, novelty, and coverage: a survey and empirical analysis of beyond-accuracy objectives in recommender systems.ACM Transactions on Interactive Intelligent Systems (TiiS)7, 1 (2016), 1–42

  67. [76]

    Li Kang, Yuhan Zhao, and Li Chen. 2025. Exploring the Potential of LLMs for Serendipity Evaluation in Recommender Systems. InProceedings of the Nineteenth ACM Conference on Recommender Systems. 746–754

  68. [77]

    Wang-Cheng Kang and Julian McAuley. 2018. Self-attentive sequential recommendation. In2018 IEEE international conference on data mining (ICDM). IEEE, 197–206

  69. [78]

    Mohammad Khosravani, Chenyang Huang, and Amine Trabelsi. 2024. Enhancing Argument Summarization: Prioritizing Exhaustiveness in Key Point Generation and Introducing an Automatic Coverage Evaluation Metric. InProceedings of the 2024 Conference of the North American Chapter of t...

  70. [79]

    Bart P Knijnenburg and Martijn C Willemsen. 2015. Evaluating recommender systems with user experiments. InRecommender systems handbook. Springer, 309–352

  71. [80]

    Zhifeng Kong, Arushi Goel, Rohan Badlani, Wei Ping, Rafael Valle, and Bryan Catanzaro. 2024. Audio Flamingo: a novel audio language model with few-shot learning and dialogue abilities. InProceedings of the 41st International Conference on Machine Learning. 25125–25148

  72. [81]

    Ngoc Luyen Le, Marie-Hélène Abel, and Philippe Gouspillou. 2023. A Constraint-based Recommender System via RDF Knowledge Graphs. In2023 26th International Conference on Computer Supported Cooperative Work in Design (CSCWD). IEEE, Rio de Janeiro, Brazil, 849–854. doi:10.1109/ C...

  73. [82]

    Harin Lee, Elif Çelen, Peter Harrison, Manuel Anglada-Tort, Pol van Rijn, Minsu Park, Marc Schönwiesner, and Nori Jacoby. 2025. GlobalMood: A cross-cultural benchmark for music emotion recognition.arXiv preprint arXiv:2505.09539(2025)

  74. [83]

    Oleg Lesota, Gustavo Escobedo, Yashar Deldjoo, Bruce Ferwerda, Simone Kopeinik, Elisabeth Lex, Navid Rekabsaz, and Markus Schedl. 2023. Computational Versus Perceived Popularity Miscalibration in Recommender Systems. InProceedings of the 46th International ACM SIGIR Conference...

  75. [84]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing sy...

  76. [85]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2021. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. arXiv:2005.1140...

  77. [86]

    Ang Li, Jennifer Thom, Praveen Chandar, Christine Hosey, Brian St Thomas, and Jean Garcia-Gathright. 2019. Search mindsets: Understanding focused and non-focused information seeking in music search. InThe World Wide Web Conference. 2971–2977

  78. [87]

    Kongmeng Liew, Vipul Mishra, Yangyang Zhou, Elena V Epure, Romain Hennequin, Shoko Wakamiya, and Eiji Aramaki. 2022. Network Analyses for Cross-Cultural Music Popularity. InIsmir 2022 Hybrid Conference

  79. [88]

    Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. InText summarization branches out. 74–81

  80. [89]

    Siyi Lin, Chongming Gao, Jiawei Chen, Sheng Zhou, Binbin Hu, Yan Feng, Chun Chen, and Can Wang. 2025. How Do Recommendation Models Amplify Popularity Bias? An Analysis from the Spectral Perspective. InProceedings of the Eighteenth ACM International Conference on Web Search and...

  81. [90]

    Meijun Liu, Xiao Hu, and Markus Schedl. 2018. The relation of culture, socio-economics, and friendship to music preferences: A large-scale, cross-country study.PLOS ONE13, 12 (Dec. 2018), e0208186. doi:10.1371/journal.pone.0208186 Publisher: Public Library of Science

  82. [91]

    Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. Lost in the middle: How language models use long contexts.Transactions of the Association for Computational Linguistics12 (2024), 157–173

  83. [92]

    Ruilun Liu and Xiao Hu. 2020. A Multimodal Music Recommendation System with Listeners’ Personality and Physiological Signals. InJCDL ’20: Proceedings of the ACM/IEEE Joint Conference on Digital Libraries in 2020, Virtual Event, China, August 1-5, 2020, Ruhua Huang, Dan Wu, Gar...

  84. [93]

    Chen, and Min-Yen Kan

    Do Xuan Long, Ngoc-Hai Nguyen, Tiviatis Sim, Hieu Dao, Shafiq Joty, Kenji Kawaguchi, Nancy F. Chen, and Min-Yen Kan. 2025. LLMs Are Biased Towards Output Formats! Systematically Evaluating and Mitigating Output Format Bias of LLMs. InProceedings of the 2025 Conference of the N...

  85. [94]

    Lonsdale and Adrian C

    Adam J. Lonsdale and Adrian C. North. 2011. Why do we listen to music? A uses and gratifications analysis.British Journal of Psychology (London, England: 1953)102, 1 (Feb. 2011), 108–134. doi:10.1348/000712610X506831

  86. [95]

    Feng Lu and Nava Tintarev. 2018. A Diversity Adjusting Strategy with Personality for Music Recommendation. InProceedings of the 5th Joint Workshop on Interfaces and Human Decision Making for Recommender Systems, IntRS 2018, co-located with ACM Conference on Recommender Systems...

  87. [96]

    Malte Ludewig and Dietmar Jannach. 2018. Evaluation of session-based recommendation algorithms.User Modeling and User-Adapted Interaction 28, 4 (2018), 331–390

  88. [97]

    Hanjia Lyu, Song Jiang, Hanqing Zeng, Yinglong Xia, Qifan Wang, Si Zhang, Ren Chen, Chris Leung, Jiajie Tang, and Jiebo Luo. 2024. LLM-Rec: Personalized Recommendation via Prompting Large Language Models. InFindings of the Association for Computational Linguistics: NAACL 2024,...

  89. [98]

    Lilian Marey, Bruno Sguerra, and Manuel Moussallam. 2024. Modeling activity-driven music listening with pace. InProceedings of the 2024 Conference on Human Information Interaction and Retrieval. 346–351

  90. [99]

    Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2020. Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation Metrics. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Nata...

  91. [100]

    Kristina Matrosova, Lilian Marey, Guillaume Salha-Galvan, Thomas Louail, Olivier Bodini, and Manuel Moussallam. 2024. Do recommender systems promote local music? a reproducibility study using music streaming data. InProceedings of the 18th ACM Conference on Recommender Systems...

  92. [101]

    McCrae and Oliver P

    Robert R. McCrae and Oliver P. John. 1992. An Introduction to the Five-Factor Model and Its Applications.Journal of Personality60, 2 (1992), 175–215. doi:10.1111/j.1467-6494.1992.tb00970.x arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.1467-6494.1992.tb00970.x

  93. [102]

    Melchiorre, Elena V

    Alessandro B. Melchiorre, Elena V. Epure, Shahed Masoudian, Gustavo Escobedo, Anna Hausberger, Manuel Moussallam, and Markus Schedl. 2025. Just Ask for Music (JAM): Multimodal and Personalized Natural Language Music Recommendation. InProceedings of the 19th ACM Conference on R...

  94. [103]

    Melchiorre and Markus Schedl

    Alessandro B. Melchiorre and Markus Schedl. 2020. Personality Correlates of Music Audio Preferences for Modelling Music Listeners. InProceedings of the 28th ACM Conference on User Modeling, Adaptation and Personalization (UMAP ’20). Association for Computing Machinery, New Yor...

  95. [104]

    MetaBrainz Foundation. [n. d.]. MusicBrainz Database. https://musicbrainz.org/doc/MusicBrainz_Database. Accessed 2025-10-11

  96. [105]

    MetaBrainz Foundation. n.d.. MusicBrainz: the open music encyclopedia. https://musicbrainz.org/. Accessed 2025-10-11

  97. [106]

    BSL METEOR. 2005. an automatic metric for MT evaluation with improved correlation with human judgments. InProceedings of the Acl Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization

  98. [107]

    Daly, Kush R

    Erik Miehling, Michael Desmond, Karthikeyan Natesan Ramamurthy, Elizabeth M. Daly, Kush R. Varshney, Eitan Farchi, Pierre Dognin, Jesus Rios, Djallel Bouneffouf, Miao Liu, and Prasanna Sattigeri. 2025. Evaluating the Prompt Steerability of Large Language Models. InProceedings ...

  99. [108]

    Sho Miyakawa and Takehito Utsuro. 2024. Emotion Classification of Lyrics through Summarization by Large Language Models. In2024 IEEE International Conference on Big Data (BigData). 2999–3006. doi:10.1109/BigData62323.2024.10825406

  100. [109]

    2015.Fundamentals of music processing: Audio, analysis, algorithms, applications

    Meinard Müller. 2015.Fundamentals of music processing: Audio, analysis, algorithms, applications. Vol. 5. Springer

  101. [110]

    Sheshera Mysore, Mahmood Jasim, Andrew McCallum, and Hamed Zamani. 2023. Editable user profiles for controllable text recommendations. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval. 993–1003

  102. [111]

    Sergio Oramas, Andres Ferraro, Alvaro Sarasua, and Fabien Gouyon. 2024. Talking to your recs: Multimodal embeddings for recommendation and retrieval. InMuRS 2024: 2nd Music Recommender Systems Workshop

  103. [112]

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. InProceedings of the 40th annual meeting of the Association for Computational Linguistics. 311–318

  104. [113]

    Emilia Parada-Cabaleiro, Anton Batliner, Marcel Zentner, and Markus Schedl. 2023. Exploring emotions in Bach chorales: a multi-modal perceptual and data-driven study.Royal Society Open Science10, 12 (Dec. 2023), 230574. doi:10.1098/rsos.230574 Publisher: Royal Society

  105. [114]

    Debjit Paul, Robert West, Antoine Bosselut, and Boi Faltings. 2024. Making reasoning matter: Measuring and improving faithfulness of chain-of- thought reasoning.arXiv preprint arXiv:2402.13950(2024)

  106. [115]

    Geoffroy Peeters et al. 2004. A large set of audio features for sound description (similarity and classification) in the CUIDADO project.CUIDADO Ist Project Report54, 0 (2004), 1–25

  107. [116]

    Andreas Peintner, Marta Moscati, Yu Kinoshita, Richard Vogl, Peter Knees, Markus Schedl, Hannah Strauss, Marcel Zentner, and Eva Zangerle

  108. [117]

    Maxime Peyrard. 2019. Studying Summarization Evaluation Metrics in the Appropriate Scoring Range. InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Anna Korhonen, David Traum, and Lluís Màrquez (Eds.). Association for Computational Ling...

  109. [118]

    Pantelis Pipergias Analytis and Philipp Hager. 2023. Collaborative filtering algorithms are prone to mainstream-taste bias. InProceedings of the 17th ACM Conference on Recommender Systems. 750–756

  110. [119]

    Liam Pond, Sichen Meng, Linnea Kirby, Simon Ngassam, Sebastien Chow, Dylan Hillerbrand, and Ichiro Fujinaga. 2025. SESEMMI for LinkedMusic: Democratizing Access to Musical Archives via Large Language Models. In1st Workshop on Large Language Models for Music{\&} Audio (LLM4MA)....

  111. [120]

    Massimo Quadrana, Paolo Cremonesi, and Dietmar Jannach. 2018. Sequence-Aware Recommender Systems. arXiv:1802.08452 [cs.IR] https: //arxiv.org/abs/1802.08452

  112. [121]

    Rashin Rahnamoun and Mehrnoush Shamsfard. 2025. Multi-Layered Evaluation Using a Fusion of Metrics and LLMs as Judges in Open-Domain Question Answering. InProceedings of the 31st International Conference on Computational Linguistics, Owen Rambow, Leo Wanner, Marianna Apidianak...

  113. [122]

    Abdallah, Mark B

    Yves Raimond, Samer A. Abdallah, Mark B. Sandler, and Frederick Giasson. 2007. The Music Ontology. InProceedings of the 8th International Conference on Music Information Retrieval, ISMIR 2007, Vienna, Austria, September 23-27, 2007, Simon Dixon, David Bainbridge, and Rainer Ty...

  114. [123]

    Thomas, Chandr Dhanush H, and Arunima C

    Rajeev Rajan, Joshua Antony, Riya Ann Joseph, Jijohn M. Thomas, Chandr Dhanush H, and Arunima C. V. 2021. Audio-Mood Classification Using Acoustic-Textual Feature Fusion. In2021 Fourth International Conference on Microelectronics, Signals & Systems (ICMSS). 1–6. doi:10.1109/ I...

  115. [124]

    Jerome Ramos, Hossein A Rahmani, Xi Wang, Xiao Fu, and Aldo Lipani. 2024. Transparent and Scrutable Recommendations Using Natural Language User Profiles. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 13971–13984

  116. [125]

    Brian Regan, Desislava Hristova, and Mariano Beguerisse-Díaz. 2023. Which Witch? Artist name disambiguation and catalog curation using audio and metadata. https://research.atspotify.com/2023/11/which-witch-artist-name-disambiguation-and-catalog-curation-using-audio-and-metadat...

  117. [126]

    Lise Regnier and Geoffroy Peeters. 2009. Singing voice detection in music tracks using direct voice vibrato detection. In2009 IEEE international conference on acoustics, speech and signal processing. IEEE, 1685–1688

  118. [127]

    Ricardo Rei, Ana C Farinha, Chrysoula Zerva, Daan van Stigt, Craig Stewart, Pedro Ramos, Taisiya Glushkova, André F. T. Martins, and Alon Lavie. 2021. Are References Really Needed? Unbabel-IST 2021 Submission for the Metrics Shared Task. InProceedings of the Sixth Conference o...

  119. [128]

    Peter J Rentfrow. 2012. The role of music in everyday life: Current directions in the social psychology of music.Social and personality psychology compass6, 5 (2012), 402–416

  120. [129]

    Rentfrow and Samuel D

    Peter J. Rentfrow and Samuel D. Gosling. 2007. The content and validity of music-genre stereotypes among college students.Psychology of Music 35, 2 (2007), 306–326. doi:10.1177/0305735607070382

  121. [130]

    Renata L Rosa, Demsteneso Z Rodriguez, and Graça Bressan. 2015. Music recommendation system based on user’s sentiments extracted from social networks.IEEE Transactions on Consumer Electronics61, 3 (2015), 359–367

  122. [131]

    Jia-Jia Ruan, Xi-Xu He, Min Zhang, and Yuan Gao. 2023. Entity generation algorithm based on reference expansion.Journal of Electronic Science and Technology21, 3 (2023), 100218

  123. [132]

    Milad Sabouri, Masoud Mansoury, Kun Lin, and Bamshad Mobasher. 2025. Towards Explainable Temporal User Profiling with LLMs. InAdjunct Proceedings of the 33rd ACM Conference on User Modeling, Adaptation and Personalization. 219–227

  124. [133]

    Sai, Tanay Dixit, Dev Yashpal Sheth, Sreyas Mohan, and Mitesh M

    Ananya B. Sai, Tanay Dixit, Dev Yashpal Sheth, Sreyas Mohan, and Mitesh M. Khapra. 2021. Perturbation CheckLists for Evaluating NLG Evaluation Metrics. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Marie-Francine Moens, Xuanjing Huan...

  125. [134]

    Oscar Sainz, Iker García-Ferrero, Rodrigo Agerri, Oier Lopez de Lacalle, German Rigau, and Eneko Agirre. 2024. GoLLIE: Annotation Guidelines improve Zero-Shot Information-Extraction. InThe Twelfth International Conference on Learning Representations. https://openreview.net/for...

  126. [135]

    Krishna Sayana, Raghavendra Vasudeva, Yuri Vasilevski, Kun Su, Liam Hebert, James Pine, Hubert Pham, Ambarish Jash, and Sukhdeep Sodhi

  127. [136]

    Markus Schedl, Emilia Gómez, Julián Urbano, et al. 2014. Music information retrieval: Recent developments and applications.Foundations and Trends®in Information Retrieval8, 2-3 (2014), 127–261

  128. [137]

    Trent, Marko Tkalčič, Hamid Eghbal-Zadeh, and Agustín Martorell

    Markus Schedl, Emilia Gómez, Erika S. Trent, Marko Tkalčič, Hamid Eghbal-Zadeh, and Agustín Martorell. 2018. On the Interrelation Between Listener Characteristics and the Perception of Emotions in Classical Orchestra Music.IEEE Transactions on Affective Computing9, 4 (Oct. 201...

  129. [138]

    InCompanion Proceedings of the ACM on Web Conference 2025(Sydney NSW, Australia)(WWW ’25)

    Beyond Retrieval: Generating Narratives in Conversational Recommender Systems. InCompanion Proceedings of the ACM on Web Conference 2025(Sydney NSW, Australia)(WWW ’25). Association for Computing Machinery, New York, NY, USA, 2411–2420. doi:10.1145/3701716.3717531

  130. [139]

    Markus Schedl, Hamed Zamani, Ching-Wei Chen, Yashar Deldjoo, and Mehdi Elahi. 2018. Current challenges and visions in music recommender systems research.International Journal of Multimedia Information Retrieval7, 2 (2018), 95–116

  131. [140]

    Thomas Schäfer, Peter Sedlmeier, Christine Städtler, and David Huron. 2013. The psychological functions of music listening.Frontiers in Psychology 4 (Aug. 2013). doi:10.3389/fpsyg.2013.00511 Publisher: Frontiers. Manuscript submitted to ACM 34 Epure et al

  132. [141]

    2021.Music Recommendation Systems: Techniques, Use Cases, and Challenges

    Markus Schedl, Peter Knees, Brian McFee, and Dmitry Bogdanov. 2021.Music Recommendation Systems: Techniques, Use Cases, and Challenges. Springer US, 927–971. doi:10.1007/978-1-0716-2197-4_24

  133. [142]

    Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020. BLEURT: Learning Robust Metrics for Text Generation. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Associa...

  134. [143]

    Bruno Sguerra, Marion Baranes, Romain Hennequin, and Manuel Moussallam. 2022. Navigational, informational or punk-rock? An exploration of search intent in the musical domain. InProceedings of the 30th ACM Conference on User Modeling, Adaptation and Personalization. 202–211

  135. [144]

    Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, and Patrick Gallinari. 2021. QuestEval: Summarization Asks for Fact-based Evaluation. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, M...

  136. [145]

    Bruno Sguerra, Viet-Anh Tran, and Romain Hennequin. 2022. Discovery dynamics: Leveraging repeated exposure for user and music characterization. InProceedings of the 16th ACM Conference on Recommender Systems. 556–561

  137. [146]

    Bruno Sguerra, Viet-Anh Tran, and Romain Hennequin. 2023. Ex2Vec: Characterizing users and items from the mere exposure effect. InProceedings of the 17th ACM Conference on Recommender Systems. 971–977

  138. [147]

    Bruno Sguerra, Elena V Epure, Harin Lee, and Manuel Moussallam. 2025. Biases in LLM-Generated Musical Taste Profiles for Recommendation. In Proceedings of the Nineteenth ACM Conference on Recommender Systems. 527–532

  139. [148]

    Lingfeng Shen, Lemao Liu, Haiyun Jiang, and Shuming Shi. 2022. On the Evaluation Metrics for Paraphrase Generation. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for...

  140. [149]

    Tiancheng Shen, Jia Jia, Yan Li, Yihui Ma, Yaohua Bu, Hanjie Wang, Bo Chen, Tat-Seng Chua, and Wendy Hall. 2020. PEIA: Personality and Emotion Integrated Attentive Model for Music Recommendation on Social Media Platforms. InThe Thirty-Fourth AAAI Conference on Artificial Intel...

  141. [150]

    Bruno Sguerra, Viet-Anh Tran, Romain Hennequin, and Manuel Moussallam. 2025. Uncertainty in Repeated Implicit Feedback as a Measure of Reliability. InProceedings of the 33rd ACM Conference on User Modeling, Adaptation and Personalization. 234–242

  142. [151]

    Harald Steck. 2018. Calibrated recommendations. InProceedings of the 12th ACM conference on recommender systems. 154–162

  143. [152]

    Lei Sun, Jinming Zhao, and Qin Jin. 2024. Revealing Personality Traits: A New Benchmark Dataset for Explainable Personality Recognition on Dialogues. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Y...

  144. [153]

    Yifan Song, Guoyin Wang, Sujian Li, and Bill Yuchen Lin. 2025. The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics:...

  145. [154]

    Noah Tekle, Alline Ayala, Jonathan Haile, Abdulla Alshabanah, Corey Baker, and Murali Annavaram. 2024. Music Recommendation through LLM Song Summary. InThe 1st Workshop on Risks, Opportunities, and Evaluation of Generative Models in Recommender Systems (ROEGEN@RECSYS’24)

  146. [155]

    Brian Thompson and Matt Post. 2020. Automatic Machine Translation Evaluation in Many Languages via Zero-Shot Paraphrasing. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds....

  147. [156]

    Tianxiang Sun, Junliang He, Xipeng Qiu, and Xuanjing Huang. 2022. BERTScore is Unfair: On Social Bias in Language Model-Based Metrics for Text Generation. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Yoav Goldberg, Zornitsa Kozareva...

  148. [157]

    Alicia Tsai, Adam Kraft, Long Jin, Chenwei Cai, Anahita Hosseini, Taibai Xu, Zemin Zhang, Lichan Hong, Ed Chi, and Xinyang Yi. 2024. Leveraging LLM Reasoning Enhances Personalized Recommender Systems. InFindings of the Association for Computational Linguistics ACL 2024. 13176–13188

  149. [158]

    Robin Ungruh, Karlijn Dinnissen, Anja Volk, Maria Soledad Pera, and Hanna Hauptmann. 2024. Putting Popularity Bias Mitigation to the Test: A User-Centric Evaluation in Music Recommenders. InProceedings of the 18th ACM Conference on Recommender Systems. 169–178

  150. [159]

    Thinh Hung Truong, Timothy Baldwin, Karin Verspoor, and Trevor Cohn. 2023. Language models are not naysayers: an analysis of language models on negation benchmarks. InProceedings of the 12th Joint Conference on Lexical and Computational Semantics (*SEM 2023), Alexis Palmer and...

  151. [160]

    Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. 2015. Cider: Consensus-based image description evaluation. InProceedings of the IEEE conference on computer vision and pattern recognition. 4566–4575

  152. [161]

    et al. Wang. 2025. Diagnostic-Guided Dynamic Profile Optimization for LLM-based User Simulation.arXiv preprint arXiv:2508.12645(2025)

  153. [162]

    Saúl Vargas and Pablo Castells. 2011. Rank and relevance in novelty and diversity metrics for recommender systems. InProceedings of the Fifth ACM Conference on Recommender Systems. ACM, Chicago, IL, USA, 109–116. doi:10.1145/2043932.2043955

  154. [163]

    Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, and Zhifang Sui. 2024. Large Language Models are not Fair Evaluators. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Vo...

  155. [164]

    Shijia Wang, Tianpei Ouyang, Yunfan Zhou, Qiang Xiao, Yintao Ren, Yifei Pan, Fangjian Li, and Chuanjiang Luo. 2025. Enhanced Emotion-aware Music Recommendation via Large Language Models. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2...

  156. [165]

    Jiarong Wang, Jiaji Wu, Mingzhou Tan, and Lingxuan Zhu. 2025. Emotion-Aware Conversational Music Recommendation With Multiagent System. IEEE Transactions on Computational Social Systems(2025), 1–14. doi:10.1109/TCSS.2025.3599008

  157. [166]

    Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models.arXiv preprint arXiv:2203.11171(2022)

  158. [167]

    Amy Beth Warriner, Victor Kuperman, and Marc Brysbaert. 2013. Norms of valence, arousal, and dominance for 13,915 English lemmas.Behavior research methods45, 4 (2013), 1191–1207

  159. [169]

    Benno Weck, Ilaria Manco, Emmanouil Benetos, Elio Quinton, György Fazekas, and Dmitry Bogdanov. 2024. MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models. InProceedings of the 25th International Society for Music Information Retrieval Conference (ISMIR)

  160. [171]

    William Webber, Alistair Moffat, and Justin Zobel. 2010. A Similarity Measure for Indefinite Rankings. InACM Transactions on Information Systems, Vol. 28. ACM, 1–38. doi:10.1145/1852102.1852106

  161. [172]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Tobias Bosma, Brian Ichter, Fei Xia, Ed Cole, Alejandro Ehinger, John Luu, and Quoc V. Le. 2022. Chain of Thought Prompting Elicits Reasoning in Large Language Models. InAdvances in Neural Information Processing Systems (NeurIPS 2022). ...

  162. [173]

    Likang Wu, Zhi Zheng, Zhaopeng Qiu, Hao Wang, Hongchao Gu, Tingjia Shen, Chuan Qin, Chen Zhu, Hengshu Zhu, Qi Liu, et al. 2024. A survey on large language models for recommendation.World Wide Web27, 5 (2024), 60

  163. [174]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems35 (2022), 24824–24837

  164. [175]

    Meng Yang, Jon McCormack, Maria Teresa Llano, and Wanchao Su. 2025. Exploring the Feasibility of LLMs for Automated Music Emotion Annotation. InIsmir 2025 Hybrid Conference

  165. [176]

    Griffiths, Yuan Cao, and Karthik Narasimhan

    Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. arXiv:2305.10601 [cs.CL] https://arxiv.org/abs/2305.10601

  166. [177]

    Wu, Yixuan Wang, Nathan Tsoi, and Julian Urbano

    Steven H. Wu, Yixuan Wang, Nathan Tsoi, and Julian Urbano. 2025. CLaMP 3: Universal Music Information Retrieval Across Multiple Languages and Modalities.arXiv preprint arXiv:2502.xxxxx(2025)

  167. [178]

    Yakun Yu, Shi-ang Qi, Baochun Li, and Di Niu. 2024. PepRec: Progressive Enhancement of Prompting for Recommendation. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association ...

  168. [179]

    Ruibin Yuan, Hanfeng Lin, Yi Wang, Zeyue Tian, Shangda Wu, Tianhao Shen, Ge Zhang, Yuhang Wu, Cong Liu, Ziya Zhou, Liumeng Xue, Ziyang Ma, Qin Liu, Tianyu Zheng, Yizhi Li, Yinghao Ma, Yiming Liang, Xiaowei Chi, Ruibo Liu, Zili Wang, Chenghua Lin, Qifeng Liu, Tao Jiang, Wenhao ...

  169. [180]

    Chao Yu, Qixin Tan, Hong Lu, Jiaxuan Gao, Xinting Yang, Yu Wang, Yi Wu, and Eugene Vinitsky. 2025. ICPL: Few-shot In-context Preference Learning via LLMs. arXiv:2410.17233 [cs.AI] https://arxiv.org/abs/2410.17233

  170. [181]

    Sojeong Yun and Youn-kyung Lim. 2025. User Experience with LLM-powered Conversational Recommendation Systems: A Case of Music Recommendation. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–15

  171. [182]

    Eva Zangerle, Martin Pichl, and Markus Schedl. 2020. User Models for Culture-Aware Music Recommendation: Fusing Acoustic and Cultural Cues. Transactions of the International Society for Music Information Retrieval3, 1 (March 2020). doi:10.5334/tismir.37

  172. [183]

    Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021. BARTSCORE: evaluating generated text as text generation. InProceedings of the 35th International Conference on Neural Information Processing Systems (NIPS ’21). Curran Associates Inc., Red Hook, NY, USA, Article 2088, 15 pages

  173. [184]

    Baiqiao Zhang, Zhifeng Liao, Xiangxian Li, Chao Zhou, Juan Liu, Xiaojuan Ma, and Yulong Bian. 2025. Rethinking Personality Assessment from Human-Agent Dialogues: Fewer Rounds May Be Better Than More. InFindings of the Association for Computational Linguistics: EMNLP 2025, Chri...

  174. [185]

    Hanlin Zhang, YiFan Zhang, Yaodong Yu, Dhruv Madeka, Dean Foster, Eric Xing, Himabindu Lakkaraju, and Sham Kakade. 2024. A Study on the Calibration of In-context Learning. InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational L...

  175. [186]

    Yuheng Zha, Yichi Yang, Ruichen Li, and Zhiting Hu. 2023. AlignScore: Evaluating Factual Consistency with A Unified Alignment Function. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Anna Rogers, Jordan Boyd-...

  176. [187]

    Tong Zhang. 2025. AdaptRec: A Self-Adaptive Framework for Sequential Recommendations with Large Language Models. arXiv:2504.08786 [cs.IR] https://arxiv.org/abs/2504.08786

  177. [188]

    Weinberger, and Yoav Artzi

    Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. BERTScore: Evaluating Text Generation with BERT. In International Conference on Learning Representations. https://openreview.net/forum?id=SkeHuCVFDr

  178. [189]

    Ruoyu Zhang, Yanzeng Li, Yongliang Ma, Ming Zhou, and Lei Zou. 2023. LLMaAA: Making Large Language Models as Active Annotators. InFindings of the Association for Computational Linguistics: EMNLP 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computatio...

  179. [190]

    Wei Zhao, Michael Strube, and Steffen Eger. 2023. DiscoScore: Evaluating Text Generation with BERT and Discourse Coherence. InProceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, Andreas Vlachos and Isabelle Augenstein (E...

  180. [191]

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P Xing, et al

  181. [192]

    Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M Meyer, and Steffen Eger. 2019. MoverScore: Text Generation Evaluating with Contextualized Embeddings and Earth Mover Distance. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and ...

  182. [193]

    Joyce Zhou, Yijia Dai, and Thorsten Joachims. 2024. Language-Based User Profiles for Recommendation.arXiv preprint arXiv:2402.15623(2024)

  183. [194]

    Jingming Zhuo, Songyang Zhang, Xinyu Fang, Haodong Duan, Dahua Lin, and Kai Chen. 2024. ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs. InFindings of the Association for Computational Linguistics: EMNLP 2024. Association for Computational Linguistics, 1950–1966

  184. [195]

    InAdvances in Neural Information Processing Systems, Vol

    Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. InAdvances in Neural Information Processing Systems, Vol. 36. 46595–46623

  185. [196]

    Han Zhou, Xingchen Wan, Lev Proleev, Diana Mincu, Jilin Chen, Katherine A Heller, and Subhrajit Roy. 2024. Batch Calibration: Rethinking Calibration for In-Context Learning and Prompt Engineering. InThe Twelfth International Conference on Learning Representations. https: //ope...

  186. [199]

    Le Zhuo, Ruibin Yuan, Jiahao Pan, Yinghao Ma, Yizhi Li, Ge Zhang, Si Liu, Roger Dannenberg, Jie Fu, Chenghua Lin, et al. 2023. Lyricwhiz: Robust multilingual zero-shot lyrics transcription by whispering to chatgpt.arXiv preprint arXiv:2306.17103(2023). Received November 2025 M...

  187. [2023]

    InProceedings of the 32nd ACM international conference on information and knowledge management

    Large language models as zero-shot conversational recommenders. InProceedings of the 32nd ACM international conference on information and knowledge management. 720–730

  188. [2025]

    doi:10.5334/tismir.235

    Nuanced Music Emotion Recognition via a Semi-Supervised Multi-Relational Graph Neural Network.Transactions of the International Society for Music Information Retrieval8, 1 (June 2025). doi:10.5334/tismir.235

Pith tools

Reviewed August 3, 2026 · model on record in the stance chip above.