REVIEW 3 major objections 6 minor 196 references
Music Recommendation with Large Language Models: Challenges, Opportunities, and Evaluation
T0 review · 3 major / 6 minor · reviewed 2026-08-03 · deepseek-v4-flash
Pith's one-line read LLM-based music recommenders cannot be judged by accuracy alone; evaluation must move to groundedness, discovery, and risk diagnostics.
desk verdict A credible, well-organized position paper that gives the MRS community a useful vocabulary for evaluating LLM-based recommenders, but its proposed new metrics are mostly sketches awaiting validation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the paper's proposed evaluation framework: a decision tree for selecting NLP-derived metrics, plus a catalog of six success dimensions (G1–G6) and eight risk diagnostics. For the natural-language recommendation scenario, the task is defined as producing a ranked list of k resolvable catalog items from a prompt that may include a user request, a natural-language profile, in-context examples, retrieved documents, or a reasoning scaffold. The framework's work is to separate faithful grounding from fluent text, and to make hallucination, bias, and instability measurable rather than hidden inside a single accuracy number.
What would settle it
A large-scale user study would settle it: collect natural-language music requests, have an LLM produce recommendation lists, then compare human judgments of quality against both standard accuracy metrics and the paper's groundedness and discovery metrics; if accuracy correlates as strongly with human preference as the proposed metrics do, the claim that accuracy metrics are inadequate collapses.
Extended reading notes
Core claim
The paper's central claim is that evaluating an LLM-based music recommender as if it were a retrieval system—scoring precision and recall against consumption logs—no longer measures whether recommendations are good. LLMs are generative, non-deterministic, and prone to hallucination, so the field must rethink evaluation from the ground up. The paper establishes that success should be assessed along six dimensions: query adherence and groundedness, discovery quality, personalization gain, profile fidelity and controllability, cultural and linguistic coverage, and classical relevance. It then adds a complementary set of risk diagnostics covering hallucination, popularity/temporal/language bias,
Load-bearing premise
The framework's prescriptions depend on the assumption that the proposed evaluation measures can be reliably computed and validated—especially LLM-based judges, whose reliability for music content is itself an open question the paper flags.
Editorial extensions
If this is right
- If the paper is right, accuracy-only benchmarks become insufficient for LLM recommenders; studies would need to report groundedness, discovery, personalization gain, and coverage alongside classical relevance.
- Entity resolution against a catalog or knowledge base should become a standard evaluation step, so hallucinated tracks and misattributed metadata are penalized explicitly.
- LLM-as-judge methods should not be adopted in music without first being validated against human judgments, because position, verbosity, and self-enhancement biases can distort scores.
- Each prompting setting needs specialized tests: shot calibration and prompt sensitivity for in-context learning, evidence grounding and document-swap tests for retrieval-augmented generation, and reasoning faithfulness plus self-consistency for chain-of-thought prompting.
- Because black-box LLMs may have been exposed to public benchmark datasets during training, offline accuracy results are hard to interpret, and evaluation should be reported with contamination awareness.
Reading between the lines
- Beyond the paper, this evaluation framework likely generalizes to other generative recommender domains such as movies, books, and e-commerce, since the core problem—ranking-based metrics failing on generative outputs—is not music-specific.
- The risk diagnostics could be operationalized into a standardized 'hallucination rate' and 'bias report card' for deployed music assistants, enabling direct comparison across systems in a way that accuracy metrics do not.
- The emphasis on profile fidelity suggests a testable interactive protocol: let users edit their natural-language profiles, then measure how quickly and consistently the recommender's outputs adapt; such a protocol is not yet standard in offline evaluation.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper argues that LLM-driven music recommender systems (MRS) require a fundamental rethinking of evaluation, because LLMs are generative, non-deterministic, prone to hallucination, and trained on opaque data, which breaks the assumptions behind traditional accuracy/IR metrics. It reviews LLM applications to user modeling, item modeling, and natural-language recommendation in music, surveys NLP evaluation methods (reference-based, reference-free, human/LLM-as-judge) and their risks, and proposes a structured set of six success dimensions (G1–G6) and eight risk dimensions, together with setting-specific bundles for in-context learning, retrieval-augmented generation, and chain-/tree-of-thought prompting. The paper is explicitly a position/perspective piece: it provides no empirical validation of the proposed metrics, but it is transparent about open questions, especially the reliability of LLM-as-judge evaluation.
Significance. If the central claim is accepted, the paper is a timely and useful synthesis. It compiles a wide range of MRS-specific challenges and connects them to evaluation practices from NLP, going beyond a simple accuracy-oriented framing. The structured catalog of success and risk dimensions provides a concrete reference point for future work, and the separation of catalog, resolver/knowledge base, and prompting components is a helpful modeling choice. Strengths include the explicit treatment of non-determinism and decoding variability, the recognition that LLM-as-judge must be validated before adoption, and the inclusion of risk dimensions such as popularity/temporal/language bias and sampling-bias amplification. As a position paper, the lack of empirical validation is not fatal, but the informal status of several proposed metrics should be clarified so that readers do not mistake a research agenda for a validated framework.
major comments (3)
- [3.2.1] The proposed success metrics for S-RAG and S-COT — EvidenceGrounding@k, Δdoc, TTFS, Day-0 Lift, ΔU_CoT — are introduced with informal hedges ('one might define', 'one might introduce') and no formal definitions, aggregation rules, or validation protocol. Since the paper's constructive contribution is precisely this catalog, the prescriptive force of G1–G6 depends on these metrics being operationalizable. Please either (a) provide precise definitions (pseudo-code or formulas) with a validation plan against human judgment, or (b) explicitly label Section 3.2 as a research agenda and add a paragraph on the validation steps needed. As written, the reader cannot tell which parts of the framework are intended as standards and which as conjectures.
- [3.1.3, 3.2.2(4)] The paper correctly states in Section 3.1.3 that LLM-as-judge must be rigorously validated against human judgments before adoption, and Section 3.2.2(4) lists evaluator bias as a risk dimension. However, several of the proposed success/risk measures (groundedness, faithfulness, reasoning faithfulness) are described as LLM-based assessments. This creates a circular dependency if the framework itself relies on the very evaluation approach whose reliability it flags as an open problem. Please state explicitly how the circularity is broken: which metrics should be computed only with human annotation, which can use LLM-as-judge after calibration, and what acceptance criteria (e.g., agreement thresholds) are proposed.
- [3.2.1–3.2.2] The framework provides separate success (G1–G6) and risk dimensions but no guidance on joint reporting or trade-offs. For example, a system may achieve high G6 accuracy while exhibiting high popularity bias (risk 2), or high G2 discovery at the cost of latency/token budget in S-RAG. A short subsection on how to report these dimensions together (e.g., a required disclosure format, or a minimum-quality threshold on risk dimensions) would make the framework actionable. Without this, the catalog remains a list of dimensions rather than a coherent evaluation protocol.
minor comments (6)
- [Figure 4 (Section 3.1)] The diagram appears to contain a stray browser/export header '20/11/2025, 16:37 nlp-eval Page 1 of 2 https://app.diagrams.net/' and the text is rendered as an image rather than typeset text. This will make the figure hard to read and contains an unintended URL/timestamp artifact. Please replace with a clean, typeset figure.
- [References] Several references are incomplete: [20] and [161] use 'et al.' with no author names, and [174] contains the placeholder 'arXiv:2502.xxxxx'. Also, [165] and [166] duplicate the same self-consistency work, and [55] lists 'Haofen Wang' twice. Please clean up the reference list.
- [Table in 3.2.1] The table's setting 'S-Prompting — In-Context Learning & Chain-of-Thought' merges ICL and CoT, while the text uses separate S-ICL and S-COT subsections. Align the nomenclature to avoid confusion.
- [3.2.3] The phrase 'most straightaway methods of evaluating' should be 'most straightforward methods'.
- [G4] The citation list in G4 includes the same reference twice as [151,151]; check and deduplicate.
- [Introduction/Abstract] The paper sometimes says 'accuracy-only metrics are questionable' but later retains G6 classical relevance as a core dimension. Make explicit that the argument is for supplementation rather than abandonment of accuracy metrics, to avoid misreading.
Circularity Check
No significant circularity: the paper is a position/survey proposing evaluation dimensions, with no derivation that reduces to its inputs.
full rationale
This is a survey/position paper, not a derivation. Its central claim—that LLM-driven MRS requires rethinking evaluation—is argued from external observations about generative models, known NLP evaluation practice, and cited empirical studies. The proposed success and risk dimensions (G1–G6, S-ICL, S-RAG, S-CoT) are explicitly framed as a catalog, with hedged definitions such as 'one might define EvidenceGrounding@k' and 'one might introduce freshness... metrics such as TTFS and Day-0 Lift.' These are proposed metrics, not fitted parameters presented as predictions, and the paper does not claim to have validated them. The self-citations (e.g., Deldjoo et al. [30] on ChatGPT bias, Sguerra et al. [144] on LLM-generated taste profiles) are used as empirical evidence from peer-reviewed venues, not as a self-referential uniqueness theorem or ansatz; they support the motivation for risk dimensions but are not the sole basis for any derived result. The paper itself flags the key open problem that LLM-as-judge must be validated before adoption ('LLMs should be rigorously validated against human judgments before being adopted as automatic evaluators'), which is an acknowledged limitation rather than a circular step. No equation or definition equates an output to an input by construction, and no fitted value is renamed as a prediction. Therefore, no significant circularity is present.
Assumptions & free parameters
assumptions (4)
- domain assumption The primary evaluation scenario is a ranked list of k resolvable catalog items produced in response to an NL request (possibly conditioned on profile, ICL, RAG, or CoT).
- domain assumption External knowledge bases (e.g., MusicBrainz) are available and reliable for entity resolution and grounding.
- domain assumption NLP evaluation metrics (reference-based, reference-free, human/LLM-as-judge) are applicable to MRS outputs.
- domain assumption Ground-truth references or human judgments are available or obtainable for evaluation.
Cite this review
Pith. "Pith review of Music Recommendation with Large Language Models: Challenges, Opportunities, and Evaluation." pith.science (2026). https://pith.science/paper/PTKXT6YT
@misc{pith2026251116478,
author = {Pith},
title = {Pith review of: Music Recommendation with Large Language Models: Challenges, Opportunities, and Evaluation},
year = {2026},
howpublished = {\url{https://pith.science/paper/PTKXT6YT}},
note = {Machine review of arXiv:2511.16478}
}
read the original abstract
Music Recommender Systems (MRSs) have long relied on an information retrieval framing, where progress is measured mainly through accuracy on retrieval-oriented subtasks. While effective, this reductionist paradigm struggles to address the deeper question of what makes a good recommendation. Attempts to broaden evaluation, through user studies or fairness analyses, have had limited impact. The emergence of Large Language Models (LLMs) disrupts this framework: LLMs are generative rather than ranking-based, making standard accuracy metrics questionable. They also introduce challenges such as hallucinations, knowledge cutoffs, non-determinism, and opaque training data, rendering traditional train or test protocols difficult to interpret. At the same time, LLMs create new opportunities, enabling natural language (NL) interaction and even allowing models to act as evaluators. This work argues that the shift toward LLM-driven MRSs requires rethinking evaluation. We first review how LLMs reshape user modeling, item modeling, and NL-based recommendation in music. We then examine evaluation practices from NLP, highlighting methodologies and open challenges relevant to MRSs. Finally, we synthesize insights, focusing on how LLM prompting applies to MRSs, to outline a structured set of success and risk dimensions. Our goal is to provide the MRSs community with an updated, pedagogical, and cross-disciplinary perspective on evaluation.
Figures
Reference graph
Works this paper leans on
-
[67]
Hui Huang, Xingyuan Bu, Hongli Zhou, Yingqi Qu, Jing Liu, Muyun Yang, Bing Xu, and Tiejun Zhao. 2025. An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4. arXiv:2403.02839 [cs.CL] https://arxiv.org/abs/2403.02839
arXiv 2025
-
[1]
Marah I Abdin, Suriya Gunasekar, Varun Chandrasekaran, Jerry Li, Mert Yuksekgonul, Rahee Ghosh Peshawaria, Ranjita Naik, and Besmira Nushi. 2024. KITAB: Evaluating LLMs on Constraint Satisfaction for Information Retrieval. InThe Twelfth International Conference on Learning Representations (ICLR). https://openreview.net/forum?id=b3kDP3IytM
2024
-
[2]
Himan Abdollahpouri, Masoud Mansoury, Robin Burke, and Bamshad Mobasher. 2019. Managing Popularity Bias in Recommender Systems with Personalized Re-ranking. InProceedings of the 29th International Florida Artificial Intelligence Research Society Conference (FLAIRS). 413–418
2019
-
[3]
Khetam Al Sharou, Zhenhao Li, and Lucia Specia. 2021. Towards a Better Understanding of Noise in Natural Language Processing. InProceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2021), Ruslan Mitkov and Galia Angelova (Eds.). INCOMA Ltd., Held Online, 53–62. https://aclanthology.org/2021.ranlp-1.7/
2021
-
[4]
Geetha Sai Aluri, Siddharth Sharma, Tarun Sharma, and Joaquin Delgado. 2024. Playlist Search Reinvented: LLMs Behind the Curtain. InProceedings of the 18th ACM Conference on Recommender Systems(Bari, Italy)(RecSys ’24). Association for Computing Machinery, New York, NY, USA, 813–815. doi:10.1145/3640457.3688047
arXiv 2024
-
[5]
Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2024. Self-rag: Learning to retrieve, generate, and critique through self-reflection. (2024)
2024
-
[6]
Yuyang Bai, Shangbin Feng, Vidhisha Balachandran, Zhaoxuan Tan, Shiqi Lou, Tianxing He, and Yulia Tsvetkov. 2024. KGQuiz: Evaluating the Generalization of Encoded Knowledge in Large Language Models. InProceedings of the ACM Web Conference 2024(Singapore, Singapore)(WWW ’24). Association for Computing Machinery, New York, NY, USA, 2226–2237. doi:10.1145/35...
arXiv 2024
-
[8]
Keqin Bao, Ming Yan, Yang Zhang, Jizhi Zhang, Wenjie Wang, Fuli Feng, and Xiangnan He. 2024. Real-Time Personalization for LLM-based Recommendation with Customized In-Context Learning.arXiv preprint arXiv:2410.23136(2024)
arXiv 2024
Show all 196 references
-
[9]
Christine Bauer and Markus Schedl. 2019. Global and country-specific mainstreaminess measures: Definitions, analysis, and usage for improving personalized music recommendation systems.PLOS ONE14, 6 (June 2019), e0217389. doi:10.1371/journal.pone.0217389 Publisher: Public Libra...
2019 doi
-
[10]
Anna Bavaresco, Raffaella Bernardi, Leonardo Bertolazzi, Desmond Elliott, Raquel Fernández, Albert Gatt, Esam Ghaleb, Mario Giulianelli, Michael Hanna, Alexander Koller, Andre Martins, Philipp Mondorf, Vera Neplenbroek, Sandro Pezzelle, Barbara Plank, David Schlangen, Alessand...
2025
-
[11]
Giovanni Maria Biancofiore, Tommaso Di Noia, Eugenio Di Sciascio, Fedelucio Narducci, and Paolo Pastore. 2022. Aspect based sentiment analysis in music: a case study with spotify. InProceedings of the 37th ACM/SIGAPP Symposium on Applied Computing(Virtual Event)(SAC ’22). Asso...
2022
-
[12]
Tony Brooke. 2014. descriptive metadata in the music industry: Why it is broken and how to fix it—part one.Journal of Digital Media Management 2, 3 (2014), 263–282
2014
-
[13]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...
2020 arXiv
-
[14]
Pedro G Campos, Fernando Díez, and Iván Cantador. 2014. Time-aware recommender systems: a comprehensive survey and analysis of existing evaluation protocols.User Modeling and User-Adapted Interaction24, 1 (2014), 67–119
2014
-
[15]
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al. 2021. Extracting Training Data from Large Language Models. In30th USENIX Security Symposium (USENIX Security 21). ...
2021
-
[16]
Brandon James Carone and Pablo Ripollés. 2024. SoundSignature: What Type of Music do you Like?. In2024 IEEE 5th International Symposium on the Internet of Sounds (IS2). IEEE, 1–10
2024
-
[17]
Elif Celen, Pol van Rijn, Harin Lee, and Nori Jacoby. 2025. Are Expressions for Music Emotions the Same Across Cultures?. InProceedings of the Annual Meeting of the Cognitive Science Society (CogSci)
2025
-
[18]
2009.Music recommendation and discovery in the long tail
Òscar Celma Herrada et al. 2009.Music recommendation and discovery in the long tail. Universitat Pompeu Fabra
2009
-
[19]
Ching-Wei Chen, Paul Lamere, Markus Schedl, and Hamed Zamani. 2018. Recsys challenge 2018: Automatic music playlist continuation. In Proceedings of the 12th ACM Conference on Recommender Systems. 527–528
2018
-
[20]
et al. Chen. 2025. MLA-Trust: Benchmarking Trustworthiness of Multimodal Large-scale Agents.arXiv preprint arXiv:2506.01616(2025)
2025 arXiv
-
[21]
Jiawei Chen, Hande Dong, Xiang Wang, Fuli Feng, Meng Wang, and Xiangnan He. 2023. Bias and debias in recommender system: A survey and future directions.ACM Transactions on Information Systems41, 3 (2023), 1–39
2023
-
[22]
Yanran Chen and Steffen Eger. 2023. MENLI: Robust Evaluation Metrics from Natural Language Inference.Transactions of the Association for Computational Linguistics11 (2023), 804–825. doi:10.1162/tacl_a_00576
2023 doi
-
[23]
Yi Chen, Rui Wang, Haiyun Jiang, Shuming Shi, and Ruifeng Xu. 2023. Exploring the Use of Large Language Models for Reference-Free Text Quality Evaluation: An Empirical Study. InFindings of the Association for Computational Linguistics: IJCNLP-AACL 2023 (Findings), Jong C. Park...
2023 doi
-
[24]
Cheng-Han Chiang and Hung-yi Lee. 2023. Can Large Language Models Be an Alternative to Human Evaluations?. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Ed...
2023 doi
-
[25]
Keunwoo Choi, Seungheon Doh, and Juhan Nam. 2025. TalkPlayData 2: An Agentic Synthetic Data Pipeline for Multimodal Conversational Music Recommendation.arXiv preprint arXiv:2509.09685(2025)
2025 arXiv
-
[26]
Pierre Colombo, Guillaume Staerman, Chloé Clavel, and Pablo Piantanida. 2021. Automatic Text Evaluation through the Lens of Wasserstein Barycenters. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Marie-Francine Moens, Xuanjing Huang, ...
2021 doi
-
[27]
Microsoft Corporation. 2025. GroundednessEvaluator — Azure AI Evaluation SDK for Python. https://learn.microsoft.com/en-us/python/api/azure- ai-evaluation/azure.ai.evaluation.groundnessevaluator?view=azure-python. Accessed: 2025-10-25
2025
-
[28]
Paolo Cremonesi and Dietmar Jannach. 2021. Progress in recommender systems research: Crisis? What crisis?AI Magazine42, 3 (2021), 43–54
2021
-
[29]
Mathieu Delcluze, Antoine Khoury, Clémence Vast, Valerio Arnaudo, Léa Briand, Walid Bendada, and Thomas Bouabça. 2025. Text2Playlist: Generating Personalized Playlists from Text on Deezer. InThe 47th European Conference on Information Retrieval (ECIR 2025). Manuscript submitte...
2025
-
[30]
Yashar Deldjoo. 2024. Understanding biases in chatgpt-based recommender systems: Provider fairness, temporal stability, and recency.ACM Transactions on Recommender Systems(2024)
2024
-
[31]
Yashar Deldjoo, Zhankui He, Julian McAuley, Anton Korikov, Scott Sanner, Arnau Ramisa, Rene Vidal, Maheswaran Sathiamoorthy, Atoosa Kasrizadeh, Silvia Milano, et al. 2024. Recommendation with Generative Models.arXiv preprint arXiv:2409.15173(2024)
2024 arXiv
-
[32]
Yashar Deldjoo, Nikhil Mehta, Maheswaran Sathiamoorthy, Shuai Zhang, Pablo Castells, and Julian McAuley. 2025. Toward Holistic Evaluation of Recommender Systems Powered by Generative Models. InProceedings of the 48th International ACM SIGIR Conference on Research and Developme...
2025
-
[33]
Yashar Deldjoo, Markus Schedl, and Peter Knees. 2024. Content-driven music recommendation: Evolution, state of the art, and challenges.Computer Science Review51 (2024), 100618. doi:10.1016/j.cosrev.2024.100618
2024
-
[34]
Zihao Deng, Yinghao Ma, Yudong Liu, Rongchen Guo, Ge Zhang, Wenhu Chen, Wenhao Huang, and Emmanouil Benetos. 2024. MusiLingo: Bridging Music and Text with Pre-trained Language Models for Music Captioning and Query Response. InFindings of the Association for Computational Lingu...
2024
-
[35]
Daniel Deutsch, Rotem Dror, and Dan Roth. 2022. On the Limitations of Reference-Free Evaluations of Generated Text. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for...
2022 doi
-
[36]
Karlijn Dinnissen and Christine Bauer. 2022. Fairness in music recommender systems: A stakeholder-centered mini review.Frontiers in big Data5 (2022), 913608
2022
-
[37]
Seungheon Doh, Keunwoo Choi, Jongpil Lee, and Juhan Nam. 2023. LP-MusicCaps: LLM-Based Pseudo Music Captioning. InIsmir 2023 Hybrid Conference
2023
-
[38]
Seungheon Doh, Junwon Lee, and Juhan Nam. 2021. Music Playlist Title Generation: A Machine-Translation Approach. InProceedings of the 2nd Workshop on NLP for Music and Spoken Audio (NLP4MusA), Sergio Oramas, Elena Epure, Luis Espinosa-Anke, Rosie Jones, Massimo Quadrana, Moham...
2021
-
[39]
SeungHeon Doh, Minhee Lee, Dasaem Jeong, and Juhan Nam. 2024. Enriching music descriptions with a finetuned-llm and metadata for text-to-music retrieval. InICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 826–830
2024
-
[40]
Elena Epure and Romain Hennequin. 2023. A human subject study of named entity recognition in conversational music recommendation queries. InProceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 1281–1296
2023
-
[41]
Epure, Anis Khlif, and Romain Hennequin
Elena V. Epure, Anis Khlif, and Romain Hennequin. 2019. Leveraging knowledge bases and parallel annotations for music genre translation. In International Society for Music Information Retrieval Conference
2019
-
[42]
Elena V Epure, Guillaume Salha, and Romain Hennequin. 2020. Multilingual music genre embeddings for effective cross-lingual music item annotation.arXiv preprint arXiv:2009.07755(2020)
2020 arXiv
-
[43]
Elena V Epure, Guillaume Salha, Manuel Moussallam, and Romain Hennequin. 2020. Modeling the Music Genre Perception across Language-Bound Cultures. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 4765–4779
2020
-
[44]
Shahul Es, Jithin James, Luis Espinosa-Anke, and Steven Schockaert. 2023. RAGAS: Automated Evaluation of Retrieval Augmented Generation. arXiv preprint arXiv:2309.15217(2023)
2023 arXiv
-
[45]
Shahul Es, Jithin James, Luis Espinosa Anke, and Steven Schockaert. 2024. RAGAs: Automated Evaluation of Retrieval Augmented Generation. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, Nikol...
2024 doi
-
[46]
Andres Ferraro. 2019. Music cold-start and long-tail recommendation: bias in deep representations. InProceedings of the 13th ACM conference on recommender systems. 586–590
2019
-
[47]
Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, George Foster, Alon Lavie, and Ondřej Bojar. 2021. Results of the WMT21 Metrics Shared Task: Evaluating Metrics with Expert-based Human Evaluations on TED and News Domain. InProceedings of the Sixth Confere...
2021
-
[48]
Jinlan Fu, See-Kiong Ng, Zhengbao Jiang, and Pengfei Liu. 2024. GPTScore: Evaluate as You Desire. InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Kevin Duh...
2024 doi
-
[49]
Yumeng Fu, Junjie Wu, Zhongjie Wang, Meishan Zhang, Lili Shan, Yulin Wu, and Bingquan Liu. 2025. LaERC-S: Improving LLM-based Emotion Recognition in Conversation with Speaker Characteristics. InProceedings of the 31st International Conference on Computational Linguistics, Owen...
2025
-
[50]
Christian Fuentes, Johan Hagberg, and Hans Kjellberg. 2019. Soundtracking: music listening practices in the digital age.European Journal of Marketing53, 3 (2019), 483–503
2019
-
[51]
Escrig, and M
Nieves Fuentes-Sánchez, Raúl Pastor, Tuomas Eerola, Miguel A. Escrig, and M. Carmen Pastor. 2022. Musical preference but not familiarity influences subjective ratings and psychophysiological correlates of music-induced emotions.Personality and Individual Differences198 (Nov. 2...
2022
-
[52]
Giovanni Gabbolini and Derek Bridge. 2021. Generating Interesting Song-to-Song Segues With Dave. InProceedings of the 29th ACM Conference on User Modeling, Adaptation and Personalization(Utrecht, Netherlands)(UMAP ’21). Association for Computing Machinery, New York, NY, USA, 9...
2021
-
[53]
Giovanni Gabbolini, Romain Hennequin, and Elena Epure. 2022. Data-Efficient Playlist Captioning With Musical and Linguistic Knowledge. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Ed...
2022 doi
-
[54]
Qingqing Gao, Jiuxin Cao, Biwei Cao, Xin Guan, and Bo Liu. 2024. CEPT: A Contrast-Enhanced Prompt-Tuning Framework for Emotion Recognition in Conversation. InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation ...
2024
-
[55]
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yixin Dai, Jiawei Sun, Haofen Wang, and Haofen Wang. 2023. Retrieval- augmented generation for large language models: A survey.arXiv preprint arXiv:2312.109972, 1 (2023)
2023 arXiv
-
[56]
Zhaolin Gao, Joyce Zhou, Yijia Dai, and Thorsten Joachims. 2024. End-to-end Training for Recommendation with Language-based User Profiles. arXiv preprint arXiv:2410.18870(2024)
2024 arXiv
-
[57]
Iker García-Ferrero, Begoña Altuna, Javier Alvez, Itziar Gonzalez-Dios, and German Rigau. 2023. This is not a Dataset: A Large Negation Benchmark to Challenge Large Language Models. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Houda...
2023 doi
-
[58]
Josh Gardner, Simon Durand, Daniel Stoller, and Rachel Bittner. 2024. LLARK: a multimodal instruction-following language model for music. In Proceedings of the 41st International Conference on Machine Learning(Vienna, Austria)(ICML’24). JMLR.org, Article 603, 46 pages
2024
- [59]
-
[60]
Yue Guo, Tal August, Gondy Leroy, Trevor Cohen, and Lucy Lu Wang. 2024. APPLS: Evaluating Evaluation Metrics for Plain Language Summarization. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung...
2024 doi
-
[61]
Simon Hachmeier and Robert Jäschke. 2025. A Benchmark and Robustness Study of In-Context-Learning with Large Language Models in Music Entity Detection. InProceedings of the 31st International Conference on Computational Linguistics. 9845–9859
2025
-
[62]
Anna Hausberger, Hannah Strauss, and Markus Schedl. 2025. ExIM: Exploring Intent of Music Listening for Retrieving User-generated Playlists. In Proceedings of the 2025 ACM SIGIR Conference on Human Information Interaction and Retrieval (CHIIR ’25). Association for Computing Ma...
2025
-
[63]
Xiangnan He, Lizi Liao, Hanwang Zhang, Liqiang Nie, Xia Hu, and Tat-Seng Chua. 2017. Neural collaborative filtering. InProceedings of the 26th international conference on world wide web. 173–182
2017
-
[64]
Zhankui He, Zhouhang Xie, Rahul Jha, Harald Steck, Dawen Liang, Yesu Feng, Bodhisattwa Prasad Majumder, Nathan Kallus, and Julian McAuley
-
[65]
Yupeng Hou, Junjie Zhang, Zihan Lin, Hongyu Lu, Ruobing Xie, Julian McAuley, and Wayne Xin Zhao. 2024. Large language models are zero-shot rankers for recommender systems. InEuropean Conference on Information Retrieval. Springer, 364–381
2024
-
[66]
Yifan Hu, Yehuda Koren, and Chris Volinsky. 2008. Collaborative filtering for implicit feedback datasets. In2008 Eighth IEEE international conference on data mining. Ieee, 263–272
2008
-
[68]
Jin Huang, Harrie Oosterhuis, Masoud Mansoury, Herke Van Hoof, and Maarten de Rijke. 2024. Going beyond popularity and positivity bias: Correcting for multifactorial bias in recommender systems. InProceedings of the 47th International ACM SIGIR Conference on Research and Devel...
2024
-
[69]
Wenyu Huang, Guancheng Zhou, Mirella Lapata, Pavlos Vougiouklis, Sebastien Montella, and Jeff Z. Pan. 2025. Prompting large language models with knowledge graphs for question answering involving long-tail facts.Knowledge-Based Systems324 (Aug. 2025), 113648. doi:10.1016/j.knos...
2025
-
[70]
Karim M Ibrahim, Elena V Epure, Geoffroy Peeters, and Gael Richard. 2020. SHOULD WE CONSIDER THE USERS IN CONTEXTUAL MUSIC AUTO-TAGGING MODELS?. In21st International Society for Music Information Retrieval Conference
2020
-
[71]
Ari Jacovi and Yoav Goldberg. 2020. Aligning Faithful Explanations: Towards Explanation as Verification.Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP 2020)(2020), 514–525. https://aclanthology.org/2020.emnlp-main.40
2020
-
[72]
Aryan Jadon and Avinash Patil. 2024. A comprehensive survey of evaluation techniques for recommendation systems. InInternational Conference on Computation of Artificial Intelligence & Machine Learning. Springer, 281–304
2024
-
[73]
Dietmar Jannach, Ahtsham Manzoor, Wanling Cai, and Li Chen. 2021. A survey on conversational recommender systems.ACM Computing Surveys (CSUR)54, 5 (2021), 1–36. Manuscript submitted to ACM Music Recommendation with Large Language Models 31
2021
-
[74]
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation.ACM computing surveys55, 12 (2023), 1–38
2023
-
[75]
Marius Kaminskas and Derek Bridge. 2016. Diversity, serendipity, novelty, and coverage: a survey and empirical analysis of beyond-accuracy objectives in recommender systems.ACM Transactions on Interactive Intelligent Systems (TiiS)7, 1 (2016), 1–42
2016
-
[76]
Li Kang, Yuhan Zhao, and Li Chen. 2025. Exploring the Potential of LLMs for Serendipity Evaluation in Recommender Systems. InProceedings of the Nineteenth ACM Conference on Recommender Systems. 746–754
2025
-
[77]
Wang-Cheng Kang and Julian McAuley. 2018. Self-attentive sequential recommendation. In2018 IEEE international conference on data mining (ICDM). IEEE, 197–206
2018
-
[78]
Mohammad Khosravani, Chenyang Huang, and Amine Trabelsi. 2024. Enhancing Argument Summarization: Prioritizing Exhaustiveness in Key Point Generation and Introducing an Automatic Coverage Evaluation Metric. InProceedings of the 2024 Conference of the North American Chapter of t...
2024
-
[79]
Bart P Knijnenburg and Martijn C Willemsen. 2015. Evaluating recommender systems with user experiments. InRecommender systems handbook. Springer, 309–352
2015
-
[80]
Zhifeng Kong, Arushi Goel, Rohan Badlani, Wei Ping, Rafael Valle, and Bryan Catanzaro. 2024. Audio Flamingo: a novel audio language model with few-shot learning and dialogue abilities. InProceedings of the 41st International Conference on Machine Learning. 25125–25148
2024
-
[81]
Ngoc Luyen Le, Marie-Hélène Abel, and Philippe Gouspillou. 2023. A Constraint-based Recommender System via RDF Knowledge Graphs. In2023 26th International Conference on Computer Supported Cooperative Work in Design (CSCWD). IEEE, Rio de Janeiro, Brazil, 849–854. doi:10.1109/ C...
2023
-
[82]
Harin Lee, Elif Çelen, Peter Harrison, Manuel Anglada-Tort, Pol van Rijn, Minsu Park, Marc Schönwiesner, and Nori Jacoby. 2025. GlobalMood: A cross-cultural benchmark for music emotion recognition.arXiv preprint arXiv:2505.09539(2025)
2025
-
[83]
Oleg Lesota, Gustavo Escobedo, Yashar Deldjoo, Bruce Ferwerda, Simone Kopeinik, Elisabeth Lex, Navid Rekabsaz, and Markus Schedl. 2023. Computational Versus Perceived Popularity Miscalibration in Recommender Systems. InProceedings of the 46th International ACM SIGIR Conference...
2023
-
[84]
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing sy...
2020
-
[85]
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2021. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. arXiv:2005.1140...
2021 arXiv
-
[86]
Ang Li, Jennifer Thom, Praveen Chandar, Christine Hosey, Brian St Thomas, and Jean Garcia-Gathright. 2019. Search mindsets: Understanding focused and non-focused information seeking in music search. InThe World Wide Web Conference. 2971–2977
2019
-
[87]
Kongmeng Liew, Vipul Mishra, Yangyang Zhou, Elena V Epure, Romain Hennequin, Shoko Wakamiya, and Eiji Aramaki. 2022. Network Analyses for Cross-Cultural Music Popularity. InIsmir 2022 Hybrid Conference
2022
-
[88]
Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. InText summarization branches out. 74–81
2004
-
[89]
Siyi Lin, Chongming Gao, Jiawei Chen, Sheng Zhou, Binbin Hu, Yan Feng, Chun Chen, and Can Wang. 2025. How Do Recommendation Models Amplify Popularity Bias? An Analysis from the Spectral Perspective. InProceedings of the Eighteenth ACM International Conference on Web Search and...
2025
-
[90]
Meijun Liu, Xiao Hu, and Markus Schedl. 2018. The relation of culture, socio-economics, and friendship to music preferences: A large-scale, cross-country study.PLOS ONE13, 12 (Dec. 2018), e0208186. doi:10.1371/journal.pone.0208186 Publisher: Public Library of Science
2018 doi
-
[91]
Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. Lost in the middle: How language models use long contexts.Transactions of the Association for Computational Linguistics12 (2024), 157–173
2024
-
[92]
Ruilun Liu and Xiao Hu. 2020. A Multimodal Music Recommendation System with Listeners’ Personality and Physiological Signals. InJCDL ’20: Proceedings of the ACM/IEEE Joint Conference on Digital Libraries in 2020, Virtual Event, China, August 1-5, 2020, Ruhua Huang, Dan Wu, Gar...
2020
-
[93]
Chen, and Min-Yen Kan
Do Xuan Long, Ngoc-Hai Nguyen, Tiviatis Sim, Hieu Dao, Shafiq Joty, Kenji Kawaguchi, Nancy F. Chen, and Min-Yen Kan. 2025. LLMs Are Biased Towards Output Formats! Systematically Evaluating and Mitigating Output Format Bias of LLMs. InProceedings of the 2025 Conference of the N...
2025
-
[94]
Lonsdale and Adrian C
Adam J. Lonsdale and Adrian C. North. 2011. Why do we listen to music? A uses and gratifications analysis.British Journal of Psychology (London, England: 1953)102, 1 (Feb. 2011), 108–134. doi:10.1348/000712610X506831
2011 doi
-
[95]
Feng Lu and Nava Tintarev. 2018. A Diversity Adjusting Strategy with Personality for Music Recommendation. InProceedings of the 5th Joint Workshop on Interfaces and Human Decision Making for Recommender Systems, IntRS 2018, co-located with ACM Conference on Recommender Systems...
2018
-
[96]
Malte Ludewig and Dietmar Jannach. 2018. Evaluation of session-based recommendation algorithms.User Modeling and User-Adapted Interaction 28, 4 (2018), 331–390
2018
-
[97]
Hanjia Lyu, Song Jiang, Hanqing Zeng, Yinglong Xia, Qifan Wang, Si Zhang, Ren Chen, Chris Leung, Jiajie Tang, and Jiebo Luo. 2024. LLM-Rec: Personalized Recommendation via Prompting Large Language Models. InFindings of the Association for Computational Linguistics: NAACL 2024,...
2024 doi
-
[98]
Lilian Marey, Bruno Sguerra, and Manuel Moussallam. 2024. Modeling activity-driven music listening with pace. InProceedings of the 2024 Conference on Human Information Interaction and Retrieval. 346–351
2024
-
[99]
Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2020. Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation Metrics. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Nata...
2020 doi
-
[100]
Kristina Matrosova, Lilian Marey, Guillaume Salha-Galvan, Thomas Louail, Olivier Bodini, and Manuel Moussallam. 2024. Do recommender systems promote local music? a reproducibility study using music streaming data. InProceedings of the 18th ACM Conference on Recommender Systems...
2024
-
[101]
McCrae and Oliver P
Robert R. McCrae and Oliver P. John. 1992. An Introduction to the Five-Factor Model and Its Applications.Journal of Personality60, 2 (1992), 175–215. doi:10.1111/j.1467-6494.1992.tb00970.x arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1111/j.1467-6494.1992.tb00970.x
1992
-
[102]
Melchiorre, Elena V
Alessandro B. Melchiorre, Elena V. Epure, Shahed Masoudian, Gustavo Escobedo, Anna Hausberger, Manuel Moussallam, and Markus Schedl. 2025. Just Ask for Music (JAM): Multimodal and Personalized Natural Language Music Recommendation. InProceedings of the 19th ACM Conference on R...
2025
-
[103]
Melchiorre and Markus Schedl
Alessandro B. Melchiorre and Markus Schedl. 2020. Personality Correlates of Music Audio Preferences for Modelling Music Listeners. InProceedings of the 28th ACM Conference on User Modeling, Adaptation and Personalization (UMAP ’20). Association for Computing Machinery, New Yor...
2020
-
[104]
MetaBrainz Foundation. [n. d.]. MusicBrainz Database. https://musicbrainz.org/doc/MusicBrainz_Database. Accessed 2025-10-11
2025
-
[105]
MetaBrainz Foundation. n.d.. MusicBrainz: the open music encyclopedia. https://musicbrainz.org/. Accessed 2025-10-11
2025
-
[106]
BSL METEOR. 2005. an automatic metric for MT evaluation with improved correlation with human judgments. InProceedings of the Acl Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization
2005
-
[107]
Daly, Kush R
Erik Miehling, Michael Desmond, Karthikeyan Natesan Ramamurthy, Elizabeth M. Daly, Kush R. Varshney, Eitan Farchi, Pierre Dognin, Jesus Rios, Djallel Bouneffouf, Miao Liu, and Prasanna Sattigeri. 2025. Evaluating the Prompt Steerability of Large Language Models. InProceedings ...
2025
-
[108]
Sho Miyakawa and Takehito Utsuro. 2024. Emotion Classification of Lyrics through Summarization by Large Language Models. In2024 IEEE International Conference on Big Data (BigData). 2999–3006. doi:10.1109/BigData62323.2024.10825406
2024
-
[109]
2015.Fundamentals of music processing: Audio, analysis, algorithms, applications
Meinard Müller. 2015.Fundamentals of music processing: Audio, analysis, algorithms, applications. Vol. 5. Springer
2015
-
[110]
Sheshera Mysore, Mahmood Jasim, Andrew McCallum, and Hamed Zamani. 2023. Editable user profiles for controllable text recommendations. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval. 993–1003
2023
-
[111]
Sergio Oramas, Andres Ferraro, Alvaro Sarasua, and Fabien Gouyon. 2024. Talking to your recs: Multimodal embeddings for recommendation and retrieval. InMuRS 2024: 2nd Music Recommender Systems Workshop
2024
-
[112]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for automatic evaluation of machine translation. InProceedings of the 40th annual meeting of the Association for Computational Linguistics. 311–318
2002
-
[113]
Emilia Parada-Cabaleiro, Anton Batliner, Marcel Zentner, and Markus Schedl. 2023. Exploring emotions in Bach chorales: a multi-modal perceptual and data-driven study.Royal Society Open Science10, 12 (Dec. 2023), 230574. doi:10.1098/rsos.230574 Publisher: Royal Society
2023 doi
-
[114]
Debjit Paul, Robert West, Antoine Bosselut, and Boi Faltings. 2024. Making reasoning matter: Measuring and improving faithfulness of chain-of- thought reasoning.arXiv preprint arXiv:2402.13950(2024)
2024 arXiv
-
[115]
Geoffroy Peeters et al. 2004. A large set of audio features for sound description (similarity and classification) in the CUIDADO project.CUIDADO Ist Project Report54, 0 (2004), 1–25
2004
-
[116]
Andreas Peintner, Marta Moscati, Yu Kinoshita, Richard Vogl, Peter Knees, Markus Schedl, Hannah Strauss, Marcel Zentner, and Eva Zangerle
-
[117]
Maxime Peyrard. 2019. Studying Summarization Evaluation Metrics in the Appropriate Scoring Range. InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Anna Korhonen, David Traum, and Lluís Màrquez (Eds.). Association for Computational Ling...
2019 doi
-
[118]
Pantelis Pipergias Analytis and Philipp Hager. 2023. Collaborative filtering algorithms are prone to mainstream-taste bias. InProceedings of the 17th ACM Conference on Recommender Systems. 750–756
2023
-
[119]
Liam Pond, Sichen Meng, Linnea Kirby, Simon Ngassam, Sebastien Chow, Dylan Hillerbrand, and Ichiro Fujinaga. 2025. SESEMMI for LinkedMusic: Democratizing Access to Musical Archives via Large Language Models. In1st Workshop on Large Language Models for Music{\&} Audio (LLM4MA)....
2025
-
[120]
Massimo Quadrana, Paolo Cremonesi, and Dietmar Jannach. 2018. Sequence-Aware Recommender Systems. arXiv:1802.08452 [cs.IR] https: //arxiv.org/abs/1802.08452
2018 arXiv
-
[121]
Rashin Rahnamoun and Mehrnoush Shamsfard. 2025. Multi-Layered Evaluation Using a Fusion of Metrics and LLMs as Judges in Open-Domain Question Answering. InProceedings of the 31st International Conference on Computational Linguistics, Owen Rambow, Leo Wanner, Marianna Apidianak...
2025
-
[122]
Abdallah, Mark B
Yves Raimond, Samer A. Abdallah, Mark B. Sandler, and Frederick Giasson. 2007. The Music Ontology. InProceedings of the 8th International Conference on Music Information Retrieval, ISMIR 2007, Vienna, Austria, September 23-27, 2007, Simon Dixon, David Bainbridge, and Rainer Ty...
2007
-
[123]
Thomas, Chandr Dhanush H, and Arunima C
Rajeev Rajan, Joshua Antony, Riya Ann Joseph, Jijohn M. Thomas, Chandr Dhanush H, and Arunima C. V. 2021. Audio-Mood Classification Using Acoustic-Textual Feature Fusion. In2021 Fourth International Conference on Microelectronics, Signals & Systems (ICMSS). 1–6. doi:10.1109/ I...
2021
-
[124]
Jerome Ramos, Hossein A Rahmani, Xi Wang, Xiao Fu, and Aldo Lipani. 2024. Transparent and Scrutable Recommendations Using Natural Language User Profiles. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 13971–13984
2024
-
[125]
Brian Regan, Desislava Hristova, and Mariano Beguerisse-Díaz. 2023. Which Witch? Artist name disambiguation and catalog curation using audio and metadata. https://research.atspotify.com/2023/11/which-witch-artist-name-disambiguation-and-catalog-curation-using-audio-and-metadat...
2023
-
[126]
Lise Regnier and Geoffroy Peeters. 2009. Singing voice detection in music tracks using direct voice vibrato detection. In2009 IEEE international conference on acoustics, speech and signal processing. IEEE, 1685–1688
2009
-
[127]
Ricardo Rei, Ana C Farinha, Chrysoula Zerva, Daan van Stigt, Craig Stewart, Pedro Ramos, Taisiya Glushkova, André F. T. Martins, and Alon Lavie. 2021. Are References Really Needed? Unbabel-IST 2021 Submission for the Metrics Shared Task. InProceedings of the Sixth Conference o...
2021
-
[128]
Peter J Rentfrow. 2012. The role of music in everyday life: Current directions in the social psychology of music.Social and personality psychology compass6, 5 (2012), 402–416
2012
-
[129]
Rentfrow and Samuel D
Peter J. Rentfrow and Samuel D. Gosling. 2007. The content and validity of music-genre stereotypes among college students.Psychology of Music 35, 2 (2007), 306–326. doi:10.1177/0305735607070382
2007 doi
-
[130]
Renata L Rosa, Demsteneso Z Rodriguez, and Graça Bressan. 2015. Music recommendation system based on user’s sentiments extracted from social networks.IEEE Transactions on Consumer Electronics61, 3 (2015), 359–367
2015
-
[131]
Jia-Jia Ruan, Xi-Xu He, Min Zhang, and Yuan Gao. 2023. Entity generation algorithm based on reference expansion.Journal of Electronic Science and Technology21, 3 (2023), 100218
2023
-
[132]
Milad Sabouri, Masoud Mansoury, Kun Lin, and Bamshad Mobasher. 2025. Towards Explainable Temporal User Profiling with LLMs. InAdjunct Proceedings of the 33rd ACM Conference on User Modeling, Adaptation and Personalization. 219–227
2025
-
[133]
Sai, Tanay Dixit, Dev Yashpal Sheth, Sreyas Mohan, and Mitesh M
Ananya B. Sai, Tanay Dixit, Dev Yashpal Sheth, Sreyas Mohan, and Mitesh M. Khapra. 2021. Perturbation CheckLists for Evaluating NLG Evaluation Metrics. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Marie-Francine Moens, Xuanjing Huan...
2021 doi
-
[134]
Oscar Sainz, Iker García-Ferrero, Rodrigo Agerri, Oier Lopez de Lacalle, German Rigau, and Eneko Agirre. 2024. GoLLIE: Annotation Guidelines improve Zero-Shot Information-Extraction. InThe Twelfth International Conference on Learning Representations. https://openreview.net/for...
2024
-
[135]
Krishna Sayana, Raghavendra Vasudeva, Yuri Vasilevski, Kun Su, Liam Hebert, James Pine, Hubert Pham, Ambarish Jash, and Sukhdeep Sodhi
-
[136]
Markus Schedl, Emilia Gómez, Julián Urbano, et al. 2014. Music information retrieval: Recent developments and applications.Foundations and Trends®in Information Retrieval8, 2-3 (2014), 127–261
2014
-
[137]
Trent, Marko Tkalčič, Hamid Eghbal-Zadeh, and Agustín Martorell
Markus Schedl, Emilia Gómez, Erika S. Trent, Marko Tkalčič, Hamid Eghbal-Zadeh, and Agustín Martorell. 2018. On the Interrelation Between Listener Characteristics and the Perception of Emotions in Classical Orchestra Music.IEEE Transactions on Affective Computing9, 4 (Oct. 201...
2018
-
[138]
InCompanion Proceedings of the ACM on Web Conference 2025(Sydney NSW, Australia)(WWW ’25)
Beyond Retrieval: Generating Narratives in Conversational Recommender Systems. InCompanion Proceedings of the ACM on Web Conference 2025(Sydney NSW, Australia)(WWW ’25). Association for Computing Machinery, New York, NY, USA, 2411–2420. doi:10.1145/3701716.3717531
2025
-
[139]
Markus Schedl, Hamed Zamani, Ching-Wei Chen, Yashar Deldjoo, and Mehdi Elahi. 2018. Current challenges and visions in music recommender systems research.International Journal of Multimedia Information Retrieval7, 2 (2018), 95–116
2018
-
[140]
Thomas Schäfer, Peter Sedlmeier, Christine Städtler, and David Huron. 2013. The psychological functions of music listening.Frontiers in Psychology 4 (Aug. 2013). doi:10.3389/fpsyg.2013.00511 Publisher: Frontiers. Manuscript submitted to ACM 34 Epure et al
2013
-
[141]
2021.Music Recommendation Systems: Techniques, Use Cases, and Challenges
Markus Schedl, Peter Knees, Brian McFee, and Dmitry Bogdanov. 2021.Music Recommendation Systems: Techniques, Use Cases, and Challenges. Springer US, 927–971. doi:10.1007/978-1-0716-2197-4_24
2021 doi
-
[142]
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020. BLEURT: Learning Robust Metrics for Text Generation. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault (Eds.). Associa...
2020 doi
-
[143]
Bruno Sguerra, Marion Baranes, Romain Hennequin, and Manuel Moussallam. 2022. Navigational, informational or punk-rock? An exploration of search intent in the musical domain. InProceedings of the 30th ACM Conference on User Modeling, Adaptation and Personalization. 202–211
2022
-
[144]
Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, Alex Wang, and Patrick Gallinari. 2021. QuestEval: Summarization Asks for Fact-based Evaluation. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, M...
2021 doi
-
[145]
Bruno Sguerra, Viet-Anh Tran, and Romain Hennequin. 2022. Discovery dynamics: Leveraging repeated exposure for user and music characterization. InProceedings of the 16th ACM Conference on Recommender Systems. 556–561
2022
-
[146]
Bruno Sguerra, Viet-Anh Tran, and Romain Hennequin. 2023. Ex2Vec: Characterizing users and items from the mere exposure effect. InProceedings of the 17th ACM Conference on Recommender Systems. 971–977
2023
-
[147]
Bruno Sguerra, Elena V Epure, Harin Lee, and Manuel Moussallam. 2025. Biases in LLM-Generated Musical Taste Profiles for Recommendation. In Proceedings of the Nineteenth ACM Conference on Recommender Systems. 527–532
2025
-
[148]
Lingfeng Shen, Lemao Liu, Haiyun Jiang, and Shuming Shi. 2022. On the Evaluation Metrics for Paraphrase Generation. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for...
2022 doi
-
[149]
Tiancheng Shen, Jia Jia, Yan Li, Yihui Ma, Yaohua Bu, Hanjie Wang, Bo Chen, Tat-Seng Chua, and Wendy Hall. 2020. PEIA: Personality and Emotion Integrated Attentive Model for Music Recommendation on Social Media Platforms. InThe Thirty-Fourth AAAI Conference on Artificial Intel...
2020
-
[150]
Bruno Sguerra, Viet-Anh Tran, Romain Hennequin, and Manuel Moussallam. 2025. Uncertainty in Repeated Implicit Feedback as a Measure of Reliability. InProceedings of the 33rd ACM Conference on User Modeling, Adaptation and Personalization. 234–242
2025
-
[151]
Harald Steck. 2018. Calibrated recommendations. InProceedings of the 12th ACM conference on recommender systems. 154–162
2018
-
[152]
Lei Sun, Jinming Zhao, and Qin Jin. 2024. Revealing Personality Traits: A New Benchmark Dataset for Explainable Personality Recognition on Dialogues. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Y...
2024 doi
-
[153]
Yifan Song, Guoyin Wang, Sujian Li, and Bill Yuchen Lin. 2025. The Good, The Bad, and The Greedy: Evaluation of LLMs Should Not Ignore Non-Determinism. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics:...
2025 doi
-
[154]
Noah Tekle, Alline Ayala, Jonathan Haile, Abdulla Alshabanah, Corey Baker, and Murali Annavaram. 2024. Music Recommendation through LLM Song Summary. InThe 1st Workshop on Risks, Opportunities, and Evaluation of Generative Models in Recommender Systems (ROEGEN@RECSYS’24)
2024
-
[155]
Brian Thompson and Matt Post. 2020. Automatic Machine Translation Evaluation in Many Languages via Zero-Shot Paraphrasing. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds....
2020 doi
-
[156]
Tianxiang Sun, Junliang He, Xipeng Qiu, and Xuanjing Huang. 2022. BERTScore is Unfair: On Social Bias in Language Model-Based Metrics for Text Generation. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Yoav Goldberg, Zornitsa Kozareva...
2022 doi
-
[157]
Alicia Tsai, Adam Kraft, Long Jin, Chenwei Cai, Anahita Hosseini, Taibai Xu, Zemin Zhang, Lichan Hong, Ed Chi, and Xinyang Yi. 2024. Leveraging LLM Reasoning Enhances Personalized Recommender Systems. InFindings of the Association for Computational Linguistics ACL 2024. 13176–13188
2024
-
[158]
Robin Ungruh, Karlijn Dinnissen, Anja Volk, Maria Soledad Pera, and Hanna Hauptmann. 2024. Putting Popularity Bias Mitigation to the Test: A User-Centric Evaluation in Music Recommenders. InProceedings of the 18th ACM Conference on Recommender Systems. 169–178
2024
-
[159]
Thinh Hung Truong, Timothy Baldwin, Karin Verspoor, and Trevor Cohn. 2023. Language models are not naysayers: an analysis of language models on negation benchmarks. InProceedings of the 12th Joint Conference on Lexical and Computational Semantics (*SEM 2023), Alexis Palmer and...
2023 doi
-
[160]
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh. 2015. Cider: Consensus-based image description evaluation. InProceedings of the IEEE conference on computer vision and pattern recognition. 4566–4575
2015
-
[161]
et al. Wang. 2025. Diagnostic-Guided Dynamic Profile Optimization for LLM-based User Simulation.arXiv preprint arXiv:2508.12645(2025)
2025
-
[162]
Saúl Vargas and Pablo Castells. 2011. Rank and relevance in novelty and diversity metrics for recommender systems. InProceedings of the Fifth ACM Conference on Recommender Systems. ACM, Chicago, IL, USA, 109–116. doi:10.1145/2043932.2043955
2011
-
[163]
Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Lingpeng Kong, Qi Liu, Tianyu Liu, and Zhifang Sui. 2024. Large Language Models are not Fair Evaluators. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Vo...
2024
-
[164]
Shijia Wang, Tianpei Ouyang, Yunfan Zhou, Qiang Xiao, Yintao Ren, Yifei Pan, Fangjian Li, and Chuanjiang Luo. 2025. Enhanced Emotion-aware Music Recommendation via Large Language Models. InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2...
2025
-
[165]
Jiarong Wang, Jiaji Wu, Mingzhou Tan, and Lingxuan Zhu. 2025. Emotion-Aware Conversational Music Recommendation With Multiagent System. IEEE Transactions on Computational Social Systems(2025), 1–14. doi:10.1109/TCSS.2025.3599008
2025
-
[166]
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc Le, Ed Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2022. Self-consistency improves chain of thought reasoning in language models.arXiv preprint arXiv:2203.11171(2022)
2022 arXiv
-
[167]
Amy Beth Warriner, Victor Kuperman, and Marc Brysbaert. 2013. Norms of valence, arousal, and dominance for 13,915 English lemmas.Behavior research methods45, 4 (2013), 1191–1207
2013
-
[169]
Benno Weck, Ilaria Manco, Emmanouil Benetos, Elio Quinton, György Fazekas, and Dmitry Bogdanov. 2024. MuChoMusic: Evaluating Music Understanding in Multimodal Audio-Language Models. InProceedings of the 25th International Society for Music Information Retrieval Conference (ISMIR)
2024
-
[171]
William Webber, Alistair Moffat, and Justin Zobel. 2010. A Similarity Measure for Indefinite Rankings. InACM Transactions on Information Systems, Vol. 28. ACM, 1–38. doi:10.1145/1852102.1852106
2010
-
[172]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Tobias Bosma, Brian Ichter, Fei Xia, Ed Cole, Alejandro Ehinger, John Luu, and Quoc V. Le. 2022. Chain of Thought Prompting Elicits Reasoning in Large Language Models. InAdvances in Neural Information Processing Systems (NeurIPS 2022). ...
2022 arXiv
-
[173]
Likang Wu, Zhi Zheng, Zhaopeng Qiu, Hao Wang, Hongchao Gu, Tingjia Shen, Chuan Qin, Chen Zhu, Hengshu Zhu, Qi Liu, et al. 2024. A survey on large language models for recommendation.World Wide Web27, 5 (2024), 60
2024
-
[174]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems35 (2022), 24824–24837
2022
-
[175]
Meng Yang, Jon McCormack, Maria Teresa Llano, and Wanchao Su. 2025. Exploring the Feasibility of LLMs for Automated Music Emotion Annotation. InIsmir 2025 Hybrid Conference
2025
-
[176]
Griffiths, Yuan Cao, and Karthik Narasimhan
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. 2023. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. arXiv:2305.10601 [cs.CL] https://arxiv.org/abs/2305.10601
2023 arXiv
-
[177]
Wu, Yixuan Wang, Nathan Tsoi, and Julian Urbano
Steven H. Wu, Yixuan Wang, Nathan Tsoi, and Julian Urbano. 2025. CLaMP 3: Universal Music Information Retrieval Across Multiple Languages and Modalities.arXiv preprint arXiv:2502.xxxxx(2025)
2025
-
[178]
Yakun Yu, Shi-ang Qi, Baochun Li, and Di Niu. 2024. PepRec: Progressive Enhancement of Prompting for Recommendation. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association ...
2024 doi
-
[179]
Ruibin Yuan, Hanfeng Lin, Yi Wang, Zeyue Tian, Shangda Wu, Tianhao Shen, Ge Zhang, Yuhang Wu, Cong Liu, Ziya Zhou, Liumeng Xue, Ziyang Ma, Qin Liu, Tianyu Zheng, Yizhi Li, Yinghao Ma, Yiming Liang, Xiaowei Chi, Ruibo Liu, Zili Wang, Chenghua Lin, Qifeng Liu, Tao Jiang, Wenhao ...
2024
-
[180]
Chao Yu, Qixin Tan, Hong Lu, Jiaxuan Gao, Xinting Yang, Yu Wang, Yi Wu, and Eugene Vinitsky. 2025. ICPL: Few-shot In-context Preference Learning via LLMs. arXiv:2410.17233 [cs.AI] https://arxiv.org/abs/2410.17233
2025 arXiv
-
[181]
Sojeong Yun and Youn-kyung Lim. 2025. User Experience with LLM-powered Conversational Recommendation Systems: A Case of Music Recommendation. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–15
2025
-
[182]
Eva Zangerle, Martin Pichl, and Markus Schedl. 2020. User Models for Culture-Aware Music Recommendation: Fusing Acoustic and Cultural Cues. Transactions of the International Society for Music Information Retrieval3, 1 (March 2020). doi:10.5334/tismir.37
2020 doi
-
[183]
Weizhe Yuan, Graham Neubig, and Pengfei Liu. 2021. BARTSCORE: evaluating generated text as text generation. InProceedings of the 35th International Conference on Neural Information Processing Systems (NIPS ’21). Curran Associates Inc., Red Hook, NY, USA, Article 2088, 15 pages
2021
-
[184]
Baiqiao Zhang, Zhifeng Liao, Xiangxian Li, Chao Zhou, Juan Liu, Xiaojuan Ma, and Yulong Bian. 2025. Rethinking Personality Assessment from Human-Agent Dialogues: Fewer Rounds May Be Better Than More. InFindings of the Association for Computational Linguistics: EMNLP 2025, Chri...
2025
-
[185]
Hanlin Zhang, YiFan Zhang, Yaodong Yu, Dhruv Madeka, Dean Foster, Eric Xing, Himabindu Lakkaraju, and Sham Kakade. 2024. A Study on the Calibration of In-context Learning. InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational L...
2024 doi
-
[186]
Yuheng Zha, Yichi Yang, Ruichen Li, and Zhiting Hu. 2023. AlignScore: Evaluating Factual Consistency with A Unified Alignment Function. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Anna Rogers, Jordan Boyd-...
2023 doi
-
[187]
Tong Zhang. 2025. AdaptRec: A Self-Adaptive Framework for Sequential Recommendations with Large Language Models. arXiv:2504.08786 [cs.IR] https://arxiv.org/abs/2504.08786
2025 arXiv
-
[188]
Weinberger, and Yoav Artzi
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020. BERTScore: Evaluating Text Generation with BERT. In International Conference on Learning Representations. https://openreview.net/forum?id=SkeHuCVFDr
2020
-
[189]
Ruoyu Zhang, Yanzeng Li, Yongliang Ma, Ming Zhou, and Lei Zou. 2023. LLMaAA: Making Large Language Models as Active Annotators. InFindings of the Association for Computational Linguistics: EMNLP 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computatio...
2023 doi
-
[190]
Wei Zhao, Michael Strube, and Steffen Eger. 2023. DiscoScore: Evaluating Text Generation with BERT and Discourse Coherence. InProceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, Andreas Vlachos and Isabelle Augenstein (E...
2023 doi
-
[191]
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P Xing, et al
-
[192]
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M Meyer, and Steffen Eger. 2019. MoverScore: Text Generation Evaluating with Contextualized Embeddings and Earth Mover Distance. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and ...
2019
-
[193]
Joyce Zhou, Yijia Dai, and Thorsten Joachims. 2024. Language-Based User Profiles for Recommendation.arXiv preprint arXiv:2402.15623(2024)
2024 arXiv
-
[194]
Jingming Zhuo, Songyang Zhang, Xinyu Fang, Haodong Duan, Dahua Lin, and Kai Chen. 2024. ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs. InFindings of the Association for Computational Linguistics: EMNLP 2024. Association for Computational Linguistics, 1950–1966
2024
-
[195]
InAdvances in Neural Information Processing Systems, Vol
Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. InAdvances in Neural Information Processing Systems, Vol. 36. 46595–46623
-
[196]
Han Zhou, Xingchen Wan, Lev Proleev, Diana Mincu, Jilin Chen, Katherine A Heller, and Subhrajit Roy. 2024. Batch Calibration: Rethinking Calibration for In-Context Learning and Prompt Engineering. InThe Twelfth International Conference on Learning Representations. https: //ope...
2024
-
[199]
Le Zhuo, Ruibin Yuan, Jiahao Pan, Yinghao Ma, Yizhi Li, Ge Zhang, Si Liu, Roger Dannenberg, Jie Fu, Chenghua Lin, et al. 2023. Lyricwhiz: Robust multilingual zero-shot lyrics transcription by whispering to chatgpt.arXiv preprint arXiv:2306.17103(2023). Received November 2025 M...
2023 arXiv
-
[2023]
InProceedings of the 32nd ACM international conference on information and knowledge management
Large language models as zero-shot conversational recommenders. InProceedings of the 32nd ACM international conference on information and knowledge management. 720–730
-
[2025]
doi:10.5334/tismir.235
Nuanced Music Emotion Recognition via a Semi-Supervised Multi-Relational Graph Neural Network.Transactions of the International Society for Music Information Retrieval8, 1 (June 2025). doi:10.5334/tismir.235
2025 doi
Reviewed August 3, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.