REVIEW 4 major objections 5 minor 131 references
Amuse: Human-AI Collaborative Songwriting with Multimodal Inspirations
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This paper claims that Amuse transforms multimodal inputs—images, text, or audio—into musically coherent, editable chord progressions by combining a multimodal LLM's noisy proposals with a unimodal chord-model filter, and that songwriters…
desk verdict Solid HCI system paper; the rejection-sampling story needs a rewrite before the technical claim holds. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the rejection-sampling acceptance ratio $P(x)/(M Q(x))$, where $P(x)$ is an LSTM learned from human-composed chord progressions (HookTheory) and $Q(x)$ is an LSTM learned from GPT-4o-generated chord progressions with keywords marginalized. Via Bayes' rule and the assumption $P(c|x)\approx Q(c|x)$, the keyword-conditional target $P(x|c)$ is replaced by this ratio, so the filter judges only musical coherence. A prompting technique that asks GPT-4o to generate 30 progressions in one batch supplies diversity, and the Chord Generator's keyword-extraction step supplies relevance and transparency.
What would settle it
Take a fixed set of keyword-conditioned proposals from GPT-4o, run the rejection filter, and have listeners (or a keyword classifier) label whether accepted progressions are more relevant to the keywords than rejected ones. If the filter is truly keyword-agnostic, relevance should be equal in the two sets, and any observed relevance is the LLM's own; if the filter is secretly changing relevance, the Bayes-ratio derivation is not what is doing the work.
Extended reading notes
Core claim
The central discovery is that noisy, keyword-conditioned chord suggestions from a multimodal LLM can be made both coherent and relevant by filtering them with a unimodal prior over real chord progressions, using a rejection-sampling ratio that drops the keyword conditioning entirely. The filtering step is keyword-agnostic because the authors assume the LLM's conditional relevance is close to the true one, so the prior only corrects musical coherence. The paper's user study claims that this pipeline, embedded in Hookpad alongside the contextual assistant Aria, enhances users' agency, creativity, and perceived efficiency without changing final satisfaction.
Load-bearing premise
The whole filter rests on the assumption that the LLM's keyword-to-chord relevance is already correct, so the prior can ignore keywords; if the LLM's relevance is off, the filter cannot repair it, and the LSTM estimate of the LLM's proposal density is also uncalibrated.
Editorial extensions
If this is right
- If the method works as claimed, a songwriter can start from a photograph or a paragraph of prose and, within one session, obtain several editable chord progressions in a chosen key and length.
- The rejection-sampling recipe becomes a general template for conditioning symbolic music models on modalities for which paired chord data do not exist, requiring only an LLM proposal and a unimodal prior.
- The user-study results imply that adding multimodal inspiration support to a contextual AI assistant shifts perceived control and creativity without sacrificing output satisfaction, and changes when and how often users query the contextual assistant.
- The keyword intermediary layer gives users a transparent handle on the AI's interpretation, making the abstract image-to-chords transformation editable and explainable.
Reading between the lines
- I would read the paper's coherence claim as inheriting all of its strength from the LSTM prior; the filter never sees the keywords, so any keyword relevance in the final chords comes entirely from GPT-4o's zero-shot reading of the prompt.
- An empirical test the authors do not report: compute keyword relevance of accepted versus rejected proposals (e.g., with a keyword classifier or human labels). If the filter is truly keyword-agnostic, accepted and rejected sets should have similar relevance, and the observed relevance should match the LLM's proposal quality.
- The same 'LLM proposes, unimodal prior disposes' scheme could be applied to other symbolic musical elements, such as drum patterns or basslines, or even to non-music creative domains where paired data are scarce but a unimodal model of the output exists.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents Amuse, a Chrome extension integrated with Hookpad that assists songwriters by transforming multimodal inputs (images, text) into chord progressions, and transcribing audio into chords. The central technical contribution is a rejection-sampling procedure that combines GPT-4o proposals, conditioned on music keywords extracted from the multimodal input, with an LSTM prior trained on the HookTheory dataset, in order to produce chord progressions that are diverse, relevant to keywords, and musically coherent without paired training data. The paper reports a technical evaluation of diversity (Self-BLEU), coherence (JSD against HookTheory), and a listening study of coherence and keyword relevance, followed by a within-subjects user study with 10 songwriters comparing Amuse+Aria against Aria alone. The user study finds that participants felt greater agency, creativity, and alignment with their creative goals when using Amuse.
Significance. If the rejection-sampling derivation were sound, the paper would contribute a practical method for conditional symbolic music generation without paired data: a generally useful recipe of using a multimodal LLM as a proposal and a unimodal prior as a filter. The user study is carefully designed and analyzed, with interaction logs, think-aloud protocols, and qualitative coding, and it provides credible evidence that a multimodal inspiration-to-chord tool can enhance perceived agency and creativity in songwriting. The paper also ships code and sound examples, which supports reproducibility. However, the technical derivation in Section 5.3 contains load-bearing unsupported equalities, and the coherence evaluation in Table 2 is substantially circular; these issues call the central technical claim into question. The HCI findings appear robust, but the paper's stated contribution (2)—a novel method for generating diverse, relevant, and coherent chord progressions—needs substantial revision.
major comments (4)
- The cancellation of the keyword conditioning from the acceptance ratio relies on two equalities: P(c)=Q(c) and P(c|x)≈Q(c|x). Neither is established. P(c) is not defined as a data distribution anywhere; the paper only defines Q(c) as the distribution over keywords used when sampling prompt keywords from a wiki. These are not the same object unless one defines P(c) to be exactly that sampling distribution, which is not a natural interpretation of Bayes' rule applied to the target P(x|c). More importantly, P(c|x)≈Q(c|x) is a strong assumption about the inverse keyword distribution of GPT-4o matching the true data inverse, and no evidence is given for it. The paper's own §6.2.1 shows that GPT-4o's forward marginal chord distribution deviates substantially from real music, making the inverse-distribution assumption especially doubtful without paired data. After this cancellation, the acceptance probability P(x)/(M Q(x)) is independent of c, and the accepted distribution is Q(x|c)·P(x)/Q(x), not P(x|c). As a result, the method as presented is a keyword-agnostic coherence filter on top of LLM proposals; any keyword relevance is inherited from the prompt, not from the rejection-sampling step. This undermines contribution (2) and the claim in §6.2 that Amuse generates keyword-conditioned progressions through the described rejection-sampling procedure.
- The constant M is set to the 95th percentile of the ratio P(x)/Q(x) over GPT-4o-generated progressions, rather than to an upper bound on that ratio. Rejection sampling is only valid when M ≥ sup_x P(x|c)/Q(x|c) (after cancellation, sup_x P(x)/Q(x)); with a 95th-percentile value, for the top 5% of proposals the ratio exceeds M, so the acceptance probability is capped at 1 instead of being the required ratio/M > 1. This changes the target distribution and invalidates the formal rejection-sampling justification. In addition, Q(x) is an LSTM density fitted to 25,000 GPT-4o samples, but no calibration is reported between this fitted density and the actual GPT-4o marginal proposal density. Both issues mean the accepted samples are not actually drawn from the claimed target distribution, even setting aside the conditioning-cancellation problem above.
- The automatic coherence evaluation compares Amuse's outputs to HookTheory, which is the same dataset used to train the filter P(x). Because the acceptance probability (after the problematic cancellation) is P(x)/Q(x), the accepted samples are biased toward P(x); the large JSD reduction in Table 2 is therefore a near-tautological consequence of filtering with the evaluation reference rather than an independent validation. The listening study in Figure 5a provides a more meaningful coherence check, but the keyword-relevance result in Figure 5b only shows a significant advantage over LSTM Prior; Amuse is not significantly preferred over GPT-4o for relevance. Taken together, the technical evaluation does not support the claim that the rejection-sampling step improves keyword relevance; it may merely preserve the relevance already present in GPT-4o's proposals while improving coherence.
- The final set of four progressions is not a pure rejection-sampling output. When fewer than four samples are accepted, the algorithm fills the remainder with the top-k rejected samples ranked by P(x)/Q(x). This means the user-facing set is a mixture of accepted and rejected samples, and is not a sample from any well-defined target distribution. The reported diversity, coherence, and relevance measurements (including the listening study) are therefore measuring the complete algorithm, which is reasonable, but the paper should acknowledge that the output is not strictly a rejection-sampling result and should evaluate the fallback path separately or at least report how often it is used.
minor comments (5)
- The abstract states that Amuse transforms multimodal image, text, and audio inputs into chord progressions, but the Chord Generator only handles image and text; audio is processed by the separate Chord Transcriber, which does not use keyword conditioning. The phrasing should be sharpened to distinguish the two pipelines.
- The listening study in §6.2.2 says it used 10 keyword sets, but Appendix B.2 lists nine keyword sets plus an attention-check set. Please clarify whether the attention check is included in the counts and how the 150 comparisons were allocated.
- The numeric values in Figure 5 are presented in a compact layout that is difficult to parse; the columns are labeled only in the caption. It would improve readability to label each column directly in the figure, for example 'vs. LSTM Prior' and 'vs. GPT-4o'.
- The implementation section says all LLM components use temperature 1.0, while §B.1 mentions a temperature of 1.7 for the LSTM distributions during rejection sampling. Clarify which temperature applies to which model and why the discrepancy exists.
- The limitations section already covers the study's small sample and controlled setting, which is good. It might also note that the Chord Transcriber's low usage was partly attributed to task constraints, and that future work could investigate the transcriber in more naturalistic settings.
Circularity Check
The coherence JSD evaluation is tautological (filter and metric share the same training distribution), and the rejection-sampling derivation cancels keyword conditioning, making the 'keyword-conditioned' filter keyword-agnostic.
-
fitted input called prediction
[Section 6.2.1 (Automatic Evaluation); Section 5.3 Implementation]
"The aim of rejection sampling is to align LLM-generated chord progressions with real music data distribution. We quantitatively assess this alignment by computing the Jensen-Shannon Divergence (JSD) between the distributions of the generated chord progressions and real music data."
The 'real music data' reference for the JSD is the HookTheory dataset [31], which is also the training data for P(x): Section 5.3 says 'For 𝑃(𝑥), we use the HookTheory dataset [31]'. The acceptance probability in Algorithm 1 is P(x)/(M Q(x)), so accepted samples are biased toward high P(x), i.e., toward the HookTheory training distribution. Measuring JSD of the accepted samples against that same distribution is a self-consistency check: the reported improvement (GPT-4o 0.42/0.57 to Amuse 0.27/0.46) is a mathematical consequence of the filter's objective, not an independent empirical finding about musical coherence.
-
self definitional
[Section 5.3, method paragraph after Bayes-rule cancellation]
"since 𝑃(𝑐) = 𝑄(𝑐) as they are the same predefined keyword distributions and we assume 𝑃(𝑐|𝑥)≈ 𝑄(𝑐|𝑥) as the generated progression 𝑥 from LLMs closely align with 𝑐. Thus, we perform rejection sampling by calculating the ratio 𝑃(𝑥)/(𝑀·𝑄(𝑥))."
After the two asserted equalities, the acceptance probability is P(x)/(M Q(x)), which is independent of the keyword c. The filter therefore cannot affect keyword relevance; all relevance comes from the LLM proposal Q(x|c). The claimed target P(x|c) is reached only if P(c|x)=Q(c|x), an assumption stated without evidence and undercut by the paper's own Table 2 showing that GPT-4o's chord distribution is far from real music (JSD 0.42/0.57). By cancelling c, the derivation defines away the very conditioning the method claims to perform, so the 'keyword-conditioned rejection sampling' reduces to a keyword-agnostic coherence filter on top of LLM proposals.
full rationale
The technical contribution has two circular or self-referential steps. First, the automatic coherence evaluation in Table 2 is a fitted-input-called-prediction: P(x) is trained on the HookTheory dataset and the acceptance rule is P(x)/Q(x), while the JSD metric measures distance to exactly that same HookTheory dataset; the reported coherence improvement is therefore forced by construction rather than independently discovered. Second, the rejection-sampling derivation in Section 5.3 asserts P(c)=Q(c) and P(c|x)≈Q(c|x), which cancels the keyword c from the acceptance ratio. This makes the filter keyword-agnostic and transfers the entire burden of relevance to the LLM's prompt-conditioned proposals, while the paper presents the filtering step as part of a keyword-conditioned system. The listening study and the user study provide independent evidence about the overall tool, but they do not validate the specific claim that the rejection-sampling step targets P(x|c). Because a central 'prediction' (coherence improvement) reduces by construction and a second central claim (keyword conditioning of the filter) is defined away, a score of 6 is appropriate; the paper is not entirely circular since the HCI evaluation and listening study have genuine independent content.
Assumptions & free parameters
free parameters (3)
- M =
7.64
- N =
30
- Temperature =
1.7
assumptions (4)
- ad hoc to paper P(c|x) ≈ Q(c|x) for all progressions x and keyword sets c
- domain assumption P(c) = Q(c), i.e., the target keyword prior equals the keyword distribution used to sample training data for Q(x)
- domain assumption The HookTheory dataset is representative of real human-composed chord progressions
- domain assumption An LSTM trained on a sample of GPT-4o outputs accurately approximates the true proposal density Q(x)
Cite this review
Pith. "Pith review of Amuse: Human-AI Collaborative Songwriting with Multimodal Inspirations." pith.science (2026). https://pith.science/paper/IXYLLYY2
@misc{pith2026241218940,
author = {Pith},
title = {Pith review of: Amuse: Human-AI Collaborative Songwriting with Multimodal Inspirations},
year = {2026},
howpublished = {\url{https://pith.science/paper/IXYLLYY2}},
note = {Machine review of arXiv:2412.18940}
}
read the original abstract
Songwriting is often driven by multimodal inspirations, such as imagery, narratives, or existing music, yet songwriters remain unsupported by current music AI systems in incorporating these multimodal inputs into their creative processes. We introduce Amuse, a songwriting assistant that transforms multimodal (image, text, or audio) inputs into chord progressions that can be seamlessly incorporated into songwriters' creative processes. A key feature of Amuse is its novel method for generating coherent chords that are relevant to music keywords in the absence of datasets with paired examples of multimodal inputs and chords. Specifically, we propose a method that leverages multimodal large language models (LLMs) to convert multimodal inputs into noisy chord suggestions and uses a unimodal chord model to filter the suggestions. A user study with songwriters shows that Amuse effectively supports transforming multimodal ideas into coherent musical suggestions, enhancing users' agency and creativity throughout the songwriting process.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
Andrea Agostinelli, Timo I. Denk, Zalán Borsos, Jesse Engel, Mauro Verzetti, An- toine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, Matt Sharifi, Neil Zeghidour, and Christian Frank. 2023. MusicLM: Generating Music From Text. arXiv:2301.11325 [cs.SD] https://arxiv.org/abs/2301.11325
arXiv 2023
-
[2]
Music AI. 2024. Music AI: AI Audio Models to Power Your Music Business. https://www.music.ai
2024
-
[3]
Suno AI. 2024. Suno AI. https://suno.com/
2024
-
[4]
Philip Alperson. 1984. On Musical Improvisation. The Journal of Aesthetics and Art Criticism 43, 1 (1984), 17–29. http://www.jstor.org/stable/430189
1984
-
[5]
Barrett R Anderson, Jash Hemant Shah, and Max Kreminski. 2024. Homog- enization Effects of Large Language Models on Human Creative Ideation. In Proceedings of the 16th Conference on Creativity & Cognition (Chicago, IL, USA) (C&C ’24). Association for Computing Machinery, New York, NY, USA, 413–425. doi:10.1145/3635636.3656204
arXiv 2024
-
[6]
Aria. 2024. Introducing Aria - Your personal AI co-creator for chords and melody. https://www.hooktheory.com/hookpad/aria
2024
-
[7]
NIC BECKER, RYAN LOUIE, JOHN THICKSTUN, and PERCY LIANG. 2024. Designing Live Human-AI Collaboration for Musical Improvisation
2024
-
[8]
Christodoulos Benetatos, Joseph VanderStel, and Zhiyao Duan. 2020. BachDuet: A Deep Learning System for Human-Machine Counterpoint Improvisation. In Proceedings of the International Conference on New Interfaces for Musical Expression, Romain Michon and Franziska Schroeder (Eds.). Birmingham City University, Birmingham, UK, 635–640. doi:10.5281/zenodo.4813234
Show all 131 references
-
[9]
Eden Bensaid, Mauro Martino, Benjamin Hoover, and Hendrik Strobelt
-
[10]
Renaud Bougueng Tchemeube, Jeffrey John Ens, and Philippe Pasquier. 2022. Calliope: A Co-creative Interface for Multi-Track Music Generation. In Proceed- ings of the 14th Conference on Creativity and Cognition (Venice, Italy) (C&C ’22). Association for Computing Machinery, New...
2022
-
[11]
Daniel Buschek, Martin Zürn, and Malin Eiband. 2021. The Impact of Multiple Parallel Phrase Suggestions on Email Input and Composition Behaviour of Na- tive and Non-Native English Writers. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokoham...
2021
-
[12]
Runze Cai, Nuwan Janaka, Yang Chen, Lucia Wang, Shengdong Zhao, and Can Liu. 2024. PANDALens: Towards AI-Assisted In-Context Writing on OHMD Dur- ing Travels. In Proceedings of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’24). Association ...
2024
-
[13]
Chang, Mihail Eric, Manolis Savva, and Christopher D
Angel X. Chang, Mihail Eric, Manolis Savva, and Christopher D. Manning. 2017. SceneSeer: 3D Scene Design with Natural Language. arXiv:1703.00050 [cs.GR] https://arxiv.org/abs/1703.00050
2017 arXiv
-
[14]
Siddhartha Chaudhuri, Evangelos Kalogerakis, Stephen Giguere, and Thomas Funkhouser. 2013. Attribit: content creation with semantic attributes. InProceed- ings of the 26th Annual ACM Symposium on User Interface Software and Technol- ogy (St. Andrews, Scotland, United Kingdom) ...
2013
-
[15]
Ke Chen, Xingjian Du, Bilei Zhu, Zejun Ma, Taylor Berg-Kirkpatrick, and Shlomo Dubnov. 2022. HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection. arXiv:2202.00874 [cs.SD] https://arxiv.org/ abs/2202.00874
2022 arXiv
-
[16]
Erin Cherry and Celine Latulipe. 2014. Quantifying the Creativity Support of Digital Tools through the Creativity Support Index. ACM Trans. Comput.-Hum. Interact. 21, 4, Article 21 (jun 2014), 25 pages. doi:10.1145/2617588
2014 doi
-
[17]
DaEun Choi, Sumin Hong, Jeongeon Park, John Joon Young Chung, and Juho Kim. 2024. CreativeConnect: Supporting Reference Recombination for Graphic Design Ideation with Generative AI. InProceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, U...
2024
-
[18]
John Joon Young Chung, Wooseok Kim, Kang Min Yoo, Hwaran Lee, Eytan Adar, and Minsuk Chang. 2022. TaleBrush: Sketching Stories with Generative Pretrained Language Models. InProceedings of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CHI...
2022
-
[19]
Elizabeth Clark, Anne Spencer Ross, Chenhao Tan, Yangfeng Ji, and Noah A. Smith. 2018. Creative Writing with a Machine in the Loop: Case Studies on Slo- gans and Stories. In Proceedings of the 23rd International Conference on Intelligent User Interfaces (Tokyo, Japan) (IUI ’18...
2018
-
[20]
WJ Conover. 1999. Practical nonparametric statistics. John Wiley & Sons, Inc, New York, NY, USA
1999
-
[21]
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez. 2024. Simple and Controllable Music Generation. arXiv:2306.05284 [cs.SD] https://arxiv.org/abs/2306.05284
2024 arXiv
-
[22]
Juliet Corbin and Anselm Strauss. 2008. Basics of Qualitative Research (3rd ed.): Techniques and Procedures for Developing Grounded Theory. doi:10.4135/ 9781452230153
2008
-
[23]
Bob Coyne and Richard Sproat. 2001. WordsEye: an automatic text-to-scene conversion system. In Proceedings of the 28th Annual Conference on Computer Graphics and Interactive Techniques (SIGGRAPH ’01). Association for Computing Machinery, New York, NY, USA, 487–496. doi:10.1145...
2001
-
[24]
Mihaly Csikszentmihalyi. 1997. Flow and Creativity. NAMTA Journal 22, 2 (1997), 60–97. https://eric.ed.gov/?id=EJ547968
1997
-
[25]
Mihaly Csikszentmihalyi. 1997. Flow and the psychology of discovery and invention. HarperPerennial, New York 39 (1997), 1–16
1997
-
[26]
Hai Dang, Sven Goller, Florian Lehmann, and Daniel Buschek. 2023. Choice Over Control: How Users Write with Large Language Models using Diegetic and Non-Diegetic Prompting. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI ...
2023
-
[27]
TheoryTab DB. 2024. TheoryTab DB: Tabs that show the theory behind songs. https://www.hooktheory.com/theorytab
2024
-
[28]
Prafulla Dhariwal, Heewoo Jun, Christine Payne, Jong Wook Kim, Alec Rad- ford, and Ilya Sutskever. 2020. Jukebox: A Generative Model for Music. arXiv:2005.00341 [eess.AS] https://arxiv.org/abs/2005.00341
2020 arXiv
-
[29]
Chris Donahue, Antoine Caillon, Adam Roberts, Ethan Manilow, Philippe Esling, Andrea Agostinelli, Mauro Verzetti, Ian Simon, Olivier Pietquin, Neil Zeghidour, and Jesse Engel. 2023. SingSong: Generating musical accompaniments from singing. arXiv:2301.12662 [cs.SD] https://arxi...
2023 arXiv
-
[30]
Cottrell, and Julian McAuley
Chris Donahue, Huanru Henry Mao, Yiting Ethan Li, Garrison W. Cottrell, and Julian McAuley. 2019. LakhNES: Improving multi-instrumental music generation with cross-domain pre-training. arXiv:1907.04868 [cs.SD] https: //arxiv.org/abs/1907.04868
2019 arXiv
-
[31]
Chris Donahue, John Thickstun, and Percy Liang. 2022. Melody transcription via generative pre-training. arXiv:2212.01884 [cs.SD] https://arxiv.org/abs/2212. 01884
2022 arXiv
-
[32]
Chris Donahue, Shih-Lun Wu, Yewon Kim, Dave Carlton, Ryan Miyakawa, and John Thickstun. 2025. Hookpad Aria: A Copilot for Songwriters. arXiv:2502.08122 [cs.SD] https://arxiv.org/abs/2502.08122
2025 arXiv
-
[33]
Doshi and Oliver P
Anil R. Doshi and Oliver P. Hauser. 2024. Generative artificial intel- ligence enhances creativity but reduces the diversity of novel content. arXiv:2312.00506 [cs.HC] https://arxiv.org/abs/2312.00506
2024 arXiv
-
[34]
Alaaeldin El-Nouby, Shikhar Sharma, Hannes Schulz, Devon Hjelm, Layla El Asri, Samira Ebrahimi Kahou, Yoshua Bengio, and Graham W. Taylor. 2019. Tell, Draw, and Repeat: Generating and Modifying Images Based on Continual Linguistic Instruction. arXiv:1811.09845 [cs.CV] https://...
2019 arXiv
-
[35]
Finke, Thomas B
Ronald A. Finke, Thomas B. Ward, and Steven M. Smith. 1992.Creative Cognition: Theory, Research, and Applications. The MIT Press, New York, NY, USA. doi:10. 7551/mitpress/7722.001.0001
1992
-
[36]
Seth* Forsgren and Hayk* Martiros. 2022. Riffusion - Stable diffusion for real- time music generation. https://riffusion.com/about
2022
-
[37]
Satoru Fukayama, Kazuyoshi Yoshii, and Masataka Goto. 2013. Chord-Sequence- Factory: A Chord Arrangement System Modifying Factorized Chord Sequence Probabilities. https://api.semanticscholar.org/CorpusID:1099764 CHI ’25, April 26-May 1, 2025, Yokohama, Japan Yewon Kim, Sung-Ju...
2013
-
[38]
Katy Ilonka Gero and Lydia B. Chilton. 2019. Metaphoria: An Algorithmic Companion for Metaphor Creation. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland Uk) (CHI ’19) . Association for Computing Machinery, New York, NY, USA, 1...
2019
-
[39]
Barney Glaser and Anselm Strauss. 1999. Discovery of Grounded Theory: Strategies for Qualitative Research (1st ed.). Routledge, Abingdon, Oxon, UK. doi:10.4324/9780203793206
1999 doi
-
[40]
Gaëtan Hadjeres and Léopold Crestel. 2021. The Piano Inpainting Application. arXiv:2107.05944 [cs.SD] https://arxiv.org/abs/2107.05944
2021 arXiv
-
[41]
Gaëtan Hadjeres, François Pachet, and Frank Nielsen. 2017. DeepBach: a Steer- able Model for Bach Chorales Generation. InProceedings of the 34th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 70), Doina Precup and Yee Whye Teh (Eds...
2017
-
[42]
Hookpad. 2024. Hookpad Songwriting Software: Create Amazing Music. https: //www.hooktheory.com/hookpad
2024
-
[43]
Hooktheory. 2024. Hooktheory: Create amazing music. https://www. hooktheory.com/
2024
-
[44]
Eric Horvitz. 1999. Principles of mixed-initiative user interfaces. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Pittsburgh, Pennsylvania, USA) (CHI ’99). Association for Computing Machinery, New York, NY, USA, 159–166. doi:10.1145/302979.303030
1999
-
[45]
Cheng-Zhi Anna Huang, Tim Cooijmans, Adam Roberts, Aaron Courville, and Douglas Eck. 2019. Counterpoint by Convolution. arXiv:1903.07227 [cs.LG] https://arxiv.org/abs/1903.07227
2019 arXiv
-
[46]
Cheng-Zhi Anna Huang, David Duvenaud, and Krzysztof Z. Gajos. 2016. Chor- dRipple: Recommending Chords to Help Novice Composers Go Beyond the Ordinary. In Proceedings of the 21st International Conference on Intelligent User Interfaces (Sonoma, California, USA) (IUI ’16). Assoc...
2016
-
[47]
Cheng-Zhi Anna Huang, Curtis Hawthorne, Adam Roberts, Monica Din- culescu, James Wexler, Leon Hong, and Jacob Howcroft. 2019. The Bach Doodle: Approachable music composition with machine learning at scale. arXiv:1907.06637 [cs.SD] https://arxiv.org/abs/1907.06637
2019 arXiv
-
[48]
Cheng-Zhi Anna Huang, Hendrik Vincent Koops, Ed Newton-Rex, Monica Dinculescu, and Carrie J. Cai. 2020. AI Song Contest: Human-AI Co-Creation in Songwriting. arXiv:2010.05388 [cs.SD] https://arxiv.org/abs/2010.05388
2020 arXiv
-
[49]
Dai, Matthew D
Cheng-Zhi Anna Huang, Ashish Vaswani, Jakob Uszkoreit, Noam Shazeer, Ian Simon, Curtis Hawthorne, Andrew M. Dai, Matthew D. Hoffman, Monica Din- culescu, and Douglas Eck. 2018. Music Transformer. arXiv:1809.04281 [cs.LG] https://arxiv.org/abs/1809.04281
2018 arXiv
-
[50]
Park, Tao Wang, Timo I
Qingqing Huang, Daniel S. Park, Tao Wang, Timo I. Denk, Andy Ly, Nanxin Chen, Zhengdong Zhang, Zhishuai Zhang, Jiahui Yu, Christian Frank, Jesse Engel, Quoc V. Le, William Chan, Zhifeng Chen, and Wei Han. 2023. Noise2Music: Text-conditioned Music Generation with Diffusion Mode...
2023 arXiv
-
[51]
Purnima Kamath, Fabio Morreale, Priambudi Lintang Bagaskara, Yize Wei, and Suranga Nanayakkara. 2024. Sound Designer-Generative AI Interactions: Towards Designing Creative Support Tools for Professional Sound Designers. In Proceedings of the 2024 CHI Conference on Human Factor...
2024
-
[52]
Hyeongcheol Kim, Shengdong Zhao, Can Liu, and Kotaro Hara. 2020. LiveS- nippets: Voice-based Live Authoring of Multimedia Articles about Experi- ences. In 22nd International Conference on Human-Computer Interaction with Mobile Devices and Services (Oldenburg, Germany) (MobileH...
2020
-
[53]
Miller, and Theresa Claire
Ziva Kunda, Dale T. Miller, and Theresa Claire. 1990. Combining Social Concepts: The Role of Causal Reasoning. Cognitive Science 14, 4 (1990), 551–577
1990
-
[54]
Tomas Lawton, Kazjon Grace, and Francisco J Ibarrola. 2023. When is a Tool a Tool? User Perceptions of System Agency in Human–AI Co-Creative Drawing. In Proceedings of the 2023 ACM Designing Interactive Systems Conference(Pittsburgh, PA, USA) (DIS ’23). Association for Computi...
2023
-
[55]
Alghamdi, Tal August, Avinash Bhat, Madiha Zahrah Choksi, Senjuti Dutta, Jin L.C
Mina Lee, Katy Ilonka Gero, John Joon Young Chung, Simon Buckingham Shum, Vipul Raheja, Hua Shen, Subhashini Venugopalan, Thiemo Wambsganss, David Zhou, Emad A. Alghamdi, Tal August, Avinash Bhat, Madiha Zahrah Choksi, Senjuti Dutta, Jin L.C. Guo, Md Naimul Hoque, Yewon Kim, S...
2024
-
[56]
Mina Lee, Percy Liang, and Qian Yang. 2022. CoAuthor: Designing a Human-AI Collaborative Writing Dataset for Exploring Language Model Capabilities. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems (New Orleans, LA, USA) (CHI ’22). Association for...
2022
-
[57]
Tuck Wah Leong, Frank Vetere, and Steve Howard. 2006. Randomness as a resource for design. In Proceedings of the 6th Conference on Designing Interac- tive Systems (University Park, PA, USA) (DIS ’06). Association for Computing Machinery, New York, NY, USA, 132–139. doi:10.1145...
2006 arXiv
-
[58]
J. Lin. 1991. Divergence measures based on the Shannon entropy. IEEE Transac- tions on Information Theory 37, 1 (1991), 145–151. doi:10.1109/18.61115
1991 doi
-
[59]
Vivian Liu, Han Qiao, and Lydia Chilton. 2022. Opal: Multimodal Image Gener- ation for News Illustration. In Proceedings of the 35th Annual ACM Symposium on User Interface Software and Technology (Bend, OR, USA) (UIST ’22). Asso- ciation for Computing Machinery, New York, NY, ...
2022
-
[60]
Vivian Liu, Jo Vermeulen, George Fitzmaurice, and Justin Matejka
-
[61]
Ryan Louie, Andy Coenen, Cheng Zhi Huang, Michael Terry, and Carrie J. Cai
-
[62]
Ryan Louie, Jesse Engel, and Cheng-Zhi Anna Huang. 2022. Expressive Commu- nication: Evaluating Developments in Generative Models and Steering Interfaces for Music Creation. In Proceedings of the 27th International Conference on Intel- ligent User Interfaces (Helsinki, Finland...
2022
-
[63]
Malandro
Martin E. Malandro. 2023. Composer’s Assistant: An Interactive Transformer for Multi-Track MIDI Infilling. arXiv:2301.12525 [cs.SD] https://arxiv.org/abs/ 2301.12525
2023 arXiv
-
[64]
Enrique Manjavacas, Folgert Karsdorp, Ben Burtenshaw, and Mike Kestemont
-
[65]
Lidia Morris, Rebecca Leger, Michele Newman, John Ashley Burgoyne, Ryan Groves, Natasha Mangal, and Jin Ha Lee. 2024. Human-AI Music Process: A Dataset of AI-Supported Songwriting Processes from the AI Song Contest
2024
-
[66]
Radford M. Neal. 2003. Slice sampling. The Annals of Statistics 31, 3 (2003), 705 –
2003
-
[67]
Michele Newman, Lidia Morris, and Jin Ha Lee 0001. 2023. Human-AI Music Creation: Understanding the Perceptions and Experiences of Music Creators for Ethical and Productive Collaboration. 80-88 pages. doi:10.5281/zenodo.10265227
2023 doi
-
[68]
D.A. Norman. 2013. The Design of Everyday Things . MIT Press, Cambridge, MA, USA. https://books.google.co.kr/books?id=heCtnQEACAAJ
2013
-
[69]
Richard L. Oliver. 1980. A Cognitive Model of the Antecedents and Consequences of Satisfaction Decisions. Journal of Marketing Research 17, 4 (1980), 460–469. http://www.jstor.org/stable/3150499
1980
-
[70]
OpenAI. 2024. GPT-4o System Card. https://cdn.openai.com/gpt-4o-system- card.pdf
2024
-
[71]
OpenAI. 2024. Hello GPT-4o. https://openai.com/index/hello-gpt-4o/
2024
-
[72]
OpenAI. 2024. OpenAI Platform. https://platform.openai.com/docs/overview
2024
-
[73]
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and...
-
[74]
Vishakh Padmakumar and He He. 2024. Does Writing with Language Models Reduce Content Diversity? https://openreview.net/forum?id=Feiz5HtCD0
2024
-
[75]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. BLEU: a method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting on Association for Computational Linguistics (Philadel- phia, Pennsylvania) (ACL ’02). Association for C...
2002
-
[76]
Christine McLeavy Payne. 2019. MuseNet. https://openai.com/blog/musenet/
2019
-
[77]
Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. 2023. SDXL: Improving Latent Dif- fusion Models for High-Resolution Image Synthesis. arXiv:2307.01952 [cs.CV] https://arxiv.org/abs/2307.01952
2023 arXiv
-
[78]
Prolific. 2024. Prolific | Quickly find research participants you can trust. https: //www.prolific.com/
2024
-
[79]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Amuse: Human-AI Collaborative Songwriting with Multimodal Inspirations CHI ’25, April 26-May 1, 2025, Yokohama, Japan Gretchen Kr...
2025 arXiv
-
[80]
Reddit. 2024. Songwriting. https://www.reddit.com/r/Songwriting/
2024
-
[81]
Mitchel Resnick, Brad Myers, Kumiyo Nakakoji, Ben Shneiderman, Randy Pausch, Ted Selker, and Mike Eisenberg. 2018. Design Principles for Tools to Support Creative Thinking. doi:10.1184/R1/6621917.v1
2018 doi
-
[82]
Flavio Schneider, Ojasv Kamal, Zhijing Jin, and Bernhard Schölkopf. 2023. Moûsai: Text-to-Music Generation with Long-Context Latent Diffusion. arXiv:2301.11757 [cs.CL] https://arxiv.org/abs/2301.11757
2023 arXiv
-
[83]
Shikhar Sharma, Dendi Suhubdy, Vincent Michalski, Samira Ebrahimi Kahou, and Yoshua Bengio. 2018. ChatPainter: Improving Text to Image Generation using Dialogue. arXiv:1802.08216 [cs.CV] https://arxiv.org/abs/1802.08216
2018 arXiv
-
[84]
Ben Shneiderman. 2022. Human-Centered AI. Oxford University Press, Oxford, United Kingdom. https://global.oup.com/academic/product/human-centered- ai-9780192845290
2022
-
[85]
2016.Designing the User Interface: Strategies for Effective Human-Computer Interaction (6th ed.)
Ben Shneiderman, Catherine Plaisant, Maxine Cohen, Steven Jacobs, Niklas Elmqvist, and Nicholas Diakopoulos. 2016.Designing the User Interface: Strategies for Effective Human-Computer Interaction (6th ed.). Pearson, London, UK
2016
-
[86]
Ian Simon and Sageev Oore. 2017. Performance RNN: Generating Music with Expressive Timing and Dynamics. https://magenta.tensorflow.org/performance- rnn
2017
-
[87]
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman. 2022. Make-A-Video: Text-to-Video Generation without Text-Video Data. arXiv:2209.14792 [cs.CV] https://arxiv.or...
2022 arXiv
-
[88]
Glassman
Nikhil Singh, Guillermo Bernal, Daria Savchenko, and Elena L. Glassman. 2023. Where to Hide a Stolen Elephant: Leaps in Creative Writing with Multimodal Machine Intelligence. ACM Trans. Comput.-Hum. Interact. 30, 5, Article 68 (sep 2023), 57 pages. doi:10.1145/3511599
2023 doi
-
[89]
Minhyang (Mia) Suh, Emily Youngblom, Michael Terry, and Carrie J Cai. 2021. AI as Social Glue: Uncovering the Roles of Deep Generative AI during Social Music Composition. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI ’21...
2021 doi
-
[90]
suno.wiki. 2024. Style and Genre List. https://www.suno.wiki/faq/style-and- lyrics/styles-and-genres/
2024
-
[91]
Ultimate Guitar Tabs. 2024. Misty Chords by Ella Fitzgerald. https://tabs. ultimate-guitar.com/tab/ella-fitzgerald/misty-chords-1206796
2024
-
[92]
John Thickstun, David Leo Wright Hall, Chris Donahue, and Percy Liang. 2024. Anticipatory Music Transformer. https://openreview.net/forum?id=EBNJ33Fcrl
2024
-
[93]
Udio. 2024. Udio | AI Music Generator - Official Website. https://udio.com/
2024
-
[94]
Unison. 2024. 7 Sad Chord Progressions That Instantly Get People Invested. https://unison.audio/sad-chord-progressions/
2024
-
[95]
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2019. Representation Learning with Contrastive Predictive Coding. arXiv:1807.03748 [cs.LG] https://arxiv.org/ abs/1807.03748
2019 arXiv
-
[96]
Ashley Walton, Auriel Washburn, Peter Langland-Hassan, Anthony Chemero, Heidi Kloos, and Michael Richardson. 2017. Creating Time: Social Collaboration in Music Improvisation. Topics in Cognitive Science 10 (11 2017). doi:10.1111/ tops.12306
2017
-
[97]
Bryan Wang, Yuliang Li, Zhaoyang Lv, Haijun Xia, Yan Xu, and Raj Sodhi. 2024. LAVE: LLM-Powered Agent Assistance and Language Augmentation for Video Editing. In Proceedings of the 29th International Conference on Intelligent User Interfaces (Greenville, SC, USA) (IUI ’24). Ass...
2024
-
[98]
Zihao Wang, Kejun Zhang, Yuxing Wang, Chen Zhang, Qihao Liang, Pengfei Yu, Yongsheng Feng, Wenbo Liu, Yikai Wang, Yuntao Bao, and Yiheng Yang
-
[99]
Thomas B Ward, Ronald A Finke, and Steven M Smith. 1995. Creativity and the mind: Discovering the genius within (1 ed.). Springer, New York, NY, USA. IX, 274 pages. doi:10.1007/978-1-4899-3330-0
1995 doi
-
[100]
Ward and E
Thomas B. Ward and E. Thomas Lawson. 2009. Creative Cognition in Science Fiction and Fantasy Writing. Cambridge University Press, Cambridge, 196–210
2009
-
[101]
Shih-Lun Wu, Chris Donahue, Shinji Watanabe, and Nicholas J. Bryan. 2024. Music ControlNet: Multiple Time-Varying Controls for Music Generation. IEEE/ACM Trans. Audio, Speech and Lang. Proc. 32 (May 2024), 2692–2703. doi:10.1109/TASLP.2024.3399026
2024
-
[102]
Tongshuang Wu, Michael Terry, and Carrie Jun Cai. 2022. AI Chains: Transpar- ent and Controllable Human-AI Interaction by Chaining Large Language Model Prompts. In Proceedings of the 2022 CHI Conference on Human Factors in Comput- ing Systems (New Orleans, LA, USA) (CHI ’22). ...
2022
-
[103]
Yusong Wu, Ke Chen, Tianyu Zhang, Yuchen Hui, Taylor Berg-Kirkpatrick, and Shlomo Dubnov. 2023. Large-Scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation. 5 pages. doi:10. 1109/ICASSP49357.2023.10095969
2023
-
[104]
Yusong Wu, Tim Cooijmans, Kyle Kastner, Adam Roberts, Ian Simon, Alexander Scarlatos, Chris Donahue, Cassie Tarakajian, Shayegan Omidshafiei, Aaron Courville, Pablo Samuel Castro, Natasha Jaques, and Cheng-Zhi Anna Huang
-
[105]
In Proceedings of the 30th ACM International Con- ference on Multimedia (Lisboa, Portugal) (MM ’22)
SongDriver: Real-time Music Accompaniment Generation without Logical Latency nor Exposure Bias. In Proceedings of the 30th ACM International Con- ference on Multimedia (Lisboa, Portugal) (MM ’22). Association for Computing Machinery, New York, NY, USA, 1057–1067. doi:10.1145/3...
-
[106]
Sihyun Yu, Weili Nie, De-An Huang, Boyi Li, Jinwoo Shin, and Anima Anandku- mar. 2024. Efficient Video Diffusion Models via Content-Frame Motion-Latent Decomposition. arXiv:2403.14148 [cs.CV] https://arxiv.org/abs/2403.14148
2024 arXiv
-
[107]
Yiming Zhang, Avi Schwarzschild, Nicholas Carlini, Zico Kolter, and Daphne Ippolito. 2024. Forcing Diffuse Distributions out of Language Models. arXiv:2404.10859 [cs.CL]
2024 arXiv
-
[108]
Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu. 2018. Texygen: A Benchmarking Platform for Text Generation Models. In The 41st International ACM SIGIR Conference on Research & Development in In- formation Retrieval (Ann Arbor, MI, USA)(SIGIR ...
2018
-
[113]
Zihan Yan, Chunxu Yang, Qihao Liang, and Xiang ’Anthony’ Chen. 2023. XCre- ation: A Graph-based Crossmodal Generative Creativity Support Tool. In Pro- ceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (San Francisco, CA, USA) (UIST ’23). Assoc...
2023
-
[117]
For text, identify the main themes, emotions, or ideas
Analyze the inputs: For an image, identify visual elements that suggest musical themes and emotions. For text, identify the main themes, emotions, or ideas. For user-written keywords, expand upon them with related music styles, genres, and types
-
[118]
Generate keywords: Use the provided keyword list, but feel free to create new, relevant keywords if they better capture the input
-
[119]
Format the output: Provide the keywords in a comma-separated list without any additional text or formatting. Keyword List: Style: dance, festive, groovy, mid-tempo, syncopated, tipsy, atmospheric, cold, dark, doom, dramatic, sinister, adjunct, art, capriccio, mellifluous, nü, ...
-
[122]
Each progression should be unique and align with the bar parameter (i.e., if bars = 4, each progression should have 4 chords)
Generate 30 Chord Progressions: Create 30 distinct chord progressions that fit the specified key and mode and match the keywords. Each progression should be unique and align with the bar parameter (i.e., if bars = 4, each progression should have 4 chords). Ensure diversity by ...
-
[129]
Ensure the chord progressions are musically coherent, stylistically appropriate, and diverse
Slash Chords: Alternate bass notes such as /E, /G#, /Bb, /Dx. Ensure the chord progressions are musically coherent, stylistically appropriate, and diverse. Include extensions, suspensions, adds, altered notes, slash chords as needed to achieve maximum diversity. Use both diato...
2025
-
[130]
Tonic (I, vi) provides resolution and stability
Analyze Chord Functions: Determine the functions of chords in the given key and mode. Tonic (I, vi) provides resolution and stability. Subdominant (IV, ii) creates movement away from the tonic. Dominant (V, vii°) creates tension that needs to resolve to the tonic
-
[131]
For example, for jazz-related keywords, consider using seventh chords, altered chords, and common jazz progressions like ii-V-I
Analyze the Keywords: Determine the chord components and progression patterns based on the keywords. For example, for jazz-related keywords, consider using seventh chords, altered chords, and common jazz progressions like ii-V-I. For keywords like ’sadness’ or ’emotional,’ use...
-
[132]
The progression should align with the bar parameter (i.e., if bars = 4, the progression should have 4 chords)
Generate a Chord Progression: Create a chord progression that fits the specified key and mode and matches the keywords. The progression should align with the bar parameter (i.e., if bars = 4, the progression should have 4 chords). Each chord text can have the following compone...
-
[133]
Root Note: A-G, with optional accidentals (#, b, x)
-
[134]
Chord Quality: maj, min, aug, dim
-
[135]
Extensions: Specific chord extensions such as 6/9, 7, 9, 11, 13
-
[136]
Suspended Chords: Suspended chords such as sus2, sus4, sus#2, sus#4
-
[137]
Added Notes: Added notes such as add2, add4, add6, add9, add11, add13
-
[138]
Altered Notes: Alterations such as b5, #5, b9, #9, #11, b13
-
[139]
Keyword Relevance
Slash Chords: Alternate bass notes such as /E, /G#, /Bb, /Dx. Ensure the chord progression is musically coherent and stylistically appropriate. Include extensions, suspensions, adds, altered notes, and slash chords as needed to achieve a rich and satisfying progression. Use bo...
2025
-
[767]
doi:10.1214/aos/1056562461
-
[2017]
Synthetic Literature: Writing Science Fiction in a Co-Creative Process. In Proceedings of the Workshop on Computational Creativity in Natural Language Generation (CC-NLG 2017) , Hugo Gonçalo Oliveira, Ben Burtenshaw, Mike Kestemont, and Tom De Smedt (Eds.). Association for Com...
2017 doi
-
[2020]
In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’20)
Novice-AI Music Co-Creation via AI-Steering Tools for Deep Generative Models. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI ’20). Association for Computing Machinery, New York, NY, USA, 1–13. doi:10.1145/3313831.3376739
2020
-
[2021]
arXiv:2108.04324 [cs.CL] https://arxiv.org/abs/2108.04324
FairyTailor: A Multimodal Generative Framework for Storytelling. arXiv:2108.04324 [cs.CL] https://arxiv.org/abs/2108.04324
-
[2022]
arXiv:2203.02155 [cs.CL] https://arxiv.org/abs/2203.02155
Training language models to follow instructions with human feedback. arXiv:2203.02155 [cs.CL] https://arxiv.org/abs/2203.02155
-
[2023]
arXiv:2210.11603 [cs.HC] https://arxiv.org/abs/2210.11603
3DALL-E: Integrating Text-to-Image AI in 3D Design Workflows. arXiv:2210.11603 [cs.HC] https://arxiv.org/abs/2210.11603
-
[2024]
In Proceedings of the 41st International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol
Adaptive Accompaniment with ReaLchords. In Proceedings of the 41st International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 235), Ruslan Salakhutdinov, Zico Kolter, Katherine Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix ...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.