Pith. sign in

REVIEW 4 major objections 5 minor 136 references

Workflow-Based Evaluation of Music Generation Systems

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Current open-source music generation systems work best as human helpers, not autonomous composers, a hands-on test of eight systems suggests.

desk verdict A useful framework paper whose central coherence claim outruns the evidence; worth peer review with requests for revision. read the letter →

arxiv 2507.01022 v1 pith:UJ47MINT submitted 2025-06-11 eess.AS cs.HCcs.LGcs.MMcs.SD

classification eess.AScs.HCcs.LGcs.MMcs.SD
keywords musicgenerationsystemsworkflow-basedevaluationhuman-AIco-creationopen-sourceAIframeworkgenerate-then-curatepromptengineeringproduction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that current open-source music generation systems (MGS), tested in a realistic, iterative production workflow, work best as complementary tools that enhance human creativity rather than as autonomous composers. Using a two-phase evaluation framework covering eight systems across composition, arrangement, and sound design, the study finds that these systems produce usable short segments and accelerate ideation, but they struggle to maintain thematic and structural coherence across full pieces, to follow precise musical instructions, and to integrate cleanly into a DAW-based workflow. The paper matters because it shifts evaluation away from isolated technical metrics toward workflow-based, generate-then-curate testing, and it proposes specific new criteria—Serendipity Support, AI Assistance Balance, Adaptation Capacity—for judging creative collaboration. The empirical basis is a single expert evaluator, so the findings are framed as an exploratory phase that establishes hypotheses for larger multi-evaluator studies rather than definitive rankings.

What carries the argument

The carrying mechanism is the evaluation framework itself, built as two aligned phases. Phase 1 (System Overview) applies descriptive 'system-level' criteria—architecture, interfaces, checkpoints, hardware—to establish what each system can do; Phase 2 (Hands-on Experimentation) applies eight 'performance' criteria (Usability, Generation Speed, Audio Quality, Stylistic Accuracy, Parameter Control, Content Generation Control, DAW Compatibility, Creative Control) scored from 1 to 5 with a standardized rubric, while the evaluator works through a cyclic generate-then-curate process that mirrors the non-linear nature of real music production. The framework's key move is to treat evaluation as workflow-embedded rather than output-isolated: the final musical track, assembled from each system's contributions, is the test, and the rubric plus qualitative notes reveal where coherence, control, and integration break down.

What would settle it

A controlled multi-evaluator study in which independent professional producers score the same eight systems with the same 1–5 rubric; if their rankings diverge substantially from the paper's Table 6, or if they do not find systematic loss of thematic and structural coherence in full-track generation, the central claim would be contradicted.

Watch

Extended reading notes

Core claim

The paper's central claim is that MGS function primarily as complementary tools in music creation, enhancing rather than replacing human expertise. Concretely, eight open-source systems—spanning symbolic and audio generation, from MusicGen and Riffusion to Magenta Studio 2.0 and DDSP-VST—were taken through a full composition-and-curation cycle to produce one track; the outputs were strongest in atomic tasks such as motif generation, sample collection, and timbral transformation, and weakest in holistic composition requiring hierarchical structural planning. The study documents systematic limitations: compositions start and end abruptly, prompts for specific instruments or grooves are frequently mismatched, isolated stems are difficult to obtain, and generation latency breaks creative flow. From this the paper concludes that human creativity remains indispensable for tasks demanding emotional depth and complex decision-making, and that the systems are best understood as catalysts and collaborators inside human-led workflows.

Load-bearing premise

All 1–5 performance scores and qualitative judgments come from a single evaluator—the first author—who also designed the criteria, so the findings stand or fall on whether those judgments represent music producers more broadly.

Editorial extensions

If this is right

  • If the central claim holds, MGS should be designed and evaluated as assistive components within human-led workflows, not as autonomous composers.
  • The effectiveness of these systems is task-dependent: strongest in atomic tasks like motif generation or timbre transformation and weakest in holistic composition, so tool design should target those strengths.
  • Prompt formulation and training-data annotation misalignment are core bottlenecks, motivating interfaces that assist prompt authoring and models with reliable instruction-following.
  • Source separation and post-processing become obligatory companions to generation, making DAW integration and stem-level control a priority for production-ready tools.
  • The three proposed new performance criteria—Serendipity Support, AI Assistance Balance, Adaptation Capacity—offer a concrete target for subsequent validation in multi-evaluator studies.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's conclusion that 'almost coherent' outputs create productive creative tension suggests a testable design principle: systems that are imperfect in structure may stimulate curation and artistic ownership, a hypothesis the paper states but does not itself test.
  • The workflow-based evaluation logic could transfer directly to other generative creative tools (image, video, text), where generate-then-curate cycles are equally dominant and where the same coherence problems appear at longer horizons.
  • Because the single evaluator is an AI music researcher and guitarist, a replication with independent professional producers using the same rubric would clarify how much of the observed ranking is system-dependent versus taste-driven.
  • The framework's professional-production standard for audio quality may systematically underrate genres that embrace artifacts and lo-fi textures; an aesthetics-inclusive scoring variant would likely shift the ranking of systems like Riffusion.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This paper proposes a workflow-based, mixed-method evaluation framework for eight open-source music generation systems (MGS), combining a system-overview phase with hands-on experimentation scored on 1–5 performance criteria. Based on observations in a single home studio and a one-evaluator methodology, it concludes that MGS serve as complementary tools rather than replacements for human expertise and that they exhibit notable limitations in maintaining thematic and structural coherence. The paper also proposes three additional evaluation dimensions—Serendipity Support, AI Assistance Balance, and Adaptation Capacity—and discusses integration challenges, prompt-engineering difficulties, and design implications for human-AI co-creation.

Significance. If the central claim were fully supported, it would have practical design implications: open-source MGS should be evaluated and developed as assistive components within human-led workflows rather than as autonomous composers. The paper has clear strengths: it documents the eight systems at a useful level of architectural and interface detail, grounds its criteria in prior HCI and co-creation literature, candidly discloses its single-evaluator limitation in Section 8, and aligns its qualitative observations with external producer reports in Section 7.5. The framework and proposed criteria are useful scaffolding for future multi-evaluator studies. However, the empirical support for the strong general claim is currently thin: the quantitative scores come from one non-independent evaluator, there is no direct coherence criterion in Table 6, and the two symbolic systems most relevant to long-form structure were not hands-on tested.

major comments (4)
  1. [Section 5, Table 6] The quantitative evaluation that is presented as the outcome of the hands-on phase does not contain a criterion that directly measures thematic or structural coherence. Table 6 lists eight criteria (Usability, Generation Speed, Audio Quality, Stylistic Accuracy, Parameter Control, Content Generation Control, DAW Compatibility, Creative Control), and none of them target coherence over the duration of a composition. The conclusion in Section 9 and the abstract that MGS 'exhibit notable limitations in maintaining thematic and structural coherence' therefore rests entirely on qualitative notes from a single evaluator (Sections 5.2.1 and 6). To make the central claim load-bearing, the authors should either add an explicit coherence criterion with a scoring rubric, or rephrase the claim as a hypothesis generated by the exploratory observations.
  2. [Section 5 and Section 4.4] MuseFormer and MuseCoco are excluded from the hands-on experimentation 'due to persistent technical impediments regarding local inference execution and the absence of accessible web-based alternatives.' Yet Section 4.4 credits MuseFormer with 'structural awareness' and 'thematic consistency,' and these are exactly the systems whose design targets long-form structural coherence. The general conclusion that MGS exhibit limitations in structural coherence is therefore not tested on the two systems most likely to contradict it. The claim should be restricted to the six systems actually tested, or the authors should obtain at least observational evidence for MuseFormer and MuseCoco through other means.
  3. [Section 3.3 and Table 6] The scoring in Table 6 was performed by the first author, who also designed the evaluation criteria and conducted the qualitative observations. The paper reports single scores per system with no inter-rater reliability, no variance, and no statistical treatment. As a result, the 'quantitative metrics' are not independently verified and the comparative rankings (e.g., DDSP-VST highest-rated) are not robust evidence. This is acknowledged in Section 8 as a limitation, but the abstract and conclusion do not carry the same caveat. At minimum, the authors should label Table 6 as single-evaluator judgments, provide the raw notes or audio artifacts for audit, and avoid comparative quantitative claims that imply reliability.
  4. [Sections 5 and 8] The paper explicitly states that the SoundCloud playlist is not provided due to the peer-review process, and no code, prompts, or generated audio samples are released. For an evaluation whose central evidence is qualitative listening judgments, this prevents readers from auditing the claims of thematic and structural incoherence. The authors should make anonymized audio examples and the exact prompt set available, or clearly state where such materials will be deposited.
minor comments (5)
  1. [Section 1.4] The heading contains a typo: 'Reasearch Questions' should be 'Research Questions.'
  2. [Section 3.2] The in-text reference to 'Tab. 13' for the performance criteria is not consistent with the table numbering in Appendix D; please harmonize the table numbering throughout.
  3. [Section 7.6] The sentence 'This study exhibited that MGS has potential...' is awkward; 'showed' or 'demonstrated' would be clearer.
  4. [Section 1.3] The capitalization of 'MuseFormer' is inconsistent (e.g., 'Museformer' appears in Section 1.3, 'MuseFormer' elsewhere); please standardize.
  5. [Section 8] The three proposed criteria—Serendipity Support, AI Assistance Balance, and Adaptation Capacity—are introduced as distinct evaluative dimensions but do not appear in Table 6 or the evaluation results; clarify whether they were used in the current evaluation or only proposed for future work.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's claims are empirical summaries of hands-on observations, not quantities derived from fitted parameters or self-citation chains.

full rationale

The paper makes no formal derivation or prediction; its central claims—that MGS serve as complementary tools and show limitations in thematic/structural coherence—are inductive summaries of qualitative notes and 1-5 scores recorded in Sections 5-6. No equation or fitted parameter is renamed as a result, and no criterion in Table 6 is defined in terms of the conclusion (indeed, thematic/structural coherence is not a scored criterion, which weakens support but does not make the claim circular). The single-evaluator design, disclosed in Sections 1.3, 3.3, and 8, is a validity limitation rather than a circularity: the same person designed criteria and assigned scores, but the scores do not mathematically force the stated findings. The exclusion of MuseFormer and MuseCoco from hands-on testing and the absence of a coherence row in Table 6 are coverage gaps, not self-referential reductions. Self-citations to Dadman et al. (2022) and Dadman & Bremdal (2024) appear in Section 7.3 as design-perspective references, but the complementary-tools conclusion is independently grounded in the evaluator's documented interactions and in artist interviews (Section 7.5); removing those citations would not collapse the argument. Accordingly, no load-bearing circular step is identifiable.

Assumptions & free parameters 0 free parameters · 3 assumptions · 3 invented entities

The central claim rests on the validity of the evaluation framework and the representativeness of one evaluator, plus the unvalidated criteria by which systems were scored. No numeric free parameters are fitted, but the 1-5 rubric and the proposed new dimensions are constructs introduced without external validation.

assumptions (3)
  • domain assumption A single expert evaluator's ratings are a valid proxy for music producer experience.
    Section 3.3 says the first author conducted the evaluation alone; Section 8 acknowledges limited generalizability but still treats the Table 6 scores as findings.
  • domain assumption The 1-5 scoring rubric is a valid measurement tool for the eight performance criteria.
    Appendix C justifies the scale from psychometric literature, but the specific descriptors are not validated against external judges or an independent panel.
  • domain assumption The selected systems are a representative sample of open-source MGS.
    Section 1.3 states systems were chosen for accessibility and architectural diversity, not through a systematic inclusion protocol; two systems were not hands-on tested.
invented entities (3)
  • Serendipity Support
    purpose: New evaluation criterion to measure a system's ability to produce unexpected, useful creative discoveries.
    Proposed in Section 8 based on the first author's notes; not externally validated.
  • AI Assistance Balance
    purpose: New evaluation criterion for the balance between automated assistance and user autonomy.
    Proposed in Section 8 based on the first author's notes; not externally validated.
  • Adaptation Capacity
    purpose: New evaluation criterion for a system's responsiveness to a user's evolving artistic style.
    Proposed in Section 8 based on the first author's notes; not externally validated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Workflow-Based Evaluation of Music Generation Systems." pith.science (2026). https://pith.science/paper/UJ47MINT

@misc{pith2026250701022,
  author       = {Pith},
  title        = {Pith review of: Workflow-Based Evaluation of Music Generation Systems},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UJ47MINT}},
  note         = {Machine review of arXiv:2507.01022}
}
read the original abstract

This study presents an exploratory evaluation of Music Generation Systems (MGS) within contemporary music production workflows by examining eight open-source systems. The evaluation framework combines technical insights with practical experimentation through criteria specifically designed to investigate the practical and creative affordances of the systems within the iterative, non-linear nature of music production. Employing a single-evaluator methodology as a preliminary phase, this research adopts a mixed approach utilizing qualitative methods to form hypotheses subsequently assessed through quantitative metrics. The selected systems represent architectural diversity across both symbolic and audio-based music generation approaches, spanning composition, arrangement, and sound design tasks. The investigation addresses limitations of current MGS in music production, challenges and opportunities for workflow integration, and development potential as collaborative tools while maintaining artistic authenticity. Findings reveal these systems function primarily as complementary tools enhancing rather than replacing human expertise. They exhibit limitations in maintaining thematic and structural coherence that emphasize the indispensable role of human creativity in tasks demanding emotional depth and complex decision-making. This study contributes a structured evaluation framework that considers the iterative nature of music creation. It identifies methodological refinements necessary for subsequent comprehensive evaluations and determines viable areas for AI integration as collaborative tools in creative workflows. The research provides empirically-grounded insights to guide future development in the field.

Figures

Figures reproduced from arXiv: 2507.01022 by the authors.

Figure 1
Figure 1. Categorization of surveys in MGS with brief description, grouped by primary focus areas. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Diagram of a two-phase evaluation framework for MGS. Phase 1 (left) presents the [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. Overview of the musical elements used in the [PITH_FULL_IMAGE:figures/full_fig_p021_3.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

136 extracted references · 43 canonical work pages

  1. [1]

    Escalona

    Miguel Civit, Javier Civit-Masot, Francisco Cuadrado, and Maria J. Escalona. A systematic review of artificial intelligence-based music generation: Scope , applications, and future trends. Expert Systems with Applications, 209: 0 118190, December 2022. ISSN 09574174. doi:10.1016/j.eswa.2022.118190. URL https://linkinghub.elsevier.com/retrieve/pii/S0957417...

  2. [2]

    A Functional Taxonomy of Music Generation Systems

    Dorien Herremans, Ching-Hua Chuan, and Elaine Chew. A Functional Taxonomy of Music Generation Systems . ACM Computing Surveys, 50 0 (5): 0 1--30, September 2018. ISSN 0360-0300, 1557-7341. doi:10.1145/3108242. URL https://dl.acm.org/doi/10.1145/3108242

  3. [3]

    Musical agents: A typology and state of the art towards Musical Metacreation

    Kıvanç Tatar and Philippe Pasquier. Musical agents: A typology and state of the art towards Musical Metacreation . Journal of New Music Research, 48 0 (1): 0 56--105, January 2019. ISSN 0929-8215, 1744-5027. doi:10.1080/09298215.2018.1511736. URL https://www.tandfonline.com/doi/full/10.1080/09298215.2018.1511736

  4. [4]

    A review of intelligent music generation systems

    Lei Wang, Ziyi Zhao, Hanwei Liu, Junwei Pang, Yi Qin, and Qidi Wu. A review of intelligent music generation systems. Neural Computing and Applications, 36 0 (12): 0 6381--6401, April 2024. ISSN 0941-0643, 1433-3058. doi:10.1007/s00521-024-09418-2. URL https://link.springer.com/10.1007/s00521-024-09418-2

  5. [5]

    Sotiroudis, Achilles D

    Lazaros Moysis, Lazaros Alexios Iliadis, Sotirios P. Sotiroudis, Achilles D. Boursianis, Maria S. Papadopoulou, Konstantinos-Iraklis D. Kokkinidis, Christos Volos, Panagiotis Sarigiannidis, Spiridon Nikolaidis, and Sotirios K. Goudos. Music Deep Learning : Deep Learning Methods for Music Signal Processing — A Review of the State -of-the- Art . IEEE Access...

  6. [6]

    A Survey of AI Music Generation Tools and Models , 2023

    Yueyue Zhu, Jared Baca, Banafsheh Rekabdar, and Reza Rawassizadeh. A Survey of AI Music Generation Tools and Models , 2023. URL https://arxiv.org/abs/2308.12982

  7. [7]

    A Survey on Deep Learning for Symbolic Music Generation : Representations , Algorithms , Evaluations , and Challenges

    Shulei Ji, Xinyu Yang, and Jing Luo. A Survey on Deep Learning for Symbolic Music Generation : Representations , Algorithms , Evaluations , and Challenges . ACM Computing Surveys, 56 0 (1): 0 7:1--7:39, August 2023. ISSN 0360-0300. doi:10.1145/3597493. URL https://doi.org/10.1145/3597493

  8. [8]

    Toward Interactive Music Generation : A Position Paper

    Shayan Dadman, Bernt Arild Bremdal, Børre Bang, and Rune Dalmo. Toward Interactive Music Generation : A Position Paper . IEEE Access, 10: 0 125679--125695, 2022. ISSN 2169-3536. doi:10.1109/ACCESS.2022.3225689. URL https://ieeexplore.ieee.org/abstract/document/9966445

Show all 136 references
  1. [9]

    From artificial neural networks to deep learning for music generation: history, concepts and trends

    Jean-Pierre Briot. From artificial neural networks to deep learning for music generation: history, concepts and trends. Neural Computing and Applications, 33 0 (1): 0 39--65, January 2021. ISSN 1433-3058. doi:10.1007/s00521-020-05399-0. URL https://doi.org/10.1007/s00521-020-05399-0

  2. [10]

    Deep Learning Techniques for Music Generation

    Jean-Pierre Briot, Gaëtan Hadjeres, and François-David Pachet. Deep Learning Techniques for Music Generation . Computational Synthesis and Creative Systems . Springer International Publishing, Cham, 2020. ISBN 9783319701622 9783319701639. doi:10.1007/978-3-319-70163-9. URL htt...

  3. [11]

    Computational Creativity and Music Generation Systems : An Introduction to the State of the Art

    Filippo Carnovalini and Antonio Rodà. Computational Creativity and Music Generation Systems : An Introduction to the State of the Art . Frontiers in Artificial Intelligence, 3: 0 14, April 2020. ISSN 2624-8212. doi:10.3389/frai.2020.00014. URL https://www.frontiersin.org/artic...

  4. [12]

    Vrahatis

    Maximos Kaliakatsos-Papakostas, Andreas Floros, and Michael N. Vrahatis. Artificial intelligence methods for music generation: a review and future perspectives. In Nature- Inspired Computation and Swarm Intelligence , pages 217--245. Elsevier, 2020. ISBN 9780128197141. doi:10....

  5. [13]

    Algoritmic music composition based on artificial intelligence: A survey

    Omar Lopez-Rincon, Oleg Starostenko, and Gerardo Ayala-San Martín. Algoritmic music composition based on artificial intelligence: A survey. In 2018 International Conference on Electronics , Communications and Computers ( CONIELECOMP ) , pages 187--193, February 2018. doi:10.11...

  6. [14]

    Computational Intelligence in Music Composition : A Survey

    Chien-Hung Liu and Chuan-Kang Ting. Computational Intelligence in Music Composition : A Survey . IEEE Transactions on Emerging Topics in Computational Intelligence, 1 0 (1): 0 2--15, February 2017. ISSN 2471-285X. doi:10.1109/TETCI.2016.2642200. URL https://ieeexplore.ieee.org...

  7. [15]

    Investigating affect in algorithmic composition systems

    Duncan Williams, Alexis Kirke, Eduardo R Miranda, Etienne Roesch, Ian Daly, and Slawomir Nasuto. Investigating affect in algorithmic composition systems. Psychology of Music, 43 0 (6): 0 831--854, November 2015. ISSN 0305-7356, 1741-3087. doi:10.1177/0305735614543282. URL http...

  8. [16]

    J. D. Fernandez and F. Vico. AI Methods in Algorithmic Composition : A Comprehensive Survey . Journal of Artificial Intelligence Research, 48: 0 513--582, November 2013. ISSN 1076-9757. doi:10.1613/jair.3908. URL https://www.jair.org/index.php/jair/article/view/10845

  9. [17]

    Alexis Kirke and Eduardo R. Miranda. An Overview of Computer Systems for Expressive Music Performance . In Alexis Kirke and Eduardo R. Miranda, editors, Guide to Computing for Expressive Music Performance , pages 1--47. Springer, London, 2013. ISBN 9781447141235. doi:10.1007/9...

  10. [18]

    Algorithmic Composition

    Gerhard Nierhaus. Algorithmic Composition . Springer Vienna, Vienna, 2009. ISBN 9783211755396 9783211755402. doi:10.1007/978-3-211-75540-2. URL http://link.springer.com/10.1007/978-3-211-75540-2

  11. [19]

    Computational Models of Expressive Music Performance : The State of the Art

    Gerhard Widmer and Werner Goebl. Computational Models of Expressive Music Performance : The State of the Art . Journal of New Music Research, 33 0 (3): 0 203--216, September 2004. ISSN 0929-8215, 1744-5027. doi:10.1080/0929821042000317804. URL http://www.tandfonline.com/doi/ab...

  12. [20]

    George Papadopoulos and Geraint A. Wiggins. Ai methods for algorithmic composition: A survey, a critical view and future prospects. 1999. URL https://api.semanticscholar.org/CorpusID:5055535

  13. [21]

    Casual creators

    Kate Compton and Michael Mteas. Casual creators. In International Conference on Innovative Computing and Cloud Computing. URL https://api.semanticscholar.org/CorpusID:1305832

  14. [22]

    From genies performing magic to sages imparting wisdom: a value-centred survey of music AI user interfaces, creative affordances and artist objectives

    Oliver Bown. From genies performing magic to sages imparting wisdom: a value-centred survey of music AI user interfaces, creative affordances and artist objectives. Journal of New Music Research, pages 1--14, January 2025. ISSN 0929-8215, 1744-5027. doi:10.1080/09298215.2024.2...

  15. [23]

    Universal music sues ai company anthropic for copyright infringement - levi's sues coperni for trade mark infringement

    Intellectual Property Helpdesk . Universal music sues ai company anthropic for copyright infringement - levi's sues coperni for trade mark infringement. 2023. URL https://intellectual-property-helpdesk.ec.europa.eu/news-events/news/universal-music-sues-ai-company-anthropic-cop...

  16. [24]

    Us record labels sue ai music generators suno and udio for copyright infringement

    Wired . Us record labels sue ai music generators suno and udio for copyright infringement. 2023. URL https://www.wired.com/story/ai-music-generators-suno-and-udio-sued-for-copyright-infringement/. Accessed: 2024-11-07

  17. [25]

    As suno and udio admit training ai with unlicensed music, record industry says: ‘there’s nothing fair about stealing an artist’s life’s work.’

    Music Business Worldwide . As suno and udio admit training ai with unlicensed music, record industry says: ‘there’s nothing fair about stealing an artist’s life’s work.’. 2023. URL https://www.musicbusinessworldwide.com/as-suno-and-udio-admit-training-ai-with-unlicensed-music-...

  18. [26]

    Yinghao Ma, Anders Øland, Anton Ragni, Bleiz MacSen Del Sette, Charalampos Saitis, Chris Donahue, Chenghua Lin, Christos Plachouras, Emmanouil Benetos, Elona Shatri, Fabio Morreale, Ge Zhang, György Fazekas, Gus Xia, Huan Zhang, Ilaria Manco, Jiawen Huang, Julien Guinot, Liwei...

  19. [27]

    Christopher T. Zirpoli. Generative artificial intelligence and copyright law. URL https://crsreports.congress.gov/product/pdf/LSB/LSB10922

  20. [28]

    Open-sourcing highly capable foundation models

    Elizabeth Seger, Noemi Dreksler, Richard Moulange, Emily Dardaman, Jonas Schuett, K Wei, Christoph Winter, Mackenzie Arnold, Se \'a n \'O h \'E igeartaigh, Anton Korinek, et al. Open-sourcing highly capable foundation models. Research paper, Centre for the Governance of AI, 2023

  21. [29]

    The Ethical Implications of Generative Audio Models : A Systematic Literature Review

    Julia Barnett. The Ethical Implications of Generative Audio Models : A Systematic Literature Review . In Proceedings of the 2023 AAAI / ACM Conference on AI , Ethics , and Society , pages 146--161, Montr '\ e\ al QC Canada, August 2023. ACM. ISBN 9798400702310. doi:10.1145/360...

  22. [30]

    Where Does the Buck Stop ? Ethical and Political Issues with AI in Music Creation

    Fabio Morreale. Where Does the Buck Stop ? Ethical and Political Issues with AI in Music Creation . Transactions of the International Society for Music Information Retrieval, 4 0 (1): 0 105--113, July 2021. ISSN 2514-3298. doi:10.5334/tismir.86. URL http://transactions.ismir.n...

  23. [31]

    The Music producer as creative agent : studio production, technology and cultural space in the work of three Finnish producers

    Tuomas Auvinen. The Music producer as creative agent : studio production, technology and cultural space in the work of three Finnish producers. Annales Universitatis Turkuensis. Turku: University of Turku, January 2019. URL https://www.utupub.fi/handle/10024/146576

  24. [32]

    The Poetics of Rock : Cutting Tracks , Making Records

    Albin Zak. The Poetics of Rock : Cutting Tracks , Making Records . University of California Press, November 2001. ISBN 9780520232242. URL http://www.jstor.org/stable/10.1525/j.ctt1ppbkt. Google-Books-ID: 5bAwDwAAQBAJ

  25. [33]

    Performing Rites : On the Value of Popular Music

    Simon Frith. Performing Rites : On the Value of Popular Music . Harvard University Press, 1996. ISBN 9780674661967. URL https://books.google.no/books?id=BPdIfT6scIoC. Google-Books-ID: BPdIfT6scIoC

  26. [34]

    The History of Music Production

    Richard James Burgess. The History of Music Production . Oxford University Press, 2014. ISBN 9780199357161. URL https://books.google.no/books?id=qMKiAwAAQBAJ. Google-Books-ID: ZeISDAAAQBAJ

  27. [35]

    Electronic and Experimental Music : Technology , Music , and Culture

    Thom Holmes. Electronic and Experimental Music : Technology , Music , and Culture . Routledge, 6 edition, March 2020. ISBN 9780429425585. doi:10.4324/9780429425585. URL https://www.taylorfrancis.com/books/9780429758447

  28. [36]

    David Moffat and Mark B. Sandler. Approaches in Intelligent Music Production . Arts, 8 0 (4): 0 125, December 2019. ISSN 2076-0752. doi:10.3390/arts8040125. URL https://www.mdpi.com/2076-0752/8/4/125

  29. [37]

    An Intermediary Between Production and Consumption : The Producer of Popular Music

    Antoine Hennion. An Intermediary Between Production and Consumption : The Producer of Popular Music . Science, Technology, & Human Values, 14 0 (4): 0 400--424, October 1989. ISSN 0162-2439, 1552-8251. doi:10.1177/016224398901400405. URL http://journals.sagepub.com/doi/10.1177...

  30. [38]

    The Art of Music Production : The Theory and Practice

    Richard James Burgess. The Art of Music Production : The Theory and Practice . Oxford University Press, September 2013. ISBN 9780199359325. URL https://books.google.no/books?id=m4dNEAAAQBAJ. Google-Books-ID: lWEUAAAAQBAJ

  31. [39]

    On the evaluation of generative models in music

    Li-Chia Yang and Alexander Lerch. On the evaluation of generative models in music. Neural Computing and Applications, 32 0 (9): 0 4773--4784, May 2020. ISSN 1433-3058. doi:10.1007/s00521-018-3849-7. URL https://doi.org/10.1007/s00521-018-3849-7

  32. [40]

    Simple and Controllable Music Generation , 2023

    Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre Défossez. Simple and Controllable Music Generation , 2023. URL https://arxiv.org/abs/2306.05284

  33. [41]

    M ^ 2 ugen: Multi-modal music understanding and generation with the power of large language models, 2024 a

    Shansong Liu, Atin Sakkeer Hussain, Chenshuo Sun, and Ying Shan. M ^ 2 ugen: Multi-modal music understanding and generation with the power of large language models, 2024 a

  34. [42]

    Riffusion - Stable diffusion for real-time music generation , 2022

    Seth* Forsgren and Hayk* Martiros. Riffusion - Stable diffusion for real-time music generation , 2022. URL https://riffusion.com

  35. [43]

    Magenta: Music and art generation with machine intelligence, 2024

    Google Magenta Team . Magenta: Music and art generation with machine intelligence, 2024. URL https://magenta.tensorflow.org/

  36. [44]

    Musika! Fast Infinite Waveform Music Generation , August 2022

    Marco Pasini and Jan Schlüter. Musika! Fast Infinite Waveform Music Generation , August 2022. URL http://arxiv.org/abs/2208.08706. arXiv:2208.08706 [cs, eess]

  37. [45]

    MuseCoco : Generating Symbolic Music from Text , May 2023

    Peiling Lu, Xin Xu, Chenfei Kang, Botao Yu, Chengyi Xing, Xu Tan, and Jiang Bian. MuseCoco : Generating Symbolic Music from Text , May 2023. URL http://arxiv.org/abs/2306.00110. arXiv:2306.00110 [cs, eess]

  38. [46]

    Museformer: Transformer with Fine- and Coarse-Grained Attention for Music Generation

    Botao Yu, Peiling Lu, Rui Wang, Wei Hu, Xu Tan, Wei Ye, Shikun Zhang, Tao Qin, and Tie-Yan Liu. Museformer: Transformer with Fine- and Coarse-Grained Attention for Music Generation . URL http://arxiv.org/abs/2210.10349

  39. [47]

    The Computer Music Tutorial

    Curtis Roads. The Computer Music Tutorial. The MIT Press, second edition edition. ISBN 978-0-262-04491-2

  40. [48]

    Mixing Secrets for the Small Studio

    Mike Senior. Mixing Secrets for the Small Studio. Sound on Sound Presents. Routledge/Taylor & Francis Group, second edition edition. ISBN 978-1-315-15001-7 978-1-351-36880-3 978-1-351-36879-7

  41. [49]

    Dance Music Manual: Tools, Toys, and Techniques

    Rick Snoman. Dance Music Manual: Tools, Toys, and Techniques. Focal Press, third edition edition. ISBN 978-0-415-82564-1

  42. [50]

    Adapting Frechet Audio Distance for Generative Music Evaluation

    Azalea Gui, Hannes Gamper, Sebastian Braun, and Dimitra Emmanouilidou. Adapting Frechet Audio Distance for Generative Music Evaluation . URL http://arxiv.org/abs/2311.01616

  43. [51]

    MusPy : A Toolkit for Symbolic Music Generation

    Hao-Wen Dong, Ke Chen, Julian McAuley, and Taylor Berg-Kirkpatrick. MusPy : A Toolkit for Symbolic Music Generation . URL http://arxiv.org/abs/2008.01951

  44. [52]

    MIR \_ EVAL : A transparent implementation of common MIR metrics

    Colin Raffel, Brian McFee, Eric J Humphrey, Justin Salamon, Oriol Nieto, Dawen Liang, Daniel PW Ellis, and C Colin Raffel. MIR \_ EVAL : A transparent implementation of common MIR metrics. In ISMIR , volume 10, page 2014, a

  45. [53]

    On the Development and Practice of AI Technology for Contemporary Popular Music Production

    Emmanuel Deruty, Maarten Grachten, Stefan Lattner, Javier Nistal, and Cyran Aouameur. On the Development and Practice of AI Technology for Contemporary Popular Music Production . Transactions of the International Society for Music Information Retrieval, 5 0 (1): 0 35, February...

  46. [54]

    A Comprehensive Survey for Evaluation Methodologies of AI-Generated Music

    Zeyu Xiong, Weitao Wang, Jing Yu, Yue Lin, and Ziyan Wang. A Comprehensive Survey for Evaluation Methodologies of AI-Generated Music . URL http://arxiv.org/abs/2308.13736

  47. [56]

    Towards efficiency of a subjective evaluation for music source separation

    Peter Kasak, Roman Jarina, and Dasa Ticha. Towards efficiency of a subjective evaluation for music source separation. In 2022 32nd International Conference Radioelektronika ( RADIOELEKTRONIKA ) , pages 01--05. IEEE . ISBN 978-1-72818-686-3. doi:10.1109/RADIOELEKTRONIKA54537.20...

  48. [57]

    Adam Linson, Chris Dobbyn, and Robin C. Laney. Critical issues in evaluating freely improvising interactive music systems. In International Conference on Innovative Computing and Cloud Computing. URL https://api.semanticscholar.org/CorpusID:95175

  49. [58]

    Sonny Rollins and the challenge of thematic improvisation

    Gunther Schuller. Sonny Rollins and the challenge of thematic improvisation. 1 0 (1): 0 6--11

  50. [59]

    Juslin and Daniel Västfjäll

    Patrik N. Juslin and Daniel Västfjäll. Emotional responses to music: The need to consider underlying mechanisms. Behavioral and Brain Sciences, 31 0 (5): 0 559--575, October 2008. ISSN 0140-525X, 1469-1825. doi:10.1017/S0140525X08005293. URL https://www.cambridge.org/core/prod...

  51. [60]

    What is beautiful is usable

    N Tractinsky, A.S Katz, and D Ikar. What is beautiful is usable. 13 0 (2): 0 127--145. ISSN 09535438. doi:10.1016/S0953-5438(00)00031-X. URL https://academic.oup.com/iwc/article-lookup/doi/10.1016/S0953-5438(00)00031-X

  52. [61]

    A Standardised Procedure for Evaluating Creative Systems : Computational Creativity Evaluation Based on What it is to be Creative

    Anna Jordanous. A Standardised Procedure for Evaluating Creative Systems : Computational Creativity Evaluation Based on What it is to be Creative . Cognitive Computation, 4 0 (3): 0 246--279, September 2012. ISSN 1866-9964. doi:10.1007/s12559-012-9156-1. URL https://doi.org/10...

  53. [62]

    Ellis, and Brian Whitman

    Adam Berenzweig, Beth Logan, Daniel P.W. Ellis, and Brian Whitman. A Large-Scale Evaluation of Acoustic and Subjective Music-Similarity Measures . 28 0 (2): 0 63--76, b . ISSN 0148-9267, 1531-5169. doi:10.1162/014892604323112257. URL https://direct.mit.edu/comj/article/28/2/63...

  54. [63]

    Are the Emotions Expressed in Music Genre-specific ? An Audio-based Evaluation of Datasets Spanning Classical , Film , Pop and Mixed Genres

    Tuomas Eerola. Are the Emotions Expressed in Music Genre-specific ? An Audio-based Evaluation of Datasets Spanning Classical , Film , Pop and Mixed Genres . 40 0 (4): 0 349--366. ISSN 0929-8215, 1744-5027. doi:10.1080/09298215.2011.602195. URL http://www.tandfonline.com/doi/ab...

  55. [64]

    The micro-and macrostructural design of improvised music

    Jeff Pressing. The micro-and macrostructural design of improvised music. 5 0 (2): 0 133--172. URL https://www.jstor.org/stable/pdf/40285390.pdf

  56. [65]

    Eric F. Clarke. Ways of Listening an Ecological Approach to the Perception of Musical Meaning. Oxford University Press. ISBN 978-0-19-028816-7

  57. [67]

    P. J. Charles Reimer and Marcelo M. Wanderley. Embracing less common evaluation strategies for studying user experience in NIME . In NIME 2021 . PubPub . doi:10.21428/92fbeb44.807a000f. URL https://nime.pubpub.org/pub/fidgs435

  58. [68]

    Wanderley and Wendy E

    Marcelo M. Wanderley and Wendy E. Mackay. HCI , Music and Art : An Interview with Wendy Mackay . In Simon Holland, Tom Mudd, Katie Wilkie-McKenna, Andrew McPherson, and Marcelo M. Wanderley, editors, New Directions in Music and Human-Computer Interaction , pages 115--120. Spri...

  59. [69]

    Cheng-Zhi Anna Huang, Hendrik Vincent Koops, Ed Newton-Rex, Monica Dinculescu, and Carrie J. Cai. AI Song Contest : Human - AI Co - Creation in Songwriting , October 2020. URL http://arxiv.org/abs/2010.05388. arXiv:2010.05388 [cs]

  60. [70]

    Cooperstock

    Dalia El-Shimy and Jeremy R. Cooperstock. User-driven techniques for the design and evaluation of new musical interfaces. 40 0 (2): 0 35--46. ISSN 0148-9267, 1531-5169. doi:10.1162/COMJ_a_00357. URL https://direct.mit.edu/comj/article/40/2/35-46/94542

  61. [71]

    Stowell, A

    D. Stowell, A. Robertson, N. Bryan-Kinns, and M.D. Plumbley. Evaluation of live human–computer music-making: Quantitative and qualitative approaches. 67 0 (11): 0 960--975, b . ISSN 10715819. doi:10.1016/j.ijhcs.2009.05.007. URL https://linkinghub.elsevier.com/retrieve/pii/S10...

  62. [72]

    Burke Johnson and Anthony J

    R. Burke Johnson and Anthony J. Onwuegbuzie. Mixed Methods Research : A Research Paradigm Whose Time Has Come . 33 0 (7): 0 14--26. ISSN 0013-189X, 1935-102X. doi:10.3102/0013189X033007014. URL https://journals.sagepub.com/doi/10.3102/0013189X033007014

  63. [73]

    Where are the mixed methods research studies? 30 0 (4): 0 311--313

    Joke Bradt. Where are the mixed methods research studies? 30 0 (4): 0 311--313. ISSN 0809-8131, 1944-8260. doi:10.1080/08098131.2021.1936771. URL https://www.tandfonline.com/doi/full/10.1080/08098131.2021.1936771

  64. [74]

    Schacher, Hanna Järveläinen, Christian Strinning, and Patrick Neff

    Jan C. Schacher, Hanna Järveläinen, Christian Strinning, and Patrick Neff. Movement Perception In Music Performance - A Mixed Methods Investigation . ISSN 2518-3672. doi:10.5281/ZENODO.851106. URL https://zenodo.org/record/851106

  65. [75]

    An Empirical Study on How People Perceive AI-generated Music

    Hyeshin Chu, Joohee Kim, Seongouk Kim, Hongkyu Lim, Hyunwook Lee, Seungmin Jin, Jongeun Lee, Taehwan Kim, and Sungahn Ko. An Empirical Study on How People Perceive AI-generated Music . In Proceedings of the 31st ACM International Conference on Information & Knowledge Managemen...

  66. [76]

    High fidelity neural audio compression

    Alexandre Défossez, Jade Copet, Gabriel Synnaeve, and Yossi Adi. High fidelity neural audio compression. URL http://arxiv.org/abs/2210.13438

  67. [77]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer, b . URL http://arxiv.org/abs/1910.10683

  68. [78]

    CLAP : Learning audio concepts from natural language supervision

    Benjamin Elizalde, Soham Deshmukh, Mahmoud Al Ismail, and Huaming Wang. CLAP : Learning audio concepts from natural language supervision. URL http://arxiv.org/abs/2206.04769

  69. [79]

    Denk, Zalán Borsos, Jesse Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, Matt Sharifi, Neil Zeghidour, and Christian Frank

    Andrea Agostinelli, Timo I. Denk, Zalán Borsos, Jesse Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, Matt Sharifi, Neil Zeghidour, and Christian Frank. Musiclm: Generating music from text, 2023

  70. [80]

    Mo usai: Text -to- Music Generation with Long - Context Latent Diffusion , October 2023

    Flavio Schneider, Ojasv Kamal, Zhijing Jin, and Bernhard Schölkopf. Mo usai: Text -to- Music Generation with Long - Context Latent Diffusion , October 2023. URL http://arxiv.org/abs/2301.11757. arXiv:2301.11757 [cs, eess]

  71. [81]

    Fr 'echet audio distance: A metric for evaluating music enhancement algorithms

    Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek, and Matthew Sharifi. Fr 'echet audio distance: A metric for evaluating music enhancement algorithms. URL http://arxiv.org/abs/1812.08466

  72. [82]

    BERT : Pre-training of deep bidirectional transformers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT : Pre-training of deep bidirectional transformers for language understanding. URL http://arxiv.org/abs/1810.04805

  73. [83]

    Transformers are RNNs : Fast autoregressive transformers with linear attention

    Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and François Fleuret. Transformers are RNNs : Fast autoregressive transformers with linear attention. URL http://arxiv.org/abs/2006.16236

  74. [84]

    MusicBERT : Symbolic Music Understanding with Large - Scale Pre - Training , June 2021

    Mingliang Zeng, Xu Tan, Rui Wang, Zeqian Ju, Tao Qin, and Tie-Yan Liu. MusicBERT : Symbolic Music Understanding with Large - Scale Pre - Training , June 2021. URL http://arxiv.org/abs/2106.05630. arXiv:2106.05630 [cs, eess]

  75. [85]

    EMOPIA : A multi-modal pop piano dataset for emotion recognition and emotion-based music generation

    Hsiao-Tzu Hung, Joann Ching, Seungheon Doh, Nabin Kim, Juhan Nam, and Yi-Hsuan Yang. EMOPIA : A multi-modal pop piano dataset for emotion recognition and emotion-based music generation. URL http://arxiv.org/abs/2108.01374

  76. [86]

    Building the MetaMIDI dataset: Linking symbolic and audio musical data

    Jeff Ens and Philippe Pasquier. Building the MetaMIDI dataset: Linking symbolic and audio musical data. In ISMIR , volume 22, pages 182--188. URL https://archives.ismir.net/ismir2021/paper/000022.pdf

  77. [87]

    Pop music transformer: Beat-based modeling and generation of expressive pop piano compositions

    Yu-Siang Huang and Yi-Hsuan Yang. Pop music transformer: Beat-based modeling and generation of expressive pop piano compositions. URL http://arxiv.org/abs/2002.00212

  78. [88]

    OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff ...

  79. [89]

    Exploring the efficacy of pre-trained checkpoints in text-to-music generation task

    Shangda Wu and Maosong Sun. Exploring the efficacy of pre-trained checkpoints in text-to-music generation task. URL http://arxiv.org/abs/2211.11216

  80. [90]

    Towards faster and stabilized GAN training for high-fidelity few-shot image synthesis

    Bingchen Liu, Yizhe Zhu, Kunpeng Song, and Ahmed Elgammal. Towards faster and stabilized GAN training for high-fidelity few-shot image synthesis. URL http://arxiv.org/abs/2101.04775

  81. [91]

    Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu

    Heiga Zen, Viet Dang, Rob Clark, Yu Zhang, Ron J. Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu. LibriTTS : A corpus derived from LibriSpeech for text-to-speech. URL http://arxiv.org/abs/1904.02882

  82. [92]

    Enabling factorized piano music modeling and generation with the MAESTRO dataset

    Curtis Hawthorne, Andriy Stasyuk, Adam Roberts, Ian Simon, Cheng-Zhi Anna Huang, Sander Dieleman, Erich Elsen, Jesse Engel, and Douglas Eck. Enabling factorized piano music modeling and generation with the MAESTRO dataset. URL http://arxiv.org/abs/1810.12247

  83. [93]

    Max W. Y. Lam, Qiao Tian, Tang Li, Zongyu Yin, Siyuan Feng, Ming Tu, Yuliang Ji, Rui Xia, Mingbo Ma, Xuchen Song, Jitong Chen, Wang Yuping, and Yuxuan Wang. Efficient neural music generation. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Adva...

  84. [94]

    An Image is Worth 16x16 Words : Transformers for Image Recognition at Scale , June 2021

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An Image is Worth 16x16 Words : Transformers for Image Recognition a...

  85. [95]

    ViViT : A Video Vision Transformer

    Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid. ViViT : A Video Vision Transformer . In Proceedings of the IEEE/CVF international conference on computer vision, pages 6836--6846, 2021. URL https://openaccess.thecvf.com/content/ICCV202...

  86. [96]

    MERT : Acoustic Music Understanding Model with Large - Scale Self -supervised Training , April 2024

    Yizhi Li, Ruibin Yuan, Ge Zhang, Yinghao Ma, Xingran Chen, Hanzhi Yin, Chenghao Xiao, Chenghua Lin, Anton Ragni, Emmanouil Benetos, Norbert Gyenge, Roger Dannenberg, Ruibo Liu, Wenhu Chen, Gus Xia, Yemin Shi, Wenhao Huang, Zili Wang, Yike Guo, and Jie Fu. MERT : Acoustic Music...

  87. [97]

    Llama 2: Open Foundation and Fine - Tuned Chat Models , July 2023

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, W...

  88. [98]

    Plumbley

    Haohe Liu, Yi Yuan, Xubo Liu, Xinhao Mei, Qiuqiang Kong, Qiao Tian, Yuping Wang, Wenwu Wang, Yuxuan Wang, and Mark D. Plumbley. AudioLDM 2: Learning Holistic Audio Generation with Self -supervised Pretraining , May 2024 b . URL http://arxiv.org/abs/2308.05734. arXiv:2308.05734...

  89. [100]

    AUDIT : Audio Editing by Following Instructions with Latent Diffusion Models

    Yuancheng Wang, Zeqian Ju, Xu Tan, Lei He, Zhizheng Wu, Jiang Bian, and Sheng Zhao. AUDIT : Audio Editing by Following Instructions with Latent Diffusion Models . Advances in Neural Information Processing Systems, 36: 0 71340--71357, December 2023. URL https://proceedings.neur...

  90. [101]

    InstructME : An Instruction Guided Music Edit And Remix Framework with Latent Diffusion Models , December 2023

    Bing Han, Junyu Dai, Weituo Hao, Xinyan He, Dong Guo, Jitong Chen, Yuxuan Wang, Yanmin Qian, and Xuchen Song. InstructME : An Instruction Guided Music Edit And Remix Framework with Latent Diffusion Models , December 2023. URL http://arxiv.org/abs/2308.14360. arXiv:2308.14360 [...

  91. [102]

    A survey of foundation models for music understanding

    Wenjun Li, Ying Cai, Ziyang Wu, Wenyi Zhang, Yifan Chen, Rundong Qi, Mengqi Dong, Peigen Chen, Xiao Dong, Fenghao Shi, Lei Guo, Junwei Han, Bao Ge, Tianming Liu, Lin Gan, and Tuo Zhang. A survey of foundation models for music understanding. URL http://arxiv.org/abs/2409.09601

  92. [103]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. URL http://arxiv.org/abs/1706.03762

  93. [104]

    A hierarchical latent vector model for learning long-term structure in music

    Adam Roberts, Jesse Engel, Colin Raffel, Curtis Hawthorne, and Douglas Eck. A hierarchical latent vector model for learning long-term structure in music. URL http://arxiv.org/abs/1803.05428

  94. [105]

    Neural audio synthesis of musical notes with WaveNet autoencoders

    Jesse Engel, Cinjon Resnick, Adam Roberts, Sander Dieleman, Douglas Eck, Karen Simonyan, and Mohammad Norouzi. Neural audio synthesis of musical notes with WaveNet autoencoders. URL http://arxiv.org/abs/1704.01279

  95. [106]

    DDSP : Differentiable Digital Signal Processing , January 2020

    Jesse Engel, Lamtharn Hantrakul, Chenjie Gu, and Adam Roberts. DDSP : Differentiable Digital Signal Processing , January 2020. URL http://arxiv.org/abs/2001.04643. arXiv:2001.04643 [cs, eess, stat]

  96. [107]

    Performance issue of m2ugen - github issue

    GitHub. Performance issue of m2ugen - github issue. URL https://github.com/shansongliu/M2UGen/issues/4. Accessed: 01 March 2024

  97. [108]

    Hybrid transformers for music source separation

    Simon Rouard, Francisco Massa, and Alexandre D \'e fossez. Hybrid transformers for music source separation. In ICASSP 23, 2023

  98. [109]

    Martínez-Ramírez, Liwei Lin, Gus Xia, Wei-Hsiang Liao, Yuki Mitsufuji, and Simon Dixon

    Yixiao Zhang, Yukara Ikemiya, Woosung Choi, Naoki Murata, Marco A. Martínez-Ramírez, Liwei Lin, Gus Xia, Wei-Hsiang Liao, Yuki Mitsufuji, and Simon Dixon. Instruct- MusicGen : Unlocking Text -to- Music Editing for Music Language Models via Instruction Tuning , May 2024. URL ht...

  99. [110]

    Max Langenkamp and Daniel N. Yue. How Open Source Machine Learning Software Shapes AI . In Proceedings of the 2022 AAAI / ACM Conference on AI , Ethics , and Society , pages 385--395, Oxford United Kingdom, July 2022. ACM. ISBN 978-1-4503-9247-1. doi:10.1145/3514094.3534167. U...

  100. [111]

    Crafting Creative Melodies : A User - Centric Approach for Symbolic Music Generation

    Shayan Dadman and Bernt Arild Bremdal. Crafting Creative Melodies : A User - Centric Approach for Symbolic Music Generation . Electronics, 13 0 (6): 0 1116, March 2024. ISSN 2079-9292. doi:10.3390/electronics13061116. URL https://www.mdpi.com/2079-9292/13/6/1116

  101. [112]

    RAVE : A variational autoencoder for fast and high-quality neural audio synthesis

    Antoine Caillon and Philippe Esling. RAVE : A variational autoencoder for fast and high-quality neural audio synthesis. URL http://arxiv.org/abs/2111.05011

  102. [113]

    How to Prompt ? Opportunities and Challenges of Zero - and Few - Shot Learning for Human - AI Interaction in Creative Applications of Generative Models , September 2022

    Hai Dang, Lukas Mecke, Florian Lehmann, Sven Goller, and Daniel Buschek. How to Prompt ? Opportunities and Challenges of Zero - and Few - Shot Learning for Human - AI Interaction in Creative Applications of Generative Models , September 2022. URL http://arxiv.org/abs/2209.0139...

  103. [114]

    A Taxonomy of Prompt Modifiers for Text - To - Image Generation

    Jonas Oppenlaender. A Taxonomy of Prompt Modifiers for Text - To - Image Generation . Behaviour & Information Technology, pages 1--14, November 2023. ISSN 0144-929X, 1362-3001. doi:10.1080/0144929X.2023.2286532. URL http://arxiv.org/abs/2204.13988. arXiv:2204.13988 [cs]

  104. [115]

    Multimodal music datasets? Challenges and future goals in music processing

    Anna-Maria Christodoulou, Olivier Lartillot, and Alexander Refsum Jensenius. Multimodal music datasets? Challenges and future goals in music processing. International Journal of Multimedia Information Retrieval, 13 0 (3): 0 37, September 2024. ISSN 2192-6611, 2192-662X. doi:10...

  105. [116]

    Pre-train, Prompt , and Predict : A Systematic Survey of Prompting Methods in Natural Language Processing , July 2021

    Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. Pre-train, Prompt , and Predict : A Systematic Survey of Prompting Methods in Natural Language Processing , July 2021. URL http://arxiv.org/abs/2107.13586. arXiv:2107.13586 [cs]

  106. [117]

    On the Open Prompt Challenge in Conditional Audio Generation

    Ernie Chang, Sidd Srinivasan, Mahi Luthra, Pin-Jie Lin, Varun Nagaraja, Forrest Iandola, Zechun Liu, Zhaoheng Ni, Changsheng Zhao, Yangyang Shi, and Vikas Chandra. On the Open Prompt Challenge in Conditional Audio Generation . In ICASSP 2024 - 2024 IEEE International Conferenc...

  107. [118]

    Translating Intercultural Creativities in Community Music , volume 1

    Pamela Burnard, Valerie Ross, Laura Hassler, and Lis Murphy. Translating Intercultural Creativities in Community Music , volume 1. Oxford University Press, February 2018. doi:10.1093/oxfordhb/9780190219505.013.6. URL https://academic.oup.com/edited-volume/34637/chapter/295100681

  108. [119]

    Learning to Answer Questions in Dynamic Audio - Visual Scenarios

    Guangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu, Ji-Rong Wen, and Di Hu. Learning to Answer Questions in Dynamic Audio - Visual Scenarios . In 2022 IEEE / CVF Conference on Computer Vision and Pattern Recognition ( CVPR ) , pages 19086--19096, New Orleans, LA, USA, June 2022....

  109. [120]

    PAGURI : a user experience study of creative interaction with text-to-music models, July 2024

    Francesca Ronchini, Luca Comanducci, Gabriele Perego, and Fabio Antonacci. PAGURI : a user experience study of creative interaction with text-to-music models, July 2024. URL http://arxiv.org/abs/2407.04333. arXiv:2407.04333 [cs, eess] version: 1

  110. [121]

    IteraTTA : An interface for exploring both text prompts and audio priors in generating music with text-to-audio models, July 2023

    Hiromu Yakura and Masataka Goto. IteraTTA : An interface for exploring both text prompts and audio priors in generating music with text-to-audio models, July 2023. URL http://arxiv.org/abs/2307.13005. arXiv:2307.13005 [cs, eess]

  111. [122]

    A survey on large language model based autonomous agents

    Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, Wayne Xin Zhao, Zhewei Wei, and Ji-Rong Wen. A survey on large language model based autonomous agents. 18 0 (6): 0 186345. ISSN 2095-2228, 2095-2236. doi:10.10...

  112. [123]

    Let AI Entertain You : Increasing User Engagement with Generative AI and Rejection Sampling , December 2023

    Jingying Zeng, Jaewon Yang, Waleed Malik, Xiao Yan, Richard Huang, and Qi He. Let AI Entertain You : Increasing User Engagement with Generative AI and Rejection Sampling , December 2023. URL http://arxiv.org/abs/2312.12457. arXiv:2312.12457 [cs]

  113. [124]

    List of questions and answers from interviews with press

    Taryn. List of questions and answers from interviews with press. Online, 2024. URL https://docs.google.com/document/d/1mTelMocJD788hk_x4Ce-bwnPmVXoogVG5tezB9lSirQ/edit. Interviews compiled and supplied by Taryn herself, including major media outlets such as Forbes, The Verge, ...

  114. [125]

    we're at the precipice of a fundamental shift in how we think about making music

    Matt Mullen. How patten used text-to-audio ai to make an entire album: "we're at the precipice of a fundamental shift in how we think about making music". MusicRadar, May 2023. URL https://www.musicradar.com/news/patten-interview

  115. [126]

    Max cooper is using ai to push the frontiers of creativity and communication

    Webb Wright. Max cooper is using ai to push the frontiers of creativity and communication. The Drum, May 2023. URL https://www.thedrum.com/news/2023/05/31/max-cooper-ai-the-future-music-and-consciousness

  116. [127]

    Doshi and Oliver P

    Anil R. Doshi and Oliver P. Hauser. Generative artificial intelligence enhances creativity but reduces the diversity of novel content. URL http://arxiv.org/abs/2312.00506

  117. [128]

    The double-edged roles of generative AI in the creative process: Experiments on design work

    Jinghui (Jove) Hou, Lei Wang, Gang Wang, Harry Wang, and Shuai Yang. The double-edged roles of generative AI in the creative process: Experiments on design work. URL https://papers.ssrn.com/abstract=4739471

  118. [129]

    Kelly, Saumya Pareek, Qiushi Zhou, and Eduardo Velloso

    Samangi Wadinambiarachchi, Ryan M. Kelly, Saumya Pareek, Qiushi Zhou, and Eduardo Velloso. The effects of generative AI on design fixation and divergent thinking. In Proceedings of the CHI Conference on Human Factors in Computing Systems , pages 1--18. doi:10.1145/3613904.3642...

  119. [130]

    Human creativity in the age of LLMs : Randomized experiments on divergent and convergent thinking

    Harsh Kumar, Jonathan Vincentius, Ewan Jordan, and Ashton Anderson. Human creativity in the age of LLMs : Randomized experiments on divergent and convergent thinking. URL http://arxiv.org/abs/2410.03703

  120. [131]

    Musicking: the meanings of performing and listening

    Christopher Small. Musicking: the meanings of performing and listening. Music/culture. University Press of New England. ISBN 9780819522566 9780819522573

  121. [132]

    Young and Dave Murphy

    Gareth W. Young and Dave Murphy. HCI Models for Digital Musical Instruments : Methodologies for Rigorous Testing of Digital Musical Instruments . URL http://arxiv.org/abs/2010.01328

  122. [133]

    Kat Agres, Jamie Forth, and Geraint A. Wiggins. Evaluation of Musical Creativity and Musical Metacreation Systems . 14 0 (3): 0 1--33. ISSN 1544-3574. doi:10.1145/2967506. URL https://dl.acm.org/doi/10.1145/2967506

  123. [134]

    Evaluating musical metacreation in a live performance context

    Arne Eigenfeldt, Adam Burnett, and Philippe Pasquier. Evaluating musical metacreation in a live performance context. Proceedings of the Third International Conference on Computational Creativity , pages 140--144

  124. [135]

    How many response categories are sufficient for Likert type scales? An empirical study based on the Item Response Theory

    Eren Can Aybek and Cetin Toraman. How many response categories are sufficient for Likert type scales? An empirical study based on the Item Response Theory . 9 0 (2): 0 534--547. ISSN 2148-7456. doi:10.21449/ijate.1132931. URL http://dergipark.org.tr/en/doi/10.21449/ijate.1132931

  125. [136]

    Optimal number of response categories in rating scales: Reliability, validity, discriminating power, and respondent preferences

    Carolyn C Preston and Andrew M Colman. Optimal number of response categories in rating scales: Reliability, validity, discriminating power, and respondent preferences. 104 0 (1): 0 1--15. ISSN 00016918. doi:10.1016/S0001-6918(99)00050-5. URL https://linkinghub.elsevier.com/ret...

  126. [137]

    Morrison

    Donald G. Morrison. Regressions with Discrete Dependent Variables : The Effect on R 2. 9 0 (3): 0 338. ISSN 00222437. doi:10.2307/3149551. URL https://www.jstor.org/stable/3149551?origin=crossref

  127. [138]

    Is a Three-Point Scale Good Enough ? URL https://measuringu.com/three-points/

    Jeff Sauro. Is a Three-Point Scale Good Enough ? URL https://measuringu.com/three-points/

  128. [139]

    Psychological Distance Between Categories in the Likert Scale : Comparing Different Numbers of Options

    Takafumi Wakita, Natsumi Ueshima, and Hiroyuki Noguchi. Psychological Distance Between Categories in the Likert Scale : Comparing Different Numbers of Options . 72 0 (4): 0 533--546. ISSN 0013-1644, 1552-3888. doi:10.1177/0013164411431162. URL https://journals.sagepub.com/doi/...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.