REVIEW 2 major objections 5 minor 56 references
Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping
T0 review · 2 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Prompted agreement measures occurrence, not instruction following; matched neutral and target-swap contrasts show which music models actually respond to key and beat requests.
desk verdict A careful, well-executed demonstration that prompted agreement can masquerade as control in text-to-music; the soft spot is model-specific auto-scorer validity, but the paper's own checks keep the conclusion credible. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the matched neutral–A–B contrast family: one neutral input that omits the scored attribute and two otherwise matched inputs that request different target values, all rendered through the system's frozen native interface with a shared seed. The argument is carried by three statistics derived from it—Acc (did the requested target occur under treatment), Δ (did treatment raise occurrence above the matched neutral output), and Margin (did swapping the requested value redirect the output toward the alternative target). A confirmatory label of effective control requires both Δ and Margin intervals to lie above zero; the paper also uses off-attribute placebos and multi-seed sentinels to rule out generic added-instruction effects and single-generation noise.
What would settle it
An independent blind human relabeling of the released 256-family benchmark per system, or a rerun with a different validated key and beat recognizer, would settle the claim: if LeVo2 shows a key Δ interval clearly above zero, or Stable Audio 3 loses its negative four-beat Δ, the paper's central contrast would not survive. The paper's own audit already shows that automatic key Δ (0.429) exceeds human key Δ (0.286) on a 28-family subset, so a full human relabeling is the direct check.
Extended reading notes
Core claim
The central empirical claim is that ACE-Step 1.5 and Stable Audio 3 Medium show large key control that is not explained by their output priors, while LeVo2 shows little attributable key response under its evaluated interface; and that for beat grouping the models redirect toward the rare three-beat target, while the apparently high four-beat agreement is a prior artifact. The central methodological claim is that agreement measures occurrence, whereas matched neutral and target-swap contrasts test instruction-attributable response. The paper reports ACE-Step key Δ = 0.612 and Stable Audio 3 key Δ = 0.646 against LeVo2's 0.026, and for beat grouping Stable Audio 3 three-beat Δ = +0.469 against four-beat Δ = −0.406, with the neutral four-beat rate at 0.969 and the explicit four-beat treatment at 0.563. The authors conclude that high prompted agreement can coexist with no improvement—or a large reversal—relative to the model's own neutral output.
Load-bearing premise
The conclusions stand or fall on the validity of the frozen automatic scorers—S-KEY for key and Beat This for beat grouping—as outcome measures; if those recognizers mislabel generated audio in a target-dependent way, the reported Δ and Margin values would not reflect what the models actually produced.
Editorial extensions
If this is right
- Prompted agreement alone should not be used as evidence of controllability for any target whose output prior is high; the matched neutral rate must be reported.
- Controllability should be described per target, not as an aggregate score, because four-beat and three-beat responses can cancel in the average.
- If the result transfers, generative-model benchmarks outside music—faces, daylight scenes, common code patterns—should adopt neutral-relative and target-swap contrasts.
- A negative four-beat Δ for Stable Audio 3 shows that a treatment can move output away from the common target; benchmarks need to allow for this rather than treating any agreement as success.
Reading between the lines
- A testable extension is to cross carrier templates with target values, since the paper's neutral contrast changes both target presence and its rendered carrier; such a factorial design could estimate carrier–value interactions rather than only end-to-end response.
- The same design could be applied to continuous controls such as BPM or loudness, but the neutral-omission construction may need a neighbouring-value reference instead of an omitted field.
- The human-audit gap in key magnitude (automatic Δ 0.429 vs human Δ 0.286) suggests that point estimates of control strength are optimistic; cross-model rankings and sign patterns are on firmer ground.
- If the benchmark convention caught on, model cards might start reporting three numbers per attribute—attainment, enhancement, differentiation—which would change how buyers compare music generators.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces a counterfactual evaluation framework for text-to-music controllability, organized around matched neutral–A–B contrast families. For each family, a neutral prompt omits the scored attribute, while two otherwise matched treatments specify different target values; all three outputs share a family seed and are rendered through frozen native-interface adapters. The framework separates three operational quantities: occurrence (treatment agreement), enhancement (change relative to the matched neutral output, Δ), and differentiation (margin between the two target arms). The method is applied to global key and three- versus four-beat grouping in ACE-Step 1.5, Stable Audio 3 Medium, and LeVo2, with 256 families per system. The main empirical findings are that ACE-Step 1.5 and Stable Audio 3 exhibit substantial key control while LeVo2 does not; for beat grouping, both responsive systems redirect toward the rare three-beat target, but high four-beat agreement largely reflects the neutral output prior rather than instruction-attributable control, with Stable Audio 3 showing a large negative four-beat Δ. The paper includes extensive validation: external recognizer benchmarks, blind expert audits, off-attribute placebos, three-seed sentinels, and a reproduction package. The central methodological claim is that agreement measures occurrence, whereas matched contrasts test instruction-attributable response.
Significance. The paper's main contribution is methodological: it makes explicit the distinction between target occurrence and instruction-attributable control, and it provides a concrete experimental template for measuring this distinction. The empirical results, if valid, change how prompted-agreement numbers should be interpreted for common target classes, and they demonstrate that a model can appear to follow a four-beat instruction while actually inheriting the four-beat grouping from its output prior. The strengths of the paper are notable: the decision rule is pre-specified, inference is performed on complete families with bootstrap methods, adapters are frozen and documented, outputs are validated against external recognizers and human audits, and the entire pipeline is packaged for reproduction. The paper also provides several falsifiable predictions, such as the sign of the four-beat reversal and the absence of LeVo2 key control, which are directly testable by other groups. If the measurement-validity caveats are resolved, the paper would be a valuable addition to the music-generation evaluation literature.
major comments (2)
- [§4.2, §7] The human validation is task-level, not model-specific. The central claim that LeVo2 shows no key control (Δ = 0.026 [0.003, 0.049]) depends on the assumption that S-KEY is as sensitive on LeVo2's generated audio distribution as on ACE-Step and Stable Audio 3 outputs. If S-KEY is less discriminative on LeVo2's typical pitch-class distributions or timbres, the near-zero Δ could be attenuation rather than absent control. The 28-family audit with a single bridge rater is explicitly described in Section 7 as not supporting per-model human rankings. Please add a model-specific human audit of LeVo2 key families (even a modest sample) or a sensitivity analysis that recalibrates S-KEY's decision threshold separately for each model and shows that LeVo2's near-zero Δ and the separation between systems persist.
- [§4.1, Appendix A.4] The confusion-matrix transport analysis addresses average measurement bias, not condition-specific error-rate differences. The headline Stable Audio 3 four-beat reversal (Δ = −0.406 [−0.516, −0.297]) could be inflated if Beat This's known tendency to make 4-to-2 half-bar errors occurs more often under the explicit '4/4 time signature' treatment than under the neutral condition. The current transport scenarios apply the same real-music confusion matrices without modeling a condition-specific increase in 4-to-2 errors. A targeted analysis that measures Beat This's error pattern separately for neutral and 4/4-generated clips (e.g., via human downbeat annotation on a sample of Stable Audio 3 outputs, or by replicating the key result with a second beat tracker) would directly address this concern and would strengthen the attribution claim.
minor comments (5)
- [§3.1] The notation for Δₜ is introduced after the family-level definitions, but the family-level formula for Δ_f is written before the target-specific version. Consider adding a short transition sentence to clarify that Δ_t is the target-specific average of the corresponding family-level quantity.
- [Figure 2] The open circles for matched neutral rates are visually light in grayscale; increasing marker size or using different point shapes for neutral and treated conditions would improve readability.
- [Table 4] The Margin rows for beat are repeated for the 3-beat and 4-beat rows. This is correct but may confuse readers into thinking there are two independent estimates; a note or a merged cell would be clearer.
- [§5.2] The sentence 'ACE-Step changes little under a four-beat instruction' is a fair summary, but the corresponding Δ interval is [0.000, 0.109]; stating 'the interval includes zero' would be more precise.
- [Appendix A.2] The reproducibility boundary for ACE-Step (no contemporaneous repository/checkpoint pin) is mentioned only in an appendix. Since reproducibility is a central claim of the paper, consider noting this caveat briefly in the main text's measurement-credibility section.
Circularity Check
No circularity: the counterfactual contrasts are measured with external frozen recognizers and independently validated, so the central claims do not reduce to their inputs.
full rationale
The paper's derivation chain is not circular. The central quantities Acc_f, Δ_f, and Margin_f are pre-specified functions of recognizer scores applied to frozen generated outputs, and the empirical claims are direct measurements of these contrasts rather than fitted parameters later relabeled as predictions. The outcome measures S-KEY and Beat This are external, frozen models, validated on real-music corpora (GTZAN, GiantSteps, RWC, Ballroom) and audited with blind expert labels; no parameter of the evaluation is fit to the generated audio in a way that reappears as the reported effect sizes. The main caveats, such as the automatic key Δ being larger than the human-audited Δ (0.429 vs. 0.286) and the small number of LeVo2 key families in the audit, concern measurement sensitivity and are explicitly acknowledged limitations, not circularity. The paper also contains no load-bearing self-citation: the cited system papers, recognizers, and behavioral-testing references are all external. Because the conclusions are not forced by the definitions or by any fitted input, the circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption S-KEY and Beat This produce valid measurements of global key and beat grouping on generated audio.
- domain assumption The canonical context matrix (four genres, instrumentation palettes, BPMs) is representative for the benchmark claims.
- domain assumption A neutral prompt that omits the target attribute is a valid counterfactual baseline.
- domain assumption The modal number of beats between detected downbeats is a faithful operationalization of three- versus four-beat grouping.
Cite this review
Pith. "Pith review of Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping." pith.science (2026). https://pith.science/paper/LF7ZZT23
@misc{pith2026260811899,
author = {Pith},
title = {Pith review of: Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping},
year = {2026},
howpublished = {\url{https://pith.science/paper/LF7ZZT23}},
note = {Machine review of arXiv:2608.11899}
}
read the original abstract
Prompted attribute agreement is widely used as evidence of text-to-music controllability, yet a requested attribute may occur simply because it is already common in the model's output distribution. We introduce a matched counterfactual evaluation that separates target occurrence from instruction-attributable control. Each family contains a neutral input that omits the scored attribute and two otherwise matched inputs that swap the requested target. All three are rendered through frozen native-interface adapters with a shared seed. Applied to global key and beat grouping in three open systems, this design changes the empirical conclusion. ACE-Step 1.5 and Stable Audio 3 Medium exhibit substantial key control, whereas LeVo2 does not. For beat grouping, the same models redirect toward the rare three-beat target, but high four-beat agreement is largely inherited from neutral outputs: Stable Audio 3 produces four-beat grouping in 0.97 of neutral cases but only 0.56 under its explicit four-beat treatment. Off-attribute placebos, external recognizer validation, blind expert annotation, and multi-seed sentinels support the attribution. When targets have unequal output priors, agreement describes what a model produced, while matched neutral and target-swap contrasts test whether the instruction changed it.
Figures
Reference graph
Works this paper leans on
-
[1]
Copet, Jade and Kreuk, Felix and Gat, Itai and Remez, Tal and Kant, David and Synnaeve, Gabriel and Adi, Yossi and D. Simple and. doi:10.48550/arXiv.2306.05284 , urldate =. arXiv , langid =:2306.05284 , primaryclass =
-
[2]
Agostinelli, Andrea and Denk, Timo I. and Borsos, Zal. doi:10.48550/arXiv.2301.11325 , urldate =. arXiv , langid =:2301.11325 , primaryclass =
-
[3]
doi:10.48550/arXiv.2506.00045 , urldate =
Gong, Junmin and Zhao, Sean and Wang, Sen and Xu, Shengyuan and Guo, Joe , year = 2025, month = may, number =. doi:10.48550/arXiv.2506.00045 , urldate =. arXiv , langid =:2506.00045 , primaryclass =
-
[4]
doi:10.48550/arXiv.2602.00744 , urldate =
Gong, Junmin and Song, Yulin and Zhao, Wenxiao and Wang, Sen and Xu, Shengyuan and Guo, Jing and Yang, Xuerui , year = 2026, month = feb, number =. doi:10.48550/arXiv.2602.00744 , urldate =. arXiv , langid =:2602.00744 , primaryclass =
-
[5]
doi:10.48550/arXiv.2407.15060 , urldate =
Lan, Yun-Han and Hsiao, Wen-Yi and Cheng, Hao-Chung and Yang, Yi-Hsuan , year = 2024, number =. doi:10.48550/arXiv.2407.15060 , urldate =. arXiv , langid =:2407.15060 , publisher =
-
[6]
doi:10.48550/arXiv.2501.10811 , urldate =
Liu, Cheng and Wang, Hui and Zhao, Jinghua and Zhao, Shiwan and Bu, Hui and Xu, Xin and Zhou, Jiaming and Sun, Haoqin and Qin, Yong , year = 2025, month = mar, number =. doi:10.48550/arXiv.2501.10811 , urldate =. arXiv , keywords =:2501.10811 , primaryclass =
-
[7]
doi:10.48550/arXiv.2506.12285 , urldate =
Ma, Yinghao and Li, Siyou and Yu, Juntao and Benetos, Emmanouil and Maezawa, Akira , year = 2025, month = jun, number =. doi:10.48550/arXiv.2506.12285 , urldate =. arXiv , keywords =:2506.12285 , primaryclass =
-
[8]
doi:10.48550/arXiv.2603.00610 , urldate =
Ma, Yinghao and Xia, Haiwen and Gao, Hewei and Chen, Weixiong and Ye, Yuxin and Yang, Yuchen and Chang, Sungkyun and Ding, Mingshuo and Li, Yizhi and Yuan, Ruibin and Dixon, Simon and Benetos, Emmanouil , year = 2026, month = jun, number =. doi:10.48550/arXiv.2603.00610 , urldate =. arXiv , keywords =:2603.00610 , primaryclass =
Show all 56 references
-
[9]
doi:10.48550/arXiv.2503.08638 , urldate =
Yuan, Ruibin and Lin, Hanfeng and Guo, Shuyue and Zhang, Ge and Pan, Jiahao and Zang, Yongyi and Liu, Haohe and Liang, Yiming and Ma, Wenye and Du, Xingjian and Du, Xinrun and Ye, Zhen and Zheng, Tianyu and Jiang, Zhengxuan and Ma, Yinghao and Liu, Minghao and Tian, Zeyue and ...
2025 doi
-
[10]
doi:10.48550/arXiv.2509.23350 , urldate =
Zhao, Jiahao and Li, Yunjia and Li, Wei and Yoshii, Kazuyoshi , year = 2025, number =. doi:10.48550/arXiv.2509.23350 , urldate =. arXiv , langid =:2509.23350 , publisher =
2025 doi
- [11]
-
[12]
doi:10.48550/arXiv.2510.22950 , urldate =
Jiang, Yuepeng and Chen, Huakang and Ning, Ziqian and Yao, Jixun and Han, Zerui and Wu, Di and Meng, Meng and Luan, Jian and Fu, Zhonghua and Xie, Lei , year = 2026, month = feb, number =. doi:10.48550/arXiv.2510.22950 , urldate =. arXiv , langid =:2510.22950 , primaryclass =
2026 doi
-
[13]
doi:10.48550/arXiv.2506.07520 , urldate =
Lei, Shun and Xu, Yaoxun and Lin, Zhiwei and Zhang, Huaicheng and Tan, Wei and Chen, Hangting and Yu, Jianwei and Zhang, Yixuan and Yang, Chenyu and Zhu, Haina and Wang, Shuai and Wu, Zhiyong and Yu, Dong , year = 2025, month = oct, number =. doi:10.48550/arXiv.2506.07520 , ur...
2025 doi
-
[14]
Zhu, Jinlong and Sakurai, Keigo and Togo, Ren and Ogawa, Takahiro and Haseyama, Miki , year = 2026, journal =
2026
- [15]
-
[16]
and Rice, Matthew and Carr, C
Evans, Zach and Parker, Julian D. and Rice, Matthew and Carr, C. J. and Zukowski, Zack and Taylor, Josiah and Pons, Jordi , year = 2026, month = may, number =. Stable. doi:10.48550/arXiv.2605.17991 , urldate =. arXiv , keywords =:2605.17991 , primaryclass =
- [17]
- [18]
- [19]
-
[20]
doi:10.48550/arXiv.2505.10793 , urldate =
Yao, Jixun and Ma, Guobin and Xue, Huixin and Chen, Huakang and Hao, Chunbo and Jiang, Yuepeng and Liu, Haohe and Yuan, Ruibin and Xu, Jin and Xue, Wei and Liu, Hao and Xie, Lei , year = 2025, month = may, number =. doi:10.48550/arXiv.2505.10793 , urldate =. arXiv , keywords =...
- [21]
-
[22]
doi:10.48550/arXiv.2606.30642 , archiveprefix =
Lei, Shun and Zhang, Huaicheng and Wu, Dapeng and Xu, Yaoxun and Zuo, Lishi and Tan, Wei and Chen, Hangting and Li, Guangzheng and Yu, Jianwei and Wu, Zhiyong and Yu, Dong , year = 2026, month = jun, number =. doi:10.48550/arXiv.2606.30642 , archiveprefix =. 2606.30642 , prima...
-
[23]
Proceedings of the International Computer Music Conference , address =
Genre-Specific Key Profiles , author =. Proceedings of the International Computer Music Conference , address =
-
[25]
Proceedings of the 7th International Conference on Music Information Retrieval , url =
Goto, Masataka , year = 2006, pages =. Proceedings of the 7th International Conference on Music Information Retrieval , url =
2006
-
[26]
and Salamon, Justin and Nieto, Oriol and Liang, Dawen and Ellis, Daniel P
Raffel, Colin and McFee, Brian and Humphrey, Eric J. and Salamon, Justin and Nieto, Oriol and Liang, Dawen and Ellis, Daniel P. W. , year = 2014, pages =. Proceedings of the 15th International Society for Music Information Retrieval Conference , address =
2014
-
[29]
Proceedings of the 16th International Society for Music Information Retrieval Conference , address =
Two Data Sets for Tempo Estimation and Key Detection in Electronic Dance Music Annotated from User Corrections , author =. Proceedings of the 16th International Society for Music Information Retrieval Conference , address =
-
[30]
Proceedings of the AES 25th International Conference on Metadata for Audio , address =
Evaluating Rhythmic Descriptors for Musical Genre Classification , author =. Proceedings of the AES 25th International Conference on Metadata for Audio , address =
-
[31]
2015 , address =
Marchand, Ugo and Fresnel, Quentin and Peeters, Geoffroy , booktitle =. 2015 , address =
2015
-
[32]
Beyond Accuracy: Behavioral Testing of
Ribeiro, Marco Tulio and Wu, Tongshuang and Guestrin, Carlos and Singh, Sameer , booktitle =. Beyond Accuracy: Behavioral Testing of
-
[33]
, year = 2026, month = may, number =
Yang, Zihao and Levy, Mosh and Goldberg, Yoav and Wallace, Byron C. , year = 2026, month = may, number =. Compared to. 2605.01048 , archiveprefix =
2026 arXiv
-
[34]
Andrea Agostinelli, Timo I. Denk, Zal \'a n Borsos, Jesse Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, Matt Sharifi, Neil Zeghidour, and Christian Frank. MusicLM : Generating Music From Text , January 2023
2023
-
[35]
Simple and Controllable Music Generation , January 2024
Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre D \'e fossez. Simple and Controllable Music Generation , January 2024
2024
-
[36]
Parker, C
Zach Evans, Julian D. Parker, C. J. Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons. Stable Audio Open , July 2024
2024
-
[37]
Parker, Matthew Rice, C
Zach Evans, Julian D. Parker, Matthew Rice, C. J. Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons. Stable Audio 3, May 2026
2026
-
[38]
Beat this! Accurate beat tracking without DBN postprocessing
Francesco Foscarin, Jan Schl \"u ter, and Gerhard Widmer. Beat this! Accurate beat tracking without DBN postprocessing. In Proceedings of the 25th International Society for Music Information Retrieval Conference, San Francisco, CA, USA, 2024. doi:10.48550/arXiv.2407.21658. URL...
-
[39]
ACE-Step 1.5: Pushing the Boundaries of Open-Source Music Generation , February 2026
Junmin Gong, Yulin Song, Wenxiao Zhao, Sen Wang, Shengyuan Xu, Jing Guo, and Xuerui Yang. ACE-Step 1.5: Pushing the Boundaries of Open-Source Music Generation , February 2026
2026
-
[40]
S-KEY : Self-supervised Learning of Major and Minor Keys from Audio
Yuexuan Kong, Gabriel Meseguer-Brocal, Vincent Lostanlen, Mathieu Lagrange, and Romain Hennequin. S-KEY : Self-supervised Learning of Major and Minor Keys from Audio . arXiv preprint arXiv:2501.12907, 2025. doi:10.48550/arXiv.2501.12907. URL https://arxiv.org/abs/2501.12907
-
[41]
MusiConGen : Rhythm and chord control for transformer-based text-to-music generation, 2024
Yun-Han Lan, Wen-Yi Hsiao, Hao-Chung Cheng, and Yi-Hsuan Yang. MusiConGen : Rhythm and chord control for transformer-based text-to-music generation, 2024
2024
-
[42]
LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training , June 2026
Shun Lei, Huaicheng Zhang, Dapeng Wu, Yaoxun Xu, Lishi Zuo, Wei Tan, Hangting Chen, Guangzheng Li, Jianwei Yu, Zhiyong Wu, and Dong Yu. LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training , June 2026
2026
-
[43]
MusicEval : A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation , March 2025
Cheng Liu, Hui Wang, Jinghua Zhao, Shiwan Zhao, Hui Bu, Xin Xu, Jiaming Zhou, Haoqin Sun, and Yong Qin. MusicEval : A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation , March 2025
2025
-
[44]
MuseCoco : Generating Symbolic Music from Text , May 2023
Peiling Lu, Xin Xu, Chenfei Kang, Botao Yu, Chengyi Xing, Xu Tan, and Jiang Bian. MuseCoco : Generating Symbolic Music from Text , May 2023
2023
-
[45]
CMI-Bench : A Comprehensive Benchmark for Evaluating Music Instruction Following , June 2025
Yinghao Ma, Siyou Li, Juntao Yu, Emmanouil Benetos, and Akira Maezawa. CMI-Bench : A Comprehensive Benchmark for Evaluating Music Instruction Following , June 2025
2025
-
[46]
GTZAN-Rhythm : Extending the GTZAN test-set with beat, downbeat and swing annotations
Ugo Marchand, Quentin Fresnel, and Geoffroy Peeters. GTZAN-Rhythm : Extending the GTZAN test-set with beat, downbeat and swing annotations. In Late-Breaking/Demo Session of the 16th International Society for Music Information Retrieval Conference, Malaga, Spain, 2015
2015
-
[47]
Mustango: Toward Controllable Text-to-Music Generation , June 2024
Jan Melechovsky, Zixun Guo, Deepanway Ghosal, Navonil Majumder, Dorien Herremans, and Soujanya Poria. Mustango: Toward Controllable Text-to-Music Generation , June 2024
2024
-
[48]
Genre-specific key profiles
Cian O'Brien and Alexander Lerch. Genre-specific key profiles. In Proceedings of the International Computer Music Conference, Denton, TX, USA, 2015. URL https://musicinformatics.gatech.edu/wp-content_nondefault/uploads/2015/09/O
2015
-
[49]
Humphrey, Justin Salamon, Oriol Nieto, Dawen Liang, and Daniel P
Colin Raffel, Brian McFee, Eric J. Humphrey, Justin Salamon, Oriol Nieto, Dawen Liang, and Daniel P. W. Ellis. mir\_eval : A Transparent Implementation of Common MIR Metrics . In Proceedings of the 15th International Society for Music Information Retrieval Conference, pages 36...
2014
-
[50]
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. Beyond accuracy: Behavioral testing of NLP models with CheckList . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4902--4912, 2020
2020
- [51]
-
[52]
FIGARO : Generating Symbolic Music with Fine-Grained Artistic Control , February 2024
Dimitri von R \"u tte, Luca Biggio, Yannic Kilcher, and Thomas Hofmann. FIGARO : Generating Symbolic Music with Fine-Grained Artistic Control , February 2024
2024
-
[53]
MuseMorphose : Full-Song and Fine-Grained Piano Music Style Transfer with One Transformer VAE , December 2022
Shih-Lun Wu and Yi-Hsuan Yang. MuseMorphose : Full-Song and Fine-Grained Piano Music Style Transfer with One Transformer VAE , December 2022
2022
-
[54]
Shih-Lun Wu, Chris Donahue, Shinji Watanabe, and Nicholas J. Bryan. Music ControlNet : Multiple time-varying controls for music generation, 2023
2023
-
[55]
Zihao Yang, Mosh Levy, Yoav Goldberg, and Byron C. Wallace. Compared to What ? Baselines and Metrics for Counterfactual Prompting , May 2026
2026
-
[56]
SongEval : A Benchmark Dataset for Song Aesthetics Evaluation , May 2025
Jixun Yao, Guobin Ma, Huixin Xue, Huakang Chen, Chunbo Hao, Yuepeng Jiang, Haohe Liu, Ruibin Yuan, Jin Xu, Wei Xue, Hao Liu, and Lei Xie. SongEval : A Benchmark Dataset for Song Aesthetics Evaluation , May 2025
2025
-
[57]
YuE : Scaling Open Foundation Models for Long-Form Music Generation , September 2025
Ruibin Yuan, Hanfeng Lin, Shuyue Guo, Ge Zhang, Jiahao Pan, Yongyi Zang, Haohe Liu, Yiming Liang, Wenye Ma, Xingjian Du, Xinrun Du, Zhen Ye, Tianyu Zheng, Zhengxuan Jiang, Yinghao Ma, Minghao Liu, Zeyue Tian, Ziya Zhou, Liumeng Xue, Xingwei Qu, Yizhi Li, Shangda Wu, Tianhao Sh...
2025
-
[58]
ABC-eval : Benchmarking large language models on symbolic music understanding and instruction following, 2025
Jiahao Zhao, Yunjia Li, Wei Li, and Kazuyoshi Yoshii. ABC-eval : Benchmarking large language models on symbolic music understanding and instruction following, 2025
2025
-
[59]
TPSMG : Text-Controllable Polyphonic Symbolic Music Generation
Jinlong Zhu, Keigo Sakurai, Ren Togo, Takahiro Ogawa, and Miki Haseyama. TPSMG : Text-Controllable Polyphonic Symbolic Music Generation . ITE Transactions on Media Technology and Applications, 14 0 (1): 0 110--118, 2026
2026
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.