Pith. sign in

REVIEW 2 major objections 5 minor 56 references

Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping

T0 review · 2 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Prompted agreement measures occurrence, not instruction following; matched neutral and target-swap contrasts show which music models actually respond to key and beat requests.

desk verdict A careful, well-executed demonstration that prompted agreement can masquerade as control in text-to-music; the soft spot is model-specific auto-scorer validity, but the paper's own checks keep the conclusion credible. read the letter →

arxiv 2608.11899 v1 pith:LF7ZZT23 submitted 2026-08-12 cs.SD

classification cs.SD
keywords text-to-musicgenerationcontrollabilityevaluationcounterfactualpromptingkeyrecognitionbeatgroupingpromptagreementoutputpriorsinstructionfollowing
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Prompted agreement—checking whether a requested attribute appears in generated audio—is the standard evidence that a text-to-music system follows instructions. This paper argues that agreement is not enough: a target may appear simply because it is already common in the model's output. It introduces matched neutral–A–B contrast families, where one prompt omits the scored attribute and two otherwise identical prompts request different values, and measures whether the instruction raises occurrence above the matched neutral output and whether swapping the target redirects the output. Applied to key and beat grouping in three open systems, this changes the conclusions: two systems genuinely control key, none of the three shows positive four-beat enhancement, and the high four-beat agreement of Stable Audio 3 is largely inherited from its default output. If the argument is right, controllability benchmarks should report occurrence, enhancement, and differentiation separately instead of a single agreement score.

What carries the argument

The central object is the matched neutral–A–B contrast family: one neutral input that omits the scored attribute and two otherwise matched inputs that request different target values, all rendered through the system's frozen native interface with a shared seed. The argument is carried by three statistics derived from it—Acc (did the requested target occur under treatment), Δ (did treatment raise occurrence above the matched neutral output), and Margin (did swapping the requested value redirect the output toward the alternative target). A confirmatory label of effective control requires both Δ and Margin intervals to lie above zero; the paper also uses off-attribute placebos and multi-seed sentinels to rule out generic added-instruction effects and single-generation noise.

What would settle it

An independent blind human relabeling of the released 256-family benchmark per system, or a rerun with a different validated key and beat recognizer, would settle the claim: if LeVo2 shows a key Δ interval clearly above zero, or Stable Audio 3 loses its negative four-beat Δ, the paper's central contrast would not survive. The paper's own audit already shows that automatic key Δ (0.429) exceeds human key Δ (0.286) on a 28-family subset, so a full human relabeling is the direct check.

Watch

Extended reading notes

Core claim

The central empirical claim is that ACE-Step 1.5 and Stable Audio 3 Medium show large key control that is not explained by their output priors, while LeVo2 shows little attributable key response under its evaluated interface; and that for beat grouping the models redirect toward the rare three-beat target, while the apparently high four-beat agreement is a prior artifact. The central methodological claim is that agreement measures occurrence, whereas matched neutral and target-swap contrasts test instruction-attributable response. The paper reports ACE-Step key Δ = 0.612 and Stable Audio 3 key Δ = 0.646 against LeVo2's 0.026, and for beat grouping Stable Audio 3 three-beat Δ = +0.469 against four-beat Δ = −0.406, with the neutral four-beat rate at 0.969 and the explicit four-beat treatment at 0.563. The authors conclude that high prompted agreement can coexist with no improvement—or a large reversal—relative to the model's own neutral output.

Load-bearing premise

The conclusions stand or fall on the validity of the frozen automatic scorers—S-KEY for key and Beat This for beat grouping—as outcome measures; if those recognizers mislabel generated audio in a target-dependent way, the reported Δ and Margin values would not reflect what the models actually produced.

Editorial extensions

If this is right

  • Prompted agreement alone should not be used as evidence of controllability for any target whose output prior is high; the matched neutral rate must be reported.
  • Controllability should be described per target, not as an aggregate score, because four-beat and three-beat responses can cancel in the average.
  • If the result transfers, generative-model benchmarks outside music—faces, daylight scenes, common code patterns—should adopt neutral-relative and target-swap contrasts.
  • A negative four-beat Δ for Stable Audio 3 shows that a treatment can move output away from the common target; benchmarks need to allow for this rather than treating any agreement as success.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension is to cross carrier templates with target values, since the paper's neutral contrast changes both target presence and its rendered carrier; such a factorial design could estimate carrier–value interactions rather than only end-to-end response.
  • The same design could be applied to continuous controls such as BPM or loudness, but the neutral-omission construction may need a neighbouring-value reference instead of an omitted field.
  • The human-audit gap in key magnitude (automatic Δ 0.429 vs human Δ 0.286) suggests that point estimates of control strength are optimistic; cross-model rankings and sign patterns are on firmer ground.
  • If the benchmark convention caught on, model cards might start reporting three numbers per attribute—attainment, enhancement, differentiation—which would change how buyers compare music generators.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper introduces a counterfactual evaluation framework for text-to-music controllability, organized around matched neutral–A–B contrast families. For each family, a neutral prompt omits the scored attribute, while two otherwise matched treatments specify different target values; all three outputs share a family seed and are rendered through frozen native-interface adapters. The framework separates three operational quantities: occurrence (treatment agreement), enhancement (change relative to the matched neutral output, Δ), and differentiation (margin between the two target arms). The method is applied to global key and three- versus four-beat grouping in ACE-Step 1.5, Stable Audio 3 Medium, and LeVo2, with 256 families per system. The main empirical findings are that ACE-Step 1.5 and Stable Audio 3 exhibit substantial key control while LeVo2 does not; for beat grouping, both responsive systems redirect toward the rare three-beat target, but high four-beat agreement largely reflects the neutral output prior rather than instruction-attributable control, with Stable Audio 3 showing a large negative four-beat Δ. The paper includes extensive validation: external recognizer benchmarks, blind expert audits, off-attribute placebos, three-seed sentinels, and a reproduction package. The central methodological claim is that agreement measures occurrence, whereas matched contrasts test instruction-attributable response.

Significance. The paper's main contribution is methodological: it makes explicit the distinction between target occurrence and instruction-attributable control, and it provides a concrete experimental template for measuring this distinction. The empirical results, if valid, change how prompted-agreement numbers should be interpreted for common target classes, and they demonstrate that a model can appear to follow a four-beat instruction while actually inheriting the four-beat grouping from its output prior. The strengths of the paper are notable: the decision rule is pre-specified, inference is performed on complete families with bootstrap methods, adapters are frozen and documented, outputs are validated against external recognizers and human audits, and the entire pipeline is packaged for reproduction. The paper also provides several falsifiable predictions, such as the sign of the four-beat reversal and the absence of LeVo2 key control, which are directly testable by other groups. If the measurement-validity caveats are resolved, the paper would be a valuable addition to the music-generation evaluation literature.

major comments (2)
  1. [§4.2, §7] The human validation is task-level, not model-specific. The central claim that LeVo2 shows no key control (Δ = 0.026 [0.003, 0.049]) depends on the assumption that S-KEY is as sensitive on LeVo2's generated audio distribution as on ACE-Step and Stable Audio 3 outputs. If S-KEY is less discriminative on LeVo2's typical pitch-class distributions or timbres, the near-zero Δ could be attenuation rather than absent control. The 28-family audit with a single bridge rater is explicitly described in Section 7 as not supporting per-model human rankings. Please add a model-specific human audit of LeVo2 key families (even a modest sample) or a sensitivity analysis that recalibrates S-KEY's decision threshold separately for each model and shows that LeVo2's near-zero Δ and the separation between systems persist.
  2. [§4.1, Appendix A.4] The confusion-matrix transport analysis addresses average measurement bias, not condition-specific error-rate differences. The headline Stable Audio 3 four-beat reversal (Δ = −0.406 [−0.516, −0.297]) could be inflated if Beat This's known tendency to make 4-to-2 half-bar errors occurs more often under the explicit '4/4 time signature' treatment than under the neutral condition. The current transport scenarios apply the same real-music confusion matrices without modeling a condition-specific increase in 4-to-2 errors. A targeted analysis that measures Beat This's error pattern separately for neutral and 4/4-generated clips (e.g., via human downbeat annotation on a sample of Stable Audio 3 outputs, or by replicating the key result with a second beat tracker) would directly address this concern and would strengthen the attribution claim.
minor comments (5)
  1. [§3.1] The notation for Δₜ is introduced after the family-level definitions, but the family-level formula for Δ_f is written before the target-specific version. Consider adding a short transition sentence to clarify that Δ_t is the target-specific average of the corresponding family-level quantity.
  2. [Figure 2] The open circles for matched neutral rates are visually light in grayscale; increasing marker size or using different point shapes for neutral and treated conditions would improve readability.
  3. [Table 4] The Margin rows for beat are repeated for the 3-beat and 4-beat rows. This is correct but may confuse readers into thinking there are two independent estimates; a note or a merged cell would be clearer.
  4. [§5.2] The sentence 'ACE-Step changes little under a four-beat instruction' is a fair summary, but the corresponding Δ interval is [0.000, 0.109]; stating 'the interval includes zero' would be more precise.
  5. [Appendix A.2] The reproducibility boundary for ACE-Step (no contemporaneous repository/checkpoint pin) is mentioned only in an appendix. Since reproducibility is a central claim of the paper, consider noting this caveat briefly in the main text's measurement-credibility section.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the counterfactual contrasts are measured with external frozen recognizers and independently validated, so the central claims do not reduce to their inputs.

full rationale

The paper's derivation chain is not circular. The central quantities Acc_f, Δ_f, and Margin_f are pre-specified functions of recognizer scores applied to frozen generated outputs, and the empirical claims are direct measurements of these contrasts rather than fitted parameters later relabeled as predictions. The outcome measures S-KEY and Beat This are external, frozen models, validated on real-music corpora (GTZAN, GiantSteps, RWC, Ballroom) and audited with blind expert labels; no parameter of the evaluation is fit to the generated audio in a way that reappears as the reported effect sizes. The main caveats, such as the automatic key Δ being larger than the human-audited Δ (0.429 vs. 0.286) and the small number of LeVo2 key families in the audit, concern measurement sensitivity and are explicitly acknowledged limitations, not circularity. The paper also contains no load-bearing self-citation: the cited system papers, recognizers, and behavioral-testing references are all external. Because the conclusions are not forced by the definitions or by any fitted input, the circularity score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper introduces no free parameters fitted to generated data and no new entities. Its claims rest on measurement validity of frozen recognizers and on the design choice that neutral prompts are meaningful baselines; both are argued with validation.

assumptions (4)
  • domain assumption S-KEY and Beat This produce valid measurements of global key and beat grouping on generated audio.
    Invoked in Section 3.3 for scoring; validation in Sections 4.1 and 4.2 shows real-music calibration and a small human audit. The audit found automatic key delta of 0.429 vs human 0.286, so magnitudes are approximate.
  • domain assumption The canonical context matrix (four genres, instrumentation palettes, BPMs) is representative for the benchmark claims.
    Defined in Section 3.2 and Appendix A.1; claims are scoped to these frozen contexts.
  • domain assumption A neutral prompt that omits the target attribute is a valid counterfactual baseline.
    The paper acknowledges in Section 7 that the neutral contrast changes both the target and its carrier; the validity of delta depends on this construction.
  • domain assumption The modal number of beats between detected downbeats is a faithful operationalization of three- versus four-beat grouping.
    Defined in Sections 3.2 and A.3.1; the paper leaves 3/4 versus 6/8 ambiguity visible.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping." pith.science (2026). https://pith.science/paper/LF7ZZT23

@misc{pith2026260811899,
  author       = {Pith},
  title        = {Pith review of: Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LF7ZZT23}},
  note         = {Machine review of arXiv:2608.11899}
}
read the original abstract

Prompted attribute agreement is widely used as evidence of text-to-music controllability, yet a requested attribute may occur simply because it is already common in the model's output distribution. We introduce a matched counterfactual evaluation that separates target occurrence from instruction-attributable control. Each family contains a neutral input that omits the scored attribute and two otherwise matched inputs that swap the requested target. All three are rendered through frozen native-interface adapters with a shared seed. Applied to global key and beat grouping in three open systems, this design changes the empirical conclusion. ACE-Step 1.5 and Stable Audio 3 Medium exhibit substantial key control, whereas LeVo2 does not. For beat grouping, the same models redirect toward the rare three-beat target, but high four-beat agreement is largely inherited from neutral outputs: Stable Audio 3 produces four-beat grouping in 0.97 of neutral cases but only 0.56 under its explicit four-beat treatment. Off-attribute placebos, external recognizer validation, blind expert annotation, and multi-seed sentinels support the attribution. When targets have unequal output priors, agreement describes what a model produced, while matched neutral and target-swap contrasts test whether the instruction changed it.

Figures

Figures reproduced from arXiv: 2608.11899 by the authors.

Figure 1
Figure 1. Counterfactual evaluation pipeline, illustrated with the open-text beat rendering. A matched neutral–A–B triplet shares one family [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Target occurrence versus instruction-attributable response. Open circles are matched neutral target rates, filled markers are [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Specificity and measurement checks. (a) Filled markers show real beat instructions, while open gray markers show off-attribute [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

56 extracted references · 38 canonical work pages

  1. [1]

    Simple and

    Copet, Jade and Kreuk, Felix and Gat, Itai and Remez, Tal and Kant, David and Synnaeve, Gabriel and Adi, Yossi and D. Simple and. doi:10.48550/arXiv.2306.05284 , urldate =. arXiv , langid =:2306.05284 , primaryclass =

  2. [2]

    and Borsos, Zal

    Agostinelli, Andrea and Denk, Timo I. and Borsos, Zal. doi:10.48550/arXiv.2301.11325 , urldate =. arXiv , langid =:2301.11325 , primaryclass =

  3. [3]

    doi:10.48550/arXiv.2506.00045 , urldate =

    Gong, Junmin and Zhao, Sean and Wang, Sen and Xu, Shengyuan and Guo, Joe , year = 2025, month = may, number =. doi:10.48550/arXiv.2506.00045 , urldate =. arXiv , langid =:2506.00045 , primaryclass =

  4. [4]

    doi:10.48550/arXiv.2602.00744 , urldate =

    Gong, Junmin and Song, Yulin and Zhao, Wenxiao and Wang, Sen and Xu, Shengyuan and Guo, Jing and Yang, Xuerui , year = 2026, month = feb, number =. doi:10.48550/arXiv.2602.00744 , urldate =. arXiv , langid =:2602.00744 , primaryclass =

  5. [5]

    doi:10.48550/arXiv.2407.15060 , urldate =

    Lan, Yun-Han and Hsiao, Wen-Yi and Cheng, Hao-Chung and Yang, Yi-Hsuan , year = 2024, number =. doi:10.48550/arXiv.2407.15060 , urldate =. arXiv , langid =:2407.15060 , publisher =

  6. [6]

    doi:10.48550/arXiv.2501.10811 , urldate =

    Liu, Cheng and Wang, Hui and Zhao, Jinghua and Zhao, Shiwan and Bu, Hui and Xu, Xin and Zhou, Jiaming and Sun, Haoqin and Qin, Yong , year = 2025, month = mar, number =. doi:10.48550/arXiv.2501.10811 , urldate =. arXiv , keywords =:2501.10811 , primaryclass =

  7. [7]

    doi:10.48550/arXiv.2506.12285 , urldate =

    Ma, Yinghao and Li, Siyou and Yu, Juntao and Benetos, Emmanouil and Maezawa, Akira , year = 2025, month = jun, number =. doi:10.48550/arXiv.2506.12285 , urldate =. arXiv , keywords =:2506.12285 , primaryclass =

  8. [8]

    doi:10.48550/arXiv.2603.00610 , urldate =

    Ma, Yinghao and Xia, Haiwen and Gao, Hewei and Chen, Weixiong and Ye, Yuxin and Yang, Yuchen and Chang, Sungkyun and Ding, Mingshuo and Li, Yizhi and Yuan, Ruibin and Dixon, Simon and Benetos, Emmanouil , year = 2026, month = jun, number =. doi:10.48550/arXiv.2603.00610 , urldate =. arXiv , keywords =:2603.00610 , primaryclass =

Show all 56 references
  1. [9]

    doi:10.48550/arXiv.2503.08638 , urldate =

    Yuan, Ruibin and Lin, Hanfeng and Guo, Shuyue and Zhang, Ge and Pan, Jiahao and Zang, Yongyi and Liu, Haohe and Liang, Yiming and Ma, Wenye and Du, Xingjian and Du, Xinrun and Ye, Zhen and Zheng, Tianyu and Jiang, Zhengxuan and Ma, Yinghao and Liu, Minghao and Tian, Zeyue and ...

  2. [10]

    doi:10.48550/arXiv.2509.23350 , urldate =

    Zhao, Jiahao and Li, Yunjia and Li, Wei and Yoshii, Kazuyoshi , year = 2025, number =. doi:10.48550/arXiv.2509.23350 , urldate =. arXiv , langid =:2509.23350 , publisher =

  3. [11]

    , year = 2023, number =

    Wu, Shih-Lun and Donahue, Chris and Watanabe, Shinji and Bryan, Nicholas J. , year = 2023, number =. Music. doi:10.48550/arXiv.2311.07069 , urldate =. arXiv , langid =:2311.07069 , publisher =

  4. [12]

    doi:10.48550/arXiv.2510.22950 , urldate =

    Jiang, Yuepeng and Chen, Huakang and Ning, Ziqian and Yao, Jixun and Han, Zerui and Wu, Di and Meng, Meng and Luan, Jian and Fu, Zhonghua and Xie, Lei , year = 2026, month = feb, number =. doi:10.48550/arXiv.2510.22950 , urldate =. arXiv , langid =:2510.22950 , primaryclass =

  5. [13]

    doi:10.48550/arXiv.2506.07520 , urldate =

    Lei, Shun and Xu, Yaoxun and Lin, Zhiwei and Zhang, Huaicheng and Tan, Wei and Chen, Hangting and Yu, Jianwei and Zhang, Yixuan and Yang, Chenyu and Zhu, Haina and Wang, Shuai and Wu, Zhiyong and Yu, Dong , year = 2025, month = oct, number =. doi:10.48550/arXiv.2506.07520 , ur...

  6. [14]

    Zhu, Jinlong and Sakurai, Keigo and Togo, Ren and Ogawa, Takahiro and Haseyama, Miki , year = 2026, journal =

  7. [15]

    and Carr, C

    Evans, Zach and Parker, Julian D. and Carr, C. J. and Zukowski, Zack and Taylor, Josiah and Pons, Jordi , year = 2024, month = jul, number =. Stable. doi:10.48550/arXiv.2407.14358 , urldate =. arXiv , keywords =:2407.14358 , primaryclass =

  8. [16]

    and Rice, Matthew and Carr, C

    Evans, Zach and Parker, Julian D. and Rice, Matthew and Carr, C. J. and Zukowski, Zack and Taylor, Josiah and Pons, Jordi , year = 2026, month = may, number =. Stable. doi:10.48550/arXiv.2605.17991 , urldate =. arXiv , keywords =:2605.17991 , primaryclass =

  9. [17]

    Mustango:

    Melechovsky, Jan and Guo, Zixun and Ghosal, Deepanway and Majumder, Navonil and Herremans, Dorien and Poria, Soujanya , year = 2024, month = jun, number =. Mustango:. doi:10.48550/arXiv.2311.08355 , urldate =. arXiv , keywords =:2311.08355 , primaryclass =

  10. [18]

    doi:10.48550/arXiv.2306.00110 , urldate =

    Lu, Peiling and Xu, Xin and Kang, Chenfei and Yu, Botao and Xing, Chengyi and Tan, Xu and Bian, Jiang , year = 2023, month = may, number =. doi:10.48550/arXiv.2306.00110 , urldate =. arXiv , keywords =:2306.00110 , primaryclass =

  11. [19]

    doi:10.48550/arXiv.2105.04090 , urldate =

    Wu, Shih-Lun and Yang, Yi-Hsuan , year = 2022, month = dec, number =. doi:10.48550/arXiv.2105.04090 , urldate =. arXiv , keywords =:2105.04090 , primaryclass =

  12. [20]

    doi:10.48550/arXiv.2505.10793 , urldate =

    Yao, Jixun and Ma, Guobin and Xue, Huixin and Chen, Huakang and Hao, Chunbo and Jiang, Yuepeng and Liu, Haohe and Yuan, Ruibin and Xu, Jin and Xue, Wei and Liu, Hao and Xie, Lei , year = 2025, month = may, number =. doi:10.48550/arXiv.2505.10793 , urldate =. arXiv , keywords =...

  13. [21]

    doi:10.48550/arXiv.2201.10936 , urldate =

    von R. doi:10.48550/arXiv.2201.10936 , urldate =. arXiv , keywords =:2201.10936 , primaryclass =

  14. [22]

    doi:10.48550/arXiv.2606.30642 , archiveprefix =

    Lei, Shun and Zhang, Huaicheng and Wu, Dapeng and Xu, Yaoxun and Zuo, Lishi and Tan, Wei and Chen, Hangting and Li, Guangzheng and Yu, Jianwei and Wu, Zhiyong and Yu, Dong , year = 2026, month = jun, number =. doi:10.48550/arXiv.2606.30642 , archiveprefix =. 2606.30642 , prima...

  15. [23]

    Proceedings of the International Computer Music Conference , address =

    Genre-Specific Key Profiles , author =. Proceedings of the International Computer Music Conference , address =

  16. [25]

    Proceedings of the 7th International Conference on Music Information Retrieval , url =

    Goto, Masataka , year = 2006, pages =. Proceedings of the 7th International Conference on Music Information Retrieval , url =

  17. [26]

    and Salamon, Justin and Nieto, Oriol and Liang, Dawen and Ellis, Daniel P

    Raffel, Colin and McFee, Brian and Humphrey, Eric J. and Salamon, Justin and Nieto, Oriol and Liang, Dawen and Ellis, Daniel P. W. , year = 2014, pages =. Proceedings of the 15th International Society for Music Information Retrieval Conference , address =

  18. [29]

    Proceedings of the 16th International Society for Music Information Retrieval Conference , address =

    Two Data Sets for Tempo Estimation and Key Detection in Electronic Dance Music Annotated from User Corrections , author =. Proceedings of the 16th International Society for Music Information Retrieval Conference , address =

  19. [30]

    Proceedings of the AES 25th International Conference on Metadata for Audio , address =

    Evaluating Rhythmic Descriptors for Musical Genre Classification , author =. Proceedings of the AES 25th International Conference on Metadata for Audio , address =

  20. [31]

    2015 , address =

    Marchand, Ugo and Fresnel, Quentin and Peeters, Geoffroy , booktitle =. 2015 , address =

  21. [32]

    Beyond Accuracy: Behavioral Testing of

    Ribeiro, Marco Tulio and Wu, Tongshuang and Guestrin, Carlos and Singh, Sameer , booktitle =. Beyond Accuracy: Behavioral Testing of

  22. [33]

    , year = 2026, month = may, number =

    Yang, Zihao and Levy, Mosh and Goldberg, Yoav and Wallace, Byron C. , year = 2026, month = may, number =. Compared to. 2605.01048 , archiveprefix =

  23. [34]

    Andrea Agostinelli, Timo I. Denk, Zal \'a n Borsos, Jesse Engel, Mauro Verzetti, Antoine Caillon, Qingqing Huang, Aren Jansen, Adam Roberts, Marco Tagliasacchi, Matt Sharifi, Neil Zeghidour, and Christian Frank. MusicLM : Generating Music From Text , January 2023

  24. [35]

    Simple and Controllable Music Generation , January 2024

    Jade Copet, Felix Kreuk, Itai Gat, Tal Remez, David Kant, Gabriel Synnaeve, Yossi Adi, and Alexandre D \'e fossez. Simple and Controllable Music Generation , January 2024

  25. [36]

    Parker, C

    Zach Evans, Julian D. Parker, C. J. Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons. Stable Audio Open , July 2024

  26. [37]

    Parker, Matthew Rice, C

    Zach Evans, Julian D. Parker, Matthew Rice, C. J. Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons. Stable Audio 3, May 2026

  27. [38]

    Beat this! Accurate beat tracking without DBN postprocessing

    Francesco Foscarin, Jan Schl \"u ter, and Gerhard Widmer. Beat this! Accurate beat tracking without DBN postprocessing. In Proceedings of the 25th International Society for Music Information Retrieval Conference, San Francisco, CA, USA, 2024. doi:10.48550/arXiv.2407.21658. URL...

  28. [39]

    ACE-Step 1.5: Pushing the Boundaries of Open-Source Music Generation , February 2026

    Junmin Gong, Yulin Song, Wenxiao Zhao, Sen Wang, Shengyuan Xu, Jing Guo, and Xuerui Yang. ACE-Step 1.5: Pushing the Boundaries of Open-Source Music Generation , February 2026

  29. [40]

    S-KEY : Self-supervised Learning of Major and Minor Keys from Audio

    Yuexuan Kong, Gabriel Meseguer-Brocal, Vincent Lostanlen, Mathieu Lagrange, and Romain Hennequin. S-KEY : Self-supervised Learning of Major and Minor Keys from Audio . arXiv preprint arXiv:2501.12907, 2025. doi:10.48550/arXiv.2501.12907. URL https://arxiv.org/abs/2501.12907

  30. [41]

    MusiConGen : Rhythm and chord control for transformer-based text-to-music generation, 2024

    Yun-Han Lan, Wen-Yi Hsiao, Hao-Chung Cheng, and Yi-Hsuan Yang. MusiConGen : Rhythm and chord control for transformer-based text-to-music generation, 2024

  31. [42]

    LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training , June 2026

    Shun Lei, Huaicheng Zhang, Dapeng Wu, Yaoxun Xu, Lishi Zuo, Wei Tan, Hangting Chen, Guangzheng Li, Jianwei Yu, Zhiyong Wu, and Dong Yu. LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training , June 2026

  32. [43]

    MusicEval : A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation , March 2025

    Cheng Liu, Hui Wang, Jinghua Zhao, Shiwan Zhao, Hui Bu, Xin Xu, Jiaming Zhou, Haoqin Sun, and Yong Qin. MusicEval : A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation , March 2025

  33. [44]

    MuseCoco : Generating Symbolic Music from Text , May 2023

    Peiling Lu, Xin Xu, Chenfei Kang, Botao Yu, Chengyi Xing, Xu Tan, and Jiang Bian. MuseCoco : Generating Symbolic Music from Text , May 2023

  34. [45]

    CMI-Bench : A Comprehensive Benchmark for Evaluating Music Instruction Following , June 2025

    Yinghao Ma, Siyou Li, Juntao Yu, Emmanouil Benetos, and Akira Maezawa. CMI-Bench : A Comprehensive Benchmark for Evaluating Music Instruction Following , June 2025

  35. [46]

    GTZAN-Rhythm : Extending the GTZAN test-set with beat, downbeat and swing annotations

    Ugo Marchand, Quentin Fresnel, and Geoffroy Peeters. GTZAN-Rhythm : Extending the GTZAN test-set with beat, downbeat and swing annotations. In Late-Breaking/Demo Session of the 16th International Society for Music Information Retrieval Conference, Malaga, Spain, 2015

  36. [47]

    Mustango: Toward Controllable Text-to-Music Generation , June 2024

    Jan Melechovsky, Zixun Guo, Deepanway Ghosal, Navonil Majumder, Dorien Herremans, and Soujanya Poria. Mustango: Toward Controllable Text-to-Music Generation , June 2024

  37. [48]

    Genre-specific key profiles

    Cian O'Brien and Alexander Lerch. Genre-specific key profiles. In Proceedings of the International Computer Music Conference, Denton, TX, USA, 2015. URL https://musicinformatics.gatech.edu/wp-content_nondefault/uploads/2015/09/O

  38. [49]

    Humphrey, Justin Salamon, Oriol Nieto, Dawen Liang, and Daniel P

    Colin Raffel, Brian McFee, Eric J. Humphrey, Justin Salamon, Oriol Nieto, Dawen Liang, and Daniel P. W. Ellis. mir\_eval : A Transparent Implementation of Common MIR Metrics . In Proceedings of the 15th International Society for Music Information Retrieval Conference, pages 36...

  39. [50]

    Beyond accuracy: Behavioral testing of NLP models with CheckList

    Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. Beyond accuracy: Behavioral testing of NLP models with CheckList . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4902--4912, 2020

  40. [51]

    Bob L. Sturm. The GTZAN dataset: Its contents, its faults, their effects on evaluation, and its future use. arXiv preprint arXiv:1306.1461, 2013. doi:10.48550/arXiv.1306.1461. URL https://arxiv.org/abs/1306.1461

  41. [52]

    FIGARO : Generating Symbolic Music with Fine-Grained Artistic Control , February 2024

    Dimitri von R \"u tte, Luca Biggio, Yannic Kilcher, and Thomas Hofmann. FIGARO : Generating Symbolic Music with Fine-Grained Artistic Control , February 2024

  42. [53]

    MuseMorphose : Full-Song and Fine-Grained Piano Music Style Transfer with One Transformer VAE , December 2022

    Shih-Lun Wu and Yi-Hsuan Yang. MuseMorphose : Full-Song and Fine-Grained Piano Music Style Transfer with One Transformer VAE , December 2022

  43. [54]

    Shih-Lun Wu, Chris Donahue, Shinji Watanabe, and Nicholas J. Bryan. Music ControlNet : Multiple time-varying controls for music generation, 2023

  44. [55]

    Zihao Yang, Mosh Levy, Yoav Goldberg, and Byron C. Wallace. Compared to What ? Baselines and Metrics for Counterfactual Prompting , May 2026

  45. [56]

    SongEval : A Benchmark Dataset for Song Aesthetics Evaluation , May 2025

    Jixun Yao, Guobin Ma, Huixin Xue, Huakang Chen, Chunbo Hao, Yuepeng Jiang, Haohe Liu, Ruibin Yuan, Jin Xu, Wei Xue, Hao Liu, and Lei Xie. SongEval : A Benchmark Dataset for Song Aesthetics Evaluation , May 2025

  46. [57]

    YuE : Scaling Open Foundation Models for Long-Form Music Generation , September 2025

    Ruibin Yuan, Hanfeng Lin, Shuyue Guo, Ge Zhang, Jiahao Pan, Yongyi Zang, Haohe Liu, Yiming Liang, Wenye Ma, Xingjian Du, Xinrun Du, Zhen Ye, Tianyu Zheng, Zhengxuan Jiang, Yinghao Ma, Minghao Liu, Zeyue Tian, Ziya Zhou, Liumeng Xue, Xingwei Qu, Yizhi Li, Shangda Wu, Tianhao Sh...

  47. [58]

    ABC-eval : Benchmarking large language models on symbolic music understanding and instruction following, 2025

    Jiahao Zhao, Yunjia Li, Wei Li, and Kazuyoshi Yoshii. ABC-eval : Benchmarking large language models on symbolic music understanding and instruction following, 2025

  48. [59]

    TPSMG : Text-Controllable Polyphonic Symbolic Music Generation

    Jinlong Zhu, Keigo Sakurai, Ren Togo, Takahiro Ogawa, and Miki Haseyama. TPSMG : Text-Controllable Polyphonic Symbolic Music Generation . ITE Transactions on Media Technology and Applications, 14 0 (1): 0 110--118, 2026

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.