{"id":"2164f689-6208-4194-9c38-f80a8981468b","arxiv_id":"2504.21815","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Across five text-to-music models, aesthetic predictor scores, pairwise human preferences, and reference-based distribution metrics produce inconsistent rankings, so the choice of evaluation metric changes the winner.","lead":"This paper compares five text-to-music generation systems using an AI aesthetics judge and two distribution-matching metrics, and finds the judges rank the systems differently. It also reports weak agreement between a human preference dataset and the AI aesthetics judge, raising questions about automated music evaluation.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim of inconsistency between human preference and automatic metrics is not directly tested: AudioBox-Aesthetics is used as a stand-in for human judgment in Section 4 and in the Section 3 'Human vs. Human' comparison, with no fresh human ratings on the target outputs.","rationale":"The reader's weakest-assumption analysis and mine converge: the load-bearing step is treating the AudioBox-Aesthetics predictor as a valid proxy for human aesthetic judgment on MusicPref pairs and on the generated benchmark clips. The paper is otherwise a useful comparative exercise: it releases a benchmark, examines five recent TTM models, checks agreement between MusicPref and an independent predictor, and includes a caveat about conditioning inputs. However, the central claim of 'significant inconsistencies' between evaluation metrics and human preference is not directly supported because no fresh human ratings are collected for the exact outputs being ranked. Section 3 shows only that MusicPref labels and AudioBox predictions disagree; Section 4 shows only AudioBox predictions; Section 5 shows only distributional distances. Each comparison is between proxies or across disjoint datasets. In addition, the small effect sizes (62.3% accuracy, r=0.258) are reported without confidence intervals or significance tests, and the KAD narrative in Section 5 contains an internal ordering inconsistency. These issues do not invalidate the comparative data, but they mean the strong human-centered conclusion is conditional on validation of the predictor on the target distribution. Since the reader already assigned CONDITIONAL, my stress-test leaves that verdict unchanged.","tokens_in":8693,"tokens_out":5155,"duration_ms":53440,"concrete_test":"Run a small listening study on the Section 4 benchmark: sample 60 prompts × 5 models (300 clips), collect pairwise preference judgments (musicality and fidelity) from at least 3 raters per pair, and also collect ratings on the four AudioBox axes for a 100-clip subset. Compute per-model human preference rankings and correlate them with (a) Table 4 AudioBox means, (b) MAD scores, and (c) KAD scores against MusicCaps GT. If the AudioBox predictor's per-model ranking does not significantly correlate with the fresh human ranking, the Section 4 leaderboard is a proxy-only result and the 'human vs. metric' inconsistency claim must be reframed; if it does correlate, the proxy transfer concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central conclusion ('significant inconsistencies across the different metrics, highlighting the limitation of the current evaluation practice') requires that at least one of the compared metrics is a valid operationalization of human preference for the exact stimuli being ranked. The paper never establishes this. Section 3 compares human pairwise labels (MusicPref) against the Meta AudioBox-Aesthetics predictor, but the predictor is trained on Meta's annotation corpus (mostly non-music audio; no evidence it covers MusicPref clips or the 10-95s TTM outputs of Section 4). Calling this 'Human vs. Human' treats a neural network trained on human ratings as equivalent to fresh human judgment. The observed 62.3% accuracy / r=0.258 may therefore measure cross-corpus transfer error of the predictor, not human-human disagreement. Section 4 then uses the same predictor to rank five TTM models, so Table 4 is a proxy leaderboard unless the predictor's transfer to these clips is validated. Section 5 adds MAD/KAD but contains no human judgments at all, so the claimed mismatch between 'human preference' and 'distributional metrics' is inferred across datasets, never measured on the same items. This is amplified by a concrete confounding factor: model outputs differ in duration (10s to 95s) and conditioning inputs, while the MusicCaps reference clips are all 10s; both the aesthetics predictor and PANNs-based embeddings can be sensitive to duration. A final internal inconsistency appears in Section 5: the text says MusicGen-Large and JASCO achieve the lowest KAD distances '(7.65 and 5.51, respectively)', but 5.51 < 7.65 implies JASCO is closest by KAD, contradicting the 'again' framing that corroborates MusicGen's lowest MAD. Without a fresh human baseline on the benchmark clips, the headline inconsistency is not yet about human preferences.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates how different evaluation procedures rank text-to-music systems, using five recent models (JASCO, Stable-Audio-Open, MusicGen-Large, YuE, DiffRhythm) and LP-MusicCaps prompts. It compares pairwise human preference labels from MusicPref against the Meta AudioBox-Aesthetics predictor across four dimensions (CE, CU, PC, PQ), reports accuracy and Spearman correlations in Section 3, and then uses the same predictor to produce a leaderboard in Section 4, including a tag-cluster bias analysis. Section 5 computes MAD and KAD distributional scores against MusicCaps ground truth. The paper reports that the rankings differ sharply across these evaluation perspectives, concluding that current evaluation practice has significant inconsistencies and advocating for more human-centered evaluation.","tokens_in":8983,"tokens_out":6128,"duration_ms":62686,"significance":"If substantiated, the descriptive pattern is a useful cautionary result for the audio generation community: model rankings differ depending on whether one uses learned aesthetic scores, MAD, or KAD, and the paper provides a public benchmark of generated samples. The cross-metric comparison and the tag-cluster analysis are constructive and go beyond a single evaluation axis. However, the central inference about human preferences rests on treating a pretrained predictor as a human surrogate, and the headline claim of 'significant inconsistencies' is not backed by significance tests or confidence intervals. The paper is therefore a suggestive comparative study rather than a rigorous demonstration of misalignment with human preference.","major_comments":[{"comment":"The abstract claims 'significant inconsistencies' across metrics, but the paper reports no significance tests, confidence intervals, or effect sizes for the accuracies and Spearman correlations. For example, with 2,049 non-tie musicality pairs, the 62.3% accuracy for CU is roughly 12 percentage points above the 50% random baseline, so deviation from chance is likely real, but the differences among CE (0.604), CU (0.623), and PQ (0.596) need paired tests or confidence intervals before the paper can claim that the metrics are inconsistent. The negative correlations for PC (-0.002 for musicality, -0.087 for fidelity) also need an explicit test against zero and against the other metrics.","section":"Abstract and §3, Tables 2–3"},{"comment":"The paper labels Section 3 'Human vs. Human' and the abstract frames the findings as revealing inconsistencies with human preference, but no fresh human ratings are collected on the exact test stimuli. The Section 3 comparison is between MusicPref human pairwise labels and the pretrained Meta AudioBox-Aesthetics predictor, not between two human annotation sources. Section 4 then uses the same predictor as the judge for the leaderboard, so Table 4 is a proxy ranking unless the predictor is validated on those exact generated clips. The AudioBox-Aesthetics training corpus, described as one-third music, is not shown to cover MusicPref clips or the 10–95 second TTM outputs, so the observed poor agreement could be cross-corpus transfer error rather than evidence about human-human or human-machine disagreement.","section":"§3 and §4, Table 4"},{"comment":"The dataset description is internally inconsistent. The paper says it uses the full LP-MusicCaps prompt set of 5,521 prompts, but then states that DiffRhythm and YuE use only 50 generated lyrics and that JASCO uses 50 drum tracks and 100 chord progressions. It is not clear how many clips were generated per model, whether all models saw the same prompt set, or how the conditioning inputs were paired with prompts. Without this information, Table 4 and the released benchmark cannot be reproduced or meaningfully interpreted as a controlled comparison.","section":"§4.2"},{"comment":"The model outputs differ in duration (JASCO 10 seconds, MusicGen-Large 20 seconds, Stable-Audio-Open 47 seconds, YuE around 50 seconds, DiffRhythm 95 seconds, with MusicCaps ground truth at 10 seconds), and the models differ in conditioning inputs such as lyrics, chords, and drum tracks. Because both the aesthetics predictor and the PANNs-wavegram-logmel embeddings used for MAD/KAD can be sensitive to duration and content composition, the cross-model comparisons in Table 4 and Figures 2–3 may partly reflect input differences rather than intrinsic generation quality. The paper acknowledges the conditioning confound in one sentence, but offers no control, sensitivity analysis, or duration-matched comparison.","section":"§4.3 and §5"},{"comment":"The central claim of inconsistency between human preference and distributional metrics is never tested on the same items. Section 5 contains no human judgments at all; it compares MAD/KAD values between generated sets and MusicCaps ground truth, while the 'human preference' side comes from the AudioBox-Aesthetics predictor in Section 4. The mismatch between the two rankings is therefore inferred across datasets rather than measured on the same clips. Additionally, the KAD sentence is ambiguous: 'MusicGen-Large and JASCO achieve the lowest distances relative to MusicCaps GT (7.65 and 5.51)' does not state which model has which value; the surrounding text and Figure 3 suggest JASCO has 5.51 and MusicGen-Large has 7.65, and this should be stated explicitly.","section":"§5"}],"minor_comments":[{"comment":"The word 'significant' in the abstract should be replaced with a precise statistical statement or qualified with confidence intervals, since no significance tests are reported.","section":"Abstract"},{"comment":"There are typos in this section: 'synthezised' should be 'synthesized', 'aethetics' should be 'aesthetics', and 'teh' in §4.1 should be 'the'.","section":"§4.3"},{"comment":"The model name is written 'YUE' in the table but 'Yue' and 'YuE' elsewhere; please standardize the capitalization.","section":"Table 4"},{"comment":"The dataset name appears as 'MusicPrefs' in §2.2 but as 'MusicPref' in Section 3 and in reference [12]; the spelling should be made consistent.","section":"§2.2 and §3"},{"comment":"The row for 'Meta-AudioBox-Aesthetics' marks 'Human Involvement' with a check mark, but this is a pretrained automatic predictor, not a human annotation process; the table should distinguish direct human involvement from models trained on human annotations.","section":"Table 1"},{"comment":"The table caption '(no low)*' is not explained in the caption itself, and the text says the subset removes recordings 'under the low quality tag or captions' while the caption says it removes recordings with the 'low quality' tag; please clarify the exact filtering criterion.","section":"§4.3"},{"comment":"The claim of being 'the first systematic study of human preference alignment in music generation' is too strong given existing work such as MusicEval and the KAD/MAD papers; please soften the novelty claim.","section":"§1 and Conclusion"},{"comment":"The number of KMeans clusters (15) appears hand-chosen, and no robustness check is reported for the tag clustering; a short sensitivity analysis or a reference to the choice would help.","section":"§4.4"},{"comment":"The paper states that a benchmark dataset is released, but no URL, repository link, license, or download instructions are provided; a data-availability statement is needed.","section":"Throughout"},{"comment":"The sentence 'The work does not relate to Huy Phan's work at Meta' appears in the main text; this kind of statement belongs in an acknowledgments or conflict-of-interest section, not in the introduction.","section":"Introduction"}],"recommendation":"major_revision","confidential_remarks":"The paper has a useful comparative dataset and an interesting descriptive finding, but the main empirical claim needs to be either reframed as a comparison among automatic proxies or supported by fresh human ratings on the exact stimuli. The absence of statistical tests and the unresolved prompt/conditioning mismatch in §4.2 are load-bearing. I would support a revision that adds confidence intervals and significance tests, validates or clearly disclaims the AudioBox-Aesthetics surrogate, and clarifies the experimental setup. The 'first systematic study' claim should also be tempered."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: the paper's central claim—automatic metrics disagree with each other and with human preference—is plausible but not actually established, because the 'human' side is a pretrained model, not fresh human judgment.\n\nWhat is new: a comparative application of the AudioBox-Aesthetics predictor to MusicPref pairs and to outputs from five text-to-music models, plus MAD and KAD scores on the same generated sets, with a benchmark dataset released. That combination is genuinely not in the prior cited work. The descriptive numbers are interesting: CU predicts musicality at 62.3% (above chance but weak), and the aesthetics leaderboard clearly differs from the distributional ranking. The paper also flags the conditioning-input confound and shows that the predictor's Production Complexity scores vary by semantic tag cluster—both fair and useful observations.\n\nThe soft spots are real and load-bearing. Section 3's 'Human vs. Human' is actually human-annotated pairs versus a predictor trained on a different corpus, so the 62.3% accuracy could just be cross-corpus transfer error, not evidence about human–human disagreement. Section 4 ranks models using the same predictor, making Table 4 a proxy leaderboard. Section 5 contains no human judgments at all, so the claimed mismatch between human preference and distributional metrics is inferred across datasets rather than measured on the same items. Duration is also a confound: outputs range from 10 to 95 seconds while MusicCaps references are all 10 seconds, and both the aesthetics model and PANNs embeddings can be sensitive to length. There is also a small but clear internal error in Section 5: the text says MusicGen-Large and JASCO achieve the lowest KAD distances '(7.65 and 5.51, respectively)', but since lower KAD means closer, JASCO is actually the closest; the sentence framing it as corroborating MusicGen's low MAD is backwards.\n\nWho is this for? Researchers working on text-to-music evaluation who want a quick, honest map of how proxies can diverge. It reads like a workshop paper with a testable hypothesis. It deserves a serious referee because the experimental scaffold is real and the released dataset could help others, but the headline conclusion needs a fresh human listening baseline on the same clips plus significance testing before it can be trusted.","headline":"The headline inconsistency is plausible but under-supported because one side of the comparison is a pretrained aesthetics predictor, not fresh human ratings; still, the paper's empirical scaffold is worth engaging.","tokens_in":9599,"tokens_out":1510,"would_cite":false,"duration_ms":15367,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Automatic metrics disagree so strongly with human preference that text-to-music model rankings depend on the chosen evaluation.","keywords":["Text-to-Music Generation","Human Preference Alignment","Evaluation Metrics","Music Quality Assessment","Generative Audio Models","MAD","KAD","Aesthetic predictors"],"falsifier":"Collect fresh human preference ratings for a random sample of the released benchmark clips, then compare the resulting model ranking with the rankings from AudioBox-Aesthetics scores and from MAD/KAD; if the automatic metrics reproduce the human ranking or agree strongly with each other on the same clips, the paper's central claim of inconsistency would not hold.","tokens_in":8506,"feed_emoji":"🎵","tokens_out":7504,"duration_ms":70480,"temperature":0.7,"pith_summary":"This paper studies whether automatic evaluation metrics can stand in for human preference when judging text-to-music systems. Using five recent generation models, it compares a learned aesthetics predictor (scoring content enjoyment, usefulness, production complexity, and production quality) against human pairwise judgments, and against reference-based distribution distances (MAD and KAD). The results show poor agreement: the best aesthetics dimension predicts human choices only about 62 percent of the time, with Spearman correlations below 0.26, and the aesthetics leaderboard disagrees with the distribution-metric leaderboard. The paper concludes that current evaluation practice is limited and calls for human-centered evaluation, releasing a benchmark dataset of generated samples to support it.","feed_headline":"Text-to-music rankings flip depending on the metric","feed_subtitle":"Five generation systems get conflicting verdicts from aesthetic scores, MAD, and KAD.","key_machinery":"The central machinery is a set of three evaluation perspectives applied to the same five text-to-music systems: (1) pairwise human preference judgments from the MusicPref dataset, treated as ground truth; (2) the AudioBox-Aesthetics neural predictor, which outputs scalar scores for content enjoyment, content usefulness, production complexity, and production quality; and (3) reference-based distribution distances, MAD and KAD, computed on audio embeddings (PANNs features) between generated sets and the human-composed LP-MusicCaps reference. MAD quantifies divergence using a KL-style measure, KAD uses a kernelized maximum mean discrepancy style distance. Comparing these perspectives on the same outputs is what exposes the inconsistency.","core_discovery":"The paper's central discovery is that different evaluation perspectives give different answers about which text-to-music model is best, and none of the automatic proxies closely tracks human preference. On MusicPref pairwise comparisons, the AudioBox-Aesthetics model's score differences agree with human preferences at best around 62.3% for musicality (Content Usefulness) and 59.6% for fidelity (Production Quality), with Spearman correlations at best 0.258, far from a reliable predictor. The aesthetics-based leaderboard ranks JASCO highest on content usefulness and production quality, while reference-based MAD and KAD rank MusicGen-Large and JASCO as closest to human-composed MusicCaps recordings, with DiffRhythm, Stable-Audio-Open, and YuE farther away. The paper interprets these inconsistencies as evidence that no single learned or distributional metric currently captures human aesthetic judgment for music generation.","pith_inferences":["If the AudioBox-Aesthetics predictor does not transfer to the MusicPref pairs and the newly generated clips (the paper does not validate it against fresh human ratings of these exact outputs), the reported agreement numbers may reflect cross-corpus transfer rather than intrinsic human-predictor agreement; a listening test on the released benchmark would separate the two.","The genre-dependent patterning of production-complexity scores suggests the aesthetics model may encode training-data or annotator biases; a testable extension is to re-rank models within matched genre clusters.","The same three-perspective comparison could be applied to other generative domains, where learned preference models and distribution metrics are also used as substitutes for human judgment without cross-validation against fresh human ratings.","If the weak agreement holds generally, current preference-optimized text-to-music systems may be optimizing a reward-model proxy that real listeners would only partly endorse; measuring the reward model's agreement with fresh human preference on generated outputs would show how much optimization signal survives."],"forward_implications":["Model rankings from text-to-music evaluations are metric-dependent; changing from aesthetics scores to MAD or KAD can change which system looks best.","Learned aesthetics scores cannot currently replace human listening tests: at roughly 62% best-case accuracy and Spearman correlations below 0.26, they capture only a weak signal of human pairwise preference.","Reference-based distribution metrics and aesthetic predictors measure different properties and should be reported side by side rather than treated as interchangeable.","The released benchmark of generated clips and evaluation scores gives other researchers a fixed corpus for testing new metrics against the same text-to-music outputs.","Because aesthetics scores vary with musical content (for example, rhythmic and electronic tags score higher on production complexity), model comparisons should control for genre and prompt content."],"supporting_citations":[{"why":"Supplies the AudioBox-Aesthetics predictor whose four dimensions (content enjoyment, content usefulness, production complexity, production quality) drive both the agreement test and the model leaderboard.","marker":"[25]"},{"why":"Supplies the MusicPref pairwise human preference dataset and the MAD metric used as human ground truth and as one distributional measure.","marker":"[12]"},{"why":"Supplies the KAD metric, the second reference-based distribution distance used to rank generated sets against human-composed audio.","marker":"[13]"},{"why":"Supplies the LP-MusicCaps prompts and the human-composed MusicCaps recordings used as generation prompts and as the reference set.","marker":"[36]"},{"why":"Provides the PANNs audio embeddings on which MAD and KAD are computed.","marker":"[37]"},{"why":"Defines MusicGen-Large, one of the five benchmark models and the closest to the reference by MAD.","marker":"[33]"},{"why":"Defines JASCO, one of the five benchmark models and the top model by several aesthetics scores.","marker":"[31]"},{"why":"Defines Stable-Audio-Open, one of the five benchmark models with contrasting aesthetics and distribution results.","marker":"[32]"},{"why":"Defines YuE, one of the five benchmark models with higher distributional distance from the reference.","marker":"[34]"},{"why":"Defines DiffRhythm, one of the five benchmark models with higher distributional distance from the reference.","marker":"[35]"}],"fun_headline_variants":["Music AI rankings shift with evaluation metric","No automatic metric tracks human music preference","Text-to-music eval metrics give conflicting verdicts","Which music model wins often depends on the metric","Automatic music evaluation diverges from human taste"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that automatic metrics diverge from human preference depends on treating the AudioBox-Aesthetics predictor's scores as a valid measure of human aesthetic judgment for the MusicPref pairs and the newly generated benchmark clips; the paper does not test that predictor against fresh human ratings of those exact outputs.","fun_headline_variants_meta":{"raw":{"variants":["Music AI rankings shift with evaluation metric","No automatic metric tracks human music preference","Text-to-music eval metrics give conflicting verdicts","Which music model wins often depends on the metric","Automatic music evaluation diverges from human taste"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000576,"raw_usage":{"total_tokens":2675,"prompt_tokens":857,"completion_tokens":1818,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":473,"completion_tokens_details":{"reasoning_tokens":1750}},"tokens_in":473,"tokens_out":1818,"duration_ms":14323,"temperature":1.0,"reasoning_tokens":1750,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:53:00.752474+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect fresh human preference ratings for a random sample of the released benchmark clips, then compare the resulting model ranking with the rankings from AudioBox-Aesthetics scores and from MAD/KAD; if the automatic metrics reproduce the human ranking or agree strongly with each other on the same clips, the paper's central claim of inconsistency would not hold.","supporting_citations":[{"cited_title":"Leveraging pre-trained audioldm for sound generation: A benchmark study,","cited_arxiv_id":null,"evidence_quote":"Supplies the AudioBox-Aesthetics predictor whose four dimensions (content enjoyment, content usefulness, production complexity, production quality) drive both the agreement test and the model leaderboard."},{"cited_title":"DSPO: Direct score preference optimization for diffusion model align- ment,","cited_arxiv_id":null,"evidence_quote":"Supplies the MusicPref pairwise human preference dataset and the MAD metric used as human ground truth and as one distributional measure."},{"cited_title":"How does the teacher rate? Observations from the NeuroPiano dataset,","cited_arxiv_id":null,"evidence_quote":"Supplies the LP-MusicCaps prompts and the human-composed MusicCaps recordings used as generation prompts and as the reference set."},{"cited_title":"Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation,","cited_arxiv_id":null,"evidence_quote":"Provides the PANNs audio embeddings on which MAD and KAD are computed."},{"cited_title":"Piano Skills Assessment,","cited_arxiv_id":null,"evidence_quote":"Defines MusicGen-Large, one of the five benchmark models and the closest to the reference by MAD."},{"cited_title":"From Audio Encoders to Piano Judges: Benchmarking Performance Under- standing for Solo Piano,","cited_arxiv_id":null,"evidence_quote":"Defines Stable-Audio-Open, one of the five benchmark models with contrasting aesthetics and distribution results."},{"cited_title":"LLaQo: Towards a Query-Based Coach in Expressive Music Performance Assessment,","cited_arxiv_id":null,"evidence_quote":"Defines YuE, one of the five benchmark models with higher distributional distance from the reference."}],"review_version":1}