{"id":"066d2b9a-fa82-4db4-b290-564249b0331f","arxiv_id":"2506.19085","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A 15,600-comparison human study ranks 12 music generation models and finds that music-trained CLAP metrics correlate best with human preference.","lead":"Researchers asked more than 2,500 people to compare 15,600 pairs of AI-generated songs. They found that the commercial model Suno sounds best to listeners and that music-trained CLAP metrics match human preference better than other automatic scores, and they released all songs and ratings.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Metric-ranking claim may be driven by noise: Fig. 5 uses only 13 model-level points without CIs; leave-one-out should be tested.","rationale":"The reader's conditional verdict is appropriate, but the reader's weakest_assumption focuses on external validity (10-second instrumental clips, narrow demographics). The more immediate load-bearing vulnerability is internal statistical robustness: the metric ranking rests on correlations over 13 model-level points with no uncertainty quantification. Even if the preference proxy were accepted, the headline conclusion about CLAP-MA could be sampling noise. This is concrete and testable with the released data. It does not overturn the paper's value, but it strengthens the need for the conditional verdict and for explicit robustness analysis in the paper.","tokens_in":6817,"tokens_out":3598,"duration_ms":39058,"concrete_test":"Recompute the Fig. 5 correlations with leave-one-model-out and bootstrap: for each metric, compute PCC/SRC on all 13 model-level points and on every 12-point subset; also resample pairwise comparisons (or Bradley-Terry fits) to propagate human-evaluation uncertainty. Then declare CLAP-MA best only if it is top-ranked in a majority of resamples and its 95% CI for the correlation excludes the second-best metric.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central conclusion (Section VI) is that music-trained CLAP embeddings, FAD-CLAP-MA and LAION-MA, best approximate human preference. This rests on Fig. 5, which reports Pearson and Spearman correlations between objective metric scores and Bradley-Terry parameters over 13 data points (12 models plus the MTG-Jamendo reference). With n=13, rank correlations have very wide sampling distributions and are highly sensitive to one observation. The paper reports no confidence intervals, significance tests, permutation tests, or leave-one-out robustness checks. Consequently, even granting the 10-second instrumental-clip setup and participant pool, the ordering of metrics in Fig. 3/4 and the 'best metric' claim are not statistically established. A single model's score could move FAD-CLAP-MA and FAD-PANN, or LAION-MA and MS-CLAP, in the correlation ranking. This is a load-bearing gap because the paper's main contribution to future benchmarking is precisely the metric ranking.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a large-scale benchmark of 12 music generation models using 6,000 generated tracks and 15,600 pairwise human preference comparisons collected from over 2,500 participants via Prolific. Prompts are selected from MTG-Jamendo tag combinations filtered for diversity with a CLAP-based cosine-similarity threshold. The authors compute Elo ratings and Bradley-Terry parameters from the human comparisons, then correlate objective metrics (FAD variants and CLAP/LAION text-audio alignment scores) with these human-derived strengths across 13 model-level data points. They conclude that Suno v3.5 is the most preferred and best-aligned model and that music-trained CLAP embeddings (FAD-CLAP-MA and LAION-MA) best approximate human preferences. All generated audio, prompts, and human response data are released openly.","tokens_in":1300,"tokens_out":1455,"duration_ms":45457,"significance":"If the conclusions are statistically robust, this would be one of the first large-scale public resources for comparing music generation models and objective metrics against human preference, and the dataset release is a valuable contribution to the community. The use of Elo and Bradley-Terry models is standard, and the bootstrapping of Elo ratings is a strength, as is the broad coverage of commercial and open-source models. However, the central metric-ranking claim rests on a small set of model-level correlation coefficients without uncertainty quantification, and the prompt-selection procedure may advantage CLAP-family metrics. These issues currently prevent the paper's headline conclusions from being fully established.","major_comments":[{"comment":"The central claim that music-trained CLAP embeddings best approximate human preference rests on Pearson and Spearman correlations computed over only 13 model-level data points (12 models plus MTG-Jamendo). With n=13, rank correlations have very wide sampling distributions and are highly sensitive to a single observation. The paper reports no confidence intervals, significance tests, permutation tests, or leave-one-out robustness checks. Please add bootstrap confidence intervals that resample both human comparisons and models, and report leave-one-model-out correlations to show that the ordering in Fig. 5 is not driven by one model. Without this, the metric ranking and the Section VI conclusion are not statistically established.","section":"Section V, Fig. 5"},{"comment":"The prompt set is selected by requiring that no two tag combinations have CLAP embedding cosine similarity above a threshold of 0.1382. The metrics evaluated later include the same CLAP family (FAD-CLAP-MA, LAION-MA, and other LAION checkpoints). This creates a selection bias: the test distribution is constructed to be well-separated in the CLAP embedding space, potentially inflating the measured performance of CLAP-based metrics relative to metrics based on other embeddings. Please test robustness by selecting prompts with an alternative diversity criterion (e.g., random selection or diversity under VGGish/PANN embeddings) and check whether the metric ranking in Fig. 5 is preserved.","section":"Section III-A"},{"comment":"The human evaluation operationalizes 'music quality' and 'text-audio alignment' through 10-second instrumental excerpts judged by fluent English speakers aged 18 to 34 who use streaming services. The abstract and conclusions state the results in terms of 'human preference' without these qualifiers. Since full-song structure, vocals, and listener demographics may change both model ranking and metric correlations, the claims should be scoped to this experimental setup, or the paper should provide evidence that the 10-second instrumental proxy is representative of full-length listening. Please add a limitations paragraph and soften the corresponding claims.","section":"Sections III-B and IV-A"}],"minor_comments":[{"comment":"Reference [21] spells the platform name as 'Prolofic'; it should be 'Prolific'.","section":"References"},{"comment":"The statement that 'Suno v3.5 achieved a significantly higher Elo than all other models' needs a significance test or confidence intervals; the current text reports only the mean over bootstraps.","section":"Section IV-B"},{"comment":"The color scale and cell values are not explained in the captions; please add a legend and define whether green always indicates 'better according to the metric' or 'better according to human evaluation'.","section":"Figures 3 and 4"},{"comment":"The numbers 15k and 6k are used loosely; the abstract later states 15,600 comparisons and 6,000 songs. Please be consistent and use exact numbers throughout.","section":"Abstract and Section I"},{"comment":"The phrase 'providing an unbiased ranking of the models' is too strong given the prompt-selection procedure and the participant filter; please replace 'unbiased' with a more cautious formulation.","section":"Section VI"}],"recommendation":"major_revision","confidential_remarks":"The dataset release and the scale of the human study are genuinely useful for the community. My main concern is statistical: the metric-ranking conclusion is based on n=13 correlations without uncertainty bounds, and the CLAP-based prompt selection may bias the comparison. These are fixable with additional analysis and robustness checks, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a genuinely useful benchmark paper: they generate 6k tracks from 12 models, collect 15,600 human pairwise judgments, and release both the audio and the responses. The model ranking itself—Suno v3.5 on top, MTG-Jamendo beaten by the best commercial models—is plausible and the survey machinery (Elo, Bradley-Terry, attention checks) is standard and competently executed. The open data alone makes it worth having.\n\nThe soft spot is the paper's central claim that music-trained CLAP embeddings best approximate human preference. That conclusion rests on Pearson/Spearman correlations between metric scores and the 13 Bradley-Terry parameters (12 models plus the reference set). With n=13, rank correlations have enormous sampling variability; one model's score can flip the ordering of FAD-CLAP-MA vs FAD-PANN or LAION-MA vs MS-CLAP. The paper reports no confidence intervals, no significance tests, no leave-one-out checks. The stress-test note is right that this is load-bearing: without error bars, the metric ranking is a set of observations, not evidence. There is also a mild circularity worry: the 500 prompts were selected to be CLAP-diverse, which could give CLAP-based alignment metrics a small head start. That is minor, not fatal, but it deserves a sentence in the limitations.\n\nTwo smaller caveats. The \"unbiased ranking\" language in the conclusion is too strong; the study used 10-second instrumental clips and a narrow participant pool (fluent English speakers 18–34 who use streaming). A broader population or full-length clips could shift the ranking. Also, several models contribute multiple checkpoints (MusicGen small/medium/large, Stable Audio v1/v2), so the 13 points aren't fully independent; that can inflate correlations.\n\nWho is this for? Anyone developing or evaluating text-to-music systems. It is a solid empirical contribution with real data, and it deserves a serious referee. My recommendation: engage with it, but push for bootstrap CIs or permutation tests on the correlations, a leave-one-out sensitivity check, and a more careful generalization caveat. With those fixes, the metric ranking would be trustworthy. I'd cite it for the dataset; I wouldn't yet cite the metric ranking as established.","headline":"A useful open benchmark for music generation evaluation, but the headline metric ranking rests on 13 points with no error bars.","tokens_in":7538,"tokens_out":3105,"would_cite":true,"duration_ms":29659,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 15,600-comparison human study ranks music generation models: Suno v3.5 wins, and music-trained CLAP embeddings best match human judgment.","keywords":["music generation","human preference","text-audio alignment","Frechet Audio Distance","CLAP embeddings","Elo rating","benchmark dataset"],"falsifier":"Re-run the same pairwise comparison survey with full-length vocal tracks and a listener pool that includes older adults and non-streamers, then check whether Suno v3.5 still leads and whether FAD-CLAP-MA and the music-trained CLAP alignment scores still correlate with human ratings; if the ranking shifts or the correlations drop materially, the 10-second instrumental-clip proxy is the load-bearing choice.","tokens_in":6610,"feed_emoji":"🎵","tokens_out":8018,"duration_ms":74818,"temperature":0.7,"pith_summary":"This paper asks which music-generation models people actually prefer and which automated metrics come closest to matching those preferences. To answer, the authors generated 6,000 songs with 12 current models, ran 15,600 pairwise listening comparisons with more than 2,500 participants, and ranked models by bootstrapped Elo ratings and Bradley-Terry strength. They find that the commercial model Suno v3.5 is preferred over all others and also matches text prompts best, beating even the human-made reference dataset. On the metric side, CLAP embedding models trained on music data—both inside Frechet Audio Distance for quality and as cosine-similarity scorers for text alignment—correlate most strongly with the human judgments. If correct, music-trained CLAP embeddings are the best cheap proxy for human taste in generated music, and the released dataset gives the field a fixed benchmark for testing new metrics.","feed_headline":"Suno v3.5 tops human music-preference test of 12 AI models","feed_subtitle":"A 15,600-comparison study finds music-trained CLAP embeddings best match listener taste and prompt fit.","key_machinery":"The engine is pairwise binary preference testing on 10-second instrumental clips, with the same tag triple used to generate both clips so that preference and text-audio alignment are judged on matched content. Human choices are converted into model strengths via bootstrapped Elo ratings and Bradley-Terry parameters. Those human strengths are then correlated, by Pearson and Spearman coefficients, against objective scores: Frechet Audio Distance computed on VGGish, PANN, EnCodec, and several CLAP embedding spaces for music quality, and CLAP cosine similarity between audio and prompt text for text-audio alignment. The decisive comparison is which embedding space makes FAD or cosine similarity track the human rank ordering.","core_discovery":"The central discovery is an empirical ranking grounded in human pairwise choices rather than in any single automated score. Among 12 generation models, Suno v3.5 earns the highest bootstrapped Elo in both music preference and text-audio alignment, with Suno v3 and Udio also surpassing the reference corpus. On the metric side, FAD computed with the music-audioset CLAP checkpoint (FAD-CLAP-MA) has the best Pearson and Spearman correlation with human music-preference Bradley-Terry parameters, and the music-trained CLAP checkpoints give the highest correlation with human text-audio alignment judgments. The authors read this as evidence that CLAP models trained on music data approximate human preferences most accurately, both as embedding models for FAD and for measuring text-audio alignment.","pith_inferences":["The paper only tests 10-second instrumental clips, so its ranking is an inference about short-form instrumental snippets; judging 30-to-60-second clips with vocals could change the model order, especially on structural coherence and vocal quality.","The listener pool is fluent English speakers aged 18 to 34 who use streaming services; broader age and cultural groups could rank models differently, particularly on genre fit.","If music-trained CLAP embeddings really track human preference this well, they could serve as a training-time reward signal or a filter for selecting generated samples, not just a post-hoc evaluation metric.","The released dataset lets future metric developers test against a fixed human ground truth; the next step is checking whether any new embedding improves on FAD-CLAP-MA at the clip level, not just at the model-aggregate level."],"forward_implications":["Suno v3.5 becomes the baseline to beat for human-preferred text-to-music generation among currently available models.","FAD-CLAP-MA can substitute for expensive listening tests when comparing music generation models on quality.","Music-trained CLAP cosine similarity can automate text-audio alignment checks during data filtering or model development.","New metrics can be validated against the released human preference data without running a new 2,500-participant study.","Within a model family, the newer and larger checkpoint generally outranked the older or smaller one, pointing to scale and iteration as reliable improvement levers."],"supporting_citations":[{"why":"Defines the Frechet Audio Distance that the paper benchmarks across embedding spaces for music quality.","marker":"[1]"},{"why":"Provides VGGish embeddings, one of the FAD variants compared against human preference.","marker":"[2]"},{"why":"Supplies the music-trained CLAP checkpoints that best approximate human preference and text-audio alignment.","marker":"[7]"},{"why":"Showed FAD can be adapted for music evaluation; the paper extends its per-song correlation finding to a 12-model benchmark.","marker":"[9]"},{"why":"Supplies the reference tracks and tag combinations used to create prompts and baseline audio.","marker":"[11]"},{"why":"MusicGen checkpoints are among the 12 models ranked, giving the open-source comparison point.","marker":"[14]"},{"why":"Suno v3 and v3.5 are the commercial models that win both human preference and text-audio alignment in this study.","marker":"[18]"},{"why":"Udio is the other commercial model ranked in the top group by human preference.","marker":"[19]"},{"why":"The Bradley-Terry model converts pairwise human choices into the model strengths correlated with objective metrics.","marker":"[25]"}],"fun_headline_variants":["Suno v3.5 wins 15k human-preference music test","Human ears rank Suno v3.5 top of 12 AI models","Music CLAP metrics match human taste in 15k study","Suno v3.5 beats 12 rivals in listener preference study"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The results assume that binary preference between 10-second instrumental clips, judged by fluent English speakers aged 18 to 34 who use streaming services, captures what music quality and text-audio alignment mean; if full-length structure, vocals, or other listener groups matter, both the model ranking and the metric correlations could change.","fun_headline_variants_meta":{"raw":{"variants":["Suno v3.5 wins 15k human-preference music test","Human ears rank Suno v3.5 top of 12 AI models","Music CLAP metrics match human taste in 15k study","Suno v3.5 beats 12 rivals in listener preference study"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000162,"raw_usage":{"total_tokens":1186,"prompt_tokens":835,"completion_tokens":351,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":451,"completion_tokens_details":{"reasoning_tokens":272}},"tokens_in":451,"tokens_out":351,"duration_ms":3622,"temperature":1.0,"reasoning_tokens":272,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:37:21.850994+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same pairwise comparison survey with full-length vocal tracks and a listener pool that includes older adults and non-streamers, then check whether Suno v3.5 still leads and whether FAD-CLAP-MA and the music-trained CLAP alignment scores still correlate with human ratings; if the ranking shifts or the correlations drop materially, the 10-second instrumental-clip proxy is the load-bearing choice.","supporting_citations":[{"cited_title":"The mtg-jamendo dataset for automatic music tagging,","cited_arxiv_id":null,"evidence_quote":"Supplies the reference tracks and tag combinations used to create prompts and baseline audio."},{"cited_title":"Simple and controllable music generation,","cited_arxiv_id":null,"evidence_quote":"MusicGen checkpoints are among the 12 models ranked, giving the open-source comparison point."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Suno v3 and v3.5 are the commercial models that win both human preference and text-audio alignment in this study."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Udio is the other commercial model ranked in the top group by human preference."}],"review_version":2}