{"id":"6fd3363e-bdf3-4e75-ba64-78f11848e09c","arxiv_id":"2608.11899","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Using matched neutral and target-swap prompts, ACE-Step 1.5 and Stable Audio 3 show real key control and partial beat control, while LeVo2 does not, and much four-beat agreement is just the models' default output.","lead":"This paper introduces a counterfactual evaluation that compares music generated with a neutral prompt against music generated with explicit key or beat instructions, to separate true instruction following from default output behavior. It finds that two text-to-music models control musical key well, while common four-beat music makes the four-beat instruction look more effective than it really is.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim is well supported, but the auto-scorer validity on each system's generated audio remains the soft spot; a model-specific human audit would settle it.","rationale":"The paper is unusually careful: matched neutral/A/B families, frozen adapters, shared seeds, off-attribute placebos, external reference validation, blind expert audits, multi-seed sentinels, and sensitivity analyses all support the central dichotomy that agreement measures occurrence while matched contrasts test instruction-attributable response. The most plausible threat to the central claim is measurement: claims about LeVo2's absence of key control and Stable Audio 3's negative four-beat effect are contrasts of automatic recognizer outputs. Real-music validation does not guarantee equal sensitivity across generated audio distributions, and the human audit is too small and too task-level to settle per-model and per-target magnitudes. However, the evidence already in the paper—family-level correlations of 0.84 and 0.80, preserved contrast signs, and confusion-matrix transport showing target signs survive—makes it unlikely that this concern reverses the central conclusion. The proposed larger blinded audit directly tests the differential-sensitivity scenario. I therefore do not alter the reader's conditional verdict; the strongest condition for accepting the paper's empirical claims is a model-specific human relabeling check, in addition to the already-needed public artifact access.","tokens_in":16619,"tokens_out":10787,"duration_ms":117942,"concrete_test":"Generate 30 additional LeVo2 key families and 30 Stable Audio 3 beat families (or reuse the frozen manifest), have at least three blinded expert raters label every clip, then compute model-specific human key Delta for LeVo2 and human 3-beat and 4-beat Deltas for Stable Audio 3 with the same family bootstrap. If human LeVo2 key Delta is indistinguishable from zero and human Stable Audio 3 4-beat Delta is negative outside its interval, the automatic-scorer concern is settled and the paper's claims stand. If human LeVo2 key Delta is positive (e.g., above 0.15) or human Stable Audio 3 4-beat Delta is near zero or positive, the central claim would need to be weakened from 'does not control key' and 'four-beat reversal' to 'not detectable with S-KEY / Beat This.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is that S-KEY and Beat This are valid outcome measures on each system's generated audio, not just on real-music corpora. The paper validates on GTZAN, GiantSteps, RWC, Ballroom, and a 28-family/one-rater audit, reporting automatic key Delta=0.429 vs human 0.286 and beat Delta=0.036 under both. But the audit is task-level: it does not report model-specific target-specific Deltas, and it has too few LeVo2 key families to rule out a sensitivity miss. If S-KEY is differentially less sensitive on LeVo2's output distribution, its near-zero key Delta could be attenuation rather than absent control. Similarly, if Beat This's known 4-to-2 half-bar errors occur more often in Stable Audio 3's 4/4 rendering than in neutral clips, the headline -0.406 four-beat reversal would be inflated. The confusion-matrix transport in Appendix A.4 addresses average bias, not condition-specific error differences. Because the central empirical separation is exactly about differences between systems and targets, this is the load-bearing point.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a counterfactual evaluation framework for text-to-music controllability, organized around matched neutral–A–B contrast families. For each family, a neutral prompt omits the scored attribute, while two otherwise matched treatments specify different target values; all three outputs share a family seed and are rendered through frozen native-interface adapters. The framework separates three operational quantities: occurrence (treatment agreement), enhancement (change relative to the matched neutral output, Δ), and differentiation (margin between the two target arms). The method is applied to global key and three- versus four-beat grouping in ACE-Step 1.5, Stable Audio 3 Medium, and LeVo2, with 256 families per system. The main empirical findings are that ACE-Step 1.5 and Stable Audio 3 exhibit substantial key control while LeVo2 does not; for beat grouping, both responsive systems redirect toward the rare three-beat target, but high four-beat agreement largely reflects the neutral output prior rather than instruction-attributable control, with Stable Audio 3 showing a large negative four-beat Δ. The paper includes extensive validation: external recognizer benchmarks, blind expert audits, off-attribute placebos, three-seed sentinels, and a reproduction package. The central methodological claim is that agreement measures occurrence, whereas matched contrasts test instruction-attributable response.","tokens_in":16834,"tokens_out":6228,"duration_ms":67458,"significance":"The paper's main contribution is methodological: it makes explicit the distinction between target occurrence and instruction-attributable control, and it provides a concrete experimental template for measuring this distinction. The empirical results, if valid, change how prompted-agreement numbers should be interpreted for common target classes, and they demonstrate that a model can appear to follow a four-beat instruction while actually inheriting the four-beat grouping from its output prior. The strengths of the paper are notable: the decision rule is pre-specified, inference is performed on complete families with bootstrap methods, adapters are frozen and documented, outputs are validated against external recognizers and human audits, and the entire pipeline is packaged for reproduction. The paper also provides several falsifiable predictions, such as the sign of the four-beat reversal and the absence of LeVo2 key control, which are directly testable by other groups. If the measurement-validity caveats are resolved, the paper would be a valuable addition to the music-generation evaluation literature.","major_comments":[{"comment":"The human validation is task-level, not model-specific. The central claim that LeVo2 shows no key control (Δ = 0.026 [0.003, 0.049]) depends on the assumption that S-KEY is as sensitive on LeVo2's generated audio distribution as on ACE-Step and Stable Audio 3 outputs. If S-KEY is less discriminative on LeVo2's typical pitch-class distributions or timbres, the near-zero Δ could be attenuation rather than absent control. The 28-family audit with a single bridge rater is explicitly described in Section 7 as not supporting per-model human rankings. Please add a model-specific human audit of LeVo2 key families (even a modest sample) or a sensitivity analysis that recalibrates S-KEY's decision threshold separately for each model and shows that LeVo2's near-zero Δ and the separation between systems persist.","section":"§4.2, §7"},{"comment":"The confusion-matrix transport analysis addresses average measurement bias, not condition-specific error-rate differences. The headline Stable Audio 3 four-beat reversal (Δ = −0.406 [−0.516, −0.297]) could be inflated if Beat This's known tendency to make 4-to-2 half-bar errors occurs more often under the explicit '4/4 time signature' treatment than under the neutral condition. The current transport scenarios apply the same real-music confusion matrices without modeling a condition-specific increase in 4-to-2 errors. A targeted analysis that measures Beat This's error pattern separately for neutral and 4/4-generated clips (e.g., via human downbeat annotation on a sample of Stable Audio 3 outputs, or by replicating the key result with a second beat tracker) would directly address this concern and would strengthen the attribution claim.","section":"§4.1, Appendix A.4"}],"minor_comments":[{"comment":"The notation for Δₜ is introduced after the family-level definitions, but the family-level formula for Δ_f is written before the target-specific version. Consider adding a short transition sentence to clarify that Δ_t is the target-specific average of the corresponding family-level quantity.","section":"§3.1"},{"comment":"The open circles for matched neutral rates are visually light in grayscale; increasing marker size or using different point shapes for neutral and treated conditions would improve readability.","section":"Figure 2"},{"comment":"The Margin rows for beat are repeated for the 3-beat and 4-beat rows. This is correct but may confuse readers into thinking there are two independent estimates; a note or a merged cell would be clearer.","section":"Table 4"},{"comment":"The sentence 'ACE-Step changes little under a four-beat instruction' is a fair summary, but the corresponding Δ interval is [0.000, 0.109]; stating 'the interval includes zero' would be more precise.","section":"§5.2"},{"comment":"The reproducibility boundary for ACE-Step (no contemporaneous repository/checkpoint pin) is mentioned only in an appendix. Since reproducibility is a central claim of the paper, consider noting this caveat briefly in the main text's measurement-credibility section.","section":"Appendix A.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is methodologically careful and the central framework is publishable. The main risk is that the empirical separation between LeVo2 and the other systems rests on the assumption that S-KEY is equally valid on LeVo2's output distribution; the current human audit does not verify model-specific sensitivity. The four-beat reversal for Stable Audio 3 is also sensitive to condition-specific beat-tracker errors. Both issues are addressable with additional targeted analyses rather than fundamental redesigns. I would be willing to look at a revision that adds a per-model human audit or a convincing sensitivity analysis, along with a condition-specific beat-error analysis. The reproduction package and pre-specified decision rule are exemplary and should be highlighted in any final version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is the first text-to-music controllability evaluation I've seen that actually separates \"the target appears\" from \"the instruction made it appear.\" The matched neutral–A–B family design is simple and correct, and the paper executes it carefully. The Stable Audio 3 result—0.97 four-beat in neutral, 0.56 under explicit four-beat instruction—is a concrete demonstration that prompted agreement can masquerade as control.\n\nThe three-property decomposition (occurrence, enhancement, differentiation) is the real contribution. It is not mathematically deep, but it is the right framing for a benchmark. The paper also does the boring work properly: frozen adapters, pre-specified decision rule, bootstrap inference on complete families, external recognizer validation, off-attribute placebos, blind expert audits, and three-seed sentinels. The limitations section is unusually honest.\n\nSoft spots are real but not fatal. The auto-scorers are validated on real music, not on each model's generated distribution. The 28-family human audit has one bridge rater and 4–5 families per system, so it supports task-level calibration but not model-specific rankings. The key delta shrinks from 0.429 automatic to 0.286 human, so magnitudes are optimistic. For LeVo2's flat key result, a differential sensitivity miss in S-KEY on its output distribution could in principle attenuate a real effect—but the placebo result (LeVo2 changes key more under an irrelevant beat instruction than under a real key instruction) makes the no-control conclusion fairly robust. For beat, the known 4-to-2 half-bar errors could inflate the reversal, but the confusion-matrix transport retains the sign pattern across scenarios.\n\nThe bigger practical issue: the reproduction package has no public URL or commit hash, and the ACE-Step checkpoint is not pinned. That should be fixed before the paper is truly reproducible.\n\nWho this is for: anyone building or benchmarking controllable music generation. It deserves a serious referee, and with the artifacts public and ideally a per-model human audit added, I would be comfortable seeing it published. My own verdict is conditional, not accept.","headline":"A careful, well-executed demonstration that prompted agreement can masquerade as control in text-to-music; the soft spot is model-specific auto-scorer validity, but the paper's own checks keep the conclusion credible.","tokens_in":17332,"tokens_out":2168,"would_cite":true,"duration_ms":20738,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Prompted agreement measures occurrence, not instruction following; matched neutral and target-swap contrasts show which music models actually respond to key and beat requests.","keywords":["text-to-music generation","controllability evaluation","counterfactual prompting","key recognition","beat grouping","prompt agreement","output priors","instruction following"],"falsifier":"An independent blind human relabeling of the released 256-family benchmark per system, or a rerun with a different validated key and beat recognizer, would settle the claim: if LeVo2 shows a key Δ interval clearly above zero, or Stable Audio 3 loses its negative four-beat Δ, the paper's central contrast would not survive. The paper's own audit already shows that automatic key Δ (0.429) exceeds human key Δ (0.286) on a 28-family subset, so a full human relabeling is the direct check.","tokens_in":16433,"feed_emoji":"🎵","tokens_out":6118,"duration_ms":56970,"temperature":0.7,"pith_summary":"Prompted agreement—checking whether a requested attribute appears in generated audio—is the standard evidence that a text-to-music system follows instructions. This paper argues that agreement is not enough: a target may appear simply because it is already common in the model's output. It introduces matched neutral–A–B contrast families, where one prompt omits the scored attribute and two otherwise identical prompts request different values, and measures whether the instruction raises occurrence above the matched neutral output and whether swapping the target redirects the output. Applied to key and beat grouping in three open systems, this changes the conclusions: two systems genuinely control key, none of the three shows positive four-beat enhancement, and the high four-beat agreement of Stable Audio 3 is largely inherited from its default output. If the argument is right, controllability benchmarks should report occurrence, enhancement, and differentiation separately instead of a single agreement score.","feed_headline":"A '4/4' prompt cut four-beat output from 97% to 56%","feed_subtitle":"Matched counterfactual tests show key control is real for two systems, but common four-beat meter is inherited from defaults.","key_machinery":"The central object is the matched neutral–A–B contrast family: one neutral input that omits the scored attribute and two otherwise matched inputs that request different target values, all rendered through the system's frozen native interface with a shared seed. The argument is carried by three statistics derived from it—Acc (did the requested target occur under treatment), Δ (did treatment raise occurrence above the matched neutral output), and Margin (did swapping the requested value redirect the output toward the alternative target). A confirmatory label of effective control requires both Δ and Margin intervals to lie above zero; the paper also uses off-attribute placebos and multi-seed sentinels to rule out generic added-instruction effects and single-generation noise.","core_discovery":"The central empirical claim is that ACE-Step 1.5 and Stable Audio 3 Medium show large key control that is not explained by their output priors, while LeVo2 shows little attributable key response under its evaluated interface; and that for beat grouping the models redirect toward the rare three-beat target, while the apparently high four-beat agreement is a prior artifact. The central methodological claim is that agreement measures occurrence, whereas matched neutral and target-swap contrasts test instruction-attributable response. The paper reports ACE-Step key Δ = 0.612 and Stable Audio 3 key Δ = 0.646 against LeVo2's 0.026, and for beat grouping Stable Audio 3 three-beat Δ = +0.469 against four-beat Δ = −0.406, with the neutral four-beat rate at 0.969 and the explicit four-beat treatment at 0.563. The authors conclude that high prompted agreement can coexist with no improvement—or a large reversal—relative to the model's own neutral output.","pith_inferences":["A testable extension is to cross carrier templates with target values, since the paper's neutral contrast changes both target presence and its rendered carrier; such a factorial design could estimate carrier–value interactions rather than only end-to-end response.","The same design could be applied to continuous controls such as BPM or loudness, but the neutral-omission construction may need a neighbouring-value reference instead of an omitted field.","The human-audit gap in key magnitude (automatic Δ 0.429 vs human Δ 0.286) suggests that point estimates of control strength are optimistic; cross-model rankings and sign patterns are on firmer ground.","If the benchmark convention caught on, model cards might start reporting three numbers per attribute—attainment, enhancement, differentiation—which would change how buyers compare music generators."],"forward_implications":["Prompted agreement alone should not be used as evidence of controllability for any target whose output prior is high; the matched neutral rate must be reported.","Controllability should be described per target, not as an aggregate score, because four-beat and three-beat responses can cancel in the average.","If the result transfers, generative-model benchmarks outside music—faces, daylight scenes, common code patterns—should adopt neutral-relative and target-swap contrasts.","A negative four-beat Δ for Stable Audio 3 shows that a treatment can move output away from the common target; benchmarks need to allow for this rather than treating any agreement as success."],"supporting_citations":[{"why":"Supplies S-KEY, the frozen key recognizer used as the primary key outcome measure.","marker":"Kong et al. [2025]"},{"why":"Supplies Beat This, the beat/downbeat detector from which beat grouping is derived.","marker":"Foscarin et al. [2024]"},{"why":"Supplies the MIREX rule for graded related-key credit, used as the secondary key score.","marker":"Raffel et al. [2014]"},{"why":"Defines ACE-Step 1.5, one of the three evaluated text-to-music systems.","marker":"Gong et al. [2026]"},{"why":"Defines Stable Audio 3 Medium, the second evaluated system whose four-beat prior drives the main beat finding.","marker":"Evans et al. [2026]"},{"why":"Defines LeVo2, the third evaluated system that shows little attributable key or beat response.","marker":"Lei et al. [2026]"},{"why":"Supplies the minimal-pairs behavioral-testing precedent for associating output changes with controlled input edits.","marker":"Ribeiro et al. [2020]"},{"why":"Supplies the counterfactual-prompting baseline insight that textual edits should be compared with benign perturbations.","marker":"Yang et al. [2026]"}],"fun_headline_variants":["4/4 prompt reduces four-beat output from 97% to 56%","Counterfactual tests reveal key control real, beat meter is prior artifact","Two systems show real key control; beat grouping inherits defaults"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusions stand or fall on the validity of the frozen automatic scorers—S-KEY for key and Beat This for beat grouping—as outcome measures; if those recognizers mislabel generated audio in a target-dependent way, the reported Δ and Margin values would not reflect what the models actually produced.","fun_headline_variants_meta":{"raw":{"variants":["4/4 prompt reduces four-beat output from 97% to 56%","Counterfactual tests reveal key control real, beat meter is prior artifact","Two systems show real key control; beat grouping inherits defaults"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000648,"raw_usage":{"total_tokens":3001,"prompt_tokens":997,"completion_tokens":2004,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":613,"completion_tokens_details":{"reasoning_tokens":1941}},"tokens_in":613,"tokens_out":2004,"duration_ms":14845,"temperature":1.0,"reasoning_tokens":1941,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:23:10.690282+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An independent blind human relabeling of the released 256-family benchmark per system, or a rerun with a different validated key and beat recognizer, would settle the claim: if LeVo2 shows a key Δ interval clearly above zero, or Stable Audio 3 loses its negative four-beat Δ, the paper's central contrast would not survive. The paper's own audit already shows that automatic key Δ (0.429) exceeds human key Δ (0.286) on a 28-family subset, so a full human relabeling is the direct check.","supporting_citations":[{"cited_title":"S-KEY: Self-supervised Learning of Major and Minor Keys from Audio","cited_arxiv_id":"2501.12907","evidence_quote":"Supplies S-KEY, the frozen key recognizer used as the primary key outcome measure."},{"cited_title":"Beyond accuracy: Behavioral testing of NLP models with CheckList","cited_arxiv_id":null,"evidence_quote":"Supplies the minimal-pairs behavioral-testing precedent for associating output changes with controlled input edits."}],"review_version":1}