{"id":"79b8d1fc-ad9f-4ec2-b12c-8582e342874d","arxiv_id":"2608.06634","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Prompts for AI music are dominated by genre and story terms, but genre words survive into perception while story-heavy prompts produce the largest semantic mismatch.","lead":"This paper asked whether the words people type into text-to-music systems match the words they use to describe the music those systems produce. It found that prompts lean heavily on genre and story language, while listeners describe instruments, emotions, and musical structure, and that story-based prompts rarely reach the listener.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's category-level results rest on LLM-based labeling validated only on 25 English Song Describer prompts; if that classifier is biased on the actual Udio prompts, on descriptions, or on Korean text, the central asymmetries could be labeler artifacts rather than linguistic facts.","rationale":"The reader's weakest_assumption is exactly the point I would stress: the central claim is a statement about category-level linguistic asymmetries, and every category label in the study comes from a single LLM whose validation is thin, English-only, and unquantified per category. This is not a manufactured concern; it is the measurement foundation for the GLMM/LMM results, the word-level associations, Table 3, and the cross-cultural comparisons. The paper itself flags the issue in Section 5.3, so the authors are aware of it, but the limitation is not resolved by reporting it. The most direct remedy is an independent human-coding check on the actual texts, especially Korean descriptions, followed by refitting the models on human labels. The reader's CONDITIONAL verdict already captures this uncertainty, so I do not recommend moving the verdict; rather, the condition should be made explicit: the category-level claims should be conditional on validation of the LLM labeling on the study's own texts. I also note secondary concerns — the recruitment-channel confound in the cross-cultural comparison and the univariate nature of the 'strongest predictor' claim — but neither is as foundational as the labeler validity issue, because the former is acknowledged as exploratory and the latter is an interpretive overreach that does not by itself overturn the main asymmetry.","tokens_in":9910,"tokens_out":3502,"duration_ms":33309,"concrete_test":"Recruit two annotators (one bilingual Korean-English) to independently apply the taxonomy from §3.1 to a stratified sample: 100 of the 200 Udio prompts, 100 English descriptions, and 100 Korean descriptions. Compute per-category Cohen's kappa between GPT-5.4 codes and the human majority label for each of the three text types. Then refit the main presence and density models from §4.1 and the language-group models from §4.4 using only human-coded labels on the sample. If agreement falls below roughly kappa = 0.6 for Story/Narrative or Mood/Emotion in any text type, or if the refitted interaction coefficients lose significance or reverse sign, the reported asymmetries cannot be attributed to the language itself rather than to labeler bias.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is that GPT-5.4's automatic taxonomic coding is accurate enough for every text it labels. The validation in §4.1 is described only as comparing LLM configurations on 'a held-out, manually-annotated subset 25 of prompts from the Song Describer dataset'; no per-category agreement, no confusion matrix, and no evaluation on the actual Udio prompts or on either population's descriptions are reported. Korean descriptions are never validated at all. Yet every category-level result — the presence and density GLMM/LMM fits (§4.1), the survival-rate and chi-square associations (§4.2), the high/low alignment comparison in Table 3 (§4.3), and the cross-cultural redistribution (§4.4) — is computed from this labeling. The taxonomy itself was derived from 100 prompts by two annotators, but the coding of 200 prompts and 2,624 descriptions is not independently checked against human annotators. The paper's own Section 5.3 admits that classification noise is 'difficult to fully quantify.' If GPT-5.4 systematically over-assigns Genre/Story to prompts and Mood/Instrumentation/Music Theory to descriptions — a plausible pattern if prompts are shorter or more noun-phrase-like, or if Korean text is less reliably parsed — then the central prompt-description asymmetry and the cross-cultural profile differences would be partly or entirely artifacts of the labeler rather than of the language. Because the central claim is defined at the category level, this unvalidated measurement step is load-bearing.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies the linguistic gap between text prompts used to generate music and free-form descriptions of the resulting audio. The authors pair 200 real-world Udio prompts with their generated audio, collect 2,624 free-form descriptions from English-speaking (n=70) and Korean-speaking (n=78) listeners, and propose a seven-category human-derived taxonomy (Genre, Mood/Emotion, Instrumentation, Music Theory, Timbre, Function, Story/Narrative). Using GPT-5.4 to label all texts with these categories, they fit mixed-effects models for category presence and density, perform word-level survival-rate and chi-square association analyses, compute Sentence-BERT prompt-description cosine similarities, and compare high- versus low-alignment prompt groups. The main findings are that prompts are dominated by Genre and Story/Narrative, descriptions are richer in Instrumentation, Mood/Emotion, and Music Theory, genre terms propagate most reliably from prompt to perception, and narrative-heavy prompts are associated with the largest semantic misalignment. A preliminary cross-cultural comparison suggests that Korean descriptions reallocate lexical space toward affective, narrative, and functional framing relative to English descriptions.","tokens_in":10205,"tokens_out":2521,"duration_ms":25587,"significance":"If the findings hold, the paper makes a useful empirical contribution to music information retrieval and human-AI interaction: it moves beyond prompt taxonomies per se and directly measures how prompting language relates to the language of perceptual evaluation, using real-world prompts and paired audio rather than synthetic or curated materials. The human-derived taxonomy, the parallel English/Korean listener populations, the mixed-effects modeling that accounts for text, stimulus, and participant non-independence, and the triangulation of category-level, word-level, and vector-level evidence are genuine strengths. The cross-cultural comparison is explicitly exploratory but opens a question that is largely absent from TTM evaluation. The central claims are, however, only as strong as the automatic taxonomy labeling on which all category-level analyses rest, and that labeling is currently validated on a very small and partially mismatched sample.","major_comments":[{"comment":"The load-bearing measurement assumption is not adequately validated. GPT-5.4 is described as selected on \"a held-out, manually-annotated subset 25 of prompts from the Song Describer dataset,\" but no per-category agreement, confusion matrix, or reliability statistic is reported, and no validation is provided for the actual Udio prompts, for listener descriptions, or for Korean text. Every category-level result in §4.1, §4.2, §4.3, and §4.4 is computed from this labeling, so a systematic labeler bias (e.g., assigning Genre/Story more readily to short noun-phrase prompts and Mood/Instrumentation to longer descriptive sentences, or handling Korean less reliably) could produce the central prompt-description asymmetry and cross-cultural profile differences as artifacts. The paper's own §5.3 admits that classification noise is \"difficult to fully quantify.\" I ask for either a full validation on human-coded material drawn from the actual corpora (including both languages and both text types, with per-category agreement) or a robustness reanalysis of a human-coded subset that directly supports the main interaction and cross-cultural effects.","section":"§4.1, §5.3, §7"},{"comment":"The claim that \"narrative-heavy prompts are the strongest predictor of semantic misalignment\" is not supported by the analysis as presented. Table 3 compares the top and bottom 25% of prompts by mean prompt-description cosine similarity on one category at a time; it does not jointly test all categories, does not control for prompt length, genre density, or other categories, and does not estimate a predictive model of alignment. The descriptive pattern (Story/Narrative density 0.452 in low-alignment vs. 0.169 in high-alignment prompts) is suggestive, but the word \"strongest\" requires a joint model, e.g., a regression with all category densities predicting alignment, or at least a model that includes the other categories as covariates. Please either provide such an analysis or soften the claim to \"narrative-heavy prompts show the largest univariate gap between high- and low-alignment groups.\"","section":"§4.3, Table 3"},{"comment":"The cross-cultural comparison is difficult to interpret as a cultural effect because the language groups differ in recruitment channel (single university psychology pool vs. volunteer personal networks plus Prolific), compensation, and number of stimuli assigned (20 vs. 10 for volunteer Korean participants). The paper acknowledges these confounds in §5.3, but the abstract and conclusion state the cross-cultural result as a substantive finding. At minimum, the authors should either report analyses restricted to the comparable Prolific-recruited Korean subgroup versus the English participants, or explicitly characterize the Korean vs. English differences as confounded by participant pool and sampling procedure. Without such a robustness check, the cross-cultural redistribution claim is not yet supported.","section":"§4.4, §5.3"}],"minor_comments":[{"comment":"The chi-square association analysis tests a large number of prompt-description word pairs and reports only positive associations, but no multiple-comparison correction or false-discovery-rate control is described. The specific threshold choices (prompt word occurring more than ten times, at least twenty paired descriptions, co-occurrence at least three times) are reasonable but should be justified or varied in a sensitivity analysis.","section":"§4.2"},{"comment":"The taxonomy is derived from prompts only and then applied to descriptions; categories such as Function or Story/Narrative may have different boundaries in perceptual descriptions. The paper should note this construct-validity limitation more explicitly, or provide a small demonstration that the taxonomy covers description language without forcing descriptions into prompt-derived categories.","section":"§3.1, Table 1"},{"comment":"The concreteness analysis excludes unmatched tokens and reports coverage of 80.3% for prompts and 91.4% for descriptions; the difference in coverage rates could itself affect the prompt-description concreteness comparison. A sensitivity analysis including all tokens or using a different norm set would strengthen the claim.","section":"§4.2"},{"comment":"Equation (1) defines the mixed-effects model, but the description of the random-effects structure for prompts is slightly confusing: prompts have no human author, so the participant random intercept is estimated from descriptions only. This is fine, but the text should state explicitly that prompt observations therefore have a structurally different random-effects design, which may affect variance estimates for the corpus comparison.","section":"§4.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for ISMIR and addresses a timely question, but the central category-level claims depend on an LLM labeler whose validation is described too briefly and on a mismatched small sample. The authors are honest about the limitation in §5.3, which is to their credit, but the current evidence does not yet rule out labeler-driven artifacts for the main prompt-description asymmetry or the cross-cultural profile. A reasonable revision path would be to add per-category validation on the actual corpora (or a human-coded subset), report a joint model for the alignment claim, and temper or condition the cross-cultural conclusion. I would not reject the paper; the empirical design is otherwise solid and the findings, if confirmed, would be a useful contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know about this paper because it's the first quantitative comparison of real text-to-music prompts with listener descriptions, and the headline result is likely right: prompting and describing are structurally different. Prompts are dominated by genre and story/narrative; descriptions carry more instrumentation, mood, and music-theory language. Genre words propagate through the audio better than narrative language. That fits an intuitive story, and the authors support it with a human-derived taxonomy, a 200-stimulus listening study with 148 participants, and mixed-effects models that account for the nested structure. The word-level survival rates and the concreteness analysis give converging evidence. I think the central direction holds up.\n\nThe soft spots are real but not fatal. The biggest is the LLM labeler. They validate GPT-5.4 against human coding on only 25 prompts from Song Describer—not on their own Udio prompts, not on descriptions, and not on any Korean text. Then every category-level result, including the cross-cultural comparison, is computed from that labeling. The paper's own Section 5.3 admits this noise is hard to quantify. If the labeler systematically over-assigns Genre and Story to prompts and Mood/Instrumentation to descriptions, the central asymmetry could be inflated. I suspect the effect sizes are large enough that the prompt-description difference would survive a better validation, but the Korean comparison is more fragile, and the authors should either validate on their own data or soften the claims.\n\nSecond, they call narrative density 'the strongest predictor' of misalignment, but Table 3 is just a high/low quartile comparison on one variable at a time. There's no joint model with all categories. That's an overclaim, and it should be reworded or tested properly.\n\nThird, no data or code are released, which limits reproducibility; the recruitment differences between language groups are acknowledged but still muddy the cross-cultural part.\n\nWho should read this: anyone building or evaluating text-to-music systems, and people working on music captioning or prompt analysis. It deserves peer review—the question is important and the empirical design is mostly sound. I'd send it out, but require the authors to either strengthen the labeler validation or temper the strongest statements.","headline":"A plausible and useful first pass at a real question, but the LLM labeling validation is too thin to support the strongest claims.","tokens_in":10724,"tokens_out":2126,"would_cite":true,"duration_ms":20143,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Prompting a music generator and describing the music you hear are structurally different acts of language: genre and story dominate prompts, while instruments, mood, and music theory dominate descriptions.","keywords":["text-to-music generation","prompt language","music description","semantic alignment","cross-cultural music perception","taxonomy of musical language","listening experiment"],"falsifier":"Have human annotators manually code every prompt and every description using the same seven-category taxonomy, then recompute the presence and density regressions; if human labels do not reproduce the genre/story dominance of prompts and the instrumentation/mood dominance of descriptions, the central asymmetry is an artifact of the automatic labeler.","tokens_in":9695,"feed_emoji":"🎵","tokens_out":6626,"duration_ms":55591,"temperature":0.7,"pith_summary":"The paper sets out to show that prompting a text-to-music system and describing heard music are not two ways of saying the same thing but structurally different communicative acts. Using 200 real prompts paired with free-form descriptions from English and Korean listeners, it builds a human-derived taxonomy of seven dimensions of musical language and finds that prompts are dominated by genre and story, while descriptions are richer in instrumentation, mood, and music theory. The paper argues this asymmetry matters because text-to-music models are trained on metadata-style labels rather than on the language listeners actually use, and because narrative-heavy prompts predict the largest mismatch between what a prompt imagines and what listeners hear.","feed_headline":"Genre drives prompts; instruments and mood drive descriptions","feed_subtitle":"Across 200 real prompts and 2,624 listener descriptions, story-heavy prompts predict the largest prompt-to-perception gap.","key_machinery":"The load-bearing instrument is a seven-category human-derived taxonomy of musical vocabulary (genre, mood/emotion, instrumentation, music theory, timbre, function, story/narrative), constructed by two annotators from a sample of real prompts. Coding every prompt and description into this taxonomy, with presence and density measures, plus word-level survival and association tests and sentence-embedding cosine similarity, lets the paper compare prompting and describing across category, word, and vector levels. Mixed-effects regressions with random intercepts for text, stimulus, and participant separate category-level differences from item and participant noise.","core_discovery":"The central discovery is a quantitative linguistic asymmetry: genre appears in about 95% of prompts but only about 61.5% of descriptions, while instrumentation, mood/emotion, and music theory are significantly more common in descriptions than in prompts. Genre terms travel best from prompt to perception, often as semantic neighbors rather than the same word, whereas narrative-heavy prompts show the weakest prompt-to-description semantic similarity, with narrative density nearly three times higher in low-alignment prompts. The paper interprets this not as failed generation but as a structural fact: narrative intent has no acoustic trace, so no amount of generative capacity can recover it. Cross-culturally, Korean and English descriptions differ in emphasis, with Korean listeners allocating more lexical space to mood/emotion, narrative, and function, and English listeners more to genre and music theory.","pith_inferences":["Left implicit in the paper: the same prompt-description asymmetry may hold for other generative modalities such as text-to-image or text-to-video, where narrative prompts are also common; testing the taxonomy in those domains would show whether the genre-versus-narrative split is music-specific.","A natural extension is to collect non-English prompts, not just non-English descriptions, to see whether the prompt-description gap itself shrinks or widens when both ends of the chain use the same non-English language.","If the cross-cultural pattern is real, text-to-music systems could be evaluated in a language-aware way, comparing model outputs against the distribution of listener descriptions in each language rather than a single aggregated English norm."],"forward_implications":["If prompts and descriptions are structurally different registers, text-to-music evaluation that relies on prompt-style metadata or automatically generated captions may not reflect how listeners actually describe music.","Genre labels are the most reliable bridge from prompt to perception, so systems aiming for semantic alignment should favor genre and concrete sonic markers over narrative framing.","Narrative-heavy prompts will keep producing low semantic alignment unless text-to-music systems learn to encode the acoustic features that support shared narrative perception.","Current English-centric training data may disadvantage users whose natural descriptive vocabulary is affective, narrative, or functional, as the English-Korean comparison suggests.","The contributed taxonomy can be reused as an annotation scheme for other prompt corpora and listening studies."],"supporting_citations":[{"why":"Supplies the real-world prompt corpus that generated the 200 audio stimuli.","marker":"[8]"},{"why":"Provides the example descriptions shown during the practice phase and the validation prompts used to pilot the automatic labeler.","marker":"[17]"},{"why":"Provides the sentence-embedding model used to compute prompt-description semantic similarity.","marker":"[24]"},{"why":"Provides the mixed-effects regression software that separates category effects from text, stimulus, and participant variation.","marker":"[21]"},{"why":"Supplies evidence that large-language-model classifiers can match human text annotations when validated, justifying the labeling pipeline.","marker":"[18]"},{"why":"Provides the prior cross-cultural finding that U.S. and East Asian listeners describe music in structural versus metaphorical terms, which the Korean/English comparison extends.","marker":"[15]"},{"why":"Supplies the concreteness norms used to compare the register of prompt and description vocabulary.","marker":"[23]"},{"why":"Supplies the earlier observation that perceptual and affective dimensions are central to listening, which the prompt-description asymmetry is consistent with.","marker":"[22]"}],"fun_headline_variants":["Prompts say genre, descriptions say mood and feel","Story-heavy prompts widen AI music perception gap","What you prompt isn't what they describe","Korean and English listeners hear AI music differently"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything rests on the automatic taxonomy labeler being accurate enough, even though it was checked against human coding on only 25 English prompts; if that labeler carries language- or category-specific bias, the prompt-description asymmetries and the Korean/English differences could be artifacts of the labeling tool rather than of the language people use.","fun_headline_variants_meta":{"raw":{"variants":["Prompts say genre, descriptions say mood and feel","Story-heavy prompts widen AI music perception gap","What you prompt isn't what they describe","Korean and English listeners hear AI music differently"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001113,"raw_usage":{"total_tokens":4607,"prompt_tokens":888,"completion_tokens":3719,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":504,"completion_tokens_details":{"reasoning_tokens":3662}},"tokens_in":504,"tokens_out":3719,"duration_ms":25256,"temperature":1.0,"reasoning_tokens":3662,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T04:10:01.250641+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have human annotators manually code every prompt and every description using the same seven-category taxonomy, then recompute the presence and density regressions; if human labels do not reproduce the genre/story dominance of prompts and the instrumentation/mood dominance of descriptions, the central asymmetry is an artifact of the automatic labeler.","supporting_citations":[{"cited_title":"Data-driven analysis of text-conditioning in ai-generated music: A case study with suno and udio,","cited_arxiv_id":null,"evidence_quote":"Supplies the real-world prompt corpus that generated the 200 audio stimuli."},{"cited_title":"A taxonomy of prompt mod- ifiers for text-to-image generation,","cited_arxiv_id":null,"evidence_quote":"Provides the example descriptions shown during the practice phase and the validation prompts used to pilot the automatic labeler."},{"cited_title":"Sentence-bert: Sentence embeddings using siamese bert-networks,","cited_arxiv_id":null,"evidence_quote":"Provides the sentence-embedding model used to compute prompt-description semantic similarity."},{"cited_title":"A prompt log analysis of text-to-image generation systems,","cited_arxiv_id":null,"evidence_quote":"Supplies evidence that large-language-model classifiers can match human text annotations when validated, justifying the labeling pipeline."},{"cited_title":"Preference responses and use of written descriptors among music and nonmusic majors in the united states, hong kong, and the people’s republic of china,","cited_arxiv_id":null,"evidence_quote":"Supplies the concreteness norms used to compare the register of prompt and description vocabulary."}],"review_version":1}