{"id":"3a459246-587f-47e6-b0af-27ad6ad3086b","arxiv_id":"2504.18799","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of multimodal music emotion recognition that organizes roughly two dozen papers into a four-stage framework and finds audio-plus-lyrics deep learning fusion to be the dominant approach.","lead":"This paper surveys how computers recognize emotion in music using multiple data types, such as audio, lyrics, video, and brain signals. It organizes the field into four stages, from data selection to emotion prediction, and identifies data scarcity and missing benchmarks as the main bottlenecks.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 6, which grounds every headline trend, has demonstrable citation mismatches and ranks incomparable accuracies; until rows are re-verified the central claims rest on an unvetted evidence table.","rationale":"The reader's weakest assumption is that Table 6 performance numbers are mutually comparable and correctly sourced. My stress-test pass confirms that this is the most load-bearing concern. The survey's value is as an organizing reference, and its central assertions about field-level trends are frequency claims over the surveyed literature. Those frequencies are computed from Table 6, so a single misattribution or an incomparable accuracy ranking directly undermines the conclusions drawn in Section 4.1.2. The citation mismatches already identified for [40] and [76] are not merely cosmetic; they show the reference list was not systematically checked, which raises the prior that other rows may also misrepresent their sources. The comparability problem is equally concrete: ranking 94.58% against numbers from PMEmo, DEAM, and self-collected datasets with different label schemes is not a valid state-of-the-art statement without matching dataset and class count. I do not think these defects require rejection, because the structural taxonomy and the qualitative observation that audio and lyrics dominate are consistent with the cited literature and with other surveys. But the survey should be accepted only conditionally, with a full re-verification of Table 6 as a prerequisite for relying on its quantitative trend claims. This matches the reader's CONDITIONAL verdict, so no change is needed.","tokens_in":24599,"tokens_out":3940,"duration_ms":43526,"concrete_test":"Re-verify every row of Table 6 against its cited paper: confirm that the cited content matches the row's method and approach label, and record dataset, label scheme, class count, and evaluation split for each reported number. Then rebuild the Section 4.1.2 ranking only within groups matched on dataset and label scheme. If [77]'s 94.58% is not the best within its matched group, or if any method or approach cell changes on re-verification, the trend claims and the 'highest accuracy to date' statement must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The survey's central empirical claims depend on Table 6: that most MMER systems use audio plus lyrics, that late fusion with deep learning is the current high-water mark, and that the four-category fusion taxonomy organizes the literature. Yet Table 6 is not a verified evidence base. Section 3.3.2 attributes LFSM to reference [40], which is a SemEval misogyny-identification paper, and Table 2 attributes Hevner's Emotional Model to [76], a circumplex-model paper. If these two entries are wrong, the table's other method, approach, and performance assignments cannot be assumed correct without rechecking each one. Furthermore, Section 4.1.2 ranks raw accuracies across MoodyLyrics, PMEmo, DEAM, self-collected sets, and other datasets with different class counts, label schemes, and evaluation splits, declaring 94.58% from [77] the highest to date. That number is not commensurable with the others unless dataset, class count, and label scheme are held fixed. Because the same table is the sole support for the survey's trend analysis, the headline empirical claim is insecure until each row is re-verified against its primary source and performance comparisons are restricted to matched settings.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript surveys multimodal music emotion recognition (MMER), proposing a four-stage framework (data selection, feature extraction, feature processing, emotion prediction) with three feature-processing approaches and four fusion categories. It reviews emotion models, modalities, datasets, and evaluation metrics, and compiles a table of 23 MMER systems (Table 6) to support claims about current trends: that audio-plus-lyrics is the dominant modality pair, that deep learning and late fusion represent the current state of the art, and that no MMER-specific benchmark exists. The survey also identifies dataset scarcity, benchmark absence, and lack of model interpretability as key gaps, and suggests future directions including more modalities, transfer learning, and real-time processing.","tokens_in":24868,"tokens_out":9499,"duration_ms":89128,"significance":"If the claims are substantiated, the survey provides a useful organizational scaffold for a growing field and a clearly articulated account of the missing MMER benchmark, which is a falsifiable and actionable gap. The four-stage framework and the fusion taxonomy are simple and reusable, and Table 3 is a convenient compact dataset reference. The paper does not present machine-checked proofs or code, but for a survey the main contribution is the synthesis; that synthesis, however, rests heavily on Table 6, whose provenance and comparability need to be established before the trend claims can be accepted.","major_comments":[{"comment":"The global performance ranking in Section 4.1.2 is not supported because Table 6 pools results from different datasets (MoodyLyrics, DEAM, PMEmo, FMA, self-collected sets) that use different emotion taxonomies, class counts, and evaluation protocols. The sentence declaring 94.58% as 'the highest accuracy of 94.58% for classification to date' from reference [77] is only meaningful if dataset, label scheme, class count, and evaluation split are held fixed, which they are not. Please restrict any accuracy comparisons to matched settings, or replace global rankings with per-dataset, per-task tables.","section":"§4.1.2, Table 6"},{"comment":"Several source attributions in the evidence tables are demonstrably wrong. Section 3.3.2 attributes late fusion subtask merging (LFSM) to reference [40], but [40] is a SemEval misogyny-identification paper, not an MMER work. Table 2 attributes Hevner's Emotional Model to reference [76], which is Posner et al.'s circumplex model of affect paper. Table 6 row [39] lists SVM as the method, although reference [39] is the bi-modal deep Boltzmann machine paper and Section 3.2.1 correctly describes it as DBM. Because Table 6 is the sole evidence base for the survey's trend claims, every row of the table and every associated attribution should be re-verified against primary sources; a supplementary provenance table or verification note should be provided.","section":"§3.3.2, Table 6, Table 2"},{"comment":"The proposed framework is described inconsistently. The Section 3 introduction states that Stage 4 'commonly involves one of two distinct approaches,' but Section 3.3 defines four fusion categories (feature-level, decision-level, model-level, cross-modal) and Figure 5 shows four such strategies. Additionally, Section 3.2.2 presents Approach 2-A and 2-B as variants of a single 'modality-specific feature processing' approach, while Figure 3 and Table 6 count them as separate approaches. The taxonomy should be enumerated consistently within the text and figures, or the text should explicitly state which level (task type vs. fusion strategy) is being counted.","section":"§3 (framework description)"},{"comment":"No literature search protocol or inclusion criteria are reported, despite the claim of a 'comprehensive overview' and a focus on recent deep learning work 'since 2022.' The reader cannot determine whether Table 6 is an exhaustive enumeration or a representative sample, what databases and query terms were used, or what criteria excluded other MMER works. This matters for the 'majority' claims in Section 4.1.2 and Section 5, which are quantitative statements about the literature. Please add a methodology paragraph covering the search strategy, screening criteria, and coverage dates, and state explicitly whether Table 6 is exhaustive or representative.","section":"§1, §3, Table 6"}],"minor_comments":[{"comment":"The abstract contains the typo 'robust, scalable, a interpretable models' (should be 'and interpretable'), and Section 5 contains 'other (con)textural data,' which appears to be a typographical error for 'contextual.'","section":"Abstract; §5"},{"comment":"The caption credits the features to 'Patrik et al. [50],' but reference [50] is the paper by Juslin and Laukka; the author name should be corrected.","section":"Table 1 caption, §2.1"},{"comment":"The text says that Thammasan [90] employed 'the MIRToolbox developed by Laurier et al. [59]'; reference [59] is by Lartillot, Toiviainen, and Eerola, so the developer name should be corrected to Lartillot et al.","section":"§3.1.5"},{"comment":"The year for reference [13] is listed as 2021 in Table 6, but reference [13] is dated 2020 and Table 4 also lists 2020; the entries should be made consistent.","section":"Table 6, row [13]"},{"comment":"The sentence reporting the 94.58% accuracy should explicitly name the dataset (MoodyLyrics) and the class/label configuration, since the bare number is not interpretable without that context, and the reference appears only in the table.","section":"§4.1.2"},{"comment":"The phrase 'late fusion subtask merging (LFSM)' is introduced without a definition; please define the acronym at first use and, if it is not a standard term in the MMER literature, state which source introduced it.","section":"§3.3.2"}],"recommendation":"major_revision","confidential_remarks":"The reference errors found in Table 2 and Table 6 go beyond ordinary typos: at least one row in the central evidence table attributes a method to a paper on a different task, and another row lists a method that contradicts the cited paper's content. I would ask for a systematic audit of all Table 6 entries against primary sources before a second round, and for the authors to state explicitly which numbers come from matched evaluation settings. This is a reliability issue for a survey whose main claims are empirical statements about the literature."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is the closest thing I have seen to a usable reference survey for multimodal music emotion recognition. The four-stage pipeline (data selection, feature extraction, feature processing, emotion prediction) and the three-plus-four taxonomy for processing and fusion are handy organizing devices, and the consolidation of datasets in Table 3 is genuinely useful. The central claim that audio plus lyrics dominates the field and that no MMER-specific benchmark exists also matches my reading of the literature. The paper is not a big conceptual leap, but it is a solid professional service to the community.\n\nThe soft spots are real, and they sit in the place where a survey lives or dies: the evidence table. I spot-checked two entries and both failed. Hevner's Emotional Model is attributed to reference [76], which is the Posner et al. circumplex paper, not Hevner. And 'late fusion subtask merging (LFSM)' in Section 3.3.2 is cited to [40], which is a SemEval misogyny-identification paper that has nothing to do with LFSM. Two verified attribution errors in a table of about twenty-five rows do not automatically sink every row, but they do mean the table cannot be trusted until each row is rechecked against its primary source.\n\nThere is also a commensurability problem. Section 4.1.2 declares 94.58% on MoodyLyrics the highest accuracy to date, but the neighboring rows come from DEAM, Indonesian songs, self-collected sets, and other datasets with different label schemes, class counts, and evaluation splits. That is an apples-to-oranges ranking. The statement should be restricted to matched settings or dropped.\n\nOne more issue: the survey claims comprehensiveness but says nothing about how papers were collected or screened. No PRISMA flow, no database list, no inclusion criteria. For a survey that wants to be cited as authoritative, that is a genuine reproducibility gap. Also minor: Section 4.1.1 reports 79.2 for Liu and Tan while Table 6 says 79.62, and the conclusion says 'textural data' where it means textual.\n\nNone of this kills the paper's organizing value. The four-stage framework is descriptive and asserted rather than proven exhaustive, but that is acceptable for a survey. The central trends would probably survive a full re-verification; they are consistent with the rest of the field as I know it. The problem is that the current version asks readers to take too much on faith.\n\nMy recommendation: send it to peer review, but with a clear requirement that the authors verify every row of Table 6, correct the citation mismatches, restrict or caveat cross-dataset performance comparisons, and describe their corpus selection method. After those fixes, this becomes a citeable reference. I would not cite its numbers before then.","headline":"A genuinely useful map of MMER with a sensible four-stage framework and fusion taxonomy, but the evidence table has two confirmed citation errors and the headline accuracy ranking is not commensurable; it deserves peer review but only with mandatory table-level verification.","tokens_in":25349,"tokens_out":2636,"would_cite":false,"duration_ms":30719,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Audio plus lyrics dominates music-emotion AI, survey finds","keywords":["multimodal music emotion recognition","music information retrieval","feature fusion","audio and lyrics","deep learning","emotion annotation","cross-modal processing","benchmark datasets"],"falsifier":"Run the surveyed systems on one held-out multimodal test set with a single emotion-annotation protocol; if the best model changes or the audio-plus-lyrics advantage disappears, the survey's core trend claim would be overturned.","tokens_in":24287,"feed_emoji":"🎵","tokens_out":4263,"duration_ms":44025,"temperature":0.7,"pith_summary":"This survey tries to establish that multimodal music emotion recognition (MMER) can be understood as a four-stage pipeline — data selection, feature extraction, feature processing, and emotion prediction — and that every published approach is a combination of one of three feature-processing options and one of four fusion strategies. Its main empirical finding is that the field has converged on audio-plus-lyrics pairs with deep-learning fusion, while other modalities remain thinly explored. It also argues that progress is gated less by model design than by data: most datasets are small, single-genre, and often copyright-restricted, and no benchmark has been built specifically for MMER. The survey matters because it maps where the field is concentrated and where capacity is missing, which is the information needed to decide what datasets and evaluation protocols should be built next.","feed_headline":"Audio plus lyrics dominates music-emotion AI, survey finds","feed_subtitle":"A four-stage framework shows most systems fuse two modalities with deep learning, and no shared benchmark exists yet.","key_machinery":"The organizing object is the four-stage MMER framework, with Stage 3 and Stage 4 carrying the taxonomy: three feature-processing approaches (feature concatenation, modality-specific processing, and cross-modal processing) plus four fusion categories (feature-level, decision-level, model-level, and cross-modal). The survey uses this grid as the axis of its literature table, and the trend claims follow from counting which cells the surveyed systems occupy. The cross-modal cell is the newest and is carried by an emotion long short-term memory (E-LSTM) cell that passes a historical emotion vector from one lyric-audio pair to the next, which the survey identifies as the mechanism that lets emotional states persist across modal interactions.","core_discovery":"The paper's central claim is that MMER research is organized by a four-stage pipeline: choose multimodal data (audio, lyrics, video, MIDI, physiological signals, text, metadata), extract features from each, process those features in one of three ways (direct concatenation, modality-specific processing, or cross-modal interaction), then predict emotion using feature-level, decision-level, model-level, or cross-modal fusion. Within that organization, the survey finds that the large majority of systems use audio and lyrics as the only modalities, that deep learning (CNN, LSTM, and BERT-family models) now dominates, and that reported accuracy varies widely across datasets and emotion models. It further claims that no MMER-specific benchmark exists, so cross-paper performance comparison is informal, and that the highest published classification accuracy in the surveyed table is 94.58%, from a CNN-BERT audio-plus-lyrics late-fusion system. The authors' conclusion is that the field's bottlenecks are the absence of a precise emotion model, limited and narrowly scoped datasets, and missing benchmarks, not the fusion architectures themselves.","pith_inferences":["The survey does not test this, but its taxonomy implies a controlled experiment that would settle the fusion question: hold the feature extractors fixed, vary only the fusion strategy on one dataset, and compare the three processing approaches directly.","A reader could infer from the table that cross-modal approaches should beat concatenation on the same dataset when the modalities carry complementary information, because cross-modal methods are the only ones that model interaction explicitly.","The survey's own comparison table suggests a testable hypothesis the authors do not pursue: a standardized MMER benchmark with one emotion model, one annotation protocol, and per-modality ablations would likely rewrite the current ranking, since several top reported numbers come from different datasets and class counts."],"forward_implications":["If the field is as concentrated on audio plus lyrics as the survey counts suggest, adding a third modality — video, MIDI, physiological signals, or metadata — is the most direct route to new performance, because the under-explored cells have the most headroom.","The absence of an MMER benchmark means that any claim of state-of-the-art status, including the 94.58% figure, is only meaningful within the dataset and annotation scheme it was reported on.","If dataset scale and diversity are the bottleneck, then unsupervised collection pipelines and transfer learning are natural next steps, as the survey itself recommends.","The four-stage framework gives future papers a shared vocabulary, so a new system can be described by which stage-3 approach and stage-4 fusion it uses rather than by ad-hoc architecture names."],"supporting_citations":[{"why":"Supplies the 94.58% audio-plus-lyrics late-fusion result that anchors the survey's trend claim about state-of-the-art classification accuracy.","marker":"[77]"},{"why":"Provides the earliest audio-and-lyrics multimodal approach, a basis for the claim that these two modalities dominate the field.","marker":"[60]"},{"why":"Reports unimodal versus multimodal accuracy comparisons that support the survey's claim that combining modalities improves emotion recognition.","marker":"[13]"},{"why":"Supplies the hierarchical cross-modal attention network example that defines the cross-modal fusion category.","marker":"[117]"},{"why":"Provides the E-LSTM mechanism for processing lyric-audio pairs, the central example of cross-modal feature processing.","marker":"[97]"},{"why":"Contributes the Hough forest model-level fusion example supporting the model-level fusion category.","marker":"[104]"},{"why":"Shows MIDI and EEG used together with an SVM, evidence for the claim that non-audio/lyric modalities remain thinly explored.","marker":"[90]"},{"why":"Demonstrates audio-plus-MIDI fusion on EMOPIA, supporting the survey's account of symbolic modalities in MMER.","marker":"[118]"},{"why":"One of the standard audio-based MER datasets cited to show that existing benchmarks are not multimodal.","marker":"[3]"},{"why":"A multimodal physiological dataset whose copyright and size limitations the survey uses to motivate the need for larger MMER datasets.","marker":"[57]"}],"fun_headline_variants":["Audio+lyrics rules multimodal emotion recognition survey","MMER survey: Two modalities dominate, benchmarks missing","Four-stage framework maps music emotion AI landscape","Deep learning and fusion key in music emotion survey","No shared benchmark stymies music emotion recognition"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's trend and ranking claims assume that the accuracy numbers collected from different papers, tested on different datasets with different emotion labels and class counts, are directly comparable and that each table entry faithfully reports its source.","fun_headline_variants_meta":{"raw":{"variants":["Audio+lyrics rules multimodal emotion recognition survey","MMER survey: Two modalities dominate, benchmarks missing","Four-stage framework maps music emotion AI landscape","Deep learning and fusion key in music emotion survey","No shared benchmark stymies music emotion recognition"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00108,"raw_usage":{"total_tokens":4500,"prompt_tokens":912,"completion_tokens":3588,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":528,"completion_tokens_details":{"reasoning_tokens":3517}},"tokens_in":528,"tokens_out":3588,"duration_ms":24194,"temperature":1.0,"reasoning_tokens":3517,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:08:49.654736+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the surveyed systems on one held-out multimodal test set with a single emotion-annotation protocol; if the best model changes or the audio-plus-lyrics advantage disappears, the survey's core trend claim would be overturned.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the earliest audio-and-lyrics multimodal approach, a basis for the claim that these two modalities dominate the field."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Reports unimodal versus multimodal accuracy comparisons that support the survey's claim that combining modalities improves emotion recognition."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the hierarchical cross-modal attention network example that defines the cross-modal fusion category."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the E-LSTM mechanism for processing lyric-audio pairs, the central example of cross-modal feature processing."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Contributes the Hough forest model-level fusion example supporting the model-level fusion category."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows MIDI and EEG used together with an SVM, evidence for the claim that non-audio/lyric modalities remain thinly explored."}],"review_version":1}