{"id":"49e7c31d-5e00-4a0c-850f-975f18a362d7","arxiv_id":"2506.08200","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A rule-based system generates retro-pop music at target levels of arousal and valence, validated by a listening study with high correspondence between target and perceived ratings.","lead":"AffectMachine-Pop is a new expert system that generates retro-pop music matched to chosen levels of emotional intensity (arousal) and pleasantness (valence). A listening study found that listeners perceived the music as close to the target emotions, supporting its use for real-time interactive and therapeutic applications.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The high R² values are computed over 13 aggregated point means, so the claim of reliable target-level control does not yet address the variance across the three instantiations at each point.","rationale":"The reader's weakest assumption already names the point-level averaging; I agree, and this is the load-bearing seam. The R² values are computed on group means, not on the ratings that a user would experience in a single generated excerpt. Since the system uses probabilistic selection at multiple levels (chord transitions, rhythmic patterns, melodic matrices), between-instantiation variation is expected. The paper's own explanation for the A=0.5 deviation acknowledges that specific instantiations can underperform the mean, which is exactly the variance that is hidden by aggregation. Without an analysis of per-stimulus responses, a high R² at the level of 13 means cannot distinguish precise affective control from a coarse monotonic mapping that happens to order the sampled regions of the V-A plane. There is no code or data release, so this concern cannot be resolved from the manuscript. I do not think this invalidates the conditional verdict: the system is described transparently and the direction of the effect is plausible, so the correct default is conditional acceptance with a required robustness check. If the proposed mixed-effects or leave-one-point-out analysis shows collapse, the verdict should move toward reject or unverified. If it confirms low within-point variance and stable prediction at unsampled points, the central claim would be substantially strengthened.","tokens_in":9892,"tokens_out":6464,"duration_ms":86353,"concrete_test":"Ask the authors to release the per-participant ratings and the 39 stimulus ratings. Fit a linear mixed-effects model, e.g. rating ~ target valence + target arousal + (1 | participant) + (1 | stimulus), report marginal R², the random-effect variance components, and the intraclass correlation among the three instantiations at each of the 13 points. Also run leave-one-point-out prediction across the 13 sampled points. If the marginal R² drops materially below 0.86/0.93, or if the three instantiations at a point disagree by more than about 0.2 normalized units on average, the current analysis does not establish reliable control at target levels.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central evidence for the system's controllability is the regression of the 13 point means (each averaged over 24 participants and 3 generated instantiations) against the target valence/arousal parameters, reported as R² = 0.93 and 0.86 in the Results and Discussion. Because the system is explicitly probabilistic, the three instantiations at each target point are not guaranteed to be similar; the paper itself concedes that the A = 0.5 stimuli 'may have deviated from the system's typical output,' but no variance decomposition or per-stimulus analysis is reported. Aggregating participants and instantiations can produce a high ecological correlation even if individual excerpts are noisy, and the 13-point R² does not reveal whether ratings match target levels or merely follow a monotonic trend. The claim that the system reliably generates music 'at target levels' therefore rests on the unverified assumption that the mean of each point is representative and stable. Slopes, intercepts, prediction intervals, and per-excerpt variability would be needed to support the claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents AffectMachine-Pop, a rule-based expert system for generating retro-pop music with specified valence and arousal values. The system uses hand-crafted probabilistic chord progressions, voice-leading matrices, tempo and velocity mappings, and rhythmic patterns to map points in Russell's two-dimensional valence-arousal space to musical parameters. The authors report a listening study with 24 participants who rated 39 musical excerpts (13 V-A target points, three instantiations each) using the Self-Assessment Manikin. They report a strong linear correspondence between target parameter settings and averaged listener ratings, with R^2 = 0.93 for valence and R^2 = 0.86 for arousal, and interpret this as confirming the system's ability to generate affective music at target levels. The paper also claims real-time generation capability and discusses applications in interactive music and biofeedback/neurofeedback systems.","tokens_in":10112,"tokens_out":3688,"duration_ms":47863,"significance":"If the central claim holds, the paper is a useful contribution to affective music generation: it offers a transparent, rule-based alternative to black-box neural models, with no reliance on copyrighted training data, and it explicitly addresses controllability along arousal and valence dimensions. The authors are candid about the system's design choices and about limitations such as sample size. The listening study, however, is currently the weakest link: the headline R^2 values are computed on a small number of aggregated points, and the paper does not report the variance across the three instantiations at each point or across participants. Because the system is stochastic and the paper itself acknowledges that particular stimuli 'may have deviated from the system's typical output,' the evidence as presented does not yet support the strong conclusion that the system reliably produces music 'at target levels' across the continuous V-A plane. With additional statistical analysis and a more complete report of per-stimulus results, the contribution would be substantially strengthened.","major_comments":[{"comment":"The central evidence for controllability is a regression fit to 13 point means, each averaged over 24 participants and 3 Monte Carlo instantiations. This aggregation discards the variance across instantiations and participants. Because the system is probabilistic, the three stimuli at each target point are not guaranteed to express the same emotion; the paper itself notes that the A = 0.5 stimuli 'may have deviated from the system's typical output.' The reported R^2 therefore represents an ecological correlation and does not establish that individual excerpts, or even the per-point means, are near target levels. Please report per-stimulus means with standard errors or prediction intervals, a variance decomposition (participant versus instantiation), and ideally a mixed-effects model with target as a fixed effect and stimulus/participant as random effects.","section":"Results and Discussion (Figure 3, regression analysis)"},{"comment":"The manuscript does not report whether the regression slope and intercept differ from the identity line (rating = target). An R^2 of 0.93 or 0.86 shows a strong linear trend but not that ratings match the intended levels, especially given the systematic deviations described in the text: the valence plateau for target values above 0.75 and the arousal dip around A = 0.5. These deviations are explained post hoc rather than quantified. Please provide the fitted slope, intercept, confidence intervals, and, if possible, a test of whether the relation is consistent with identity; otherwise the claim of generating music 'at target levels' is overstated.","section":"Results and Discussion (comparative analysis)"},{"comment":"The generality of the validation is limited by the sampling of both stimuli and listeners. The 13 V-A points are not specified in this paper; the reader is referred to prior work, which makes the validation not self-contained. Only 13 points are used to represent the continuous valence-arousal plane, and the three instantiations per point are too few to estimate generation variability. The participant sample (N = 24, mean age 23.0, university students) is small and homogeneous. The conclusion that the system controls emotion across the entire V-A plane is not supported by the current data. Please list the exact 13 coordinates, justify their coverage of the space, and report per-point variability across instantiations and participants.","section":"Musical Stimuli Generation and Experimental Procedure"},{"comment":"The title and abstract claim 'real-time' music generation, and the introduction states that the system features 'near real-time adaptability,' but no latency measurement, timing benchmark, or demonstration of continuous re-parameterization during playback is provided. If real-time operation is part of the contribution, please report a technical measurement or at least an architecture-level latency estimate; otherwise, the real-time claim should be softened or explicitly labeled as a design goal rather than a validated property.","section":"Abstract, Conclusion, and System Description (real-time claims)"}],"minor_comments":[{"comment":"The text reads 'To access whether our system accurately expresses the desired emotions'; this should be 'To assess whether.' Also, 'plataeu' should be 'plateau' in the Results section.","section":"Experimental Procedure"},{"comment":"Please clarify whether the normalization in Eq. (2) was applied to each participant's ratings individually or to the pooled raw ratings, and state the exact mapping from the 9-point SAM scale to the normalized [0, 1] range in the figure axes.","section":"Equation (2) and Figure 3"},{"comment":"The manuscript states an average stimulus duration of 32.6 seconds but later describes low-arousal stimuli as 4 bars ranging from 27 to 34 seconds; please clarify whether the 32.6-second average is computed over all 39 excerpts and explain the duration variability more precisely.","section":"Musical Stimuli Generation"},{"comment":"The tempo range is reported as [36, 130] bpm with a logarithmic mapping to arousal, but no equation is given. Since tempo is a primary manipulation, please provide the exact mapping or a reference to where it is specified.","section":"System Description (Tempo)"},{"comment":"The paper focuses on perceived emotion and explicitly defers induced emotion to future work. The introduction and conclusion would benefit from clearer wording that the validation concerns perceived emotion only, so that readers do not infer that the system has been shown to reliably induce emotional states in listeners.","section":"Conclusion / Future Work"}],"recommendation":"major_revision","confidential_remarks":"The paper is a system description with a validation that is currently under-powered for the strength of the claims. The main issue is statistical: the headline R^2 values are computed on 13 aggregated points, which can be high even when individual excerpts are noisy. The authors should be asked to provide per-excerpt variance, prediction intervals, and a more careful treatment of the deviations they explain post hoc. If those analyses are supplied and are favorable, the paper could be a solid contribution to the affective music generation literature. I also note that the paper is formatted with an AAAI copyright notice, which may be unusual for the venue; the editor may wish to check policy."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuine, clearly-described extension of the authors' AffectMachine-Classical system into retro-pop, and the listening study gives real evidence that listeners perceive the intended valence/arousal trends. The high R² values (0.93 valence, 0.86 arousal) are computed on 13 aggregated point-means, which flatters stability, so the 'reliably at target levels' claim is a bit stronger than the data show. But that is a caution, not a fatal flaw.\n\nWhat's new: the system itself—new chord progression graph, rhythm and instrumentation rules, low/high valence melodic matrices, tempo mapping. It's a rule-based alternative to black-box neural generators, and the paper is refreshingly transparent about the design. The validation approach is reasonable for a first pass: 13 points spanning the V-A plane, 3 instantiations each, 24 raters, external SAM ratings. The regression on the 13 means shows a clear monotonic correspondence, and the authors honestly report the valence plateau and the arousal dip at A=0.5 rather than hiding them.\n\nSoft spots. The main one: aggregating over 24 participants and 3 instantiations per point can produce an ecological correlation that overstates per-excerpt accuracy. The paper gives no variance decomposition, no per-stimulus analysis, no prediction intervals. The authors themselves say the A=0.5 stimuli 'may have deviated from the system's typical output'—that is exactly the kind of variance they should have quantified. Second, the 'real-time' capability is asserted but never demonstrated; there are no latency or throughput measurements. Third, no code, audio examples, or data are released, which makes independent validation harder than it should be for a rule-based system. The sample is small and homogeneous (N=24, mean age 23), but the paper acknowledges that.\n\nProportion: none of these sink the central claim. The system plausibly does what it says—generate pop music whose perceived valence and arousal track the input parameters on average. The paper is an incremental but honest step in a niche line of research, and the authors are clear about what they have and haven't shown.\n\nWho should read it: people working on affective music generation, biofeedback/BCI music, or controllable generative music. It deserves a real peer review; the referee should ask for per-excerpt variability and ideally released stimuli. I'd send it to review.","headline":"A transparent rule-based pop music generator with credible but provisional validation; the R²s are real but rest on aggregated point-means.","tokens_in":10599,"tokens_out":2671,"would_cite":true,"duration_ms":30433,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AffectMachine-Pop, a rule-based expert system, composes retro-pop music at specified arousal and valence levels, and a 24-participant listening study found listeners' mean ratings track the target settings closely ($R^2 = 0.93$ for…","keywords":["affective music generation","expert system","real-time music generation","valence-arousal model","pop music","rule-based music generation","emotion control","listening study"],"falsifier":"Generate music at a held-out point on the valence-arousal plane not among the 13 sampled (for example, valence 0.65 and arousal 0.35), have a fresh group of listeners rate it on the same 9-point Self-Assessment Manikin scales, and compare the mean normalized ratings with the target; the system's claim is falsified if the linear correspondence found at the sampled points does not extend, for instance if the mean ratings at such held-out points are systematically offset by more than the error observed in the original study.","tokens_in":9693,"feed_emoji":"🎵","tokens_out":6119,"duration_ms":67990,"temperature":0.7,"pith_summary":"This paper claims that a rule-based expert system called AffectMachine-Pop can compose retro-pop music in real time at specified levels of arousal and valence, the two dimensions of Russell's circumplex model of emotion. Unlike black-box generative models, the system is built from hand-tuned musical rules and does not train on recordings, making its output deterministic and transparent enough to control. In a listening study with 24 participants rating 39 excerpts generated from 13 target points on the valence-arousal plane, mean perceived valence and arousal tracked the target settings closely, with linear-fit $R^2 = 0.93$ for valence and $R^2 = 0.86$ for arousal. If these results hold, the system offers a practical path to interactive music that adapts to a listener's emotional state, including in therapeutic or biofeedback settings.","feed_headline":"Real-time pop composer hits emotion targets in listener test","feed_subtitle":"Rule-based system maps valence and arousal settings onto retro-pop; listener ratings match target levels.","key_machinery":"The load-bearing mechanism is a parametric mapping from Russell's valence-arousal plane onto musical features: a directed probabilistic chord graph selects chord sequences driven by a per-bar valence array; tempo follows a logarithmic relation with arousal over the range $[36,130]$ bpm; MIDI attack velocity is set by $60 + 15 \\times \\text{arousal}$; separate voice-leading transition matrices for low ($V \\le 0.5$) and high ($V > 0.5$) valence shape the melody; rhythmic patterns and note density are chosen by arousal region; and pitch-register ascent is governed by valence with a probability $p = \\text{val}$ of inversion increase. Together these rules translate any (valence, arousal) pair into a deterministic MIDI score realized with five virtual instruments (percussion, bass guitar, electric guitar, violin section, French horn), and they can be re-evaluated bar by bar to adapt the music in near real time.","core_discovery":"The paper reports that AffectMachine-Pop generates continuous, non-repetitive pop music whose emotional character is set by two input numbers, valence and arousal, and that listeners perceive roughly the intended emotion: averaged ratings across the 13 sampled points regress on the target settings with $R^2 = 0.93$ (valence) and $R^2 = 0.86$ (arousal). The authors also describe an asymmetric crossover effect: arousal settings influence perceived valence at extreme combinations, while perceived arousal is largely independent of valence settings, consistent with earlier findings in rule-based affective music systems. This is a validation of perceived emotion, not of emotion induction; the paper states that inducing emotions in listeners is left to future work.","pith_inferences":["An untested extension: because the mapping rules are explicit functions of arousal and valence, the system can be probed at any unsampled point in the valence-arousal plane without retraining; a natural next experiment is to confirm linearity on a denser grid of points.","The logarithmic tempo law implies that equal steps in arousal are more perceptually distinct at low tempos, which could let designers choose arousal quantization steps that sound perceptually uniform.","The paper validates perceived emotion only; the stronger claim for therapy is emotion induction, so a plausible next study is to measure skin conductance or self-reported mood before and after listening to test whether the target emotion transfers to the listener's own state.","The component-wise architecture suggests the crossover effect could be modeled by treating perceived valence as a function of both valence and arousal parameters, then inverting that function at the input stage to pre-compensate for the observed bias."],"forward_implications":["If the system reliably hits target valence and arousal, it can serve as an interactive tool that adapts music in real time to user input or physiological signals such as heart rate or EEG.","Because generation is rule-based and deterministic, a designer can trace why a given excerpt sounds a certain way, an auditable control that black-box generative models do not offer.","The system avoids reliance on pre-existing recordings for training, sidestepping copyright concerns while still producing stylistically coherent retro-pop.","The pop style and continuous valence-arousal traversal make it suitable for emotion-regulation applications aimed at general audiences, including middle-aged and older adults who favor 1960s-70s pop.","The empirically observed crossover between arousal settings and perceived valence provides a specific target for future versions to compensate for, improving accuracy at extreme corners of the valence-arousal plane."],"supporting_citations":[{"why":"Supplies the valence-arousal circumplex model that serves as the system's emotion control space.","marker":"(Russell 1980)"},{"why":"Describes the predecessor AffectMachine-Classical, from which this system adapts the V-A point sampling, voice-leading dissimilarity details, and overall design.","marker":"(Agres, Dash, and Chua 2023)"},{"why":"Provides the rule-based generative music approach controlled by valence and arousal, the voice-leading heuristic, and the asymmetric crossover effect the paper compares against.","marker":"(Wallis et al. 2011)"},{"why":"Demonstrates a closed-loop brain-computer interface generating MIDI events modulated by arousal and valence, a direct precursor for real-time emotion-driven generation.","marker":"(Ehrlich et al. 2019)"},{"why":"Empirically grounds the mapping of musical variables such as loudness to emotions, supporting the velocity-arousal mapping used here.","marker":"(Bresin and Friberg 2011)"},{"why":"Reviews AI-based affective music generation systems and methods, motivating the controllability and transparency requirements this system addresses.","marker":"(Dash and Agres 2024)"},{"why":"Explains the psychometric tendency for extreme scale values to receive fewer responses, used to interpret the valence rating plateau at high settings.","marker":"(Leung 2011)"}],"fun_headline_variants":["Pop music generator hits emotion targets in listener study","Turn dials for mood: real-time pop composer works as intended","Valence and arousal dials set real-time pop mood, study confirms","Two emotion numbers drive pop music, listeners verify effect","Real-time pop composer matches listener mood from two knobs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim rests on the assumption that averaging ratings from 24 listeners and three musical instantiations at each of 13 sampled points gives a reliable estimate of the emotion the system expresses at that point, so that the high $R^2$ values on these aggregated points would carry over to the rest of the continuous valence-arousal plane.","fun_headline_variants_meta":{"raw":{"variants":["Pop music generator hits emotion targets in listener study","Turn dials for mood: real-time pop composer works as intended","Valence and arousal dials set real-time pop mood, study confirms","Two emotion numbers drive pop music, listeners verify effect","Real-time pop composer matches listener mood from two knobs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0005,"raw_usage":{"total_tokens":2391,"prompt_tokens":831,"completion_tokens":1560,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":447,"completion_tokens_details":{"reasoning_tokens":1477}},"tokens_in":447,"tokens_out":1560,"duration_ms":14282,"temperature":1.0,"reasoning_tokens":1477,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:16:19.321315+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate music at a held-out point on the valence-arousal plane not among the 13 sampled (for example, valence 0.65 and arousal 0.35), have a fresh group of listeners rate it on the same 9-point Self-Assessment Manikin scales, and compare the mean normalized ratings with the target; the system's claim is falsified if the linear correspondence found at the sampled points does not extend, for instance if the mean ratings at such held-out points are systematically offset by more than the error observed in the original study.","supporting_citations":[{"cited_title":"R.; Dash, A.; and Chua, P","cited_arxiv_id":null,"evidence_quote":"Describes the predecessor AffectMachine-Classical, from which this system adapts the V-A point sampling, voice-leading dissimilarity details, and overall design."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the rule-based generative music approach controlled by valence and arousal, the voice-leading heuristic, and the asymmetric crossover effect the paper compares against."},{"cited_title":"K.; Agres, K","cited_arxiv_id":null,"evidence_quote":"Demonstrates a closed-loop brain-computer interface generating MIDI events modulated by arousal and valence, a direct precursor for real-time emotion-driven generation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Empirically grounds the mapping of musical variables such as loudness to emotions, supporting the velocity-arousal mapping used here."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Reviews AI-based affective music generation systems and methods, motivating the controllability and transparency requirements this system addresses."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Explains the psychometric tendency for extreme scale values to receive fewer responses, used to interpret the valence rating plateau at high settings."}],"review_version":1}