{"id":"43ae8bd8-8c96-4a3e-817b-ea85da33ecf7","arxiv_id":"2412.13786","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A language-model song generator adapted for multi-task editing, supporting segment-wise and track-wise modifications alongside full-song generation.","lead":"SongEditor turns a song-generation language model into an editor that can rewrite specific sections, adjust vocals or accompaniment, and compose songs from scratch. This matters because it moves AI song tools from one-shot generation to practical, controllable music production.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Track-wise editing is trained and evaluated on BS-RNN-separated stems with only a 40 dB white-noise safeguard; no separation-quality metric is reported, so independent vocal/accompaniment editing is not established.","rationale":"The reader's weakest assumption (BS-RNN separation accuracy) is also the most load-bearing for the paper's central claim. Both headline capabilities—track-wise editing and editing a segment of a single track—depend on the conditioning track being clean. The paper's only mitigation is white noise at 40 dB SNR added to separated vocals during training. White noise is not a proxy for structured leakage: BS-RNN errors are typically coherent (e.g., vocal harmonics bleeding into the accompaniment stem, drum transients in the vocal stem), and such structured leakage is exactly the kind of cue an autoregressive model can exploit to 'copy' the target from the conditioning input. Because training and evaluation both use BS-RNN, Table 3 cannot distinguish true generation from leakage reconstruction. The subjective intelligibility score is also blind to this failure mode. I considered other concerns—such as the fairness of the VoiceCraft baseline, the lack of significance testing, and the proprietary data—but those mainly affect persuasiveness and scope, whereas separation leakage goes to whether the track-wise capability exists as claimed. A clean-stem or separation-quality test would settle it. This reinforces the reader's conditional verdict rather than changing it.","tokens_in":13739,"tokens_out":7095,"duration_ms":71001,"concrete_test":"Assemble a held-out test set with ground-truth vocal and accompaniment stems (e.g., studio multitracks or clean source-separation references). For each song, run SongEditor+ track-wise editing under three input conditions: (i) clean stems, (ii) BS-RNN-separated stems, and (iii) BS-RNN stems plus the paper's σ=0.01 noise. Compute SDR/SI-SNR and FAD of the generated target stem against the ground-truth target stem, and report BS-RNN SDR on the same inputs. If clean-stem inputs substantially outperform BS-RNN inputs, or if the generated target correlates with leaked content in the conditioning stem, the current evaluation overstates track-wise editing independence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Track-wise editing—one of the two headline capabilities—is trained and evaluated entirely on BS-RNN separated stems. The only explicit defense against leakage is adding white noise (σ=0.01, ≈40 dB SNR) to separated vocals (Section 5). White noise does not model structured leakage: residual vocals in the accompaniment stem, or drum/pad bleed in the vocal stem, are correlated with the target and can be exploited by the autoregressive model. No SDR/SI-SNR or similar separation metric is reported for BS-RNN on the test set, so the degree of leakage is unknown. Because the same separator is used for both training and evaluation, any systematic leak is present on both sides; Table 3 cannot reveal whether the model is generating the target track from the conditioning track or reconstructing it from leaked cues. The intelligibility MOS checks vocal clarity and lyric match, not independence from leaked content. Thus the central claim of independent track-wise editing is not demonstrated by the current evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes SongEditor, an extension of a zero-shot song generation language model (SongLM) that adds segment-wise editing (regenerating a specific lyric sentence or section) and track-wise editing (regenerating the vocal or accompaniment track given the other). The method rearranges semantic token sequences in the style of VoiceCraft, introduces a lyric context-free editing strategy, a force-smoothing training objective, a score-based candidate selection step at inference, and a multi-source encoder plus BS-RNN source separation for track-wise editing. The authors train on roughly 700K songs (50K hours) and evaluate with PER, FAD, and MOS listening tests, reporting improvements over a SongLM baseline and over VoiceCraft on several metrics. The paper claims to be the first song editing paradigm that introduces editing capabilities into language-modeling song generation approaches.","tokens_in":13963,"tokens_out":3558,"duration_ms":32485,"significance":"If the evaluation were fully convincing, this would be a useful contribution to the relatively underexplored area of song editing, and the combination of context-free editing, force-smoothing, and candidate selection is a sensible and well-motivated design. The authors also provide a demo page with audio samples. However, the current empirical support is not sufficient for the strong claims in the abstract: the objective results are mixed, the subjective evaluation is small and lacks statistical grounding, and the track-wise editing capability depends on an unmeasured source-separation quality. The central architecture and training ideas are plausible and worth publishing after the load-bearing evaluation gaps are addressed.","major_comments":[{"comment":"Track-wise editing is trained and evaluated entirely on BS-RNN-separated stems, and the only explicit leakage safeguard is adding white noise with sigma=0.01 (approximately 40 dB SNR). White noise does not model structured leakage such as residual vocals in the accompaniment stem or drum/pad bleed in the vocal stem, which is correlated with the target track. No separation quality metric (e.g., SDR, SI-SNR) is reported for BS-RNN on the test set. Because the same separator is used in both training and evaluation, Table 3 cannot rule out that the model is partly reconstructing the target from leaked cues rather than generating it independently from the conditioning track. Please report separation metrics on the evaluation set and, if possible, include an analysis with oracle stems or with separators of varying quality to demonstrate that the editing is robust to separation error.","section":"Multi-Source Track-Wise Editing, 'Source Separation' paragraph"},{"comment":"The smoothness MOS of SongEditor (3.53) is numerically below VoiceCraft (3.63), and the force-smoothing plus candidate-selection variant (3.43) is lower still. This contradicts the paper's assertion that force-smoothing and candidate selection improve transition naturalness. The text attributes the decline to 'potential inaccuracies in the annotation of temporal boundaries,' but no evidence is provided for that explanation. Please either present a controlled analysis of boundary-alignment errors or revise the claim that the proposed smoothing techniques improve smoothness.","section":"Results and Analysis, 'Segment-Wise Editing' (Table 2)"},{"comment":"The subjective evaluation uses only 15 samples and 30 listeners, with no confidence intervals, inter-rater agreement, or significance tests reported. Several differences that are used to support the main claims are small (e.g., quality 3.39 vs. 3.17, smoothness 3.53 vs. 3.63), so the MOS results do not support the abstract's 'exceptional performance' claim. Please provide per-item confidence intervals or significance tests, and temper the wording of the claims to match the statistical strength of the evidence.","section":"Evaluation Metrics, subjective evaluation"},{"comment":"The text states that incorporating segment-wise editing capability 'does not degrade the performance' of SongLM, but Table 1 shows FAD increases from 1.99 (SongLM) to 2.24 (SongEditor). This is an objective degradation on an automatically computed metric. Please correct the statement, report statistical significance, or offer a concrete explanation for why this FAD increase is acceptable.","section":"Results and Analysis, 'Song Generation' (Table 1)"}],"minor_comments":[{"comment":"The summation in Eq. (2) is typeset incorrectly as 'KX'; it should be a sum over k=1,...,K.","section":"Base Model: SongLM, Eq. (2)"},{"comment":"The text says 'which are then assessed by expert musicians' but the sentence structure is unclear; also 'expert musicians' should be 'expert musicians'. More importantly, the allocation of the 30 listeners across the 15 samples is not described (e.g., how many ratings per sample).","section":"Experimental Settings, 'Subjective evaluation'"},{"comment":"The rows of Table 3 use '-' in the V and A columns, but the reader must infer which rows correspond to 'neither source', 'vocal only', 'accompaniment only', and 'both'. Please add explicit row labels or a legend.","section":"Results and Analysis, Table 3"},{"comment":"The number of candidates N is not specified in the main text. Please state the value used in the experiments.","section":"SongEditor, 'Score-Based Candidate Selection'"},{"comment":"The sentence 'Each model contains approximately 800M parameters' is ambiguous because SongLM/SongEditor have 16 decoder layers while SongEditor+ has 12 layers plus a multi-source encoder; the parameter counts must differ. Please report parameter counts separately.","section":"Appendix B, 'Configuration Details'"},{"comment":"The caption reads 'the below corresponds to the vocal'; this should be 'the lower one corresponds to the vocal' for grammatical correctness.","section":"Figure 4 caption"},{"comment":"The related work cites MusicMagus and VampNet for music editing, but the paper's claim of being the first song editing paradigm for language-modeling song generation could be better positioned by explicitly discussing why these prior editing methods do not address song-level vocals-plus-accompaniment editing in a language-modeling setup.","section":"Introduction, 'Related Work'"}],"recommendation":"major_revision","confidential_remarks":"The manuscript reports results on a private 700K-song dataset and does not release code or checkpoints, which makes independent verification difficult. The demo page is helpful but does not substitute for a reproducible evaluation protocol. I would encourage the editor to request that the authors release at least the evaluation set, the exact inference settings, and the listening-test materials so that the PER/FAD/MOS numbers can be independently reproduced."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"SongEditor is a reasonable systems paper for LM-based song editing. The genuinely new pieces are the context-free lyric strategy for long-content editing, the force-smoothing training and score-based candidate selection for transition quality, and the multi-source encoder that lets vocal/accompaniment tracks condition each other. The ablations support force-smoothing and candidate selection, and the PER gain over VoiceCraft (20.82 vs 35.23) is real. Segment-wise editing is a plausible, well-engineered contribution.\n\nThe soft spots are concentrated in two places. First, the track-wise editing capability—one of the two headline claims—rests on BS-RNN-separated stems without any reported separation-quality metric. The white-noise safeguard (σ=0.01) does not model structured leakage. Since the same separator is used in training and evaluation, the model might be reconstructing leaked cues rather than truly generating from the conditioning track. The stress-test note lands. The paper needs at least an SDR/SI-SNR measurement on the test set and ideally a leakage-controlled experiment before the independence claim is credible.\n\nSecond, the evaluation is thinner than the abstract implies. The subjective test uses 15 samples and 30 listeners with no confidence intervals or significance tests; the claim of “exceptional performance” is not supported. The objective metrics are mixed: FAD worsens in full-song generation (2.24 vs 1.99) and smoothness is below the VoiceCraft baseline (3.53 vs 3.63). These do not sink the paper, but they should be discussed honestly rather than glossed over.\n\nMinor issues: no code or data, and the FAD metric uses MERT-95M while the tokenizer uses MERT-330M—a mild overlap, not circular. The limitations section is candid about what the model cannot do.\n\nWho should read this: people doing music generation, editing, or singing voice synthesis. It is worth a close read for the editing-specific techniques. The paper deserves a serious peer review, but a reviewer should push for separation-quality evaluation and significance testing.","headline":"Reasonable framework paper for LM-based song editing; segment-wise effects hold up, but track-wise independence is not demonstrated because separation leakage is never measured.","tokens_in":14492,"tokens_out":2707,"would_cite":true,"duration_ms":23524,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SongEditor claims to be the first song-editing paradigm built into a language-model song generator, letting users rewrite a chosen segment or an entire vocal or accompaniment track while keeping the rest coherent.","keywords":["song editing","language model","semantic tokens","source separation","music infilling","zero-shot generation","track-wise editing","segment-wise editing"],"falsifier":"Feed a song through SongEditor's track-wise mode, then run a standard music source-separation model on the output and compare the leaked residual to the 40dB white-noise floor. If the regenerated 'vocal-only' track contains recognizable accompaniment content, or if the separated output's signal-to-distortion ratio is substantially worse than the training-time masking level, the claimed independence of track-wise editing fails.","tokens_in":13550,"feed_emoji":"🎵","tokens_out":8714,"duration_ms":72497,"temperature":0.7,"pith_summary":"SongEditor claims to be the first song-editing system built on a language-model song generator. It can rewrite a chosen span of an existing song (segment-wise editing) or regenerate an entire vocal or accompaniment track from the other track (track-wise editing), while also generating full songs from scratch. The value would be in targeted studio fixes: correcting one bad chorus or swapping a vocal without regenerating the whole piece. The authors support the claim with objective metrics and listening tests, and with a design that keeps generation quality comparable to the base generator.","feed_headline":"Song editor rewrites any segment or single track of a song","feed_subtitle":"A zero-shot language model handles both targeted fixes and full vocal or accompaniment swaps in one framework.","key_machinery":"The key mechanism is token-sequence rearrangement with a delayed pattern. The segment being edited is removed from the middle of the token sequence and appended after a separator token, making infilling a continuation problem; the delayed pattern stacks multiple residual-quantizer codebooks so the model predicts all codebooks coherently. Force-smoothing training copies a few frames from the following context as extra prediction targets, and a score-based candidate selection chooses among regenerated endings by the likelihood with which they lead into the real following audio. Track-wise editing adds a gated multi-source encoder that conditions the decoder via cross-attention, with 40dB white noise added to the separated-vocal condition during training to mask residual leakage from the separator.","core_discovery":"The paper's central claim is that editing can be made an intrinsic capability of an autoregressive song language model, not a separate task added after generation. The editing segment is cut out of the semantic-token sequence and moved to the end, so the model learns to fill the gap by continuing from the surrounding audio context; only the target lyrics are supplied for the edited region, which the paper calls a context-free strategy. For track-wise editing, vocals and accompaniment are separated with BS-RNN, the given track is tokenized and injected through a multi-source encoder with cross-attention, and the model completes the missing track. The paper reports that this single framework handles generation, segment-wise editing, track-wise editing, and iterative long-song story-mode generation, and that the removed lyric context reduces word errors at edit boundaries.","pith_inferences":["Because the tokenizer and diffusion generator are not intrinsically tied to songs, the same rearrangement-plus-source-conditioning recipe should transfer to long-form speech editing, where context preservation and speaker consistency are the same problems.","The fixed 40dB white-noise masking implies that track-wise editing quality is bounded by source-separation quality, so any future improvement in separation should carry over to editing without retraining the language model.","A testable extension would be to evaluate track-wise edits with an objective separation metric on the output, such as signal-to-distortion ratio against the original separated tracks, which would expose whether the generated track contains leakage from the condition track."],"forward_implications":["A targeted segment of a song can be regenerated with new lyrics while the rest of the song stays untouched, reducing the cost of fixing one bad verse or chorus.","Full vocal or accompaniment tracks can be resynthesized from the other track, enabling vocal swaps or instrumental re-records without retraining the model.","Segment-wise and track-wise editing can be combined, so a short portion of the vocal can be edited independently of the accompaniment.","Because the lyric context is dropped, the model can generate songs longer than its training length by iterating over sections without manual prefix annotation (Story Mode).","The same system also synthesizes complete songs from lyrics and a 10-second prompt, so generation and editing share one architecture."],"supporting_citations":[{"why":"Supplies the token-rearrangement editing method that SongEditor adapts and uses as the segment-wise editing baseline.","marker":"Peng et al. 2024"},{"why":"Provides BS-RNN, the source separator that extracts vocals and accompaniment for track-wise editing and for PER evaluation.","marker":"Luo and Yu 2023"},{"why":"Supplies the delayed pattern used to rearrange the multiple residual-quantizer token streams.","marker":"Copet et al. 2023"},{"why":"Supplies the score-based candidate selection that re-ranks regenerated endings for smoother transitions.","marker":"Jiang et al. 2023"},{"why":"Motivates adding white noise to separated vocals during training to hide source-separation leakage.","marker":"Donahue et al. 2023"},{"why":"MERT encoder provides the semantic tokens for the accompaniment branch and the features used for FAD computation.","marker":"Li et al. 2024"},{"why":"HuBERT encoder provides the semantic tokens for the vocal branch.","marker":"Hsu et al. 2021"},{"why":"Establishes the semantic-token language-modeling paradigm and the iterative story-mode generation approach SongEditor extends.","marker":"Agostinelli et al. 2023"},{"why":"SongComposer is the closest song-composition language model that edits via lyrics/MIDI but discards vocal timbre, defining the gap this work fills.","marker":"Ding et al. 2024"}],"fun_headline_variants":["Zero-shot song editor tweaks lyrics, vocals, or full tracks","One language model edits any song portion or stem","SongEditor: edit segments, tracks, or generate from scratch","Language model learns to fill any song gap context-free","Song editing unified for generation, fixes, and track swaps"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The weakest load-bearing premise is that the BS-RNN source separator leaves no significant cross-track leakage, because the track-wise editor is trained on vocals and accompaniments separated by this model, and only a fixed 40dB white-noise masking is used to hide leakage.","fun_headline_variants_meta":{"raw":{"variants":["Zero-shot song editor tweaks lyrics, vocals, or full tracks","One language model edits any song portion or stem","SongEditor: edit segments, tracks, or generate from scratch","Language model learns to fill any song gap context-free","Song editing unified for generation, fixes, and track swaps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000162,"raw_usage":{"total_tokens":1213,"prompt_tokens":894,"completion_tokens":319,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":510,"completion_tokens_details":{"reasoning_tokens":238}},"tokens_in":510,"tokens_out":319,"duration_ms":3404,"temperature":1.0,"reasoning_tokens":238,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:47:27.835459+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Feed a song through SongEditor's track-wise mode, then run a standard music source-separation model on the output and compare the leaked residual to the 40dB white-noise floor. If the regenerated 'vocal-only' track contains recognizable accompaniment content, or if the separated output's signal-to-distortion ratio is substantially worse than the training-time masking level, the claimed independence of track-wise editing fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the token-rearrangement editing method that SongEditor adapts and uses as the segment-wise editing baseline."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides BS-RNN, the source separator that extracts vocals and accompaniment for track-wise editing and for PER evaluation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the delayed pattern used to rearrange the multiple residual-quantizer token streams."}],"review_version":1}