{"id":"b6ddc1f2-3fc2-4c7d-83ec-355fc339bb2e","arxiv_id":"2608.09571","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"SonicWeave routes chunks of audio through specialized experts, using a learned gate between text-derived prior and local evidence, improving compositional fidelity in unified text-to-audio scene generation.","lead":"SonicWeave is a single audio model that generates speech, music, sound effects, singing, and mixtures from text using a new chunk-level mixture-of-experts router. A generalist might read it because it shows how to balance global scene intent and local acoustic detail in one unified generator.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Base-MoE control changes both routing granularity and router inputs, so the reported gains do not isolate chunk-level routing; a token-level CPE-MoE (C=1) control is missing.","rationale":"The reader's verdict (CONDITIONAL) is fair, and the reader's concern about missing significance is valid. I identify a more specific gap: the paper never isolates the chunking mechanism from the router design. The central claim in the conclusion explicitly names 'chunk-level prior–evidence routing' and 'intermediate granularity' as the contribution. A scientific claim about granularity requires comparing at least two granularities under an otherwise identical router. The paper's Base-MoE baseline is token-level but lacks the prior–evidence gate, so it is confounded. The C=8 ablation is coarser than C=4, which supports the claim that C=4 is better than C=8, but it does not establish that C=4 is better than C=1. In fact, the paper's own rationale for chunking—'local acoustic states are unreliable at early noisy diffusion phases'—could be handled by the conflict gate without any chunking. A C=1 variant is the direct, minimal experiment that would settle the causal role of chunk granularity. If C=1 matches C=4, the central novelty is not the chunking but the prior–evidence fusion, and the paper's framing and conclusion would need revision. If C=4 beats C=1, the granularity claim is supported. I therefore keep the reader's CONDITIONAL verdict and recommend adding this control (and the associated significance testing) before claims of intermediate granularity are treated as established.","tokens_in":21186,"tokens_out":11063,"duration_ms":103464,"concrete_test":"Train and evaluate a C=1 variant of SonicWeave that keeps every CPE-MoE component (prior logits, evidence logits, conflict gate, shared expert, top-2, load balancing) but routes each latent frame independently, using the same data, backbone, optimizer, and 800k-update budget as the reported model. Evaluate it on the same TTS/TTA/TTM and Complex-Scene protocols, including MOS-R with per-listener significance testing. If C=4 does not beat C=1 on the primary compositional metrics (Complex-Scene MOS-R and AI-Sem, plus TTS WER/CER), the 'chunk' dimension is not load-bearing and the conclusion should be revised to credit prior-evidence fusion rather than chunk granularity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central conclusion is that chunk-level prior-evidence routing is a useful intermediate granularity. The evidence for this specific claim is incomplete. In Section 4.1, Base-MoE is described as the parameter-matched control that 'isolates the effect of the routing design,' but it differs from CPE-MoE in two coupled ways: it routes individual tokens instead of chunks, and it omits the conflict-gated prior–evidence fusion. A comparison between CPE-MoE and Base-MoE therefore cannot attribute the observed gains (Tables 2–5) to chunk granularity; they could come from the prior–evidence router, the shared text/time expert, or the learned conflict gate. The ablations in Appendix E do not close this gap: the fixed-gate control (E.2) removes the learned gate but keeps C=4, and the chunk-size study (E.3) compares C=8 with C=4, both within the chunked family. There is no C=1 variant (token-level routing that retains the same prior, evidence, and conflict gate) and no clip-level (C=T) variant. Without these, the phrase 'intermediate granularity' is asserted rather than demonstrated: C=4 might outperform C=8 simply because finer routing is better, and a plain token-level router with the same prior-evidence fusion might match or beat C=4. This is an internal control gap, not a disagreement with external consensus.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes SonicWeave, a unified flow-matching model for audio scene generation, with a core CPE-MoE module that routes contiguous audio chunks through a mixture of experts using a global text-and-diffusion-phase prior, local acoustic evidence, and a learned conflict gate. The method is evaluated on public TTS, TTA, and TTM benchmarks, on a curated 100-prompt complex-scene suite with both a reference-free LLM judge and human listening, and through routing analyses. The central claim is that chunk-level prior-evidence routing offers a useful intermediate granularity between domain-level and token-level conditional computation for unified audio generation.","tokens_in":21533,"tokens_out":7021,"duration_ms":66952,"significance":"If the central claim is established, SonicWeave would be a useful contribution to unified audio generation: the chunking idea is well motivated by acoustic continuity, the structured-caption interface is practical, and the controlled comparisons against matched Dense and Base-MoE baselines on public tasks are a genuine strength. The paper is also unusually transparent about protocols in the appendix, including sampling settings, the structured-prompt conversion rules, and the routing analysis machinery. However, the specific claim of an 'intermediate granularity' is not directly tested by the reported experiments, and the evaluation lacks variance or significance information on the objective metrics. The contribution is therefore promising but currently conditional on additional controls.","major_comments":[{"comment":"The central claim that chunk-level routing is a 'useful intermediate granularity' is not tested by the reported controls. The Base-MoE baseline changes two factors at once relative to CPE-MoE: it routes individual tokens rather than chunks, and it omits the conflict-gated prior-evidence fusion. The ablations in Appendix E hold the router design fixed at C=4 and C=8, so no comparison isolates granularity. A token-level CPE-MoE variant with C=1 that retains the same prior, evidence, and conflict gate, and ideally a clip-level variant with C=T, are needed to attribute the observed gains to chunking. Without them, the improvements could come from the prior-evidence router or the learned gate, and C=4 could outperform C=8 simply because finer routing is better. The paper itself concedes 'non-exhaustive routing ablations' in the conclusion, but this is the ablation that the paper's title and conclusion depend on.","section":"§4.1, §E.3, §5"},{"comment":"The controlled comparisons are reported as point estimates from a single training run with no variance or significance information. The TTS gains over Base-MoE are 0.4, 0.3, and 0.8 percentage points in Table 2, and MOS-Q is tied at 4.57 in Table 5, so the statement that SonicWeave 'consistently improves' over Base-MoE would be much more convincing with multiple seeds, confidence intervals, or significance tests. Section C.1 describes the shared data, backbone, and training budget but does not state that the controlled models were trained with multiple seeds, nor does it document the hyperparameter search budget for each baseline, so the claim that the baselines are exactly matched controls is not fully verifiable.","section":"Tables 2–5, §C.1"},{"comment":"The Complex-Scene headline metrics, AI-Tech and AI-Sem in Table 5, come from a reference-free Gemini judge, but the paper reports no validation of that judge against the human ratings and no inter-rater agreement statistics, despite the detailed seven-dimensional rubric in Appendix D.3. The human study covers only 25 of the 100 prompts with 25 listeners, and the MOS numbers are reported as means with standard deviations across listeners only; the MOS-R difference between SonicWeave and Base-MoE is 4.49 versus 4.31 with overlapping standard deviations. Given the central role of complex-scene compositional fidelity in the conclusions, the evaluation needs either judge-human correlation and agreement statistics, or an explicit statement that the reference-free scores are exploratory.","section":"§4.3, §D.3, §D.4"},{"comment":"The comparison against external systems in the Complex-Scene suite is confounded by the prompt interface. Public systems receive natural-language prompts while SonicWeave, Dense, and Base-MoE receive structured captions; since structured conditioning is itself a contribution of the paper (Section 3.2), the MOS-R gains over Higgs Audio V2 and Dasheng AudioGen in Table 5 cannot be attributed to CPE-MoE. The conclusion's statement of improvement 'over the strongest external baseline' should be framed as a full-system comparison, not as evidence for the routing contribution, and the text should acknowledge that the external systems were not given the structured interface.","section":"§D.2, Table 5"}],"minor_comments":[{"comment":"Please report the exact parameter counts for Dense, Base-MoE, and SonicWeave; the term 'parameter-matched' is load-bearing for the controlled comparison and should be quantified.","section":"§4.1"},{"comment":"The APG guidance uses eta=0.85 for speech-only generation and eta=0.5 for scenes with sound effects or music; state whether these values were selected per task or benchmark and consider a sensitivity analysis over w and eta, since inference-time hyperparameters can affect the comparison.","section":"§C.3"},{"comment":"The chunk-pooling operation is simple mean pooling; a brief discussion of why mean pooling is preferred over max or attention pooling for the routing decision would help the reader assess the design.","section":"Equation (5)"},{"comment":"The mixed-scene routing trace is clearly labeled as illustrative, but the caption's claim that speech regions 'show stronger reliance on the text prior' should be accompanied by an uncertainty or cross-example summary rather than a single demo prompt.","section":"Figure 5, §F.4"},{"comment":"Several references carry 2026 dates and some are arXiv preprints; please verify all bibliographic details at the final submission stage, especially for Dasheng AudioGen and UniMoE-Audio, since these are moving targets.","section":"References"},{"comment":"The paper states that all reported results should use the fixed sampling protocol 'rather than selecting a schedule per example'; it would be helpful to also state explicitly that no per-benchmark selection of the guidance scale w was performed, or to list the w values used if they varied.","section":"Appendix C.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of a serious speech/audio generation venue and the overall direction is sound. My main concern is the missing C=1 control, which is required to support the 'intermediate granularity' claim; I would be willing to accept after that ablation is added and after variance or significance reporting is included. I saw no evidence of a citation or novelty disclosure problem; the issue is evidence completeness rather than correctness of the core derivation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read SonicWeave. The CPE-MoE mechanism is a real addition, but the paper's own headline claim — that chunk-level routing is the useful intermediate granularity — is not actually isolated by the experiments. The Base-MoE control differs from CPE-MoE in two ways at once: token vs chunk routing and no conflict-gated prior–evidence fusion. So the gains in Tables 2–5 over Base-MoE can't be attributed to the chunking. The authors do call the ablations non-exhaustive in the conclusion, which is honest, but the sentence about intermediate granularity is asserted rather than shown. A C=1 variant (token-level routing with the same prior, evidence, and gate) and a clip-level variant would close the gap. That's a concrete internal control issue, not a disagreement with external consensus.\n\nWhat's genuinely new: the prior–evidence router with a learned conflict gate that interpolates a global text-plus-phase prior and local chunk evidence, applied in a flow-matching DiT for unified audio generation. The structured-caption interface is sensible and carefully constrained. The controlled comparisons against matched Dense and Base-MoE, public benchmarks across TTS/TTA/TTM, the 100-prompt complex-scene suite with human listening, and the routing analyses all represent real work. The routing analysis showing content-dependent expert specialization and phase-dependent gate behavior is a good addition and supports the mechanism's plausibility.\n\nSoft spots, in proportion. First, the missing C=1 control is the main one. Second, no error bars or significance tests on the objective metrics; several Base-MoE deltas are small (TTS 0.3–0.8 points). Third, complex-scene scores rely on a reference-free Gemini judge; the human MOS subset shows the same direction, but the judge's reliability is not fully established against human scores. Fourth, no code, model, or evaluation artifacts are released, so the internal 20,000-hour corpus and the prompt-conversion pipeline make independent reproduction hard. None of these kill the paper; they just set the bar for revision.\n\nWho this is for: people working on unified audio generation or MoE routing for diffusion models. It deserves a serious referee. I'd send it out, with a request for the C=1 control, error bars, and artifacts before acceptance.","headline":"CPE-MoE is a real new routing mechanism, but the paper's central claim about chunk granularity needs the missing token-level control before it fully holds.","tokens_in":22039,"tokens_out":2678,"would_cite":true,"duration_ms":23418,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"One flow-matching model generates speech, music, singing, and sound effects together by routing contiguous audio chunks through a conflict-gated prior–evidence mixture-of-experts.","keywords":["unified audio generation","mixture-of-experts","chunk-level routing","conflict-gated routing","flow matching","text-to-audio","complex scene generation","compositional fidelity"],"falsifier":"Retrain Base-MoE and Dense with the same per-model hyperparameter sweep and several random seeds; if the reported WER, CER, FAD, KL, CLAP, and MOS-R differences collapse to within seed-to-seed noise on the same benchmarks, then the central claim that CPE-MoE routing causes the improvement is refuted.","tokens_in":21012,"feed_emoji":"🎧","tokens_out":5728,"duration_ms":48445,"temperature":0.7,"pith_summary":"SonicWeave sets out to show that a single flow-matching generator can compose speech, singing, music, sound effects, and their mixtures, and that the right conditional-computation design—routing contiguous chunks of audio through experts under a learned balance between a global text-and-phase prior and local acoustic evidence—is what makes unified generation work. The paper claims that this chunk-routed, conflict-gated mixture-of-experts (CPE-MoE) consistently beats matched dense and token-routed MoE baselines on TTS, TTA, and TTM benchmarks, and that its largest gains appear in complex scenes where components overlap or alternate. A sympathetic reader would care because unified audio generation has been blocked by the conflicting structural demands of speech versus music versus effects, and this is a concrete recipe for allocating computation flexibly within a single scene without sacrificing quality.","feed_headline":"Chunk-routed experts sharpen mixed-audio composition","feed_subtitle":"One flow-matching model blends speech, music, and SFX by letting a learned gate balance text prior against local audio evidence.","key_machinery":"The central object is CPE-MoE (Conflict-gated Prior–Evidence Mixture-of-Experts), which replaces the dense feed-forward network in the final four transformer layers. For each contiguous chunk of acoustic frames it fuses two routing signals—a global prior from the structured caption and diffusion-time embedding, and a local evidence vector pooled from the emerging audio state—through a learned conflict gate, then performs top-2 expert dispatch at chunk granularity while a shared expert handles text and time tokens. The chunk is the routing unit, not the representation unit: expert selection is shared within a chunk, but experts transform the original frame states, preserving frame-level detail while enforcing local computational continuity. A Switch-Transformer-style auxiliary loss prevents expert collapse, and the gate is trained by the generation objective rather than as a calibrated probability of evidence correctness.","core_discovery":"On its own terms, the discovery is that prior–evidence chunk routing is a useful intermediate granularity for unified audio generation. Concretely, CPE-MoE partitions audio latents into chunks of four frames, computes a global prior logit vector from the structured text summary and diffusion phase, computes a local evidence logit vector from the pooled chunk state, and lets a learned conflict gate interpolate between them: $\\ell_j = (1-g_j)\\ell_{\\text{prior}} + g_j\\ell_{\\text{evid}}$. Text and time tokens bypass the routing gate through a shared expert, and each selected expert still transforms the original frame states, so pooling controls selection without discarding frame-level detail. The paper reports content-dependent expert specialization across layers and diffusion phases, and credits this mechanism for improved compositional quality—higher semantic adherence and request realization—while keeping perceptual quality on par with the token-MoE control.","pith_inferences":["If the matched-baseline comparison is taken at face value, the same chunking and gate design could transfer to other conditional diffusion transformers, including image and video generation, wherever a global prompt and a local state coexist.","The paper's routing analyses suggest that gate values could double as a per-region confidence signal; using them to switch between prior-driven and evidence-driven generation at inference is a testable extension the paper does not pursue.","The reported gains on complex scenes, if robust, imply that evaluation of unified audio models should include compositional metrics, since single-task benchmarks understate the benefit of this routing design.","Because the paper reports no variance or significance testing, a multi-seed re-run with per-model hyperparameter tuning is the direct check on whether chunk-routed gating causes the reported improvements."],"forward_implications":["A single set of weights can generate speech, music, singing, sound effects, and mixtures from structured captions, so model count need not scale with task count.","Chunk size four emerges as the trade-off point: fine enough to react to short events and speaker turns, coarse enough to keep locally coherent audio on one expert path.","The conflict gate yields a phase-adaptive prior-to-evidence trajectory, meaning early denoising leans on the text prior and later refinement leans on acoustic evidence.","Compositional fidelity and perceptual quality are separable: the routing module improves the former while leaving the latter close to a matched token-MoE.","Structured captions that separate foreground, background, and texture provide a stable conditioning pathway that the routing prior can exploit."],"supporting_citations":[{"why":"Supplies the flow-matching objective that SonicWeave trains, regressing the linear-path velocity x1-x0.","marker":"[1]"},{"why":"Provides the strongest external complex-scene baseline for the 100-prompt suite.","marker":"[8]"},{"why":"Defines the prior domain-level audio MoE approach that CPE-MoE is designed to go beyond.","marker":"[9]"},{"why":"Supplies the auxiliary load-balancing loss computed at chunk granularity to prevent expert collapse.","marker":"[15]"},{"why":"Supplies the symmetric in-batch contrastive loss that aligns audio and text states for the prior and evidence signals.","marker":"[21]"},{"why":"Supplies the pretrained stereo VAE architecture used to encode and decode audio latents.","marker":"[27]"},{"why":"Supplies the pairwise matching features (elementwise product and absolute difference) used by the conflict gate.","marker":"[31]"},{"why":"Supplies the adaptive projected guidance used at inference to damp the parallel component of the guidance direction.","marker":"[33]"},{"why":"Supplies the TTS benchmark used for WER and CER comparisons.","marker":"[34]"},{"why":"Supplies the AudioCaps benchmark used for FAD, KL, and CLAP comparisons.","marker":"[36]"}],"fun_headline_variants":["Chunk-routed MoE blends audio scenes from text","Conflict-gated routing fuses text and audio cues","Prior-evidence chunk routing sharpens mixed audio","One model, chunk-wise routing for audio scenes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The Dense and Base-MoE controls are exactly matched to SonicWeave—same data, backbone, and training budget—so the measured gaps must be attributable to the chunk-routed gating design itself.","fun_headline_variants_meta":{"raw":{"variants":["Chunk-routed MoE blends audio scenes from text","Conflict-gated routing fuses text and audio cues","Prior-evidence chunk routing sharpens mixed audio","One model, chunk-wise routing for audio scenes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000308,"raw_usage":{"total_tokens":1806,"prompt_tokens":1036,"completion_tokens":770,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":652,"completion_tokens_details":{"reasoning_tokens":708}},"tokens_in":652,"tokens_out":770,"duration_ms":7086,"temperature":1.0,"reasoning_tokens":708,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:50:23.207777+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain Base-MoE and Dense with the same per-model hyperparameter sweep and several random seeds; if the reported WER, CER, FAD, KL, CLAP, and MOS-R differences collapse to within seed-to-seed noise on the same benchmarks, then the central claim that CPE-MoE routing causes the improvement is refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the flow-matching objective that SonicWeave trains, regressing the linear-path velocity x1-x0."},{"cited_title":"Dasheng AudioGen: A Unified Model for Generating Coherent Audio Scenes from Text","cited_arxiv_id":"2605.27838","evidence_quote":"Provides the strongest external complex-scene baseline for the 100-prompt suite."},{"cited_title":"UniMoE-Audio: Unified speech and music generation with dynamic- capacity mixture-of-experts","cited_arxiv_id":null,"evidence_quote":"Defines the prior domain-level audio MoE approach that CPE-MoE is designed to go beyond."},{"cited_title":"Learning transferable visual models from natural language supervision","cited_arxiv_id":null,"evidence_quote":"Supplies the symmetric in-batch contrastive loss that aligns audio and text states for the prior and evidence signals."},{"cited_title":"Parker, CJ Carr, Zack Zukowski, Josiah Taylor, and Jordi Pons","cited_arxiv_id":null,"evidence_quote":"Supplies the pretrained stereo VAE architecture used to encode and decode audio latents."},{"cited_title":"Eliminating oversaturation and artifacts of high guidance scales in diffusion models","cited_arxiv_id":null,"evidence_quote":"Supplies the adaptive projected guidance used at inference to damp the parallel component of the guidance direction."},{"cited_title":"AudioCaps: Generating captions for audios in the wild","cited_arxiv_id":null,"evidence_quote":"Supplies the AudioCaps benchmark used for FAD, KL, and CLAP comparisons."}],"review_version":1}