{"id":"69e260a2-0720-40cb-9640-b7a3f0d3668a","arxiv_id":"2412.09789","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A caption-augmentation method that adds DSP-derived acoustic descriptors to text prompts gives a text-to-audio diffusion model controllable loudness, pitch, reverb, noise, brightness, fade, and duration.","lead":"SILA is a method that appends acoustic labels, such as loudness and reverb, to text captions used for training a text-to-audio model. This lets users steer generated sound effects with prompt tags like loudness: soft or reverb: very wet, and the authors report better alignment with user intent than several baselines.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claimed disentangled control is not yet established: Table II reuses the Section III-B estimators to score noise/brightness/pitch, and no counterfactual test varies one descriptor while holding all others fixed.","rationale":"The reader's verdict is CONDITIONAL with high confidence, and I agree. The strongest claim is that SILA gives fine-grained, disentangled control while preserving quality. I see independent support: a same-dataset baseline with and without SILA, a subjective test with 22 listeners, and comparable FAD. The most load-bearing gap is not the absence of significance bars per se, but that the objective evidence for noise, brightness, and pitch control is generated by the same estimators that wrote the training labels (explicitly stated in Section V), and the subjective table omits brightness. This creates a circularity for three of the seven claimed descriptors. A second, related gap is that no experiment fixes the semantic caption and all other descriptors while varying exactly one descriptor; aggregate means per category cannot establish disentanglement. The reader's weakest assumption touched the first point (label faithfulness and estimator reuse) but framed it mainly as classifier reliability; my concern extends to the missing counterfactual independence test, which is needed even if the classifiers are perfect. Both gaps are addressable with a counterfactual listening and measurement protocol, so the correct verdict stays CONDITIONAL; my pass does not move it. I would add the counterfactual experiment and an independent, human-based brightness evaluation as acceptance conditions. The paper's limitations section does not acknowledge these evaluation gaps, which strengthens the need for the conditions.","tokens_in":8699,"tokens_out":5828,"duration_ms":66218,"concrete_test":"Construct 50 base captions (e.g., \"dog bark,\" \"footsteps,\" \"engine hum\") and for each generate outputs for pairs that differ in exactly one descriptor (e.g., \"& brightness: dull\" vs \"& brightness: bright\") while all other descriptor tokens and semantic text are identical. Have listeners (n>=20) perform forced-choice on the target attribute (e.g., \"which is brighter?\") and also measure the non-target attributes to confirm they stay fixed. Compute per-pair agreement above chance and effect size. If separation is absent or non-target attributes drift, the disentanglement claim fails; if it passes, the circularity and confound concerns are substantially answered.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Section III-B, SILA's training descriptors are produced by hand-chosen estimators: SNR from Mel-spectrogram frame contrast (thresholds <=2 and >=6), brightness from spectral centroid (45/65), pitch from CREPE octave bands (1.5/3.5). Section V then evaluates \"noise, brightness, and pitch\" with \"the metrics discussed in Section III-B\" (Table II). This is circular: the objective evidence for those three descriptors verifies that the model can reproduce the estimator that wrote its labels, not that the descriptors are perceptually meaningful or disentangled. Brightness is especially exposed because it appears in Table II but has no column in the subjective study (Table III), so its entire support rests on the circular metric. Independently, the paper reports only aggregate mean values per descriptor category. There is no counterfactual condition in which one descriptor token is changed while the semantic caption and all other descriptor tokens are held fixed. The \"soft explosion\" example in Section V is not such a test: changing \"loudness: soft\" to \"loud\" also changes the expected event size/distance, so the observed difference could be a semantically sensible rendition rather than independent acoustic control. Without within-caption, single-descriptor manipulations, the central \"disentangled representations\" claim is not actually measured.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes SILA, a training-time augmentation for text-to-audio models: acoustic descriptors (loudness, pitch, reverb, noise, brightness, fade, duration) are computed from training audio and appended to the text caption as structured tokens such as '& loudness: soft'. A DiT-based text-to-audio model is trained on these augmented captions, and at inference the user writes the same descriptor format to control acoustic attributes. The method is evaluated on the AuditionSFX dataset against Stable Audio Open, AudioGen, and Tango 2, using CLAP score, FAD, objective acoustic descriptor values (Table II), and a subjective listening test with 22 participants (Table III). The paper reports the highest CLAP score, comparable FAD, and strong subjective preference for SILA.","tokens_in":8971,"tokens_out":4619,"duration_ms":49341,"significance":"If the central claim holds, SILA is a simple, model-agnostic recipe for adding fine-grained acoustic control to text-to-audio models without architectural changes, which would be useful for sound design and creative production. The paper has several concrete strengths: the audio examples are available on a project page; the subjective test covers six descriptors plus overall alignment; reverb and duration are evaluated with metrics independent of the training-label estimators; and the CLAP/FAD results indicate that control does not obviously degrade generation quality. However, the load-bearing evidence for disentangled control is currently incomplete: three of the seven descriptors are evaluated with the same estimators used to build the training labels, no single-descriptor counterfactual experiment is reported, and no statistical uncertainty is provided for the subjective or objective comparisons. These issues are fixable and do not undermine the plausibility of the method, but they must be addressed before the central claims are established.","major_comments":[{"comment":"The objective evaluation for noise, brightness, and pitch is circular. Table II reports these three rows using 'the metrics discussed in Section III-B,' i.e., the same SNR frame-contrast, spectral-centroid, and CREPE-octave estimators that generated the SILA training labels. The comparison therefore shows that the model reproduces its own labeler, not that these descriptors are perceptually meaningful or disentangled. Brightness is especially exposed because it has no column in the subjective study of Table III, so its entire support rests on this circular metric. Consequently, the only non-circular objective evidence is reverb (RT60) and duration, while loudness and fade have no objective evaluation at all.","section":"Section V, Table II; Section III-B"},{"comment":"No counterfactual single-descriptor manipulation is reported. Disentangled control requires varying one descriptor token while holding the semantic caption and all other descriptor tokens fixed; Table II instead compares aggregate per-category means over different captions. The soft/loud explosion example is not such a test, because changing loudness also changes the plausible source size and distance ('soft explosion ... in the distance' vs 'loud explosion ... very close'), which could be a semantically sensible rendition rather than independent acoustic control. Please add within-caption, single-descriptor paired comparisons for each of the seven descriptors.","section":"Section V, soft explosion example"},{"comment":"No error bars, confidence intervals, or statistical tests are reported anywhere. The subjective results come from 22 participants and 30 items, and proportions such as SILA 0.36 vs AudioGen 0.22 for duration or SILA 0.50 vs Stable Audio 0.23 for pitch need paired significance tests before preference claims are warranted. Table II reports single average values (e.g., baseline 4.31 vs SILA 6.78 for silent-background SNR), and the CLAP/FAD gaps (0.29 vs 0.27; 0.84 vs 0.81) also lack variance estimates, so the reliability of all headline comparisons is unquantified.","section":"Tables II and III"},{"comment":"The 'model-agnostic' claim is not empirically supported by the experiments. All training and evaluation use a single DiT-based text-to-audio model; the paper does not instantiate SILA with Stable Audio Open, AudioGen, Tango 2, or any other text-conditioned backbone. If the claimed contribution is that any text-conditioned model can be made controllable without architectural changes, at least one additional architecture should be tested.","section":"Abstract and Section IV-C"}],"minor_comments":[{"comment":"The loudness classes leave an unlabeled gap between -40 and -30 LKFS; please state the rule for audio files falling in that gap, and whether they are excluded from the descriptor-labeled training subset.","section":"Section III-B.1"},{"comment":"The duration descriptor is said to be 'probabilistically appended' to the metadata, but the probability is not specified; please state the value and confirm whether the same probability is used at inference.","section":"Section III-B.7"},{"comment":"The model used for caption refinement is called both 'Mistral-7B' and 'Mixtral-7B'; please make the naming consistent.","section":"Section III-A"},{"comment":"The paper defers dataset statistics to the project page; including the number of training samples, the class distribution for each descriptor, and the augmentation proportions would make the training setup self-contained and reproducible.","section":"Section IV, Training Datasets"},{"comment":"Please clarify whether each participant made one choice per category per trial or one overall choice, and how ties or incomplete responses were handled in the reported proportions.","section":"Section IV, User-Evaluation"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read. SILA's recipe is simple and genuinely useful: take an LLM-generated caption, append DSP-derived descriptor tokens (loudness, pitch, reverb, fade, brightness, noise, duration), and train any text-conditioned audio model on that. The paper does the right controlled comparison—same DiT architecture and data, with and without SILA—and the subjective test shows a clear listener preference across most categories. That preference is independent evidence and should not be discounted.\n\nWhat's new is the specific packaging: GAMA/Mistral coarse captions plus appended acoustic descriptor tags. That combination is not in the cited prior work. It is an incremental extension of caption augmentation and prompt conditioning, but a clean one. The authors also honestly list limitations (compositional audio, TTS not tested).\n\nThe soft spots are real, and the stress-test note gets the big one right. Table II evaluates noise, brightness, and pitch using the same estimators from Section III-B that wrote the training labels. That's circular for those three descriptors: it shows the model can mimic the estimator, not that the descriptors are perceptually meaningful or disentangled. Brightness is the worst case—it appears in Table II but is absent from the subjective study, so its only support is the circular metric. Loudness and fade have no objective metric; the subjective numbers carry those claims. No error bars or significance tests are reported anywhere. And the 'soft explosion' example is not a counterfactual: changing 'loudness: soft' to 'loud' also changes the implied distance and event size, so you cannot isolate loudness from semantics.\n\nThese are gaps, not refutations. The subjective result is strong, and reverb and duration have non-circular objective support (timbral RT60 and clock time). So the practical claim—appending descriptors gives users control that listeners notice—probably holds. The stronger wording about 'disentangled representations' goes beyond the evidence.\n\nWho should read it: people building or evaluating TTA systems. Practitioners get a cheap, model-agnostic control recipe. Researchers get a nice example of why controllability evaluation needs counterfactual prompts and metrics that are not the label generator.\n\nMy recommendation: send it to peer review. The recipe is reproducible, the subjective evidence is worth referee time, and the gaps are addressable. Reviewers should ask for (1) objective metrics for brightness, loudness, and fade that are not the training label estimators, (2) error bars or significance, and (3) at least one within-caption single-descriptor counterfactual. If the authors add those, the disentanglement claim would actually be tested.","headline":"SILA is a clean caption-augmentation recipe for TTA control; the subjective evidence is real, but the disentanglement claim outruns the metrics, and the objective check for three descriptors is circular.","tokens_in":9487,"tokens_out":2987,"would_cite":true,"duration_ms":29703,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that appending automatically extracted acoustic tags to text captions during training can give any text-conditioned audio generation model fine-grained control over loudness, pitch, reverb, fade, brightness, noise, and…","keywords":["text-to-audio generation","acoustic descriptors","fine-grained control","signal-to-language augmentation","diffusion models","sound effects","disentangled representations","caption augmentation"],"falsifier":"Generate the same prompt with only the loudness descriptor changed between 'very soft' and 'very loud', measure the integrated loudness of the outputs with an independent loudness meter, and have listeners rank them; if there is no consistent, perceptible level gap across many prompts, the control claim fails.","tokens_in":8517,"feed_emoji":"🎛️","tokens_out":9681,"duration_ms":96292,"temperature":0.7,"pith_summary":"SILA is a training-time recipe for text-to-audio models: before training, every sound in the dataset gets a caption, and attached to that caption are short labels describing measurable acoustic properties such as loudness band, pitch range, reverb level, fade direction, brightness, background noise, and duration. At inference the user can write or edit those same labels, and the model learns to treat them as separate control knobs rather than as part of the scene description. The paper's central claim is that this augmentation teaches the model disentangled representations of acoustic characteristics, so generated audio follows the requested property without the overall semantic alignment or perceptual quality collapsing. On their own text-to-audio diffusion transformer, SILA improves the CLAP text-audio alignment score over the no-augmentation baseline and over Stable Audio Open, AudioGen, and Tango 2, with a competitive FAD, and listeners preferred SILA on every evaluated attribute. The method is deliberately model-agnostic, so the same trick should carry over to other text-conditioned audio generators.","feed_headline":"Seven acoustic parameters become controllable via text prompts","feed_subtitle":"A training-time trick appends signal-derived tags to captions, improving semantic alignment without hurting audio quality.","key_machinery":"The machinery is the SILA caption: a normal semantic caption followed by a block of '& descriptor: value' tags, such as '& loudness: soft, & pitch: low, & reverb: very wet'. The descriptor values are produced by small signal estimators—loudness in LKFS bands, pitch in octave ranges via a neural pitch tracker, brightness via spectral centroid, noise via an SNR-style frame comparison—together with reverb and fade classes created by data augmentation and a duration label. Training a text-conditioned diffusion transformer on these concatenated strings is what forces the language conditioning to carry explicit acoustic information as separable dimensions, so that at inference a user can edit a tag and the generated audio changes only that property.","core_discovery":"On the paper's own terms, the discovery is that fine-grained acoustic control in text-to-audio generation can be achieved without new architectures, auxiliary networks, or inference-time guidance: the model just needs to see the signal-level truth during training. The authors append descriptor strings such as '& loudness: soft' and '& reverb: very wet' to captions, where the descriptors come from signal estimates, simulated reverb and fade via data augmentation, and a duration label. Because the same label appears across many different semantic events, the text encoder can separate what the sound is from how loud, how reverberant, or how bright it is. The result claimed is higher CLAP alignment than all baselines, comparable FAD, and subjective preference across loudness, pitch, reverb, noise, fade, duration, and overall alignment.","pith_inferences":["Because the descriptor vocabulary is categorical and estimate-based, SILA should be compatible with editor-style controls such as sliders, presets, or partial prompt editing, turning text prompts into a parametric surface without extra model machinery; the paper does not build such an interface.","The same training-time augmentation could extend to attributes the paper only lists as future work, such as stereo width, panning, and apparent source motion, since those are also measurable signal properties that could be phrased as descriptor tags.","An independent test of the descriptors with a second estimator or human labels would separate real perceptual control from control over artifacts of the chosen estimators; the paper's objective evaluation of noise, brightness, and pitch reuses the estimators that created the labels.","If the disentanglement story holds, editing or deleting one descriptor tag at inference should leave the other acoustic properties and the semantic content intact, which would make SILA a natural fit for prompt-to-prompt audio editing; this is a direct corollary the paper does not demonstrate."],"forward_implications":["Retraining an existing text-to-audio model on SILA-style captions should transfer the same descriptor vocabulary to new datasets, because the descriptors are computed from the audio itself rather than requiring human annotations.","A user at inference can specify 'very loud', 'very wet', 'bright', or 'fade out' in the prompt and expect the generated sound to land in the corresponding measured range, as the paper's objective results indicate.","The soft versus loud distinction is learned as a contextual difference, not a volume knob: a soft explosion sounds distant while a loud one sounds close.","The method is confined to single-event sound effects; compositional scenes and text-to-speech are explicitly left as limitations."],"supporting_citations":[{"why":"Closest diffusion-based text-to-audio baseline; SILA is compared against it on CLAP and FAD.","marker":"[1]"},{"why":"Supplies the CLAP score used as the objective alignment metric for the central claim of better text-audio alignment.","marker":"[8]"},{"why":"Supplies the FAD metric used to show that control does not come at the cost of generation quality.","marker":"[9]"},{"why":"Autoregressive text-to-audio baseline used in objective and subjective comparisons.","marker":"[13]"},{"why":"Generates the coarse semantic captions that SILA augments with acoustic descriptors.","marker":"[19]"},{"why":"Rewrites metadata and coarse captions into diverse detailed captions, providing the variety of training text SILA needs.","marker":"[20]"},{"why":"Provides the psychoacoustic basis for the loudness descriptor categories.","marker":"[21]"},{"why":"Supplies the neural pitch estimation used to create the low and high pitch descriptors.","marker":"[22]"},{"why":"Preference-tuned text-to-audio baseline used in the same objective and subjective comparisons.","marker":"[29]"}],"fun_headline_variants":["Training-time tags unlock fine acoustic control","Acoustic knobs via text: signal-to-language trick","Model-agnostic SILA enhances audio prompt control","Append signal tags to captions for better audio control","Control loudness, pitch, reverb with prompt words"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The approach assumes that the automatic labels describing each training sound—how loud, how high-pitched, how bright, how noisy—truly match what listeners perceive, since the model can only learn to control what its labels actually measure.","fun_headline_variants_meta":{"raw":{"variants":["Training-time tags unlock fine acoustic control","Acoustic knobs via text: signal-to-language trick","Model-agnostic SILA enhances audio prompt control","Append signal tags to captions for better audio control","Control loudness, pitch, reverb with prompt words"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000457,"raw_usage":{"total_tokens":2264,"prompt_tokens":887,"completion_tokens":1377,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":503,"completion_tokens_details":{"reasoning_tokens":1301}},"tokens_in":503,"tokens_out":1377,"duration_ms":10333,"temperature":1.0,"reasoning_tokens":1301,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T16:43:36.211748+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate the same prompt with only the loudness descriptor changed between 'very soft' and 'very loud', measure the integrated loudness of the outputs with an independent loudness meter, and have listeners rank them; if there is no consistent, perceptible level gap across many prompts, the control claim fails.","supporting_citations":[{"cited_title":"Large-scale contrastive language- audio pretraining with feature fusion and keyword-to-caption augmen- tation,","cited_arxiv_id":null,"evidence_quote":"Supplies the CLAP score used as the objective alignment metric for the central claim of better text-audio alignment."},{"cited_title":"GAMA: A large audio-language model with advanced audio understanding and complex reasoning abilities,","cited_arxiv_id":null,"evidence_quote":"Generates the coarse semantic captions that SILA augments with acoustic descriptors."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the psychoacoustic basis for the loudness descriptor categories."},{"cited_title":"Crepe: A convolutional representation for pitch estimation,","cited_arxiv_id":null,"evidence_quote":"Supplies the neural pitch estimation used to create the low and high pitch descriptors."}],"review_version":1}