{"id":"6b576936-e65f-40b0-a591-2e5a7d568dbf","arxiv_id":"2608.02035","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"AcoustiTrace introduces eight physics-grounded evaluation dimensions and shows that leading audio-video generators consistently fail acoustic realism even when their sound appears plausible and synchronized.","lead":"AcoustiTrace is a new benchmark that checks whether sounds generated by AI video-audio models obey basic physics, such as loudness matching motion and echoes matching rooms. It finds that even the best current generators produce plausible sound that systematically violates measurable acoustic rules.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"RT60 Consistency's visual estimator rests on sparse real-world validation and a non-scene-disjoint split; the headline RT60 failure scores could be estimator artifacts, not physics violations.","rationale":"The reader's weakest assumption correctly identifies the RT60 visual estimator as the least secure component of the benchmark. The central claim that generators consistently violate acoustic relations is supported by several evaluators, but the RT60 Consistency dimension is both a headline result and the one whose visual side is least validated. The simulated-label training, non-scene-disjoint split, and sparse real-world evidence (26 STARSS23 proxy pairs and one BRAS room) mean that the RT60 scores in Table 1 could reflect estimator bias on generated content rather than genuine physical inconsistency. The proposed test directly addresses this by retraining on a scene-disjoint split and checking whether the headline RT60 results are stable. I do not see a need to change the conditional verdict: the concern is real but scoped, and the benchmark's other dimensions and validation experiments remain credible. The intervention's circularity is a secondary issue that the paper partially acknowledges by presenting it as a proof of concept, but the RT60 validity concern is more load-bearing for the benchmark's primary claim.","tokens_in":38007,"tokens_out":5689,"duration_ms":56003,"concrete_test":"Retrain the visual RT60 estimator on a scene-disjoint split of Matterport3D (e.g., hold out 20 of the 83 scenes), then recompute the T2AV and I2AV RT60 Consistency rows of Table 1 with the same audio-side estimator and validity gates. If any model's score crosses the factor-1.5 score boundary or the cross-model ordering changes materially, the current non-scene-disjoint split is insufficient to support the reported RT60 Consistency findings.","verdict_should_be":"UNCHANGED","load_bearing_attack":"AcoustiTrace's RT60 Consistency dimension depends on a visual RT60 estimator (S2.2) trained on SoundSpaces 2.0 simulated labels from 83 Matterport3D scenes with a non-scene-disjoint train/test split. The real-world support is limited to 26 valid STARSS23 proxy pairs (S6.3) and one illustrative BRAS CR3 case (S6.4). When applied to generated videos, the estimator may be outside its training distribution; a biased visual RT60 prediction would directly distort the Table 1 RT60 Consistency scores, which are among the lowest and most attention-grabbing results (e.g., Seedance 2.0 T2AV 24.28, Wan 2.7 T2AV 33.20). The paper itself warns in S2.2 that the estimator is 'not presented as a measurement-grade estimator for arbitrary real rooms,' and the non-scene-disjoint split makes the held-out simulated MAE of 0.0642 s optimistic. This concern does not invalidate the other seven dimensions, but RT60 Consistency is a headline dimension, and if the visual side is unreliable, the claim of 'consistent violations' loses one of its strongest pieces of evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"AcoustiTrace is a diagnostic benchmark for acoustic physical realism in joint text-to-audio-video (T2AV) and image-to-audio-video (I2AV) generation. It defines eight evaluation dimensions spanning sound generation, propagation environment, and acoustic reception; constructs a dataset of 11,296 real-world audio-video clips and 82,828 RGB-D observations with simulated 500 Hz RT60 labels; builds targeted prompt suites (605 T2AV and 748 I2AV prompts); validates evaluators through real-world relation recovery, controlled perturbations, and human agreement; and evaluates nine joint audio-video generators. The authors report that plausible, synchronized sound does not guarantee faithful modeling of acoustic relations, with RT60 Consistency, Range Attenuation, and Log Attack Time being particularly challenging. A final intervention converts the diagnosed range-attenuation residual into a differentiable audio-sampling objective, improving the targeted metric in 80.16% of valid samples while largely preserving non-target quality metrics.","tokens_in":38239,"tokens_out":9524,"duration_ms":88390,"significance":"If the identified limitations are addressed, AcoustiTrace would be a valuable community benchmark: it organizes evaluation around interpretable acoustic mechanisms, explicitly gates validity, and validates evaluators with controlled perturbations, human judgments, bootstrap confidence intervals, and matched-valid analyses. These strengths are genuinely above the norm for audio-video generation benchmarks. The headline conclusion is credible for several dimensions, but the RT60 Consistency evidence is weakened by limited external validation of the visual estimator, and the intervention is more an optimization feasibility study than a demonstrated general refinement method. With a re-scoped RT60 dimension and a softened intervention claim, the paper could become a reference point for physics-aware evaluation of joint audio-video models.","major_comments":[{"comment":"The visual RT60 estimator is trained and evaluated on SoundSpaces 2.0 simulated labels from the same 83 Matterport3D scenes with a non-scene-disjoint split, and its real-world support is limited to 26 valid STARSS23 proxy pairs and one illustrative BRAS CR3 case. The paper itself states in S2.2 that the estimator is 'not presented as a measurement-grade estimator for arbitrary real rooms.' Since RT60 Consistency is one of the headline, lowest scores in Table 1 (e.g., Seedance 2.0 T2AV 24.28 and Wan 2.7 T2AV 33.20), the current evidence cannot rule out that these results partly reflect estimator distribution shift on generated videos rather than genuine acoustic violations. Please re-validate with a scene-disjoint split and substantially more measured-room evidence, or explicitly demote the RT60 dimension to an exploratory sub-diagnostic and remove it from the paper's central claims.","section":"S2.2, S6.3, S6.4, Table 1"},{"comment":"Table 1 reports conditional means over valid outputs, and RT60 Consistency has the lowest and most variable validity rates (40.5% to 82.1%). The matched-valid analysis covers only three models per task, and within that analysis RT60 is the only dimension whose ordering changes (JavisDiT++ and Ovi swap in T2AV). No matched-valid evidence is provided for Seedance 2.0 and Wan 2.7, the two models with the lowest T2AV RT60 scores. The cross-model RT60 rankings in Table 1 are therefore not robust to validity-coverage differences and should be either fully matched-validated across all reported models or excluded from the headline comparisons.","section":"S5.2, Tables S6, S10, S11"},{"comment":"The range-guided intervention optimizes a loss that is a differentiable surrogate of the Range Attenuation evaluator's own residual (log audio envelope versus log visual range). Improving mean R2 from 0.7138 to 0.8580 and winning in 80.16% of samples is therefore expected if the optimization is successful; it does not by itself demonstrate that AcoustiTrace diagnostics transfer to 'model refinement' beyond optimizing the same measurement. The non-target metrics (CLAP, LUFS, PQ) show the effect is not a global loudness adjustment, and the CLAP improvement is encouraging, but the authors should either add held-out prompts/models or human perceptual judgments, or explicitly reframe the study as a feasibility demonstration of converting a diagnostic residual into an optimization signal.","section":"S8.1, Eq. (S17), Table 2"}],"minor_comments":[{"comment":"The term 'conditional scores' is used in the main text without definition; please state explicitly that all scores are means over valid outputs, with validity rates reported separately.","section":"Table 1, S5.2"},{"comment":"The unique T2AV prompt-count formula (84+222+110+111+84-6=605) is confusing because the text lists 222 each for Approach Gain and Lateral Stability and separately lists Causality Violation; clarify that Approach and Lateral share a common receiver-motion pool and that Causality Violation reuses prompts from other pools.","section":"S4"},{"comment":"The human evaluation reports agreement between evaluator preferences and human judgments, but not inter-rater agreement; please report a chance-corrected agreement statistic such as Fleiss' kappa to show that raters agree with each other.","section":"S6.5"},{"comment":"The RT60 consistency score saturates for all ratios within a factor of 1.5, which can compress meaningful differences between generators; please state this saturation behavior and its implication for cross-model comparisons in the main text.","section":"S3.2, Eq. (S10)"},{"comment":"The paper should include an artifact availability statement in the main text, since the reproducibility of a benchmark depends on public release of code, evaluator weights, and the data manifest.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The RT60 external-validity gap and the partly by-construction intervention are the main barriers to acceptance. The other six or seven dimensions are carefully validated, and the authors are unusually transparent about limitations, which supports a constructive revision. I would not reject: the benchmark framework itself is a solid contribution, and the required fixes are re-scoping RT60 claims or adding scene-disjoint/measured validation, plus softening the intervention language."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"AcoustiTrace is worth a serious referee. It is the most complete attempt I know to turn acoustic physical realism in audio-video generation into measurable, per-output diagnostics. The eight-dimension organization around sound generation, propagation, and reception is sensible, and the authors put real work into validating the evaluators: real-world anchors, controlled perturbations, human agreement, bootstrap confidence intervals, and matched-valid robustness checks. The large RGB-D dataset with absorption maps and RT60 labels is a genuine resource, even if the labels are simulated. They are also refreshingly honest about scoping, for example the visual RT60 estimator is explicitly not measurement-grade and the train/test split is not scene-disjoint. That transparency earns credit.\n\nThe main soft spot is exactly where the stress-test note points: RT60 Consistency. The visual estimator is trained on SoundSpaces 2.0 simulated labels from 83 Matterport3D scenes with a non-scene-disjoint split, and real-world validation is thin: 26 STARSS23 proxy pairs and one BRAS case. The held-out MAE of 0.064 s is therefore optimistic, and the eye-catching low RT60 scores for Seedance and Wan could be partly estimator artifact. The authors caveat this in the supplement, but the main text leans on those numbers for the 'plausible sound violates physics' narrative. That claim still holds from the other seven dimensions, but the RT60 entry should be treated as provisional until the visual estimator gets scene-disjoint and broader real-room validation.\n\nThe other soft spot is the intervention study. Optimizing the same range-attenuation relation that the evaluator measures makes the 80.16% improvement nearly a self-fulfilling prophecy. The authors call it a proof of concept, which is fair, but it should not be read as strong evidence that AcoustiTrace can drive model refinement. The non-target quality metrics (CLAP, LUFS, PQ) are a good-faith addition.\n\nNo code or full data release is promised, which limits reproducibility and adoption. That is a concrete weakness for a benchmark paper.\n\nBottom line: the benchmark is solid, the validation is above average for this area, and the RT60 concern is real but disclosed and non-fatal. I would send this to a serious referee. I would probably cite it if I worked on A/V generation evaluation, and it is a good reading-group paper for the discussion of validation methodology.","headline":"A serious, unusually transparent benchmark for acoustic physical realism in A/V generation; the RT60 visual estimator is the main soft spot and the refinement study is only a proof of concept.","tokens_in":38775,"tokens_out":3394,"would_cite":true,"duration_ms":28526,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AcoustiTrace shows that plausible, synchronized sound in generated video does not guarantee faithful modeling of the tested acoustic relations.","keywords":["audio-video generation","acoustic physical realism","diagnostic benchmark","range attenuation","RT60 consistency","joint audio-video models","physics-guided evaluation","diffusion guidance"],"falsifier":"Measure the visual RT60 estimator's predictions on a held-out set of rooms never seen in training, using independently measured reverberation times rather than simulated labels; if the visual estimates disagree with the measured values by more than the benchmark's tolerance for many cases, the RT60-based failure rankings would not stand.","tokens_in":37811,"feed_emoji":"🔊","tokens_out":7183,"duration_ms":60889,"temperature":0.7,"pith_summary":"AcoustiTrace tests joint audio-video generators against eight acoustic relations that pair visible evidence with measurable audio quantities, spanning sound generation (motion-loudness, log attack time, impact decay), propagation (RT60 consistency, causality), and reception (range attenuation, approach gain, lateral stability). The benchmark's central claim is that generators can produce semantically plausible, well-synchronized sound while systematically violating these tested relations. Evaluated across nine generators, the paper finds consistent weak spots in Log Attack Time, RT60 Consistency, and Range Attenuation even when models score well on causality and local event relations. A proof-of-concept intervention on one diagnosed failure, range attenuation, converts the residual into a differentiable guidance signal and improves the targeted relation in 80.16% of valid samples while holding video, prompt, and model fixed.","feed_headline":"Plausible audio in generated video flouts acoustic physics","feed_subtitle":"Nine leading generators clear sync tests yet fail decay, reverb, and distance rules.","key_machinery":"The load-bearing mechanism is the relation-level diagnostic contract written as $D_k=(V_k,A_k,R_k,G_k,S_k)$. For each dimension it extracts required visual evidence, measures the matched audio quantity, states the expected acoustic relation, checks validity, and maps residual to a score. Example relations include the inverse-square law for range attenuation, Sabine's formula $T_{60}=0.161V/A$ for RT60 consistency, and exponential energy decay for impact decay. The same residual used for scoring can be turned into a differentiable guidance objective.","core_discovery":"The paper's central discovery is that perceptual plausibility and synchronization are not proxies for acoustic physical fidelity. When nine joint audio-video generators are evaluated on eight relation-level tests, no model consistently satisfies the expected relations; scores are frequently near chance for Log Attack Time, RT60 Consistency, and Range Attenuation, while Causality Violation and local event relations are generally high. The paper states this as: 'plausible and synchronized sound events do not guarantee faithful modeling of the tested acoustic relations.' It further shows that the diagnosed range-attenuation residual can be used as a differentiable guidance objective, improving the targeted relation in 80.16% of valid samples without retraining.","pith_inferences":["The paper leaves implicit that the same diagnostic contract could be reused for reward modeling in reinforcement-learning fine-tuning, because a per-output relation residual is exactly a reward signal.","If the range-attenuation result generalizes, other continuous acoustic relations such as approach gain and lateral stability may also be correctable by gradient guidance on the decoded audio envelope, without retraining.","A scene-disjoint validation of the visual RT60 estimator would separate estimator error from generator error; this is a natural next experiment the paper does not run.","The weak correlation with embedding-transition scores suggests future benchmarks should report both types of evidence, since each can miss what the other catches."],"forward_implications":["No evaluated generator dominates across all eight dimensions; leadership is split, implying the models are not uniformly physical.","Relations that are locally explicit in training clips (causality, lateral stability, motion-loudness, impact decay) score higher than relations requiring longer geometric or environmental consistency (log attack time, RT60 consistency, range attenuation).","A diagnosed residual can guide inference: range-guided audio resampling raises mean $R^2$ from 0.714 to 0.858, improves text-audio alignment, and preserves loudness, without retraining.","Embedding-transition metrics and relation-specific acoustic diagnostics are complementary: a high embedding-based score can coexist with a low attenuation score and vice versa."],"supporting_citations":[{"why":"Supplies the classical acoustic relations (inverse-square law, Sabine's formula, exponential decay) that define the expected relations.","marker":"Kinsler et al. 2000"},{"why":"Provides the energy-decay integration method used to estimate apparent RT60 from audio.","marker":"Schroeder 1965"},{"why":"Supplies SoundSpaces 2.0 simulated room impulse responses used to generate RT60 labels for the RGB-D observations.","marker":"Chen et al. 2022"},{"why":"Supplies Matterport3D RGB-D scenes used for the acoustically annotated RGB-D collection.","marker":"Chang et al. 2017"},{"why":"Supplies paired conditioning image and reference audio used for the Log Attack Time evaluation.","marker":"Owens et al. 2016"},{"why":"Provides open-vocabulary audio-visual event localization used by several evaluators.","marker":"Zhou et al. 2025"},{"why":"Provides open-vocabulary sound event detection used to localize audio events in the evaluators.","marker":"Hai et al. 2025"},{"why":"Offers the embedding-transition CPRS baseline that the paper compares against its range-attenuation diagnosis.","marker":"Xie et al. 2025"}],"fun_headline_variants":["Plausible generated audio routinely breaks acoustic physics","Sync tests pass, but acoustic physics tests fail for top generators","AcoustiTrace: leading generators fail decay, reverb, and distance tests","Audio-visual generators: plausible sound, impossible acoustics"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The RT60 Consistency dimension assumes that a visual estimator trained on simulated room labels from the same room set it is tested on accurately estimates apparent reverberation in real and generated scenes, with supporting real-world evidence limited to 26 proxy pairs and one illustrative measured-room case.","fun_headline_variants_meta":{"raw":{"variants":["Plausible generated audio routinely breaks acoustic physics","Sync tests pass, but acoustic physics tests fail for top generators","AcoustiTrace: leading generators fail decay, reverb, and distance tests","Audio-visual generators: plausible sound, impossible acoustics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000548,"raw_usage":{"total_tokens":2580,"prompt_tokens":869,"completion_tokens":1711,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":485,"completion_tokens_details":{"reasoning_tokens":1640}},"tokens_in":485,"tokens_out":1711,"duration_ms":12222,"temperature":1.0,"reasoning_tokens":1640,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:02:11.524834+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the visual RT60 estimator's predictions on a held-out set of rooms never seen in training, using independently measured reverberation times rather than simulated labels; if the visual estimates disagree with the measured values by more than the benchmark's tolerance for many cases, the RT60-based failure rankings would not stand.","supporting_citations":[{"cited_title":"and Dai, Angela and Funkhouser, Thomas and Halber, Maciej and Niessner, Matthias and Savva, Manolis and Song, Shuran and Zeng, Andy and Zhang, Yinda , booktitle =","cited_arxiv_id":null,"evidence_quote":"Supplies Matterport3D RGB-D scenes used for the acoustically annotated RGB-D collection."},{"cited_title":"and Torralba, Antonio and Adelson, Edward H","cited_arxiv_id":null,"evidence_quote":"Supplies paired conditioning image and reference audio used for the Log Attack Time evaluation."},{"cited_title":"2025 , doi =","cited_arxiv_id":null,"evidence_quote":"Provides open-vocabulary sound event detection used to localize audio events in the evaluators."}],"review_version":1}