Pith. sign in

REVIEW 3 major objections 5 minor 46 references

AcoustiTrace: When Plausible Sound Violates Physics

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read AcoustiTrace shows that plausible, synchronized sound in generated video does not guarantee faithful modeling of the tested acoustic relations.

desk verdict A serious, unusually transparent benchmark for acoustic physical realism in A/V generation; the RT60 visual estimator is the main soft spot and the refinement study is only a proof of concept. read the letter →

arxiv 2608.02035 v1 pith:EM26T47E submitted 2026-08-03 cs.MM cs.SD

classification cs.MMcs.SD
keywords audio-videogenerationacousticphysicalrealismdiagnosticbenchmarkrangeattenuationRT60consistencyjointmodelsphysics-guidedevaluationdiffusionguidance
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

AcoustiTrace tests joint audio-video generators against eight acoustic relations that pair visible evidence with measurable audio quantities, spanning sound generation (motion-loudness, log attack time, impact decay), propagation (RT60 consistency, causality), and reception (range attenuation, approach gain, lateral stability). The benchmark's central claim is that generators can produce semantically plausible, well-synchronized sound while systematically violating these tested relations. Evaluated across nine generators, the paper finds consistent weak spots in Log Attack Time, RT60 Consistency, and Range Attenuation even when models score well on causality and local event relations. A proof-of-concept intervention on one diagnosed failure, range attenuation, converts the residual into a differentiable guidance signal and improves the targeted relation in 80.16% of valid samples while holding video, prompt, and model fixed.

What carries the argument

The load-bearing mechanism is the relation-level diagnostic contract written as $D_k=(V_k,A_k,R_k,G_k,S_k)$. For each dimension it extracts required visual evidence, measures the matched audio quantity, states the expected acoustic relation, checks validity, and maps residual to a score. Example relations include the inverse-square law for range attenuation, Sabine's formula $T_{60}=0.161V/A$ for RT60 consistency, and exponential energy decay for impact decay. The same residual used for scoring can be turned into a differentiable guidance objective.

What would settle it

Measure the visual RT60 estimator's predictions on a held-out set of rooms never seen in training, using independently measured reverberation times rather than simulated labels; if the visual estimates disagree with the measured values by more than the benchmark's tolerance for many cases, the RT60-based failure rankings would not stand.

Watch

Extended reading notes

Core claim

The paper's central discovery is that perceptual plausibility and synchronization are not proxies for acoustic physical fidelity. When nine joint audio-video generators are evaluated on eight relation-level tests, no model consistently satisfies the expected relations; scores are frequently near chance for Log Attack Time, RT60 Consistency, and Range Attenuation, while Causality Violation and local event relations are generally high. The paper states this as: 'plausible and synchronized sound events do not guarantee faithful modeling of the tested acoustic relations.' It further shows that the diagnosed range-attenuation residual can be used as a differentiable guidance objective, improving the targeted relation in 80.16% of valid samples without retraining.

Load-bearing premise

The RT60 Consistency dimension assumes that a visual estimator trained on simulated room labels from the same room set it is tested on accurately estimates apparent reverberation in real and generated scenes, with supporting real-world evidence limited to 26 proxy pairs and one illustrative measured-room case.

Editorial extensions

If this is right

  • No evaluated generator dominates across all eight dimensions; leadership is split, implying the models are not uniformly physical.
  • Relations that are locally explicit in training clips (causality, lateral stability, motion-loudness, impact decay) score higher than relations requiring longer geometric or environmental consistency (log attack time, RT60 consistency, range attenuation).
  • A diagnosed residual can guide inference: range-guided audio resampling raises mean $R^2$ from 0.714 to 0.858, improves text-audio alignment, and preserves loudness, without retraining.
  • Embedding-transition metrics and relation-specific acoustic diagnostics are complementary: a high embedding-based score can coexist with a low attenuation score and vice versa.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper leaves implicit that the same diagnostic contract could be reused for reward modeling in reinforcement-learning fine-tuning, because a per-output relation residual is exactly a reward signal.
  • If the range-attenuation result generalizes, other continuous acoustic relations such as approach gain and lateral stability may also be correctable by gradient guidance on the decoded audio envelope, without retraining.
  • A scene-disjoint validation of the visual RT60 estimator would separate estimator error from generator error; this is a natural next experiment the paper does not run.
  • The weak correlation with embedding-transition scores suggests future benchmarks should report both types of evidence, since each can miss what the other catches.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. AcoustiTrace is a diagnostic benchmark for acoustic physical realism in joint text-to-audio-video (T2AV) and image-to-audio-video (I2AV) generation. It defines eight evaluation dimensions spanning sound generation, propagation environment, and acoustic reception; constructs a dataset of 11,296 real-world audio-video clips and 82,828 RGB-D observations with simulated 500 Hz RT60 labels; builds targeted prompt suites (605 T2AV and 748 I2AV prompts); validates evaluators through real-world relation recovery, controlled perturbations, and human agreement; and evaluates nine joint audio-video generators. The authors report that plausible, synchronized sound does not guarantee faithful modeling of acoustic relations, with RT60 Consistency, Range Attenuation, and Log Attack Time being particularly challenging. A final intervention converts the diagnosed range-attenuation residual into a differentiable audio-sampling objective, improving the targeted metric in 80.16% of valid samples while largely preserving non-target quality metrics.

Significance. If the identified limitations are addressed, AcoustiTrace would be a valuable community benchmark: it organizes evaluation around interpretable acoustic mechanisms, explicitly gates validity, and validates evaluators with controlled perturbations, human judgments, bootstrap confidence intervals, and matched-valid analyses. These strengths are genuinely above the norm for audio-video generation benchmarks. The headline conclusion is credible for several dimensions, but the RT60 Consistency evidence is weakened by limited external validation of the visual estimator, and the intervention is more an optimization feasibility study than a demonstrated general refinement method. With a re-scoped RT60 dimension and a softened intervention claim, the paper could become a reference point for physics-aware evaluation of joint audio-video models.

major comments (3)
  1. [S2.2, S6.3, S6.4, Table 1] The visual RT60 estimator is trained and evaluated on SoundSpaces 2.0 simulated labels from the same 83 Matterport3D scenes with a non-scene-disjoint split, and its real-world support is limited to 26 valid STARSS23 proxy pairs and one illustrative BRAS CR3 case. The paper itself states in S2.2 that the estimator is 'not presented as a measurement-grade estimator for arbitrary real rooms.' Since RT60 Consistency is one of the headline, lowest scores in Table 1 (e.g., Seedance 2.0 T2AV 24.28 and Wan 2.7 T2AV 33.20), the current evidence cannot rule out that these results partly reflect estimator distribution shift on generated videos rather than genuine acoustic violations. Please re-validate with a scene-disjoint split and substantially more measured-room evidence, or explicitly demote the RT60 dimension to an exploratory sub-diagnostic and remove it from the paper's central claims.
  2. [S5.2, Tables S6, S10, S11] Table 1 reports conditional means over valid outputs, and RT60 Consistency has the lowest and most variable validity rates (40.5% to 82.1%). The matched-valid analysis covers only three models per task, and within that analysis RT60 is the only dimension whose ordering changes (JavisDiT++ and Ovi swap in T2AV). No matched-valid evidence is provided for Seedance 2.0 and Wan 2.7, the two models with the lowest T2AV RT60 scores. The cross-model RT60 rankings in Table 1 are therefore not robust to validity-coverage differences and should be either fully matched-validated across all reported models or excluded from the headline comparisons.
  3. [S8.1, Eq. (S17), Table 2] The range-guided intervention optimizes a loss that is a differentiable surrogate of the Range Attenuation evaluator's own residual (log audio envelope versus log visual range). Improving mean R2 from 0.7138 to 0.8580 and winning in 80.16% of samples is therefore expected if the optimization is successful; it does not by itself demonstrate that AcoustiTrace diagnostics transfer to 'model refinement' beyond optimizing the same measurement. The non-target metrics (CLAP, LUFS, PQ) show the effect is not a global loudness adjustment, and the CLAP improvement is encouraging, but the authors should either add held-out prompts/models or human perceptual judgments, or explicitly reframe the study as a feasibility demonstration of converting a diagnostic residual into an optimization signal.
minor comments (5)
  1. [Table 1, S5.2] The term 'conditional scores' is used in the main text without definition; please state explicitly that all scores are means over valid outputs, with validity rates reported separately.
  2. [S4] The unique T2AV prompt-count formula (84+222+110+111+84-6=605) is confusing because the text lists 222 each for Approach Gain and Lateral Stability and separately lists Causality Violation; clarify that Approach and Lateral share a common receiver-motion pool and that Causality Violation reuses prompts from other pools.
  3. [S6.5] The human evaluation reports agreement between evaluator preferences and human judgments, but not inter-rater agreement; please report a chance-corrected agreement statistic such as Fleiss' kappa to show that raters agree with each other.
  4. [S3.2, Eq. (S10)] The RT60 consistency score saturates for all ratios within a factor of 1.5, which can compress meaningful differences between generators; please state this saturation behavior and its implication for cross-model comparisons in the main text.
  5. [General] The paper should include an artifact availability statement in the main text, since the reproducibility of a benchmark depends on public release of code, evaluator weights, and the data manifest.

Circularity Check

1 steps flagged · score 4.0 of 10

Benchmark evaluation is not circular; only the range-guidance intervention improves its target by construction.

  1. self definitional [Main paper, 'Range-Guided Audio Sampling'; Supplementary Eq. S17 and S8.1]
    "We target the identified Range Attenuation failure and formulate the expected attenuation relation used by the evaluator as a differentiable guidance objective for audio sampling. ... Distance guidance increases the mean Range Attenuation R2 from 0.7138 to 0.8580 and outperforms unguided resampling in 80.16% of cases."

    The guidance loss in Eq. S17 is a log-ratio envelope objective with gamma=1, i.e., log(e_i/e_j) = log(d_j/d_i), which is exactly the inverse-distance amplitude relation underlying the Range Attenuation evaluator (Eq. S12: Delta L* = -20 log10(d(t)/d(t0))). Optimizing this objective on the fixed video directly maximizes the evaluator's R2, so the reported 80.16% improvement is a manipulation check rather than an independent discovery. The paper's non-target metrics (CLAP, LUFS, PQ) are independent and show the intervention is not wholly vacuous, but the headline 'improvement in the targeted relation' reduces by construction to optimizing the measured relation.

full rationale

AcoustiTrace's benchmark evaluation is not circular: each of the eight dimensions tests an externally specified acoustic relation (inverse-square attenuation, Sabine reverberation, exponential decay, causality) against independently measured visual and audio evidence, and the evaluators are validated with controlled perturbations and human agreement. The RT60 visual-estimator limitation, including the non-scene-disjoint train/test split and sparse real-world proxy validation, is a validity and robustness concern rather than a circularity: the estimator is honestly scoped and not used to define the physical relation. No load-bearing self-citation chain is present. The only construction-redundant element is the intervention study, where Eq. S17 optimizes the same log-ratio range-loudness relation that the Range Attenuation evaluator scores, making the reported R2 and 80.16% win expected; the independent CLAP, LUFS, and Audiobox PQ metrics provide partial independent content. Overall, the central benchmark derivation is self-contained, so the score is moderate rather than severe.

Assumptions & free parameters 6 free parameters · 7 assumptions · 0 invented entities

The paper introduces no new physical particles, forces, or conservation laws. Its invented artifacts are evaluation constructs and data representations, chiefly the Acoustic Alpha Map and the eight diagnostic dimensions. The free parameters listed above are hand-set evaluator thresholds and score mappings that the benchmark's rankings depend on; they are disclosed but not derived from first principles.

free parameters (6)
  • RT60 consistency tolerance factor = score max at ratio 1.5, zero at ratio 3.0
    Hand-set score mapping in Equation S10 determines which audio-visual RT60 ratios count as consistent; changing these thresholds would change model rankings on RT60 Consistency.
  • Log Attack Time exponential scale = 0.35 in Equation S3
    Hand-set normalization width converts log attack time difference into a 0-100 score; it sets the severity scale for the I2AV onset dimension.
  • Causality violation scoring margin = 1 ms
    Hand-set strict precedence threshold in Equation S11; the authors note it is a timestamp increment rather than event-localization accuracy, but it determines violation counts.
  • Impact decay fit window and tail penalty = fit from 0.02 to 0.50 s after peak, tail residual normalized by dynamic range
    Hand-set decay window and residual penalty in Equation S4 define what counts as a plausible exponential decay shape.
  • Range attenuation sliding-window configuration = 0.40 s windows, 0.05 s stride; minimum 1.50 s dominant segment; relative-depth thresholds
    These hand-set thresholds gate validity and determine which range trajectories are scored; the authors state they do not imply metric distances.
  • Range-guided sampling gamma exponent = gamma = 1 (inverse-distance amplitude attenuation)
    Chosen target attenuation exponent for the intervention loss in Equation S17; calibration of the decoded-mel proxy recovered rhat = 0.997, but gamma is still a modeling choice rather than a derived constant.
assumptions (7)
  • domain assumption Ideal free-field inverse-square law applies to range attenuation with approximately constant source power, stable directivity, and negligible reflections.
    Used in the Range Attenuation and Approach Gain evaluators and Equation S12; the authors flag it as a modeling assumption in Section S3.3, but departures from it affect scores.
  • domain assumption Sabine's diffuse-field relation T60 = 0.161 V / A approximates real room reverberation.
    Guides the visual RT60 estimator and the RGB-D dataset labels; assumes a diffuse sound field and roughly uniform absorption across surfaces.
  • domain assumption Simulated SoundSpaces 2.0 room impulse responses at 500 Hz provide valid RT60 supervision for real and generated audio.
    The 82,828 RT60 labels are simulated rather than measured in situ; real-room validation is limited to 26 STARSS23 proxy pairs and one illustrative BRAS case.
  • domain assumption Semantic-material matching of PTB absorption coefficients to RGB-D segmentation yields surface absorption maps accurate enough for visual RT60 estimation.
    The Acoustic Alpha Map is a proxy, not an in-situ absorption measurement; segmentation errors and material-retrieval mismatches propagate into visual RT60 predictions.
  • domain assumption Apparent RT60 estimated from generated audio via a Schroeder decay curve is comparable to the room RT60 implied by visual cues.
    Used by RT60 Consistency; the authors distinguish the two concepts but still compare them through Equation S10.
  • domain assumption The decoded-mel envelope extracted from the Audio VAE is an amplitude-like proxy whose logarithm scales linearly with distance attenuation.
    Range-guided sampling uses Equations S18 and S19; calibration on 40 held-out clips supports the proxy, but the support is itself a fitted validation.
  • domain assumption Two-reviewer screening and automated MLLM/detector pipelines reliably identify when a target acoustic relation is observable and measurable.
    Dataset construction and evaluator validation rely on human and automated screening to confirm that the relation is present in real-world anchors.

how reviews work

0 comments
Cite this review

Pith. "Pith review of AcoustiTrace: When Plausible Sound Violates Physics." pith.science (2026). https://pith.science/paper/EM26T47E

@misc{pith2026260802035,
  author       = {Pith},
  title        = {Pith review of: AcoustiTrace: When Plausible Sound Violates Physics},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EM26T47E}},
  note         = {Machine review of arXiv:2608.02035}
}
read the original abstract

Recent audio-video generators can produce semantically plausible and apparently synchronized sound, yet may still violate the acoustic processes implied by visible events and environments. Existing benchmarks provide limited support for attributing such violations to particular acoustic processes and quantifying their severity. We introduce AcoustiTrace, a diagnostic benchmark that formalizes acoustic physical realism in audio-video generation. AcoustiTrace organizes text-to-audio-video (T2AV) and image-to-audio-video (I2AV) evaluation around the acoustic process, covering sound generation, propagation environment, and acoustic reception through eight dimensions grounded in measurable acoustic quantities. Based on these evaluation dimensions, we construct a large-scale dataset organized around acoustic mechanisms, comprising real-world audio-video recordings and acoustically annotated RGB-D observations, and use it to develop targeted prompt suites and validated evaluators. Experiments reveal that even leading generators still struggle with fundamental acoustic processes despite producing plausible sound events. Finally, we show that the diagnostics AcoustiTrace provides for specific acoustic relations can guide model refinement toward more physically faithful audio and open new directions for incorporating acoustic principles into training objectives, reward modeling, and candidate selection.

Figures

Figures reproduced from arXiv: 2608.02035 by the authors.

Figure 1
Figure 1. Representative cases where plausible sound vio [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Framework of AcoustiTrace. The benchmark organizes acoustic physical realism into sound generation, propagation [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Distribution of the AcoustiTrace prompt suites. The [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Construction of the AcoustiTrace dataset from real [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: From diagnosis to intervention. (a) The Range Attenuation residual from a fixed visual trajectory guides decoded-mel audio denoising. (b) A qualitative receding-source example visualizes how guidance reshapes the audio envelope toward the expected attenuation trend. (c…
Figure 6
Figure 6. Figure 6: (a) Responses of the eight evaluators to controlled audio perturbations that selectively violate their targeted acoustic relations. The unperturbed point in each sweep corresponds to the original, unmodified real-world A/V samples. For RT60 Consistency, perturbation se…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

46 extracted references · 34 canonical work pages

  1. [1]

    A Benchmark for Room Acoustical Simulation: Concept and Database , journal =

    Brinkmann, Fabian and Asp. A Benchmark for Room Acoustical Simulation: Concept and Database , journal =. 2021 , doi =

  2. [2]

    2020 , howpublished =

    Asp. 2020 , howpublished =. doi:10.14279/depositonce-6726.3 , note =

  3. [3]

    2026 , eprint =

    Do Joint Audio-Video Generation Models Understand Physics? , author =. 2026 , eprint =

  4. [4]

    2023 IEEE International Conference on Acoustics, Speech and Signal Processing , pages =

    Large-Scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation , author =. 2023 IEEE International Conference on Acoustics, Speech and Signal Processing , pages =. 2023 , doi =

  5. [5]

    Girdhar, Rohit and El-Nouby, Alaaeldin and Liu, Zhuang and Singh, Mannat and Alwala, Kalyan Vasudev and Joulin, Armand and Misra, Ishan , booktitle =

  6. [6]

    2024 , doi =

    Iashin, Vladimir and Xie, Weidi and Rahtu, Esa and Zisserman, Andrew , booktitle =. 2024 , doi =

  7. [7]

    2026 , doi =

    Zhang, Yiming and Gu, Yicheng and Zeng, Yanhong and Xing, Zhening and Wang, Yuancheng and Wu, Zhizheng and Liu, Bin and Chen, Kai , journal =. 2026 , doi =

  8. [8]

    Proceedings of the AAAI Conference on Artificial Intelligence , volume =

    Read, Watch and Scream! Sound Generation from Text and Video , author =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2025 , doi =

Show all 46 references
  1. [9]

    Proceedings of the AAAI Conference on Artificial Intelligence , volume =

    Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model Adaptation , author =. Proceedings of the AAAI Conference on Artificial Intelligence , volume =. 2024 , doi =

  2. [10]

    Cheng, Ho Kei and Ishii, Masato and Hayakawa, Akio and Shibuya, Takashi and Schwing, Alexander and Mitsufuji, Yuki , booktitle =

  3. [11]

    2508.16930 , archivePrefix =

    Shan, Sizhe and Li, Qiulin and Cui, Yutao and Yang, Miles and Wang, Yuehai and Yang, Qun and Zhou, Jin and Zhong, Zhao , year =. 2508.16930 , archivePrefix =

  4. [12]

    2506.21448 , archivePrefix =

    Liu, Huadai and Luo, Kaicheng and Wang, Jialei and Wang, Wen and Chen, Qian and Zhao, Zhou and Xue, Wei , year =. 2506.21448 , archivePrefix =

  5. [13]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =

    Seeing and Hearing: Open-Domain Visual-Audio Generation with Diffusion Latent Aligners , author =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =

  6. [14]

    2406.07686 , archivePrefix =

    Wang, Kai and Deng, Shijian and Shi, Jing and Hatzinakos, Dimitrios and Tian, Yapeng , year =. 2406.07686 , archivePrefix =

  7. [15]

    and Shi, Yangyang and Chandra, Vikas , year =

    Liu, Haohe and Le Lan, Gael and Mei, Xinhao and Ni, Zhaoheng and Kumar, Anurag and Nagaraja, Varun and Wang, Wenwu and Plumbley, Mark D. and Shi, Yangyang and Chandra, Vikas , year =. 2412.15220 , archivePrefix =

  8. [16]

    2026 , url =

    Liu, Kai and Li, Wei and Chen, Lai and Wu, Shengqiong and Zheng, Yanhao and Ji, Jiayi and Zhou, Fan and Luo, Jiebo and Liu, Ziwei and Fei, Hao and Chua, Tat-Seng , booktitle =. 2026 , url =

  9. [17]

    2509.06155 , archivePrefix =

    Wang, Duomin and Zuo, Wei and Li, Aojie and Chen, Ling-Hao and Liao, Xinyao and Zhou, Deyu and Yin, Zixin and Dai, Xili and Jiang, Daxin and Yu, Gang , year =. 2509.06155 , archivePrefix =

  10. [18]

    2510.01284 , archivePrefix =

    Low, Chetwin and Wang, Weimin and Katyal, Calder , year =. 2510.01284 , archivePrefix =

  11. [19]

    2601.03233 , archivePrefix =

    HaCohen, Yoav and Brazowski, Benny and Chiprut, Nisan and Bitterman, Yaki and Kvochko, Andrew and Berkowitz, Avishai and Shalem, Daniel and Lifschitz, Daphna and Moshe, Dudu and Porat, Eitan and Richardson, Eitan and Shiran, Guy and Chachy, Itay and Chetboun, Jonathan and Fink...

  12. [20]

    2026 , eprint =

    Native Audio-Visual Alignment for Generation , author =. 2026 , eprint =

  13. [21]

    2026 , url =

    Liu, Kai and Zheng, Yanhao and Wang, Kai and Wu, Shengqiong and Zhang, Rongjunchen and Luo, Jiebo and Hatzinakos, Dimitrios and Liu, Ziwei and Fei, Hao and Chua, Tat-Seng , booktitle =. 2026 , url =

  14. [22]

    Huang, Ziqi and He, Yinan and Yu, Jiashuo and Zhang, Fan and Si, Chenyang and Jiang, Yuming and Zhang, Yuanhan and Wu, Tianxing and Jin, Qingyang and Chanpaisit, Nattapol and Wang, Yaohui and Chen, Xinyuan and Wang, Limin and Lin, Dahua and Qiao, Yu and Liu, Ziwei , booktitle =

  15. [23]

    2024 , doi =

    Mao, Yuxin and Shen, Xuyang and Zhang, Jing and Qin, Zhen and Zhou, Jinxing and Xiang, Mochu and Zhong, Yiran and Dai, Yuchao , booktitle =. 2024 , doi =

  16. [24]

    2026 , doi =

    Shimada, Kazuki and Simon, Christian and Shibuya, Takashi and Takahashi, Shusuke and Mitsufuji, Yuki , booktitle =. 2026 , doi =

  17. [25]

    Hua, Daili and Wang, Xizhi and Zeng, Bohan and Huang, Xinyi and Liang, Hao and Niu, Junbo and Chen, Xinlong and Xu, Quanqing and Zhang, Wentao , booktitle =

  18. [26]

    2512.21094 , archivePrefix =

    Cao, Zhe and Wang, Tao and Wang, Jiaming and Wang, Yanghai and Zhang, Yuanxing and Wang, Jiahao and Chen, Jialu and Deng, Miao and Guo, Yubin and Liao, Chenxi and Zhang, Yize and Zhang, Zhaoxiang and Liu, Jiaheng , year =. 2512.21094 , archivePrefix =

  19. [27]

    2512.23994 , archivePrefix =

    Xie, Tianxin and Lei, Wentao and Jiang, Kai and Huang, Guanjie and Zhang, Pengfei and Zhang, Chunhui and Ma, Fengji and He, Haoyu and Zhang, Han and He, Jiangshan and Wang, Jinting and Fang, Linghan and Gao, Lufei and Ablet, Orkesh and Zhang, Peihua and Hu, Ruolin and Li, Shen...

  20. [28]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =

    Benchmarking Single-Factor Physical Video-to-Audio Generation , author =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =

  21. [29]

    Oh, Hyun-Bin and Takida, Yuhta and Uesaka, Toshimitsu and Oh, Tae-Hyun and Mitsufuji, Yuki , booktitle =

  22. [30]

    The Journal of the Acoustical Society of America , volume =

    The Timbre Toolbox: Extracting Audio Descriptors from Musical Signals , author =. The Journal of the Acoustical Society of America , volume =. 2011 , doi =

  23. [31]

    The Journal of the Acoustical Society of America , volume =

    New Method of Measuring Reverberation Time , author =. The Journal of the Acoustical Society of America , volume =. 1965 , doi =

  24. [32]

    2000 , isbn =

    Fundamentals of Acoustics , author =. 2000 , isbn =

  25. [33]

    and Dai, Angela and Funkhouser, Thomas and Halber, Maciej and Niessner, Matthias and Savva, Manolis and Song, Shuran and Zeng, Andy and Zhang, Yinda , booktitle =

    Chang, Angel X. and Dai, Angela and Funkhouser, Thomas and Halber, Maciej and Niessner, Matthias and Savva, Manolis and Song, Shuran and Zeng, Andy and Zhang, Yinda , booktitle =. 2017 , doi =

  26. [34]

    Chen, Changan and Schissler, Carl and Garg, Sanchit and Kobernik, Philip and Clegg, Alexander and Calamia, Paul and Batra, Dhruv and Robinson, Philip and Grauman, Kristen , booktitle =

  27. [35]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =

    Real Acoustic Fields: An Audio-Visual Room Acoustics Dataset and Benchmark , author =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =

  28. [36]

    The Room Acoustics Absorption Coefficient Database , howpublished =. n.d. , note =

  29. [37]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =

    Towards Open-Vocabulary Audio-Visual Event Localization , author =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages =

  30. [38]

    2025 , doi =

    Hai, Jiarui and Wang, Helin and Guo, Weizhe and Elhilali, Mounya , booktitle =. 2025 , doi =

  31. [39]

    2026 , howpublished =

  32. [40]

    2025 , howpublished =

  33. [41]

    2026 , month = apr, howpublished =

    Alibaba Unveils. 2026 , month = apr, howpublished =

  34. [42]

    2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages =

    Elizalde, Benjamin and Deshmukh, Soham and Al Ismail, Mahmoud and Wang, Huaming , title =. 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages =. 2023 , publisher =

  35. [43]

    2023 , month = nov, note =

    Algorithms to Measure Audio Programme Loudness and True-Peak Audio Level , institution =. 2023 , month = nov, note =

  36. [44]

    2023 , month = nov, note =

    Loudness Normalisation and Permitted Maximum Level of Audio Signals , institution =. 2023 , month = nov, note =

  37. [45]

    2025 , eprint =

    Tjandra, Andros and Wu, Yi-Chiao and Guo, Baishan and Hoffman, John and Ellis, Brian and Vyas, Apoorv and Shi, Bowen and Chen, Sanyuan and Le, Matt and Zacharov, Nick and Wood, Carleigh and Lee, Ann and Hsu, Wei-Ning , title =. 2025 , eprint =

  38. [46]

    and Torralba, Antonio and Adelson, Edward H

    Owens, Andrew and Isola, Phillip and McDermott, Josh H. and Torralba, Antonio and Adelson, Edward H. and Freeman, William T. , title =. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages =. 2016 , doi =

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.