REVIEW 4 major objections 5 minor 93 references
A Survey on World Models Grounded in Acoustic Physical Information
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This survey argues that acoustic signals—the radiated mechanical energy of physical events—should be a primary modality for AI world models, enabling causal physical understanding and predictive simulation through sound.
desk verdict Useful framing for a survey, but the citation mismatches and a dimensionally wrong KZK term need fixing before it can be trusted as a field map. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by a chain of physics-to-signal encodings. The elastodynamic wave equation $$(\$\lambda$+\mu)\nabla(\nabla\cdot u)+\mu\$nabla^{2}$ u=\rho\,\$partial^{2}$ u/\partial $t^{2}$$$ with modal analysis links an object's natural frequencies $\omega_n$ to its material and geometry; the acoustic analogy recasts the Navier-Stokes equations into an inhomogeneous wave equation for flow-generated sound; the room impulse response $h(t)$, with standard reverberation-time estimates $T_{60}$, encodes a space's geometry and absorption; and the Westervelt and KZK equations describe nonlinear high-amplitude propagation. On the machine-learning side, the composite PINN loss $$\mathcal{L}(\$\theta$)=w_{\mathrm{data}}\mathcal{L}_{\mathrm{data}}+w_{\mathrm{phys}}\mathcal{L}_{\mathrm{phys}}+w_{\mathrm{bc}}\mathcal{L}_{\mathrm{bc}}+w_{\mathrm{ic}}\mathcal{L}_{\mathrm{ic}}$$ is the gray-box device that enforces these partial differential equations as regularizers, while differentiable simulators, neural acoustic fields, and contrastive audio-visual models supply forward prediction and scalable representation learning. Each element does a specific job: physics supplies the acoustic fingerprint, physics-informed networks invert it with sparse data, generative models predict and synthesize, and self-supervised learning scales the representations.
What would settle it
A concrete falsifier: train a model to infer material and shape from impact sounds, then evaluate it on objects whose geometry, excitation type, and recording conditions are randomized and never seen during training; if accuracy falls to chance, the modal-inversion route to physical properties is not doing the work claimed. A second check: on a held-out physical-prediction task, if an acoustic-only model adds no predictive power over a vision-only model, the claimed complementarity of sound for physical intuition is unsupported.
Extended reading notes
Core claim
The paper's central claim is that the physical laws governing sound—elastodynamics, aeroacoustics, room acoustics, and nonlinear wave propagation—leave measurable acoustic signatures that can be inverted for physical knowledge and used forward for prediction. Vibrations of struck objects encode material and modal properties such as elastic modulus, density, and damping; contact and flow sounds encode collision, friction, and fluid dynamics; room impulse responses encode geometry and surface absorption; and nonlinear equations such as the Westervelt and KZK equations extend the picture to high-intensity fields. The survey's thesis is that these signatures make audio complementary to vision precisely where vision is weak: seeing surfaces but not masses, internal defects, contact forces, or occluded structure. An AI that learns these mappings can simulate events in latent space and reason about physical interventions, which the paper calls an internal 'intuitive physics' engine through sound.
Load-bearing premise
The load-bearing premise is that acoustic signals carry enough latent physical information—and the cited methods can extract enough of it—that sound alone can ground causal physical understanding; the survey's evidence base also has to be a fair sample of the field for its roadmap to be trustworthy.
Editorial extensions
If this is right
- If the acoustic world-model program works, robots can localize and map in darkness, smoke, or dust where vision fails, using Acoustic SLAM built from echoes and passive sound sources.
- Contact and friction sounds give robotic manipulators haptic-like feedback, letting them detect incipient slip, adjust grip, and infer hidden object properties such as fill level by shaking a container.
- Vehicles gain a 360-degree safety layer: siren detection beyond line of sight and real-time road-surface classification from tire-road noise, complementing vision and LiDAR where those sensors are ambiguous.
- Non-invasive health screening can scale through cough, breathing, and voice analysis on ubiquitous microphones, with vocal biomarkers tracking respiratory, neurological, and mental-health conditions.
- Financial applications can extract non-consensus signals from paralinguistic cues in earnings calls and use acoustic stress markers for fraud detection and compliance.
Reading between the lines
- Editorial inference: if audio truly encodes physical properties as the survey argues, sound can serve as a cheap self-supervision signal for vision-based models, correcting visual errors about mass, material, and internal state without new hardware.
- Editorial inference: the survey's own hybrid scenario—self-supervised pretraining, generative synthesis, and PINN regularization—can be tested directly by comparing that combined architecture against each pillar alone on a physical audio prediction benchmark.
- Editorial note grounded in the paper's stated limitations: the authors concede that sensor design and classical array processing are treated only as background, so the deployment roadmap is stronger on algorithms than on the hardware that would make those algorithms reliable in the field.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a survey of 'world models grounded in acoustic physical information.' Its central thesis is that acoustic signals, as direct carriers of mechanical wave energy from physical events, encode latent information about material properties, internal structure, geometry, and interaction dynamics, and that this makes sound a privileged modality for AI systems that must build causal, physics-grounded world models. The paper reviews the physical theory (elastodynamics, aeroacoustics, room acoustics, nonlinear acoustics), organizes methodological work into three pillars (physics-informed neural networks, generative forward models, and self-supervised multimodal learning), surveys applications in robotics, autonomous driving, healthcare, and finance, and closes with an ethics discussion and a five-pathway research roadmap.
Significance. If its map of the literature were reliable, this survey would provide a useful synthesis of an emerging interdisciplinary area and a concrete agenda for future work. The paper is clearly organized, and the comparative analysis in Table 1, the application taxonomy in Table 2, the explicit limitations statement, and the resource list in Appendix A are useful editorial features. However, the survey's central claim depends on the cited literature substantiating specific acoustic-inference capabilities, and many load-bearing citations do not support the claims they are attached to (Sections 2.1, 2.2, 3.2, and 4). In addition, Eq. (9) contains a dimensional inconsistency in the KZK nonlinear term. Because a survey's contribution is precisely its map of the evidence base, these problems materially weaken the paper. They are, however, correctable within the manuscript's scope.
major comments (4)
- [2.1] The sentence 'Recent works have demonstrated that deep learning models can leverage these rich acoustic signatures to classify object materials with high accuracy [16], estimate the viscosity of liquids ... [17], or even infer the particle size of granular materials from the statistical properties of many small impacts during shaking [18]' is not supported by the cited papers: [16] is a review of ultrasonic non-destructive evaluation of composites, [17] is a study of porosity evaluation of additively manufactured parts using ultrasonic NDT, and [18] is a study of acoustic-emission damage classification in composites. None of these works reports impact-sound material classification, liquid-viscosity estimation, or granular particle-size inference. The same paragraph's earlier claim that material identification from sound rests on a 'large body of work' also cites [13], which is a study of natural reverberation statistics, alongside the relevant [14]. This passage is load-bearing for the survey's central thesis that acoustic signals encode material and fluid properties, so it should be rewritten with correct references or the claims should be removed.
- [2.2] The cyclostationary-analysis paragraph states that deep learning models trained on spectral correlation and envelope analysis 'can be trained on these processed representations to perform highly sensitive fault diagnosis and prognostics, effectively creating a detailed acoustic physical model of a machine's health state [34–36]'. References [34], [35], and [36] are, respectively, a review of sound source localization, an acoustic SLAM paper, and a survey of underwater acoustic SLAM; none of them addresses cyclostationary machinery fault diagnosis. This leaves the paragraph's specific technical claim without supporting evidence. Please replace these citations with the actual cyclostationary/condition-monitoring literature (for example, the Randall monograph already in the reference list as [27]) or delete the unsupported claim.
- [2.4, Eqs. (9)-(10)] Equation (9) and its axisymmetric counterpart Eq. (10) write the nonlinear term of the KZK equation as (beta_NL/(2 rho0 c^3)) partial^3(p^2)/partial tau^3. The standard KZK/Westervelt nonlinearity is (beta/(2 rho0 c^3)) partial^2(p^2)/partial tau^2; the third-order derivative makes the term dimensionally inconsistent with the left-hand side (it introduces an extra factor of 1/s) and does not correspond to the canonical KZK equation. Since the paper explicitly uses the KZK equation as part of its theoretical foundation and references it in the authors' own framework [44], this is a technical error in a load-bearing equation and should be corrected.
- [3.2 and 4] The citation mismatches are not confined to Section 2. In Section 3.2.1, [50] (MUGEN) does not support the claim about differentiable physics simulation; in Section 3.2.2, [52] is an audio-visual anomaly-detection paper, not the Neural Acoustic Fields approach described; and in Section 3.2.3, [55] and [56] are, respectively, a heart-sound review and a COVID chest-X-ray study, not controllable sound synthesis. Similar mismatches recur in Section 4 (for example, [32] and [33] are not tire-road noise papers, and [34] and [35] are not vehicle health-monitoring papers). Because these are the papers that anchor the methodological pillars and application claims, the survey's map of the literature is unreliable as it stands; a systematic reference audit and re-citation is needed.
minor comments (5)
- [1] The Introduction's phrase 'deeply rooted in principles of auditory scene analysis [11, 12]' cites [12] as Computational Physics (Vesely), which is not an auditory scene analysis reference; the citation should be corrected.
- [6 and Appendix B] Section 6, 'Next Steps in Our Research', is unusual for a survey and functions as self-promotion of the authors' own paper [44] and the company-affiliated GitHub repository in Appendix B; if these are kept, they should be clearly separated from the survey content and explicitly identified as author-affiliated resources.
- [4] The Section 4 heading reads 'Deploying Acoustic World Modes in Real World'; 'Modes' should be 'Models'.
- [Throughout] Several grammatical errors remain (e.g., 'a acoustic world model', 'a important'), and the manuscript would benefit from a careful language edit.
- [6] The 'Limitations in Our Survey' passage appears after the Conclusions; integrating it before the Conclusions would improve the paper's structure.
Circularity Check
No circular derivation: the survey's central content is independent of the authors' self-citations; only minor self-promotional self-citation appears in Section 6 and Appendix B.
full rationale
This paper is a survey rather than a derivation: it proposes no fitted parameter that is later renamed a prediction, imports no uniqueness theorem from self-citations, and does not define its central concept in terms of its conclusion. The closest issue is self-referential promotion: Section 6 ('Next Steps in Our Research') highlights the authors' own [44] as a practical foundation, [44] is also cited in Section 2.4 for KZK applications and in Section 4.1 for human-robot collaboration, and Appendix B points to a company-affiliated GitHub repository. These are self-citations, but they are not load-bearing for the central claim that acoustic signals encode physical information and that PINNs, generative models, and self-supervised learning can exploit that information. That claim rests on standard acoustics and a broad external literature. Separately, the skeptical concerns about misassigned citations in Sections 2.1-2.2 and the dimensional inconsistency in Eq. (9) are important evidence-quality and correctness issues, but they are not circularity under the rubric: a wrong equation or a poorly matched citation does not make the argument equivalent to its inputs. No step in the paper exhibits the required reduction-by-construction, so the appropriate finding is a low score reflecting only minor self-promotional self-citation.
Assumptions & free parameters
assumptions (3)
- domain assumption Linear, isotropic, homogeneous elastodynamics with small deformations governs the vibration of struck objects (Section 2.1, Eq. 1).
- domain assumption Lighthill's acoustic analogy and the FW-H equation adequately describe the sound generated by flows and moving surfaces (Section 2.2, Eq. 3).
- domain assumption Sabine and Eyring reverberation-time formulas are valid for the rooms discussed (Section 2.3, Eqs. 5 and 6).
Cite this review
Pith. "Pith review of A Survey on World Models Grounded in Acoustic Physical Information." pith.science (2026). https://pith.science/paper/Q2MWQPKQ
@misc{pith2026250613833,
author = {Pith},
title = {Pith review of: A Survey on World Models Grounded in Acoustic Physical Information},
year = {2026},
howpublished = {\url{https://pith.science/paper/Q2MWQPKQ}},
note = {Machine review of arXiv:2506.13833}
}
read the original abstract
This survey provides a comprehensive overview of the emerging field of world models grounded in the foundation of acoustic physical information. It examines the theoretical underpinnings, essential methodological frameworks, and recent technological advancements in leveraging acoustic signals for high-fidelity environmental perception, causal physical reasoning, and predictive simulation of dynamic events. The survey explains how acoustic signals, as direct carriers of mechanical wave energy from physical events, encode rich, latent information about material properties, internal geometric structures, and complex interaction dynamics. Specifically, this survey establishes the theoretical foundation by explaining how fundamental physical laws govern the encoding of physical information within acoustic signals. It then reviews the core methodological pillars, including Physics-Informed Neural Networks (PINNs), generative models, and self-supervised multimodal learning frameworks. Furthermore, the survey details the significant applications of acoustic world models in robotics, autonomous driving, healthcare, and finance. Finally, it systematically outlines the important technical and ethical challenges while proposing a concrete roadmap for future research directions toward robust, causal, uncertainty-aware, and responsible acoustic intelligence. These elements collectively point to a research pathway towards embodied active acoustic intelligence, empowering AI systems to construct an internal "intuitive physics" engine through sound.
Figures
Reference graph
Works this paper leans on
-
[16]
Ultrasonic non-destructive evaluation of composites: A review,
J. Jodhani, A. Handa, A. Gautam, R. Ranaet al., “Ultrasonic non-destructive evaluation of composites: A review,”Materials Today: Proceedings, vol. 78, pp. 627–632, 2023. 23
2023
-
[17]
Porosity evaluation of additively manufactured components using deep learning-based ultrasonic nondestructive testing,
S.-H. Park, S. Choi, and K.-Y. Jhang, “Porosity evaluation of additively manufactured components using deep learning-based ultrasonic nondestructive testing,”International Journal of Precision Engineering and Manufacturing-Green Technology, vol. 9, no. 2, pp. 395–407, 2022
2022
-
[18]
Deep learning approach for damage classification based on acoustic emission data in composite materials,
F. Guo, W. Li, P. Jiang, F. Chen, and Y. Liu, “Deep learning approach for damage classification based on acoustic emission data in composite materials,”Materials, vol. 15, no. 12, p. 4270, 2022
2022
-
[12]
F. J. Veselyet al.,Computational Physics. Springer, 1994
1994
-
[13]
Statistics of natural reverberation enable perceptual separation of sound and space,
J. Traer and J. H. McDermott, “Statistics of natural reverberation enable perceptual separation of sound and space,”Proceedings of the National Academy of Sciences, vol. 113, no. 48, pp. E7856–E7865, 2016
2016
-
[14]
Material categorization and hardness scaling in real and synthetic impact sounds,
B. L. Giordano, “Material categorization and hardness scaling in real and synthetic impact sounds,” The sounding object, pp. 73–94, 2003
2003
-
[34]
A review on sound source localization systems,
D. Desai and N. Mehendale, “A review on sound source localization systems,”Archives of Computational Methods in Engineering, vol. 29, no. 7, pp. 4631–4642, 2022
work page 2022
-
[35]
C. Evers and P. A. Naylor, “Acoustic slam,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 26, no. 9, pp. 1484–1498, 2018. 24
work page 2018
-
[36]
A survey of underwater acoustic slam system,
M. Jiang, S. Song, Y. Li, W. Jin, J. Liu, and X. Feng, “A survey of underwater acoustic slam system,” inIntelligent Robotics and Applications: 12th International Conference, ICIRA 2019, Shenyang, China, August 8–11, 2019, Proceedings, Part II 12. Springer, 2019, pp. 159–170
work page 2019
-
[27]
R. B. Randall,Vibration-based condition monitoring: industrial, automotive and aerospace applications. John Wiley & Sons, 2021
work page 2021
-
[44]
X. Chen, X. Yu, L. Chang, Y. Huang, J. He, S. Zhang, J. Li, L. Lin, Z. Zeng, X. Tuet al., “A synergistic framework of nonlinear acoustic computing and reinforcement learning for real-world human- robot interaction,”arXiv preprint arXiv:2505.01998, 2025
arXiv 2025
-
[50]
Mugen: A playground for video-audio-text multimodal understanding and generation,
T. Hayes, S. Zhang, X. Yin, G. Pang, S. Sheng, H. Yang, S. Ge, Q. Hu, and D. Parikh, “Mugen: A playground for video-audio-text multimodal understanding and generation,” inEuropean Conference on Computer Vision. Springer, 2022, pp. 431–449
work page 2022
-
[52]
Audio–visualrepresentationlearningforanomalyeventsdetection in crowds,
J.Gao, H.Yang, M.Gong, andX.Li, “Audio–visualrepresentationlearningforanomalyeventsdetection in crowds,”Neurocomputing, vol. 582, p. 127489, 2024
work page 2024
-
[55]
Algorithms for automatic analysis and classi- fication of heart sounds–a systematic review,
A. K. Dwivedi, S. A. Imtiaz, and E. Rodriguez-Villegas, “Algorithms for automatic analysis and classi- fication of heart sounds–a systematic review,”IEEE Access, vol. 7, pp. 8316–8345, 2018
work page 2018
-
[56]
Deep learning approaches for covid-19 detection based on chest x-ray images,
A. M. Ismael and A. Şengür, “Deep learning approaches for covid-19 detection based on chest x-ray images,”Expert Systems with Applications, vol. 164, p. 114054, 2021
work page 2021
-
[32]
Deep room recognition using inaudible echos,
Q. Song, C. Gu, and R. Tan, “Deep room recognition using inaudible echos,”Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, vol. 2, no. 3, pp. 1–28, 2018
work page 2018
-
[33]
Springer Science & Business Media, 2001
M.BrandsteinandD.Ward,Microphone arrays: signal processing techniques and applications. Springer Science & Business Media, 2001
work page 2001
Show all 93 references
-
[1]
Language models are few-shot learners,
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askellet al., “Language models are few-shot learners,”Advances in neural information processing systems, vol. 33, pp. 1877–1901, 2020
1901
-
[2]
Open foundation and fine-tuned chat models,
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaeiet al., “Open foundation and fine-tuned chat models,”arXiv preprint arXiv:2307.09288, 2023
2023 arXiv
-
[3]
Climbing towards nlu: On meaning, form, and understanding in the age of data,
E. M. Bender and A. Koller, “Climbing towards nlu: On meaning, form, and understanding in the age of data,” inProceedings of the 58th annual meeting of the association for computational linguistics, 2020, pp. 5185–5198
2020
-
[4]
World models,
D. Ha and J. Schmidhuber, “World models,”arXiv preprint arXiv:1803.10122, 2018
2018 arXiv
-
[5]
Dream to control: Learning behaviors by latent imagination,
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi, “Dream to control: Learning behaviors by latent imagination,”arXiv preprint arXiv:1912.01603, 2019
1912 arXiv
-
[6]
A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27,
Y. LeCun, “A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27,”Open Review, vol. 62, no. 1, pp. 1–62, 2022
2022
-
[7]
Deep learning, reinforcement learning, and world models,
Y. Matsuo, Y. LeCun, M. Sahani, D. Precup, D. Silver, M. Sugiyama, E. Uchibe, and J. Morimoto, “Deep learning, reinforcement learning, and world models,”Neural Networks, vol. 152, pp. 267–275, 2022
2022
-
[8]
Imagenet classification with deep convolutional neural networks,
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,”Communications of the ACM, vol. 60, no. 6, pp. 84–90, 2017
2017
-
[9]
An image is worth 16x16 words: Transformers for image recognition at scale,
A.Dosovitskiy, L.Beyer, A.Kolesnikov, D.Weissenborn, X.Zhai, T.Unterthiner, M.Dehghani, M.Min- derer, G. Heigold, S. Gellyet al., “An image is worth 16x16 words: Transformers for image recognition at scale,”arXiv preprint arXiv:2010.11929, 2020
2010 arXiv
-
[10]
Benchmarking neural network robustness to common corruptions and perturbations,
D. Hendrycks and T. Dietterich, “Benchmarking neural network robustness to common corruptions and perturbations,”arXiv preprint arXiv:1903.12261, 2019
1903 arXiv
-
[11]
A. S. Bregman,Auditory scene analysis: The perceptual organization of sound. MIT press, 1994
1994
-
[15]
Resonant ultrasound spectroscopy: applications, current status and limitations,
R. Schwarz and J. Vuorinen, “Resonant ultrasound spectroscopy: applications, current status and limitations,”Journal of Alloys and Compounds, vol. 310, no. 1-2, pp. 243–250, 2000
2000
-
[19]
Sound production and modeling,
P. R. Cook, “Sound production and modeling,”IEEE Computer Graphics and applications, vol. 22, no. 4, pp. 23–27, 2002
2002
-
[20]
Controlling material properties in physical models of sounding ob- jects,
F. Avanzini, D. Rocchessoet al., “Controlling material properties in physical models of sounding ob- jects,” inICMC, 2001
2001
-
[21]
Auditory scene analysis: Hearing in complex environments,
A. S. Bregnian, “Auditory scene analysis: Hearing in complex environments,” 1993
1993
-
[22]
Multimodal human–robot interaction for human-centric smart manufacturing: a survey,
T. Wang, P. Zheng, S. Li, and L. Wang, “Multimodal human–robot interaction for human-centric smart manufacturing: a survey,”Advanced Intelligent Systems, vol. 6, no. 3, p. 2300359, 2024
2024
-
[23]
Robot sound interpretation: Combining sight and sound in learning-based control,
P. Chang, S. Liu, H. Chen, and K. Driggs-Campbell, “Robot sound interpretation: Combining sight and sound in learning-based control,” in2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2020, pp. 5580–5587
2020
-
[24]
Leak detection in real water distribution networks based on acoustic emission and machine learning,
A. Fares, I. Tijani, Z. Rui, and T. Zayed, “Leak detection in real water distribution networks based on acoustic emission and machine learning,”Environmental Technology, vol. 44, no. 25, pp. 3850–3866, 2023
2023
-
[25]
Estimating rainfall intensity based on surveillance audio and deep-learning,
M. Wang, M. Chen, Z. Wang, Y. Guo, Y. Wu, W. Zhao, and X. Liu, “Estimating rainfall intensity based on surveillance audio and deep-learning,”Environmental Science and Ecotechnology, vol. 22, p. 100450, 2024
2024
-
[26]
Road type classificationusingdeeplearningfortire-pavementinteractionnoisedatainautonomousdrivingvehicle,
S.-K. Lee, J. Yoo, C.-H. Lee, K. An, Y.-S. Yoon, J. Lee, G.-H. Yeom, and S.-U. Hwang, “Road type classificationusingdeeplearningfortire-pavementinteractionnoisedatainautonomousdrivingvehicle,” Applied Acoustics, vol. 212, p. 109597, 2023
2023
-
[28]
Onsoundgeneratedaerodynamicallyi.generaltheory,
M.J.Lighthill, “Onsoundgeneratedaerodynamicallyi.generaltheory,”Proceedings of the Royal Society of London. Series A. Mathematical and Physical Sciences, vol. 211, no. 1107, pp. 564–587, 1952
1952
-
[29]
A novel deep autoencoder feature learning method for rotating machinery fault diagnosis,
H. Shao, H. Jiang, H. Zhao, and F. Wang, “A novel deep autoencoder feature learning method for rotating machinery fault diagnosis,”Mechanical Systems and Signal Processing, vol. 95, pp. 187–204, 2017
2017
-
[30]
A review of data-driven machinery fault diagnosis using machine learning algorithms,
J. Cen, Z. Yang, X. Liu, J. Xiong, and H. Chen, “A review of data-driven machinery fault diagnosis using machine learning algorithms,”Journal of Vibration Engineering & Technologies, vol. 10, no. 7, pp. 2481–2507, 2022
2022
-
[31]
Kuttruff and M
H. Kuttruff and M. Vorländer,Room acoustics. Crc Press, 2024
2024
-
[37]
Acoustic echoes reveal room shape,
I. Dokmanić, R. Parhizkar, A. Walther, Y. M. Lu, and M. Vetterli, “Acoustic echoes reveal room shape,” Proceedings of the National Academy of Sciences, vol. 110, no. 30, pp. 12186–12191, 2013
2013
-
[38]
Batslam: Simultaneous localization and mapping using biomimetic sonar,
J. Steckel and H. Peremans, “Batslam: Simultaneous localization and mapping using biomimetic sonar,” PloS one, vol. 8, no. 1, p. e54076, 2013
2013
-
[39]
Neural representation of three-dimensional acoustic space in the human temporal lobe,
X. Zhang, Q. Zhang, X. Hu, and B. Zhang, “Neural representation of three-dimensional acoustic space in the human temporal lobe,”Frontiers in Human Neuroscience, vol. 9, p. 203, 2015
2015
-
[40]
Overview of geometrical room acoustic modeling techniques,
L. Savioja and U. P. Svensson, “Overview of geometrical room acoustic modeling techniques,”The Journal of the Acoustical Society of America, vol. 138, no. 2, pp. 708–730, 2015
2015
-
[41]
A system for data-driven concatenative sound synthesis,
D. Schwarz, “A system for data-driven concatenative sound synthesis,” inDigital Audio Effects (DAFx), 2000, pp. 97–102
2000
-
[42]
Pertilä,Acoustic source localization in a room environment and at moderate distances
P. Pertilä,Acoustic source localization in a room environment and at moderate distances. Tampere University of Technology, 2009
2009
-
[43]
Nonlinear acoustics in china,
Q. Zuwen, “Nonlinear acoustics in china,”WULI-BEIJING-, vol. 28, no. 10, pp. 593–599, 1999
1999
-
[45]
Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations,
M. Raissi, P. Perdikaris, and G. E. Karniadakis, “Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations,” Journal of Computational physics, vol. 378, pp. 686–707, 2019
2019
-
[46]
Physics-informed machine learning,
G. E. Karniadakis, I. G. Kevrekidis, L. Lu, P. Perdikaris, S. Wang, and L. Yang, “Physics-informed machine learning,”Nature Reviews Physics, vol. 3, no. 6, pp. 422–440, 2021
2021
-
[47]
Physics-informed deep learning for structural vibration identification and its application on a benchmark structure,
M. Zhang, T. Guo, G. Zhang, Z. Liu, and W. Xu, “Physics-informed deep learning for structural vibration identification and its application on a benchmark structure,”Philosophical Transactions of the Royal Society A, vol. 382, no. 2264, p. 20220400, 2024
2024
-
[48]
A study on the auralization system using flexible rendering for virtual reality-based acoustic simulation,
Y. Lee and J. Ryu, “A study on the auralization system using flexible rendering for virtual reality-based acoustic simulation,”Journal of Advanced Mechanical Design, Systems, and Manufacturing, vol. 13, no. 5, pp. JAMDSM0094–JAMDSM0094, 2019
2019
-
[49]
Physics- informed neural network for volumetric sound field reconstruction of speech signals,
M. Olivieri, X. Karakonstantis, M. Pezzoli, F. Antonacci, A. Sarti, and E. Fernandez-Grande, “Physics- informed neural network for volumetric sound field reconstruction of speech signals,”EURASIP Journal on Audio, Speech, and Music Processing, vol. 2024, no. 1, p. 42, 2024
2024
-
[51]
Nerf: Repre- senting scenes as neural radiance fields for view synthesis,
B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “Nerf: Repre- senting scenes as neural radiance fields for view synthesis,”Communications of the ACM, vol. 65, no. 1, pp. 99–106, 2021
2021
-
[53]
Equivariant neural rendering,
E. Dupont, M. B. Martin, A. Colburn, A. Sankar, J. Susskind, and Q. Shan, “Equivariant neural rendering,” inInternational Conference on Machine Learning. PMLR, 2020, pp. 2761–2770. 25
2020
-
[54]
Diffwave: A versatile diffusion model for audio synthesis,
Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, “Diffwave: A versatile diffusion model for audio synthesis,”arXiv preprint arXiv:2009.09761, 2020
2009 arXiv
-
[57]
Look, listen and learn,
R. Arandjelovic and A. Zisserman, “Look, listen and learn,” inProceedings of the IEEE international conference on computer vision, 2017, pp. 609–617
2017
-
[58]
Audio-visual scene analysis with self-supervised multisensory features,
A. Owens and A. A. Efros, “Audio-visual scene analysis with self-supervised multisensory features,” in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 631–648
2018
-
[59]
Learning transferable visual models from natural language supervision,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clarket al., “Learning transferable visual models from natural language supervision,” inInternational conference on machine learning. PMLR, 2021, pp. 8748–8763
2021
-
[60]
Flamingo: a visual language model for few-shot learning,
J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynoldset al., “Flamingo: a visual language model for few-shot learning,”Advances in neural information processing systems, vol. 35, pp. 23716–23736, 2022
2022
-
[61]
Learning to localize sound source in visual scenes,
A. Senocak, T.-H. Oh, J. Kim, M.-H. Yang, and I. S. Kweon, “Learning to localize sound source in visual scenes,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 4358–4366
2018
-
[62]
Co-separating sounds of visual objects,
R. Gao and K. Grauman, “Co-separating sounds of visual objects,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 3879–3888
2019
-
[63]
Object category recognition by a humanoid robot using behavior- grounded relational learning,
J. Sinapov and A. Stoytchev, “Object category recognition by a humanoid robot using behavior- grounded relational learning,” in2011 IEEE International Conference on Robotics and Automation. IEEE, 2011, pp. 184–190
2011
-
[64]
Communication in human-robot interaction,
A. Bonarini, “Communication in human-robot interaction,”Current Robotics Reports, vol. 1, no. 4, pp. 279–285, 2020
2020
-
[65]
A comprehensive review of polyphonic sound event detection,
T. K. Chan and C. S. Chin, “A comprehensive review of polyphonic sound event detection,”IEEE Access, vol. 8, pp. 103339–103373, 2020
2020
-
[66]
Acoustic based emergency vehicle detection using ensemble of deep learning models,
U. Mittal and P. Chawla, “Acoustic based emergency vehicle detection using ensemble of deep learning models,”Procedia Computer Science, vol. 218, pp. 227–234, 2023
2023
-
[67]
Audio surveillance: A systematic review,
M. Crocco, M. Cristani, A. Trucco, and V. Murino, “Audio surveillance: A systematic review,”ACM Computing Surveys (CSUR), vol. 48, no. 4, pp. 1–46, 2016
2016
-
[68]
A review of depression and suicide risk assessment using speech analysis,
N. Cummins, S. Scherer, J. Krajewski, S. Schnieder, J. Epps, and T. F. Quatieri, “A review of depression and suicide risk assessment using speech analysis,”Speech communication, vol. 71, pp. 10–49, 2015
2015
-
[69]
Covid-19 artificial intelligence diagnosis using only cough recordings,
J. Laguarta, F. Hueto, and B. Subirana, “Covid-19 artificial intelligence diagnosis using only cough recordings,”IEEE Open Journal of Engineering in Medicine and Biology, vol. 1, pp. 275–281, 2020
2020
-
[70]
Adaptive, personalized closed-loop therapy for parkinson’s disease: Biochemical, neurophysiological, and wearable sensing systems,
L. di Biase, G. Tinkhauser, E. Martin Moraud, M. L. Caminiti, P. M. Pecoraro, and V. Di Lazzaro, “Adaptive, personalized closed-loop therapy for parkinson’s disease: Biochemical, neurophysiological, and wearable sensing systems,”Expert review of neurotherapeutics, vol. 21, no....
2021
-
[71]
Emotion recognition from speech: a review,
S. G. Koolagudi and K. S. Rao, “Emotion recognition from speech: a review,”International journal of speech technology, vol. 15, pp. 99–117, 2012
2012
-
[72]
A comprehensive review of speech emotion recognition systems,
T. M. Wani, T. S. Gunawan, S. A. A. Qadri, M. Kartiwi, and E. Ambikairajah, “A comprehensive review of speech emotion recognition systems,”IEEE access, vol. 9, pp. 47795–47814, 2021. 26
2021
-
[73]
Multimodal machine learning: A survey and taxonomy,
T. Baltrušaitis, C. Ahuja, and L.-P. Morency, “Multimodal machine learning: A survey and taxonomy,” IEEE transactions on pattern analysis and machine intelligence, vol. 41, no. 2, pp. 423–443, 2018
2018
-
[74]
The power of voice: Managerial affective states and future firm performance,
W. J. Mayew and M. Venkatachalam, “The power of voice: Managerial affective states and future firm performance,”The Journal of Finance, vol. 67, no. 1, pp. 1–43, 2012
2012
-
[75]
Analyzingspeechtodetectfinancialmisreporting,
J.L.Hobson, W.J.Mayew, andM.Venkatachalam, “Analyzingspeechtodetectfinancialmisreporting,” Journal of Accounting Research, vol. 50, no. 2, pp. 349–392, 2012
2012
-
[76]
Predicting user satisfaction from turn-taking in spoken conversations
S. A. Chowdhury, E. A. Stepanov, G. Riccardiet al., “Predicting user satisfaction from turn-taking in spoken conversations.” inInterspeech, 2016, pp. 2910–2914
2016
-
[77]
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding,
S. Han, H. Mao, and W. J. Dally, “Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding,”arXiv preprint arXiv:1510.00149, 2015
2015 arXiv
-
[78]
Supervisedspeechseparationbasedondeeplearning: Anoverview,
D.WangandJ.Chen, “Supervisedspeechseparationbasedondeeplearning: Anoverview,”IEEE/ACM transactions on audio, speech, and language processing, vol. 26, no. 10, pp. 1702–1726, 2018
2018
-
[79]
Federated learning: Strategies for improving communication efficiency,
J. Konečn` y, H. B. McMahan, F. X. Yu, P. Richtárik, A. T. Suresh, and D. Bacon, “Federated learning: Strategies for improving communication efficiency,”arXiv preprint arXiv:1610.05492, 2016
2016 arXiv
-
[80]
Privacy-preserving human activity sensing: A survey,
Y. Yang, P. Hu, J. Shen, H. Cheng, Z. An, and X. Liu, “Privacy-preserving human activity sensing: A survey,”High-Confidence Computing, vol. 4, no. 1, p. 100204, 2024
2024
-
[81]
Transfer learning from speaker verification to multispeaker text-to-speech synthesis,
Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, P. Nguyen, R. Pang, I. Lopez Moreno, Y. Wu et al., “Transfer learning from speaker verification to multispeaker text-to-speech synthesis,”Advances in neural information processing systems, vol. 31, 2018
2018
-
[82]
Explaining deep neural networks and beyond: A review of methods and applications,
W. Samek, G. Montavon, S. Lapuschkin, C. J. Anders, and K.-R. Müller, “Explaining deep neural networks and beyond: A review of methods and applications,”Proceedings of the IEEE, vol. 109, no. 3, pp. 247–278, 2021
2021
-
[83]
Audio set: An ontology and human-labeled dataset for audio events,
J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Rit- ter, “Audio set: An ontology and human-labeled dataset for audio events,” in2017 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 2017, ...
2017
-
[84]
Fsd50k: an open dataset of human-labeled sound events,
E. Fonseca, X. Favory, J. Pons, F. Font, and X. Serra, “Fsd50k: an open dataset of human-labeled sound events,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 30, pp. 829–852, 2021
2021
-
[85]
Wham!: Extending speech separation to noisy environments,
G. Wichern, J. Antognini, M. Flynn, L. R. Zhu, E. McQuinn, D. Crow, E. Manilow, and J. L. Roux, “Wham!: Extending speech separation to noisy environments,”arXiv preprint arXiv:1907.01160, 2019
1907 arXiv
-
[86]
Tut database for acoustic scene classification and sound event detection,
A. Mesaros, T. Heittola, and T. Virtanen, “Tut database for acoustic scene classification and sound event detection,” in2016 24th European Signal Processing Conference (EUSIPCO). IEEE, 2016, pp. 1128–1132
2016
-
[87]
Habitat: A platform for embodied ai research,
M. Savva, A. Kadian, O. Maksymets, Y. Zhao, E. Wijmans, B. Jain, J. Straub, J. Liu, V. Koltun, J. Maliket al., “Habitat: A platform for embodied ai research,” inProceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 9339–9347
2019
-
[88]
Soundspaces: Audio-visual navigation in 3d environments,
C. Chen, U. Jain, C. Schissler, S. V. A. Gari, Z. Al-Halah, V. K. Ithapu, P. Robinson, and K. Grauman, “Soundspaces: Audio-visual navigation in 3d environments,” inComputer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part VI 16. Sp...
2020
-
[89]
Isaac gym: High performance gpu-based physics simulation for robot learning,
V. Makoviychuk, L. Wawrzyniak, Y. Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. All- shire, A. Handaet al., “Isaac gym: High performance gpu-based physics simulation for robot learning,” arXiv preprint arXiv:2108.10470, 2021. 27
2021 arXiv
-
[90]
Threedworld: A platform for interactive multi-modal physical simulation,
C. Gan, J. Schwartz, S. Alter, D. Mrowca, M. Schrimpf, J. Traer, J. De Freitas, J. Kubilius, A. Bhand- waldar, N. Haberet al., “Threedworld: A platform for interactive multi-modal physical simulation,” arXiv preprint arXiv:2007.04954, 2020
2007 arXiv
-
[91]
Sapien: A simulatedpart-basedinteractiveenvironment,
F. Xiang, Y. Qin, K. Mo, Y. Xia, H. Zhu, F. Liu, M. Liu, H. Jiang, Y. Yuan, H. Wanget al., “Sapien: A simulatedpart-basedinteractiveenvironment,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 11097–11107
2020
-
[92]
Adecadeofdcase: Achievements, practices, evaluations and future challenges,
A.Mesaros, R.Serizel, T.Heittola, T.Virtanen, andM.D.Plumbley, “Adecadeofdcase: Achievements, practices, evaluations and future challenges,” inICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025, pp. 1–5
2025
-
[93]
Hear: Holistic evaluation of audio representations,
J. Turian, J. Shier, H. R. Khan, B. Raj, B. W. Schuller, C. J. Steinmetz, C. Malloy, G. Tzanetakis, G. Velarde, K. McNallyet al., “Hear: Holistic evaluation of audio representations,” inNeurIPS 2021 Competitions and Demonstrations Track. PMLR, 2022, pp. 125–145. 28
2021
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.