Pith. sign in

REVIEW 4 major objections 5 minor 93 references

A Survey on World Models Grounded in Acoustic Physical Information

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This survey argues that acoustic signals—the radiated mechanical energy of physical events—should be a primary modality for AI world models, enabling causal physical understanding and predictive simulation through sound.

desk verdict Useful framing for a survey, but the citation mismatches and a dimensionally wrong KZK term need fixing before it can be trusted as a field map. read the letter →

arxiv 2506.13833 v1 pith:Q2MWQPKQ submitted 2025-06-16 cs.SD cs.AIcs.ROeess.ASphysics.app-ph

classification cs.SDcs.AIcs.ROeess.ASphysics.app-ph
keywords worldmodelsacousticphysicalinformationphysics-informedneuralnetworksperceptionmultimodallearningdynamicpredictionembodiedintelligencecausalinference
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey makes the case that sound should be treated as a primary physical modality for AI world models, not a side channel for event detection. Its central claim is that acoustic signals are radiated mechanical energy from physical events, so they carry latent information about material properties, internal structure, contact and fluid dynamics, and spatial geometry. The paper organizes the field around three methodological pillars—physics-informed neural networks, generative forward models, and self-supervised multimodal learning—and argues that combining them lets an AI build an internal 'intuitive physics' engine by listening. It then maps applications in robotics, autonomous driving, healthcare, and finance, and closes with a research roadmap toward causal, uncertainty-aware, embodied, and responsible acoustic intelligence.

What carries the argument

The argument is carried by a chain of physics-to-signal encodings. The elastodynamic wave equation $$(\$\lambda$+\mu)\nabla(\nabla\cdot u)+\mu\$nabla^{2}$ u=\rho\,\$partial^{2}$ u/\partial $t^{2}$$$ with modal analysis links an object's natural frequencies $\omega_n$ to its material and geometry; the acoustic analogy recasts the Navier-Stokes equations into an inhomogeneous wave equation for flow-generated sound; the room impulse response $h(t)$, with standard reverberation-time estimates $T_{60}$, encodes a space's geometry and absorption; and the Westervelt and KZK equations describe nonlinear high-amplitude propagation. On the machine-learning side, the composite PINN loss $$\mathcal{L}(\$\theta$)=w_{\mathrm{data}}\mathcal{L}_{\mathrm{data}}+w_{\mathrm{phys}}\mathcal{L}_{\mathrm{phys}}+w_{\mathrm{bc}}\mathcal{L}_{\mathrm{bc}}+w_{\mathrm{ic}}\mathcal{L}_{\mathrm{ic}}$$ is the gray-box device that enforces these partial differential equations as regularizers, while differentiable simulators, neural acoustic fields, and contrastive audio-visual models supply forward prediction and scalable representation learning. Each element does a specific job: physics supplies the acoustic fingerprint, physics-informed networks invert it with sparse data, generative models predict and synthesize, and self-supervised learning scales the representations.

What would settle it

A concrete falsifier: train a model to infer material and shape from impact sounds, then evaluate it on objects whose geometry, excitation type, and recording conditions are randomized and never seen during training; if accuracy falls to chance, the modal-inversion route to physical properties is not doing the work claimed. A second check: on a held-out physical-prediction task, if an acoustic-only model adds no predictive power over a vision-only model, the claimed complementarity of sound for physical intuition is unsupported.

Watch

Extended reading notes

Core claim

The paper's central claim is that the physical laws governing sound—elastodynamics, aeroacoustics, room acoustics, and nonlinear wave propagation—leave measurable acoustic signatures that can be inverted for physical knowledge and used forward for prediction. Vibrations of struck objects encode material and modal properties such as elastic modulus, density, and damping; contact and flow sounds encode collision, friction, and fluid dynamics; room impulse responses encode geometry and surface absorption; and nonlinear equations such as the Westervelt and KZK equations extend the picture to high-intensity fields. The survey's thesis is that these signatures make audio complementary to vision precisely where vision is weak: seeing surfaces but not masses, internal defects, contact forces, or occluded structure. An AI that learns these mappings can simulate events in latent space and reason about physical interventions, which the paper calls an internal 'intuitive physics' engine through sound.

Load-bearing premise

The load-bearing premise is that acoustic signals carry enough latent physical information—and the cited methods can extract enough of it—that sound alone can ground causal physical understanding; the survey's evidence base also has to be a fair sample of the field for its roadmap to be trustworthy.

Editorial extensions

If this is right

  • If the acoustic world-model program works, robots can localize and map in darkness, smoke, or dust where vision fails, using Acoustic SLAM built from echoes and passive sound sources.
  • Contact and friction sounds give robotic manipulators haptic-like feedback, letting them detect incipient slip, adjust grip, and infer hidden object properties such as fill level by shaking a container.
  • Vehicles gain a 360-degree safety layer: siren detection beyond line of sight and real-time road-surface classification from tire-road noise, complementing vision and LiDAR where those sensors are ambiguous.
  • Non-invasive health screening can scale through cough, breathing, and voice analysis on ubiquitous microphones, with vocal biomarkers tracking respiratory, neurological, and mental-health conditions.
  • Financial applications can extract non-consensus signals from paralinguistic cues in earnings calls and use acoustic stress markers for fraud detection and compliance.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: if audio truly encodes physical properties as the survey argues, sound can serve as a cheap self-supervision signal for vision-based models, correcting visual errors about mass, material, and internal state without new hardware.
  • Editorial inference: the survey's own hybrid scenario—self-supervised pretraining, generative synthesis, and PINN regularization—can be tested directly by comparing that combined architecture against each pillar alone on a physical audio prediction benchmark.
  • Editorial note grounded in the paper's stated limitations: the authors concede that sensor design and classical array processing are treated only as background, so the deployment roadmap is stronger on algorithms than on the hardware that would make those algorithms reliable in the field.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This manuscript is a survey of 'world models grounded in acoustic physical information.' Its central thesis is that acoustic signals, as direct carriers of mechanical wave energy from physical events, encode latent information about material properties, internal structure, geometry, and interaction dynamics, and that this makes sound a privileged modality for AI systems that must build causal, physics-grounded world models. The paper reviews the physical theory (elastodynamics, aeroacoustics, room acoustics, nonlinear acoustics), organizes methodological work into three pillars (physics-informed neural networks, generative forward models, and self-supervised multimodal learning), surveys applications in robotics, autonomous driving, healthcare, and finance, and closes with an ethics discussion and a five-pathway research roadmap.

Significance. If its map of the literature were reliable, this survey would provide a useful synthesis of an emerging interdisciplinary area and a concrete agenda for future work. The paper is clearly organized, and the comparative analysis in Table 1, the application taxonomy in Table 2, the explicit limitations statement, and the resource list in Appendix A are useful editorial features. However, the survey's central claim depends on the cited literature substantiating specific acoustic-inference capabilities, and many load-bearing citations do not support the claims they are attached to (Sections 2.1, 2.2, 3.2, and 4). In addition, Eq. (9) contains a dimensional inconsistency in the KZK nonlinear term. Because a survey's contribution is precisely its map of the evidence base, these problems materially weaken the paper. They are, however, correctable within the manuscript's scope.

major comments (4)
  1. [2.1] The sentence 'Recent works have demonstrated that deep learning models can leverage these rich acoustic signatures to classify object materials with high accuracy [16], estimate the viscosity of liquids ... [17], or even infer the particle size of granular materials from the statistical properties of many small impacts during shaking [18]' is not supported by the cited papers: [16] is a review of ultrasonic non-destructive evaluation of composites, [17] is a study of porosity evaluation of additively manufactured parts using ultrasonic NDT, and [18] is a study of acoustic-emission damage classification in composites. None of these works reports impact-sound material classification, liquid-viscosity estimation, or granular particle-size inference. The same paragraph's earlier claim that material identification from sound rests on a 'large body of work' also cites [13], which is a study of natural reverberation statistics, alongside the relevant [14]. This passage is load-bearing for the survey's central thesis that acoustic signals encode material and fluid properties, so it should be rewritten with correct references or the claims should be removed.
  2. [2.2] The cyclostationary-analysis paragraph states that deep learning models trained on spectral correlation and envelope analysis 'can be trained on these processed representations to perform highly sensitive fault diagnosis and prognostics, effectively creating a detailed acoustic physical model of a machine's health state [34–36]'. References [34], [35], and [36] are, respectively, a review of sound source localization, an acoustic SLAM paper, and a survey of underwater acoustic SLAM; none of them addresses cyclostationary machinery fault diagnosis. This leaves the paragraph's specific technical claim without supporting evidence. Please replace these citations with the actual cyclostationary/condition-monitoring literature (for example, the Randall monograph already in the reference list as [27]) or delete the unsupported claim.
  3. [2.4, Eqs. (9)-(10)] Equation (9) and its axisymmetric counterpart Eq. (10) write the nonlinear term of the KZK equation as (beta_NL/(2 rho0 c^3)) partial^3(p^2)/partial tau^3. The standard KZK/Westervelt nonlinearity is (beta/(2 rho0 c^3)) partial^2(p^2)/partial tau^2; the third-order derivative makes the term dimensionally inconsistent with the left-hand side (it introduces an extra factor of 1/s) and does not correspond to the canonical KZK equation. Since the paper explicitly uses the KZK equation as part of its theoretical foundation and references it in the authors' own framework [44], this is a technical error in a load-bearing equation and should be corrected.
  4. [3.2 and 4] The citation mismatches are not confined to Section 2. In Section 3.2.1, [50] (MUGEN) does not support the claim about differentiable physics simulation; in Section 3.2.2, [52] is an audio-visual anomaly-detection paper, not the Neural Acoustic Fields approach described; and in Section 3.2.3, [55] and [56] are, respectively, a heart-sound review and a COVID chest-X-ray study, not controllable sound synthesis. Similar mismatches recur in Section 4 (for example, [32] and [33] are not tire-road noise papers, and [34] and [35] are not vehicle health-monitoring papers). Because these are the papers that anchor the methodological pillars and application claims, the survey's map of the literature is unreliable as it stands; a systematic reference audit and re-citation is needed.
minor comments (5)
  1. [1] The Introduction's phrase 'deeply rooted in principles of auditory scene analysis [11, 12]' cites [12] as Computational Physics (Vesely), which is not an auditory scene analysis reference; the citation should be corrected.
  2. [6 and Appendix B] Section 6, 'Next Steps in Our Research', is unusual for a survey and functions as self-promotion of the authors' own paper [44] and the company-affiliated GitHub repository in Appendix B; if these are kept, they should be clearly separated from the survey content and explicitly identified as author-affiliated resources.
  3. [4] The Section 4 heading reads 'Deploying Acoustic World Modes in Real World'; 'Modes' should be 'Models'.
  4. [Throughout] Several grammatical errors remain (e.g., 'a acoustic world model', 'a important'), and the manuscript would benefit from a careful language edit.
  5. [6] The 'Limitations in Our Survey' passage appears after the Conclusions; integrating it before the Conclusions would improve the paper's structure.

Circularity Check

0 steps flagged · score 2.0 of 10

No circular derivation: the survey's central content is independent of the authors' self-citations; only minor self-promotional self-citation appears in Section 6 and Appendix B.

full rationale

This paper is a survey rather than a derivation: it proposes no fitted parameter that is later renamed a prediction, imports no uniqueness theorem from self-citations, and does not define its central concept in terms of its conclusion. The closest issue is self-referential promotion: Section 6 ('Next Steps in Our Research') highlights the authors' own [44] as a practical foundation, [44] is also cited in Section 2.4 for KZK applications and in Section 4.1 for human-robot collaboration, and Appendix B points to a company-affiliated GitHub repository. These are self-citations, but they are not load-bearing for the central claim that acoustic signals encode physical information and that PINNs, generative models, and self-supervised learning can exploit that information. That claim rests on standard acoustics and a broad external literature. Separately, the skeptical concerns about misassigned citations in Sections 2.1-2.2 and the dimensional inconsistency in Eq. (9) are important evidence-quality and correctness issues, but they are not circularity under the rubric: a wrong equation or a poorly matched citation does not make the argument equivalent to its inputs. No step in the paper exhibits the required reduction-by-construction, so the appropriate finding is a low score reflecting only minor self-promotional self-citation.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The survey introduces no free parameters and no invented entities; it is a review. Its central claims rest on standard physics equations (elastodynamics, Lighthill's analogy, KZK) and on domain assumptions about their validity in real-world settings. The main burden is the accuracy of the surveyed literature: several citations are mismatched, which weakens the evidence base for the specific capabilities claimed.

assumptions (3)
  • domain assumption Linear, isotropic, homogeneous elastodynamics with small deformations governs the vibration of struck objects (Section 2.1, Eq. 1).
    The survey uses this to argue that natural frequencies encode material properties, but real objects often violate isotropy, homogeneity, and small-strain assumptions.
  • domain assumption Lighthill's acoustic analogy and the FW-H equation adequately describe the sound generated by flows and moving surfaces (Section 2.2, Eq. 3).
    These are standard results, but their applicability to real-world leak and flow monitoring depends on conditions the survey does not discuss.
  • domain assumption Sabine and Eyring reverberation-time formulas are valid for the rooms discussed (Section 2.3, Eqs. 5 and 6).
    The formulas assume diffuse sound fields; the survey applies them without stating the diffuse-field limitation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Survey on World Models Grounded in Acoustic Physical Information." pith.science (2026). https://pith.science/paper/Q2MWQPKQ

@misc{pith2026250613833,
  author       = {Pith},
  title        = {Pith review of: A Survey on World Models Grounded in Acoustic Physical Information},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/Q2MWQPKQ}},
  note         = {Machine review of arXiv:2506.13833}
}
read the original abstract

This survey provides a comprehensive overview of the emerging field of world models grounded in the foundation of acoustic physical information. It examines the theoretical underpinnings, essential methodological frameworks, and recent technological advancements in leveraging acoustic signals for high-fidelity environmental perception, causal physical reasoning, and predictive simulation of dynamic events. The survey explains how acoustic signals, as direct carriers of mechanical wave energy from physical events, encode rich, latent information about material properties, internal geometric structures, and complex interaction dynamics. Specifically, this survey establishes the theoretical foundation by explaining how fundamental physical laws govern the encoding of physical information within acoustic signals. It then reviews the core methodological pillars, including Physics-Informed Neural Networks (PINNs), generative models, and self-supervised multimodal learning frameworks. Furthermore, the survey details the significant applications of acoustic world models in robotics, autonomous driving, healthcare, and finance. Finally, it systematically outlines the important technical and ethical challenges while proposing a concrete roadmap for future research directions toward robust, causal, uncertainty-aware, and responsible acoustic intelligence. These elements collectively point to a research pathway towards embodied active acoustic intelligence, empowering AI systems to construct an internal "intuitive physics" engine through sound.

Figures

Figures reproduced from arXiv: 2506.13833 by the authors.

Figure 1
Figure 1. A conceptual diagram illustrating the methodological framework for acoustic world models. [PITH_FULL_IMAGE:figures/full_fig_p010_1.png] view at source ↗
Figure 2
Figure 2. A proposed research roadmap for the next generation of acoustic world models. [PITH_FULL_IMAGE:figures/full_fig_p019_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

93 extracted references · 63 canonical work pages

  1. [16]

    Ultrasonic non-destructive evaluation of composites: A review,

    J. Jodhani, A. Handa, A. Gautam, R. Ranaet al., “Ultrasonic non-destructive evaluation of composites: A review,”Materials Today: Proceedings, vol. 78, pp. 627–632, 2023. 23

  2. [17]

    Porosity evaluation of additively manufactured components using deep learning-based ultrasonic nondestructive testing,

    S.-H. Park, S. Choi, and K.-Y. Jhang, “Porosity evaluation of additively manufactured components using deep learning-based ultrasonic nondestructive testing,”International Journal of Precision Engineering and Manufacturing-Green Technology, vol. 9, no. 2, pp. 395–407, 2022

  3. [18]

    Deep learning approach for damage classification based on acoustic emission data in composite materials,

    F. Guo, W. Li, P. Jiang, F. Chen, and Y. Liu, “Deep learning approach for damage classification based on acoustic emission data in composite materials,”Materials, vol. 15, no. 12, p. 4270, 2022

  4. [12]

    F. J. Veselyet al.,Computational Physics. Springer, 1994

  5. [13]

    Statistics of natural reverberation enable perceptual separation of sound and space,

    J. Traer and J. H. McDermott, “Statistics of natural reverberation enable perceptual separation of sound and space,”Proceedings of the National Academy of Sciences, vol. 113, no. 48, pp. E7856–E7865, 2016

  6. [14]

    Material categorization and hardness scaling in real and synthetic impact sounds,

    B. L. Giordano, “Material categorization and hardness scaling in real and synthetic impact sounds,” The sounding object, pp. 73–94, 2003

  7. [34]

    A review on sound source localization systems,

    D. Desai and N. Mehendale, “A review on sound source localization systems,”Archives of Computational Methods in Engineering, vol. 29, no. 7, pp. 4631–4642, 2022

  8. [35]

    Acoustic slam,

    C. Evers and P. A. Naylor, “Acoustic slam,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 26, no. 9, pp. 1484–1498, 2018. 24

  9. [36]

    A survey of underwater acoustic slam system,

    M. Jiang, S. Song, Y. Li, W. Jin, J. Liu, and X. Feng, “A survey of underwater acoustic slam system,” inIntelligent Robotics and Applications: 12th International Conference, ICIRA 2019, Shenyang, China, August 8–11, 2019, Proceedings, Part II 12. Springer, 2019, pp. 159–170

  10. [27]

    R. B. Randall,Vibration-based condition monitoring: industrial, automotive and aerospace applications. John Wiley & Sons, 2021

  11. [44]

    A synergistic framework of nonlinear acoustic computing and reinforcement learning for real-world human- robot interaction,

    X. Chen, X. Yu, L. Chang, Y. Huang, J. He, S. Zhang, J. Li, L. Lin, Z. Zeng, X. Tuet al., “A synergistic framework of nonlinear acoustic computing and reinforcement learning for real-world human- robot interaction,”arXiv preprint arXiv:2505.01998, 2025

  12. [50]

    Mugen: A playground for video-audio-text multimodal understanding and generation,

    T. Hayes, S. Zhang, X. Yin, G. Pang, S. Sheng, H. Yang, S. Ge, Q. Hu, and D. Parikh, “Mugen: A playground for video-audio-text multimodal understanding and generation,” inEuropean Conference on Computer Vision. Springer, 2022, pp. 431–449

  13. [52]

    Audio–visualrepresentationlearningforanomalyeventsdetection in crowds,

    J.Gao, H.Yang, M.Gong, andX.Li, “Audio–visualrepresentationlearningforanomalyeventsdetection in crowds,”Neurocomputing, vol. 582, p. 127489, 2024

  14. [55]

    Algorithms for automatic analysis and classi- fication of heart sounds–a systematic review,

    A. K. Dwivedi, S. A. Imtiaz, and E. Rodriguez-Villegas, “Algorithms for automatic analysis and classi- fication of heart sounds–a systematic review,”IEEE Access, vol. 7, pp. 8316–8345, 2018

  15. [56]

    Deep learning approaches for covid-19 detection based on chest x-ray images,

    A. M. Ismael and A. Şengür, “Deep learning approaches for covid-19 detection based on chest x-ray images,”Expert Systems with Applications, vol. 164, p. 114054, 2021

  16. [32]

    Deep room recognition using inaudible echos,

    Q. Song, C. Gu, and R. Tan, “Deep room recognition using inaudible echos,”Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, vol. 2, no. 3, pp. 1–28, 2018

  17. [33]

    Springer Science & Business Media, 2001

    M.BrandsteinandD.Ward,Microphone arrays: signal processing techniques and applications. Springer Science & Business Media, 2001

Show all 93 references
  1. [1]

    Language models are few-shot learners,

    T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askellet al., “Language models are few-shot learners,”Advances in neural information processing systems, vol. 33, pp. 1877–1901, 2020

  2. [2]

    Open foundation and fine-tuned chat models,

    H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaeiet al., “Open foundation and fine-tuned chat models,”arXiv preprint arXiv:2307.09288, 2023

  3. [3]

    Climbing towards nlu: On meaning, form, and understanding in the age of data,

    E. M. Bender and A. Koller, “Climbing towards nlu: On meaning, form, and understanding in the age of data,” inProceedings of the 58th annual meeting of the association for computational linguistics, 2020, pp. 5185–5198

  4. [4]

    World models,

    D. Ha and J. Schmidhuber, “World models,”arXiv preprint arXiv:1803.10122, 2018

  5. [5]

    Dream to control: Learning behaviors by latent imagination,

    D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi, “Dream to control: Learning behaviors by latent imagination,”arXiv preprint arXiv:1912.01603, 2019

  6. [6]

    A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27,

    Y. LeCun, “A path towards autonomous machine intelligence version 0.9. 2, 2022-06-27,”Open Review, vol. 62, no. 1, pp. 1–62, 2022

  7. [7]

    Deep learning, reinforcement learning, and world models,

    Y. Matsuo, Y. LeCun, M. Sahani, D. Precup, D. Silver, M. Sugiyama, E. Uchibe, and J. Morimoto, “Deep learning, reinforcement learning, and world models,”Neural Networks, vol. 152, pp. 267–275, 2022

  8. [8]

    Imagenet classification with deep convolutional neural networks,

    A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,”Communications of the ACM, vol. 60, no. 6, pp. 84–90, 2017

  9. [9]

    An image is worth 16x16 words: Transformers for image recognition at scale,

    A.Dosovitskiy, L.Beyer, A.Kolesnikov, D.Weissenborn, X.Zhai, T.Unterthiner, M.Dehghani, M.Min- derer, G. Heigold, S. Gellyet al., “An image is worth 16x16 words: Transformers for image recognition at scale,”arXiv preprint arXiv:2010.11929, 2020

  10. [10]

    Benchmarking neural network robustness to common corruptions and perturbations,

    D. Hendrycks and T. Dietterich, “Benchmarking neural network robustness to common corruptions and perturbations,”arXiv preprint arXiv:1903.12261, 2019

  11. [11]

    A. S. Bregman,Auditory scene analysis: The perceptual organization of sound. MIT press, 1994

  12. [15]

    Resonant ultrasound spectroscopy: applications, current status and limitations,

    R. Schwarz and J. Vuorinen, “Resonant ultrasound spectroscopy: applications, current status and limitations,”Journal of Alloys and Compounds, vol. 310, no. 1-2, pp. 243–250, 2000

  13. [19]

    Sound production and modeling,

    P. R. Cook, “Sound production and modeling,”IEEE Computer Graphics and applications, vol. 22, no. 4, pp. 23–27, 2002

  14. [20]

    Controlling material properties in physical models of sounding ob- jects,

    F. Avanzini, D. Rocchessoet al., “Controlling material properties in physical models of sounding ob- jects,” inICMC, 2001

  15. [21]

    Auditory scene analysis: Hearing in complex environments,

    A. S. Bregnian, “Auditory scene analysis: Hearing in complex environments,” 1993

  16. [22]

    Multimodal human–robot interaction for human-centric smart manufacturing: a survey,

    T. Wang, P. Zheng, S. Li, and L. Wang, “Multimodal human–robot interaction for human-centric smart manufacturing: a survey,”Advanced Intelligent Systems, vol. 6, no. 3, p. 2300359, 2024

  17. [23]

    Robot sound interpretation: Combining sight and sound in learning-based control,

    P. Chang, S. Liu, H. Chen, and K. Driggs-Campbell, “Robot sound interpretation: Combining sight and sound in learning-based control,” in2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2020, pp. 5580–5587

  18. [24]

    Leak detection in real water distribution networks based on acoustic emission and machine learning,

    A. Fares, I. Tijani, Z. Rui, and T. Zayed, “Leak detection in real water distribution networks based on acoustic emission and machine learning,”Environmental Technology, vol. 44, no. 25, pp. 3850–3866, 2023

  19. [25]

    Estimating rainfall intensity based on surveillance audio and deep-learning,

    M. Wang, M. Chen, Z. Wang, Y. Guo, Y. Wu, W. Zhao, and X. Liu, “Estimating rainfall intensity based on surveillance audio and deep-learning,”Environmental Science and Ecotechnology, vol. 22, p. 100450, 2024

  20. [26]

    Road type classificationusingdeeplearningfortire-pavementinteractionnoisedatainautonomousdrivingvehicle,

    S.-K. Lee, J. Yoo, C.-H. Lee, K. An, Y.-S. Yoon, J. Lee, G.-H. Yeom, and S.-U. Hwang, “Road type classificationusingdeeplearningfortire-pavementinteractionnoisedatainautonomousdrivingvehicle,” Applied Acoustics, vol. 212, p. 109597, 2023

  21. [28]

    Onsoundgeneratedaerodynamicallyi.generaltheory,

    M.J.Lighthill, “Onsoundgeneratedaerodynamicallyi.generaltheory,”Proceedings of the Royal Society of London. Series A. Mathematical and Physical Sciences, vol. 211, no. 1107, pp. 564–587, 1952

  22. [29]

    A novel deep autoencoder feature learning method for rotating machinery fault diagnosis,

    H. Shao, H. Jiang, H. Zhao, and F. Wang, “A novel deep autoencoder feature learning method for rotating machinery fault diagnosis,”Mechanical Systems and Signal Processing, vol. 95, pp. 187–204, 2017

  23. [30]

    A review of data-driven machinery fault diagnosis using machine learning algorithms,

    J. Cen, Z. Yang, X. Liu, J. Xiong, and H. Chen, “A review of data-driven machinery fault diagnosis using machine learning algorithms,”Journal of Vibration Engineering & Technologies, vol. 10, no. 7, pp. 2481–2507, 2022

  24. [31]

    Kuttruff and M

    H. Kuttruff and M. Vorländer,Room acoustics. Crc Press, 2024

  25. [37]

    Acoustic echoes reveal room shape,

    I. Dokmanić, R. Parhizkar, A. Walther, Y. M. Lu, and M. Vetterli, “Acoustic echoes reveal room shape,” Proceedings of the National Academy of Sciences, vol. 110, no. 30, pp. 12186–12191, 2013

  26. [38]

    Batslam: Simultaneous localization and mapping using biomimetic sonar,

    J. Steckel and H. Peremans, “Batslam: Simultaneous localization and mapping using biomimetic sonar,” PloS one, vol. 8, no. 1, p. e54076, 2013

  27. [39]

    Neural representation of three-dimensional acoustic space in the human temporal lobe,

    X. Zhang, Q. Zhang, X. Hu, and B. Zhang, “Neural representation of three-dimensional acoustic space in the human temporal lobe,”Frontiers in Human Neuroscience, vol. 9, p. 203, 2015

  28. [40]

    Overview of geometrical room acoustic modeling techniques,

    L. Savioja and U. P. Svensson, “Overview of geometrical room acoustic modeling techniques,”The Journal of the Acoustical Society of America, vol. 138, no. 2, pp. 708–730, 2015

  29. [41]

    A system for data-driven concatenative sound synthesis,

    D. Schwarz, “A system for data-driven concatenative sound synthesis,” inDigital Audio Effects (DAFx), 2000, pp. 97–102

  30. [42]

    Pertilä,Acoustic source localization in a room environment and at moderate distances

    P. Pertilä,Acoustic source localization in a room environment and at moderate distances. Tampere University of Technology, 2009

  31. [43]

    Nonlinear acoustics in china,

    Q. Zuwen, “Nonlinear acoustics in china,”WULI-BEIJING-, vol. 28, no. 10, pp. 593–599, 1999

  32. [45]

    Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations,

    M. Raissi, P. Perdikaris, and G. E. Karniadakis, “Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations,” Journal of Computational physics, vol. 378, pp. 686–707, 2019

  33. [46]

    Physics-informed machine learning,

    G. E. Karniadakis, I. G. Kevrekidis, L. Lu, P. Perdikaris, S. Wang, and L. Yang, “Physics-informed machine learning,”Nature Reviews Physics, vol. 3, no. 6, pp. 422–440, 2021

  34. [47]

    Physics-informed deep learning for structural vibration identification and its application on a benchmark structure,

    M. Zhang, T. Guo, G. Zhang, Z. Liu, and W. Xu, “Physics-informed deep learning for structural vibration identification and its application on a benchmark structure,”Philosophical Transactions of the Royal Society A, vol. 382, no. 2264, p. 20220400, 2024

  35. [48]

    A study on the auralization system using flexible rendering for virtual reality-based acoustic simulation,

    Y. Lee and J. Ryu, “A study on the auralization system using flexible rendering for virtual reality-based acoustic simulation,”Journal of Advanced Mechanical Design, Systems, and Manufacturing, vol. 13, no. 5, pp. JAMDSM0094–JAMDSM0094, 2019

  36. [49]

    Physics- informed neural network for volumetric sound field reconstruction of speech signals,

    M. Olivieri, X. Karakonstantis, M. Pezzoli, F. Antonacci, A. Sarti, and E. Fernandez-Grande, “Physics- informed neural network for volumetric sound field reconstruction of speech signals,”EURASIP Journal on Audio, Speech, and Music Processing, vol. 2024, no. 1, p. 42, 2024

  37. [51]

    Nerf: Repre- senting scenes as neural radiance fields for view synthesis,

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “Nerf: Repre- senting scenes as neural radiance fields for view synthesis,”Communications of the ACM, vol. 65, no. 1, pp. 99–106, 2021

  38. [53]

    Equivariant neural rendering,

    E. Dupont, M. B. Martin, A. Colburn, A. Sankar, J. Susskind, and Q. Shan, “Equivariant neural rendering,” inInternational Conference on Machine Learning. PMLR, 2020, pp. 2761–2770. 25

  39. [54]

    Diffwave: A versatile diffusion model for audio synthesis,

    Z. Kong, W. Ping, J. Huang, K. Zhao, and B. Catanzaro, “Diffwave: A versatile diffusion model for audio synthesis,”arXiv preprint arXiv:2009.09761, 2020

  40. [57]

    Look, listen and learn,

    R. Arandjelovic and A. Zisserman, “Look, listen and learn,” inProceedings of the IEEE international conference on computer vision, 2017, pp. 609–617

  41. [58]

    Audio-visual scene analysis with self-supervised multisensory features,

    A. Owens and A. A. Efros, “Audio-visual scene analysis with self-supervised multisensory features,” in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 631–648

  42. [59]

    Learning transferable visual models from natural language supervision,

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clarket al., “Learning transferable visual models from natural language supervision,” inInternational conference on machine learning. PMLR, 2021, pp. 8748–8763

  43. [60]

    Flamingo: a visual language model for few-shot learning,

    J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynoldset al., “Flamingo: a visual language model for few-shot learning,”Advances in neural information processing systems, vol. 35, pp. 23716–23736, 2022

  44. [61]

    Learning to localize sound source in visual scenes,

    A. Senocak, T.-H. Oh, J. Kim, M.-H. Yang, and I. S. Kweon, “Learning to localize sound source in visual scenes,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 4358–4366

  45. [62]

    Co-separating sounds of visual objects,

    R. Gao and K. Grauman, “Co-separating sounds of visual objects,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 3879–3888

  46. [63]

    Object category recognition by a humanoid robot using behavior- grounded relational learning,

    J. Sinapov and A. Stoytchev, “Object category recognition by a humanoid robot using behavior- grounded relational learning,” in2011 IEEE International Conference on Robotics and Automation. IEEE, 2011, pp. 184–190

  47. [64]

    Communication in human-robot interaction,

    A. Bonarini, “Communication in human-robot interaction,”Current Robotics Reports, vol. 1, no. 4, pp. 279–285, 2020

  48. [65]

    A comprehensive review of polyphonic sound event detection,

    T. K. Chan and C. S. Chin, “A comprehensive review of polyphonic sound event detection,”IEEE Access, vol. 8, pp. 103339–103373, 2020

  49. [66]

    Acoustic based emergency vehicle detection using ensemble of deep learning models,

    U. Mittal and P. Chawla, “Acoustic based emergency vehicle detection using ensemble of deep learning models,”Procedia Computer Science, vol. 218, pp. 227–234, 2023

  50. [67]

    Audio surveillance: A systematic review,

    M. Crocco, M. Cristani, A. Trucco, and V. Murino, “Audio surveillance: A systematic review,”ACM Computing Surveys (CSUR), vol. 48, no. 4, pp. 1–46, 2016

  51. [68]

    A review of depression and suicide risk assessment using speech analysis,

    N. Cummins, S. Scherer, J. Krajewski, S. Schnieder, J. Epps, and T. F. Quatieri, “A review of depression and suicide risk assessment using speech analysis,”Speech communication, vol. 71, pp. 10–49, 2015

  52. [69]

    Covid-19 artificial intelligence diagnosis using only cough recordings,

    J. Laguarta, F. Hueto, and B. Subirana, “Covid-19 artificial intelligence diagnosis using only cough recordings,”IEEE Open Journal of Engineering in Medicine and Biology, vol. 1, pp. 275–281, 2020

  53. [70]

    Adaptive, personalized closed-loop therapy for parkinson’s disease: Biochemical, neurophysiological, and wearable sensing systems,

    L. di Biase, G. Tinkhauser, E. Martin Moraud, M. L. Caminiti, P. M. Pecoraro, and V. Di Lazzaro, “Adaptive, personalized closed-loop therapy for parkinson’s disease: Biochemical, neurophysiological, and wearable sensing systems,”Expert review of neurotherapeutics, vol. 21, no....

  54. [71]

    Emotion recognition from speech: a review,

    S. G. Koolagudi and K. S. Rao, “Emotion recognition from speech: a review,”International journal of speech technology, vol. 15, pp. 99–117, 2012

  55. [72]

    A comprehensive review of speech emotion recognition systems,

    T. M. Wani, T. S. Gunawan, S. A. A. Qadri, M. Kartiwi, and E. Ambikairajah, “A comprehensive review of speech emotion recognition systems,”IEEE access, vol. 9, pp. 47795–47814, 2021. 26

  56. [73]

    Multimodal machine learning: A survey and taxonomy,

    T. Baltrušaitis, C. Ahuja, and L.-P. Morency, “Multimodal machine learning: A survey and taxonomy,” IEEE transactions on pattern analysis and machine intelligence, vol. 41, no. 2, pp. 423–443, 2018

  57. [74]

    The power of voice: Managerial affective states and future firm performance,

    W. J. Mayew and M. Venkatachalam, “The power of voice: Managerial affective states and future firm performance,”The Journal of Finance, vol. 67, no. 1, pp. 1–43, 2012

  58. [75]

    Analyzingspeechtodetectfinancialmisreporting,

    J.L.Hobson, W.J.Mayew, andM.Venkatachalam, “Analyzingspeechtodetectfinancialmisreporting,” Journal of Accounting Research, vol. 50, no. 2, pp. 349–392, 2012

  59. [76]

    Predicting user satisfaction from turn-taking in spoken conversations

    S. A. Chowdhury, E. A. Stepanov, G. Riccardiet al., “Predicting user satisfaction from turn-taking in spoken conversations.” inInterspeech, 2016, pp. 2910–2914

  60. [77]

    Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding,

    S. Han, H. Mao, and W. J. Dally, “Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding,”arXiv preprint arXiv:1510.00149, 2015

  61. [78]

    Supervisedspeechseparationbasedondeeplearning: Anoverview,

    D.WangandJ.Chen, “Supervisedspeechseparationbasedondeeplearning: Anoverview,”IEEE/ACM transactions on audio, speech, and language processing, vol. 26, no. 10, pp. 1702–1726, 2018

  62. [79]

    Federated learning: Strategies for improving communication efficiency,

    J. Konečn` y, H. B. McMahan, F. X. Yu, P. Richtárik, A. T. Suresh, and D. Bacon, “Federated learning: Strategies for improving communication efficiency,”arXiv preprint arXiv:1610.05492, 2016

  63. [80]

    Privacy-preserving human activity sensing: A survey,

    Y. Yang, P. Hu, J. Shen, H. Cheng, Z. An, and X. Liu, “Privacy-preserving human activity sensing: A survey,”High-Confidence Computing, vol. 4, no. 1, p. 100204, 2024

  64. [81]

    Transfer learning from speaker verification to multispeaker text-to-speech synthesis,

    Y. Jia, Y. Zhang, R. Weiss, Q. Wang, J. Shen, F. Ren, P. Nguyen, R. Pang, I. Lopez Moreno, Y. Wu et al., “Transfer learning from speaker verification to multispeaker text-to-speech synthesis,”Advances in neural information processing systems, vol. 31, 2018

  65. [82]

    Explaining deep neural networks and beyond: A review of methods and applications,

    W. Samek, G. Montavon, S. Lapuschkin, C. J. Anders, and K.-R. Müller, “Explaining deep neural networks and beyond: A review of methods and applications,”Proceedings of the IEEE, vol. 109, no. 3, pp. 247–278, 2021

  66. [83]

    Audio set: An ontology and human-labeled dataset for audio events,

    J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Rit- ter, “Audio set: An ontology and human-labeled dataset for audio events,” in2017 IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE, 2017, ...

  67. [84]

    Fsd50k: an open dataset of human-labeled sound events,

    E. Fonseca, X. Favory, J. Pons, F. Font, and X. Serra, “Fsd50k: an open dataset of human-labeled sound events,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 30, pp. 829–852, 2021

  68. [85]

    Wham!: Extending speech separation to noisy environments,

    G. Wichern, J. Antognini, M. Flynn, L. R. Zhu, E. McQuinn, D. Crow, E. Manilow, and J. L. Roux, “Wham!: Extending speech separation to noisy environments,”arXiv preprint arXiv:1907.01160, 2019

  69. [86]

    Tut database for acoustic scene classification and sound event detection,

    A. Mesaros, T. Heittola, and T. Virtanen, “Tut database for acoustic scene classification and sound event detection,” in2016 24th European Signal Processing Conference (EUSIPCO). IEEE, 2016, pp. 1128–1132

  70. [87]

    Habitat: A platform for embodied ai research,

    M. Savva, A. Kadian, O. Maksymets, Y. Zhao, E. Wijmans, B. Jain, J. Straub, J. Liu, V. Koltun, J. Maliket al., “Habitat: A platform for embodied ai research,” inProceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 9339–9347

  71. [88]

    Soundspaces: Audio-visual navigation in 3d environments,

    C. Chen, U. Jain, C. Schissler, S. V. A. Gari, Z. Al-Halah, V. K. Ithapu, P. Robinson, and K. Grauman, “Soundspaces: Audio-visual navigation in 3d environments,” inComputer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part VI 16. Sp...

  72. [89]

    Isaac gym: High performance gpu-based physics simulation for robot learning,

    V. Makoviychuk, L. Wawrzyniak, Y. Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. All- shire, A. Handaet al., “Isaac gym: High performance gpu-based physics simulation for robot learning,” arXiv preprint arXiv:2108.10470, 2021. 27

  73. [90]

    Threedworld: A platform for interactive multi-modal physical simulation,

    C. Gan, J. Schwartz, S. Alter, D. Mrowca, M. Schrimpf, J. Traer, J. De Freitas, J. Kubilius, A. Bhand- waldar, N. Haberet al., “Threedworld: A platform for interactive multi-modal physical simulation,” arXiv preprint arXiv:2007.04954, 2020

  74. [91]

    Sapien: A simulatedpart-basedinteractiveenvironment,

    F. Xiang, Y. Qin, K. Mo, Y. Xia, H. Zhu, F. Liu, M. Liu, H. Jiang, Y. Yuan, H. Wanget al., “Sapien: A simulatedpart-basedinteractiveenvironment,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 11097–11107

  75. [92]

    Adecadeofdcase: Achievements, practices, evaluations and future challenges,

    A.Mesaros, R.Serizel, T.Heittola, T.Virtanen, andM.D.Plumbley, “Adecadeofdcase: Achievements, practices, evaluations and future challenges,” inICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025, pp. 1–5

  76. [93]

    Hear: Holistic evaluation of audio representations,

    J. Turian, J. Shier, H. R. Khan, B. Raj, B. W. Schuller, C. J. Steinmetz, C. Malloy, G. Tzanetakis, G. Velarde, K. McNallyet al., “Hear: Holistic evaluation of audio representations,” inNeurIPS 2021 Competitions and Demonstrations Track. PMLR, 2022, pp. 125–145. 28

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.