{"id":"2636a2de-2838-4f74-8816-0c1f0ad84c97","arxiv_id":"2506.13833","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey that frames acoustic signals as physical information for world models and reviews PINNs, generative models, and self-supervised learning toward that end.","lead":"This survey reviews and organizes the emerging field of 'acoustic world models', where AI systems use sound as physical evidence to perceive, reason about, and predict their environment. It is a broad literature review with a proposed research roadmap, not a new experimental result.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim's evidence base fails: §2.1 and §2.2 cite unrelated works for material/viscosity/fault-inference capabilities, and Eq. (9)'s nonlinear term is dimensionally wrong.","rationale":"The reader's weakest assumption identified the same load-bearing concern: the survey's evidence base must actually substantiate the capabilities it attributes to cited works. My analysis finds concrete confirming instances, so the concern lands. The central thesis is plausible and there is a real body of acoustic-perception research, but the survey as written does not yet reliably map it: §2.1 and §2.2 attach experimental results to references that report different tasks, and §2.4 contains a dimensionally inconsistent version of the KZK equation. Because this is a survey, its primary claim is about what the field has shown; unreliable citation support directly weakens that claim. The errors are correctable, however, and do not negate the underlying idea, so a conditional acceptance with mandatory citation audit and equation fix is the appropriate outcome. The reader's CONDITIONAL verdict stands; no adjustment is needed.","tokens_in":23508,"tokens_out":7367,"duration_ms":70809,"concrete_test":"Audit every citation in §2.1 and §2.2: for each claimed capability (material classification from impact sound, liquid-viscosity estimation from container damping, granular particle-size inference from shaking, cyclostationary fault detection), read the cited paper and record whether it actually reports that result. In particular, check [16-18] for the material/viscosity/particle-size claims and [34-36] for the cyclostationary paragraph. If the stated experiments are absent, the survey's evidence base is unsupported and the text must be revised or the claims qualified. Separately, re-derive Eq. (9) from the standard KZK equation (e.g., from the authors' own reference [44] or a nonlinear acoustics textbook) and confirm the order of the τ-derivative on the p² term; if the term is ∂³(p²)/∂τ³ rather than ∂²(p²)/∂τ², the equation is dimensionally incorrect and §2.4 needs a correction.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The survey's central thesis is that acoustic signals are a privileged modality for physical world models because the literature demonstrates acoustic inference of material properties, fluid dynamics, and geometry. That demonstration is not reliable as written. In §2.1, the sentence claiming deep learning models 'classify object materials with high accuracy [16], estimate the viscosity of liquids [17], or infer the particle size of granular materials [18]' cites papers on ultrasonic NDE of composites, porosity evaluation of additively manufactured parts, and acoustic-emission damage classification; none of these report the stated impact-sound or shaking-sound inference tasks. In §2.2, the cyclostationary machinery-health paragraph cites sound-source-localization and acoustic-SLAM surveys [34-36], not cyclostationary fault-diagnosis work. Separately, Eq. (9) writes the KZK nonlinear term as ∂³(p²)/∂τ³; this is dimensionally inconsistent with the left-hand side by a factor of 1/s, and the standard KZK/Westervelt nonlinearity is ∂²(p²)/∂τ². These are not isolated typos: they indicate that the theoretical and empirical scaffolding for the field map has not been checked. If the cited works do not substantiate the assigned capabilities, the survey's map of acoustic world models—the object under review—is misleading and overstates the strength of the evidence for the central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of 'world models grounded in acoustic physical information.' Its central thesis is that acoustic signals, as direct carriers of mechanical wave energy from physical events, encode latent information about material properties, internal structure, geometry, and interaction dynamics, and that this makes sound a privileged modality for AI systems that must build causal, physics-grounded world models. The paper reviews the physical theory (elastodynamics, aeroacoustics, room acoustics, nonlinear acoustics), organizes methodological work into three pillars (physics-informed neural networks, generative forward models, and self-supervised multimodal learning), surveys applications in robotics, autonomous driving, healthcare, and finance, and closes with an ethics discussion and a five-pathway research roadmap.","tokens_in":23686,"tokens_out":12922,"duration_ms":124782,"significance":"If its map of the literature were reliable, this survey would provide a useful synthesis of an emerging interdisciplinary area and a concrete agenda for future work. The paper is clearly organized, and the comparative analysis in Table 1, the application taxonomy in Table 2, the explicit limitations statement, and the resource list in Appendix A are useful editorial features. However, the survey's central claim depends on the cited literature substantiating specific acoustic-inference capabilities, and many load-bearing citations do not support the claims they are attached to (Sections 2.1, 2.2, 3.2, and 4). In addition, Eq. (9) contains a dimensional inconsistency in the KZK nonlinear term. Because a survey's contribution is precisely its map of the evidence base, these problems materially weaken the paper. They are, however, correctable within the manuscript's scope.","major_comments":[{"comment":"The sentence 'Recent works have demonstrated that deep learning models can leverage these rich acoustic signatures to classify object materials with high accuracy [16], estimate the viscosity of liquids ... [17], or even infer the particle size of granular materials from the statistical properties of many small impacts during shaking [18]' is not supported by the cited papers: [16] is a review of ultrasonic non-destructive evaluation of composites, [17] is a study of porosity evaluation of additively manufactured parts using ultrasonic NDT, and [18] is a study of acoustic-emission damage classification in composites. None of these works reports impact-sound material classification, liquid-viscosity estimation, or granular particle-size inference. The same paragraph's earlier claim that material identification from sound rests on a 'large body of work' also cites [13], which is a study of natural reverberation statistics, alongside the relevant [14]. This passage is load-bearing for the survey's central thesis that acoustic signals encode material and fluid properties, so it should be rewritten with correct references or the claims should be removed.","section":"2.1"},{"comment":"The cyclostationary-analysis paragraph states that deep learning models trained on spectral correlation and envelope analysis 'can be trained on these processed representations to perform highly sensitive fault diagnosis and prognostics, effectively creating a detailed acoustic physical model of a machine's health state [34–36]'. References [34], [35], and [36] are, respectively, a review of sound source localization, an acoustic SLAM paper, and a survey of underwater acoustic SLAM; none of them addresses cyclostationary machinery fault diagnosis. This leaves the paragraph's specific technical claim without supporting evidence. Please replace these citations with the actual cyclostationary/condition-monitoring literature (for example, the Randall monograph already in the reference list as [27]) or delete the unsupported claim.","section":"2.2"},{"comment":"Equation (9) and its axisymmetric counterpart Eq. (10) write the nonlinear term of the KZK equation as (beta_NL/(2 rho0 c^3)) partial^3(p^2)/partial tau^3. The standard KZK/Westervelt nonlinearity is (beta/(2 rho0 c^3)) partial^2(p^2)/partial tau^2; the third-order derivative makes the term dimensionally inconsistent with the left-hand side (it introduces an extra factor of 1/s) and does not correspond to the canonical KZK equation. Since the paper explicitly uses the KZK equation as part of its theoretical foundation and references it in the authors' own framework [44], this is a technical error in a load-bearing equation and should be corrected.","section":"2.4, Eqs. (9)-(10)"},{"comment":"The citation mismatches are not confined to Section 2. In Section 3.2.1, [50] (MUGEN) does not support the claim about differentiable physics simulation; in Section 3.2.2, [52] is an audio-visual anomaly-detection paper, not the Neural Acoustic Fields approach described; and in Section 3.2.3, [55] and [56] are, respectively, a heart-sound review and a COVID chest-X-ray study, not controllable sound synthesis. Similar mismatches recur in Section 4 (for example, [32] and [33] are not tire-road noise papers, and [34] and [35] are not vehicle health-monitoring papers). Because these are the papers that anchor the methodological pillars and application claims, the survey's map of the literature is unreliable as it stands; a systematic reference audit and re-citation is needed.","section":"3.2 and 4"}],"minor_comments":[{"comment":"The Introduction's phrase 'deeply rooted in principles of auditory scene analysis [11, 12]' cites [12] as Computational Physics (Vesely), which is not an auditory scene analysis reference; the citation should be corrected.","section":"1"},{"comment":"Section 6, 'Next Steps in Our Research', is unusual for a survey and functions as self-promotion of the authors' own paper [44] and the company-affiliated GitHub repository in Appendix B; if these are kept, they should be clearly separated from the survey content and explicitly identified as author-affiliated resources.","section":"6 and Appendix B"},{"comment":"The Section 4 heading reads 'Deploying Acoustic World Modes in Real World'; 'Modes' should be 'Models'.","section":"4"},{"comment":"Several grammatical errors remain (e.g., 'a acoustic world model', 'a important'), and the manuscript would benefit from a careful language edit.","section":"Throughout"},{"comment":"The 'Limitations in Our Survey' passage appears after the Conclusions; integrating it before the Conclusions would improve the paper's structure.","section":"6"}],"recommendation":"major_revision","confidential_remarks":"The citation mismatches are pervasive enough that I would treat the reference list as unverified until a systematic audit is completed; some entries appear to be attached to claims unrelated to their content. I would also flag Section 6 and the GitHub repository in Appendix B as potential conflict-of-interest issues, though they are disclosed in the author affiliations. This is not a fit issue for the journal, but the revision should be checked carefully."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The 'acoustic world model' label is a genuinely useful organizing device, and the survey is clearly written and well structured. The taxonomy of PINNs, generative models, and self-supervised learning is sensible, and the roadmap in Section 5.3 is concrete enough to guide newcomers. The applications chapters give a good sense of the breadth, even if they are mostly aspirational. This deserves to exist as a survey.\n\nThe soft spots are real, though. The stress-test note is right: the claims in Section 2.1 about deep learning estimating material, viscosity, and particle size are attached to references that do not say those things. [16] is a review of ultrasonic NDE of composites, [17] is about porosity in additively manufactured parts, and [18] is acoustic emission damage classification. None of them support the stated impact-sound or shaking-sound inference tasks. Similarly, the cyclostationary machinery-health paragraph cites sound-source localization and acoustic SLAM surveys rather than fault-diagnosis work. And [12] is a computational physics textbook sitting in for auditory scene analysis. These are not minor reference typos; they are load-bearing for the claim that the literature already demonstrates acoustic inference of physical properties, and they need to be corrected.\n\nThe KZK equation in (9) also has the nonlinear term written as ∂³(p²)/∂τ³, which is dimensionally inconsistent. The standard form is ∂²(p²)/∂τ². That is a straightforward typo, but it is exactly the kind of thing that makes a reader doubt the rest of the physics.\n\nThe self-promotional 'Next Steps in Our Research' section and the company-affiliated GitHub repository are unusual in a survey. They are transparently labeled, which helps, but they do raise a fair question about whether the literature selection is independent. The authors should disclose the affiliation more prominently and, ideally, separate the roadmap from their own agenda.\n\nThe central thesis—that sound carries rich physical information and is a valuable modality for world models—is plausible and survives these problems. The survey fails as a reliable map of the field in its current form because the evidence base has not been checked carefully. It is still worth a serious referee: the errors are correctable, and the framing is useful to the community. I would suggest sending it to peer review with a required revision round, not a desk reject. Who is this for? Anyone looking for an entry point into acoustic physics plus generative modeling, or an instructor building a course module. I would not cite it in my own work until the citations and equations are cleaned up.","headline":"Useful framing for a survey, but the citation mismatches and a dimensionally wrong KZK term need fixing before it can be trusted as a field map.","tokens_in":24302,"tokens_out":1687,"would_cite":false,"duration_ms":21721,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey argues that acoustic signals—the radiated mechanical energy of physical events—should be a primary modality for AI world models, enabling causal physical understanding and predictive simulation through sound.","keywords":["world models","acoustic physical information","physics-informed neural networks","acoustic perception","multimodal learning","dynamic prediction","embodied intelligence","causal inference"],"falsifier":"A concrete falsifier: train a model to infer material and shape from impact sounds, then evaluate it on objects whose geometry, excitation type, and recording conditions are randomized and never seen during training; if accuracy falls to chance, the modal-inversion route to physical properties is not doing the work claimed. A second check: on a held-out physical-prediction task, if an acoustic-only model adds no predictive power over a vision-only model, the claimed complementarity of sound for physical intuition is unsupported.","tokens_in":23225,"feed_emoji":"🔊","tokens_out":12144,"duration_ms":129179,"temperature":0.7,"pith_summary":"This survey makes the case that sound should be treated as a primary physical modality for AI world models, not a side channel for event detection. Its central claim is that acoustic signals are radiated mechanical energy from physical events, so they carry latent information about material properties, internal structure, contact and fluid dynamics, and spatial geometry. The paper organizes the field around three methodological pillars—physics-informed neural networks, generative forward models, and self-supervised multimodal learning—and argues that combining them lets an AI build an internal 'intuitive physics' engine by listening. It then maps applications in robotics, autonomous driving, healthcare, and finance, and closes with a research roadmap toward causal, uncertainty-aware, embodied, and responsible acoustic intelligence.","feed_headline":"Sound can give AI a physical intuition of the world","feed_subtitle":"A new survey maps how acoustic physics—vibration, flow, echo—can power perception, prediction, and causal reasoning.","key_machinery":"The argument is carried by a chain of physics-to-signal encodings. The elastodynamic wave equation $$(\\$\\lambda$+\\mu)\\nabla(\\nabla\\cdot u)+\\mu\\$nabla^{2}$ u=\\rho\\,\\$partial^{2}$ u/\\partial $t^{2}$$$ with modal analysis links an object's natural frequencies $\\omega_n$ to its material and geometry; the acoustic analogy recasts the Navier-Stokes equations into an inhomogeneous wave equation for flow-generated sound; the room impulse response $h(t)$, with standard reverberation-time estimates $T_{60}$, encodes a space's geometry and absorption; and the Westervelt and KZK equations describe nonlinear high-amplitude propagation. On the machine-learning side, the composite PINN loss $$\\mathcal{L}(\\$\\theta$)=w_{\\mathrm{data}}\\mathcal{L}_{\\mathrm{data}}+w_{\\mathrm{phys}}\\mathcal{L}_{\\mathrm{phys}}+w_{\\mathrm{bc}}\\mathcal{L}_{\\mathrm{bc}}+w_{\\mathrm{ic}}\\mathcal{L}_{\\mathrm{ic}}$$ is the gray-box device that enforces these partial differential equations as regularizers, while differentiable simulators, neural acoustic fields, and contrastive audio-visual models supply forward prediction and scalable representation learning. Each element does a specific job: physics supplies the acoustic fingerprint, physics-informed networks invert it with sparse data, generative models predict and synthesize, and self-supervised learning scales the representations.","core_discovery":"The paper's central claim is that the physical laws governing sound—elastodynamics, aeroacoustics, room acoustics, and nonlinear wave propagation—leave measurable acoustic signatures that can be inverted for physical knowledge and used forward for prediction. Vibrations of struck objects encode material and modal properties such as elastic modulus, density, and damping; contact and flow sounds encode collision, friction, and fluid dynamics; room impulse responses encode geometry and surface absorption; and nonlinear equations such as the Westervelt and KZK equations extend the picture to high-intensity fields. The survey's thesis is that these signatures make audio complementary to vision precisely where vision is weak: seeing surfaces but not masses, internal defects, contact forces, or occluded structure. An AI that learns these mappings can simulate events in latent space and reason about physical interventions, which the paper calls an internal 'intuitive physics' engine through sound.","pith_inferences":["Editorial inference: if audio truly encodes physical properties as the survey argues, sound can serve as a cheap self-supervision signal for vision-based models, correcting visual errors about mass, material, and internal state without new hardware.","Editorial inference: the survey's own hybrid scenario—self-supervised pretraining, generative synthesis, and PINN regularization—can be tested directly by comparing that combined architecture against each pillar alone on a physical audio prediction benchmark.","Editorial note grounded in the paper's stated limitations: the authors concede that sensor design and classical array processing are treated only as background, so the deployment roadmap is stronger on algorithms than on the hardware that would make those algorithms reliable in the field."],"forward_implications":["If the acoustic world-model program works, robots can localize and map in darkness, smoke, or dust where vision fails, using Acoustic SLAM built from echoes and passive sound sources.","Contact and friction sounds give robotic manipulators haptic-like feedback, letting them detect incipient slip, adjust grip, and infer hidden object properties such as fill level by shaking a container.","Vehicles gain a 360-degree safety layer: siren detection beyond line of sight and real-time road-surface classification from tire-road noise, complementing vision and LiDAR where those sensors are ambiguous.","Non-invasive health screening can scale through cough, breathing, and voice analysis on ubiquitous microphones, with vocal biomarkers tracking respiratory, neurological, and mental-health conditions.","Financial applications can extract non-consensus signals from paralinguistic cues in earnings calls and use acoustic stress markers for fraud detection and compliance."],"supporting_citations":[{"why":"Supplies the definition of world models as learnable generative internal representations that the survey builds on.","marker":"[4]"},{"why":"Provides the action-conditioned latent-space predictive learning framework used for embodied agents.","marker":"[5]"},{"why":"Introduces the physics-informed neural network framework that anchors the physics-informed methodological pillar.","marker":"[45]"},{"why":"Gives the acoustic analogy for flow-generated sound, the basis for treating aeroacoustic signals as physical encodings.","marker":"[28]"},{"why":"Introduces neural radiance fields, the template for neural acoustic fields as implicit scene representations.","marker":"[51]"},{"why":"Establishes audio-visual contrastive learning, the foundation of the self-supervised multimodal pillar.","marker":"[57]"},{"why":"Defines Acoustic SLAM, the core technique for constructing spatial maps from sound.","marker":"[35]"},{"why":"Frames auditory scene analysis, the perceptual foundation for reading object and scene properties from sound.","marker":"[11]"}],"fun_headline_variants":["Sound gives AI an intuitive physics engine","Acoustic world models let AI hear physics","AI learns physical laws from acoustic signals","Echoes and vibrations: AI's physics teacher","Hearing the world: AI's path to intuitive physics"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that acoustic signals carry enough latent physical information—and the cited methods can extract enough of it—that sound alone can ground causal physical understanding; the survey's evidence base also has to be a fair sample of the field for its roadmap to be trustworthy.","fun_headline_variants_meta":{"raw":{"variants":["Sound gives AI an intuitive physics engine","Acoustic world models let AI hear physics","AI learns physical laws from acoustic signals","Echoes and vibrations: AI's physics teacher","Hearing the world: AI's path to intuitive physics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000854,"raw_usage":{"total_tokens":3700,"prompt_tokens":925,"completion_tokens":2775,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":541,"completion_tokens_details":{"reasoning_tokens":2706}},"tokens_in":541,"tokens_out":2775,"duration_ms":21991,"temperature":1.0,"reasoning_tokens":2706,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:36:48.909610+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete falsifier: train a model to infer material and shape from impact sounds, then evaluate it on objects whose geometry, excitation type, and recording conditions are randomized and never seen during training; if accuracy falls to chance, the modal-inversion route to physical properties is not doing the work claimed. A second check: on a held-out physical-prediction task, if an acoustic-only model adds no predictive power over a vision-only model, the claimed complementarity of sound for physical intuition is unsupported.","supporting_citations":[{"cited_title":"Onsoundgeneratedaerodynamicallyi.generaltheory,","cited_arxiv_id":null,"evidence_quote":"Gives the acoustic analogy for flow-generated sound, the basis for treating aeroacoustic signals as physical encodings."},{"cited_title":"Nerf: Repre- senting scenes as neural radiance fields for view synthesis,","cited_arxiv_id":null,"evidence_quote":"Introduces neural radiance fields, the template for neural acoustic fields as implicit scene representations."},{"cited_title":"Look, listen and learn,","cited_arxiv_id":null,"evidence_quote":"Establishes audio-visual contrastive learning, the foundation of the self-supervised multimodal pillar."},{"cited_title":"Acoustic slam,","cited_arxiv_id":null,"evidence_quote":"Defines Acoustic SLAM, the core technique for constructing spatial maps from sound."}],"review_version":1}