Pith. sign in

REVIEW 2 major objections 4 minor 1 cited by

Deep Learning for Personalized Binaural Audio Reproduction

T0 review · 2 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read This survey claims that deep learning is transforming personalized binaural audio through two pathways—explicit HRTF filtering and end-to-end synthesis—and organizes the field around that split.

desk verdict A solid, genuinely useful survey that organizes DL binaural audio into explicit HRTF personalization and end-to-end synthesis; the gap claim is plausible but rests on an undocumented literature selection, so treat the 'first comprehensive' language with caution. read the letter →

arxiv 2509.00400 v1 pith:2EISULC4 submitted 2025-08-30 eess.AS cs.SD

classification eess.AScs.SD
keywords binauralaudiohead-relatedtransferfunctionHRTFpersonalizationend-to-endsynthesismulti-modalspatialimplicitneuralrepresentationsevaluationdeeplearningsurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey claims that deep learning now drives personalized binaural audio through two complementary paradigms: explicit personalized filtering, where a model predicts a listener-specific head-related transfer function (HRTF) and then renders with it, and end-to-end rendering, where a model maps source audio plus visual, text, or parametric guidance directly to binaural signals. The paper organizes recent methods into this split, covering HRTF prediction from anthropometry, photos, and 3D scans; interpolation and dataset fusion; and single- or multi-modal binaural synthesis. It also catalogs public datasets, evaluation metrics, applications, and open challenges. A sympathetic reader would take the central contribution to be the organizing structure itself: a field map that lets new work be positioned and compared, rather than any single technical result.

What carries the argument

The organizing device is a two-paradigm taxonomy. Paradigm one, explicit personalized filtering, keeps the classic rendering pipeline: a deep model predicts the head-related transfer function (the filter that encodes how a listener's head, torso, and pinna shape sound directionally) and then convolution renders binaural audio. Paradigm two, end-to-end rendering, discards the explicit filter and learns the whole mapping from source audio plus guidance (visual, textual, or parametric) to binaural output. Within this split, the survey uses data representation (time-domain, frequency-domain, sparse, and implicit neural representations), datasets, and metrics as secondary organizing axes.

What would settle it

A reader could run a structured bibliographic search for deep-learning binaural audio papers from 2019-2025 with explicit inclusion criteria. If a substantial number of well-cited methods fall outside both paradigms, or if important datasets and metrics are absent from the tables, the coverage claim would be falsified.

Watch

Extended reading notes

Core claim

The central claim is that deep learning is not just improving but reshaping both fundamental pathways to binaural audio. On the explicit side, models predict personalized HRTFs from sparse measurements, morphological features, 3D geometry, or even ambient listening cues, replacing costly lab measurement. On the end-to-end side, models synthesize binaural audio directly from mono audio, video, text, or scene geometry, learning personalization implicitly. The survey's thesis is that these two paradigms, plus shared datasets and metrics, now define the field, and that this is the first structured overview to treat end-to-end synthesis as a first-class area alongside HRTF modeling.

Load-bearing premise

The survey's usefulness depends on its reviewed set of methods, datasets, and metrics fairly representing the whole field, but the paper gives no systematic search strategy or inclusion criteria.

Editorial extensions

If this is right

  • Researchers can position any new method as either explicit HRTF personalization or end-to-end synthesis, and compare it against the tables of representative approaches.
  • End-to-end binaural synthesis is a recognized first-class research area, with text- and image-guided generation as active directions.
  • Implicit neural representations appear to be the leading mechanism for continuous HRTF interpolation and cross-dataset fusion, easing the data heterogeneity bottleneck.
  • Perceptual validity, not raw signal error, is the field's main evaluation gap; objective metrics need stronger correlation with listening tests.
  • Applications in VR/AR, hearing aids, telepresence, audio-visual navigation, and scene understanding follow directly from the two paradigms' progress.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The authors do not say this, but the taxonomy suggests the two paradigms are converging: end-to-end models learn implicit HRTFs, while explicit methods are adopting neural fields, so the split may soon be a matter of rendering control rather than architecture.
  • An inference beyond the survey: if in-the-wild and text-guided personalization mature, personalized binaural audio could be generated on demand from a photo of a pinna or a sentence describing a scene, without any lab measurement.
  • The survey's emphasis on dataset fusion implies that the next large gains may come from data, not new architectures; readers should expect cross-database training to become standard.
  • Because the review is narrative rather than systematic, quantitative claims about trends should be read as illustrative until a formal literature search confirms the coverage.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper is a survey of deep learning for personalized binaural audio reproduction. It organizes recent work into two paradigms: explicit HRTF-based personalization/filtering and end-to-end binaural synthesis. The survey covers HRTF data representations, personalization from morphological and environmental cues, spatial interpolation, dataset fusion, and evaluation metrics; it also reviews end-to-end synthesis from audio, visual, text, and joint multimodal inputs, along with relevant datasets, applications, and open challenges. The paper claims to fill a gap in the literature by being one of the first comprehensive overviews covering both pathways in a unified structure.

Significance. If the coverage is accepted, the survey provides a useful organizing reference for a rapidly growing field. Its strengths include a clear taxonomy separating explicit and end-to-end paradigms, extensive summary tables for methods, datasets, and metrics, and a number of useful code links. Spot-checks of the metric formulas in Tables IV and VIII indicate they are consistent with standard usage, and the textual descriptions of methods align with the cited literature. The paper does not introduce new derivations, so internal consistency is high. However, the central claim of being a comprehensive, gap-filling overview rests on an undocumented literature selection, which limits confidence in the representativeness of the reviewed set and in the conclusions derived from it.

major comments (2)
  1. [Section I-C] The claimed contribution, 'one of the first comprehensive overview of DL-based end-to-end spatial audio synthesis', is not supported by a reproducible literature selection. No search strategy, database list, inclusion/exclusion criteria, date cutoff, or screening process is reported. The reference set appears to be a convenience sample with many 2024-2025 arXiv preprints and a notable number of works from the authors' own network (e.g., [55], [76], [89]); those works are summarized like all others, but without a protocol the reader cannot verify that the selected set fairly represents the field or that a prior survey focused on binaural synthesis does not already fill part of the claimed gap. This is load-bearing for the survey's contribution. Please add a methodology subsection describing the search and selection protocol, and use it to substantiate the gap assertion by comparing with t
  2. [Section II-D, Eq. (1)] The notation fw : (θ, φ) → H(θ, φ, f) is ambiguous: if f is not an input, then H(θ, φ, f) must be a vector-valued output over frequency, but this is not stated. Clarify whether the network outputs a full spectrum or whether frequency is also an input coordinate. This is a presentation issue, but it affects the interpretation of the continuous-domain interpolation methods summarized in Table II.
minor comments (4)
  1. [Section I-C] The phrase 'one of the first comprehensive overview' mixes singular and plural; suggest 'one of the first comprehensive overviews'.
  2. [Section III-D] The text refers to 'SonicSet [202]' in the paragraph after Table VII, but the table row and reference [202] are for 'SonicSim'. Please align the naming.
  3. [Section II-F] In Table IV, the LMD formula is rendered as '20 log10 |Hhat/H|' and described as 'mean absolute log-magnitude difference'; if the metric is the mean absolute difference, the outer absolute value should appear explicitly around the log term. Clarify the typesetting.
  4. [Section V] The open challenges in Section V-A through V-E are reasonable but somewhat generic; tying each challenge back to specific gaps visible in Tables I-VIII would strengthen the survey's contribution.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is an organizing survey whose claims are descriptive, not derived from its cited inputs.

full rationale

This is a survey paper; it makes no quantitative claim, prediction, or derivation whose conclusion could be encoded in its inputs. The central assertion is that a structured overview of DL for personalized binaural reproduction is missing and that this paper provides one (Section I-C). That is a coverage claim, not a reduction: the two-paradigm taxonomy (explicit HRTF-based filtering vs. end-to-end synthesis) is a classification choice applied to the cited literature, not a result forced by any fitted parameter or self-citation. The paper does cite several works by its own authors ([50], [55], [76], [89]), but these are summarized in the same manner as all other cited methods and are not used to justify the taxonomy, the gap claim, or any evaluation conclusion. No uniqueness theorem, ansatz, or known result is imported via self-citation. The absence of a stated literature search protocol is a legitimate completeness/coverage limitation, but it is not circularity: an incomplete survey could still be a non-circular survey. The 'open challenges' discussion is explicitly framed as the authors' summary and does not masquerade as an inference from the reviewed set. Therefore the derivation chain, such as it is, is self-contained as a literature organization and no step reduces by construction to its inputs.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

As a survey, the paper introduces no new entities or fitted constants. Its conclusions rest on the validity of its organizing taxonomy and on the representativeness of the cited literature, both of which are domain assumptions rather than mathematical axioms.

assumptions (2)
  • domain assumption The two-paradigm taxonomy (explicit personalized filtering vs. end-to-end rendering) is a valid and exhaustive organization of the surveyed literature.
    Section I-C declares this as the structure; if this partition misrepresents some methods, the survey's overview would be distorted.
  • domain assumption The selected papers are representative of the field and are accurately summarized.
    No systematic selection protocol is described in Sections II and III, so the survey's coverage claim rests on the completeness and fairness of the authors' literature search.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Deep Learning for Personalized Binaural Audio Reproduction." pith.science (2026). https://pith.science/paper/2EISULC4

@misc{pith2026250900400,
  author       = {Pith},
  title        = {Pith review of: Deep Learning for Personalized Binaural Audio Reproduction},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2EISULC4}},
  note         = {Machine review of arXiv:2509.00400}
}
read the original abstract

Personalized binaural audio reproduction is the basis of realistic spatial localization, sound externalization, and immersive listening, directly shaping user experience and listening effort. This survey reviews recent advances in deep learning for this task and organizes them by generation mechanism into two paradigms: explicit personalized filtering and end-to-end rendering. Explicit methods predict personalized head-related transfer functions (HRTFs) from sparse measurements, morphological features, or environmental cues, and then use them in the conventional rendering pipeline. End-to-end methods map source signals directly to binaural signals, aided by other inputs such as visual, textual, or parametric guidance, and they learn personalization within the model. We also summarize the field's main datasets and evaluation metrics to support fair and repeatable comparison. Finally, we conclude with a discussion of key applications enabled by these technologies, current technical limitations, and potential research directions for deep learning-based spatial audio systems.

Figures

Figures reproduced from arXiv: 2509.00400 by the authors.

Figure 1
Figure 1. Comparison of the two main paradigms in binaural audio reproduction: (a) HRTF-based filtering and (b) end-to-end binaural synthesis. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Overview of DL applications in HRTF personalized modeling, illustrating approaches for (a) personalization using morphological data and (b) spatial [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Conceptual illustration of DL-based binaural audio synthesis paradigms: (a) single-modal synthesis driven by audio inputs and spatial parameters, and [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HRTFformer: A Spatially-Aware Transformer for Individual HRTF Upsampling in Immersive Audio Rendering

    cs.SD 2025-10 conditional novelty 5.0 of 10

    HRTFformer reconstructs high-resolution head-related transfer functions from as few as three measured directions using a transformer in the spherical harmonic domain, beating prior methods in accuracy.

Reference graph

Works this paper leans on

260 extracted references · 70 canonical work pages · cited by 1 Pith paper

  1. [55]

    Modelling individual head-related transfer function (HRTF) based on anthropometric parameters and generic HRTF amplitudes,

    R. Zhang, R. Meng, J. Sang, Y . Hu, X. Li, and C. Zheng, “Modelling individual head-related transfer function (HRTF) based on anthropometric parameters and generic HRTF amplitudes,” CAAI Transactions on Intelligence Technology , vol. 8, no. 2, pp. 364–378, 2023

  2. [76]

    Modeling individual head-related transfer functions from sparse measurements using a convolutional neural network,

    Z. Jiang, J. Sang, C. Zheng, A. Li, and X. Li, “Modeling individual head-related transfer functions from sparse measurements using a convolutional neural network,” The Journal of the Acoustical Society of America , vol. 153, no. 1, pp. 248–259, 2023

  3. [89]

    BiCG: binaural cue generation from unified HRTF datasets,

    X. Lu, Y . Wang, J. Sang, and C. Zheng, “BiCG: binaural cue generation from unified HRTF datasets,” in ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2025, pp. 1–5

  4. [1]

    B. C. Moore, An introduction to the psychology of hearing. Leiden: Brill, 2012

  5. [2]

    Blauert, Spatial hearing: the psychophysics of human sound localization

    J. Blauert, Spatial hearing: the psychophysics of human sound localization. Cambridge: MIT press, 1997

  6. [3]

    Spatial sound- scapes and virtual worlds: Challenges and opportuni- ties,

    C. Rajguru, M. Obrist, and G. Memoli, “Spatial sound- scapes and virtual worlds: Challenges and opportuni- ties,” Frontiers in Psychology, vol. 11, p. 569056, 2020

  7. [4]

    Spatial sound-history, principle, progress and challenge,

    B. Xie, “Spatial sound-history, principle, progress and challenge,” Chinese Journal of Electronics , vol. 29, no. 3, pp. 397–416, 2020

  8. [5]

    A computationally-efficient and perceptually-plausible al- gorithm for binaural room impulse response simula- tion,

    T. Wendt, S. Van De Par, and S. D. Ewert, “A computationally-efficient and perceptually-plausible al- gorithm for binaural room impulse response simula- tion,” Journal of the Audio Engineering Society, vol. 62, no. 11, pp. 748–766, 2014

Show all 260 references
  1. [6]

    Low-order filter approxima- tion of diffraction for virtual acoustics,

    C. Kirsch and S. D. Ewert, “Low-order filter approxima- tion of diffraction for virtual acoustics,” in 2021 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA). IEEE, 2021, pp. 341–345

  2. [7]

    A filter representation of diffraction at infinite and finite wedges,

    S. D. Ewert, “A filter representation of diffraction at infinite and finite wedges,” JASA Express Letters, vol. 2, no. 9, 2022

  3. [8]

    A universal filter approx- imation of edge diffraction for geometrical acoustics,

    C. Kirsch and S. D. Ewert, “A universal filter approx- imation of edge diffraction for geometrical acoustics,” IEEE/ACM Transactions on Audio, Speech, and Lan- guage Processing, vol. 31, pp. 1636–1651, 2023

  4. [9]

    Machine-learning-based estimation and rendering of scattering in virtual reality,

    V . Pulkki and U. P. Svensson, “Machine-learning-based estimation and rendering of scattering in virtual reality,” The Journal of the Acoustical Society of America , vol. 145, no. 4, pp. 2664–2676, 2019

  5. [10]

    Computationally-efficient simulation of late reverberation for inhomogeneous boundary conditions and coupled rooms,

    C. Kirsch, T. Wendt, S. Van De Par, H. Hu, and S. D. Ewert, “Computationally-efficient simulation of late reverberation for inhomogeneous boundary conditions and coupled rooms,” Journal of the Audio Engineering Society, vol. 71, no. 4, pp. 186–201, 2023

  6. [11]

    Integrating real-time room acoustics simulation into a cad modeling software to enhance the architectural design process,

    S. Pelzer, L. Asp ¨ock, D. Schr ¨oder, and M. V orl ¨ander, “Integrating real-time room acoustics simulation into a cad modeling software to enhance the architectural design process,” Buildings, vol. 4, no. 2, pp. 113–138, 2014

  7. [12]

    A round robin on room acoustical simulation and auralization,

    F. Brinkmann, L. Asp ¨ock, D. Ackermann, S. Lepa, M. V orl¨ander, and S. Weinzierl, “A round robin on room acoustical simulation and auralization,” The Journal of the Acoustical Society of America , vol. 145, no. 4, pp. 2746–2760, 2019

  8. [13]

    Overview of geomet- rical room acoustic modeling techniques,

    L. Savioja and U. P. Svensson, “Overview of geomet- rical room acoustic modeling techniques,” The Journal of the Acoustical Society of America , vol. 138, no. 2, pp. 708–730, 2015

  9. [14]

    Interactive simulation and free-field auralization of acoustic space with the rtSOFE,

    B. U. Seeber and S. W. Clapp, “Interactive simulation and free-field auralization of acoustic space with the rtSOFE,” The Journal of the Acoustical Society of America, vol. 141, no. 5 Supplement, p. 3974, 2017

  10. [15]

    Schr ¨oder, Physically based real-time auralization of interactive virtual environments

    D. Schr ¨oder, Physically based real-time auralization of interactive virtual environments. Berlin: Logos Verlag Berlin GmbH, 2011, vol. 11

  11. [16]

    Efficient HRTF-based spatial audio for area and volumetric sources,

    C. Schissler, A. Nicholls, and R. Mehra, “Efficient HRTF-based spatial audio for area and volumetric sources,” IEEE Transactions on Visualization and Com- puter Graphics, vol. 22, no. 4, pp. 1356–1366, 2016

  12. [17]

    Natural listening over head- phones in augmented reality using adaptive filtering techniques,

    R. Ranjan and W.-S. Gan, “Natural listening over head- phones in augmented reality using adaptive filtering techniques,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 23, no. 11, pp. 1988– 2002, 2015

  13. [18]

    Larsson, A

    P. Larsson, A. V ¨aljam¨ae, D. V ¨astfj¨all, A. Tajadura- Jim´enez, and M. Kleiner, Auditory-Induced Presence in Mixed Reality Environments and Related Technology . London: Springer London, 2010, pp. 143–163

  14. [19]

    Acoustic control by wave field synthesis,

    A. J. Berkhout, D. de Vries, and P. V ogel, “Acoustic control by wave field synthesis,” The Journal of the Acoustical Society of America, vol. 93, no. 5, pp. 2764– 2778, 1993

  15. [20]

    Further investiga- tions of high-order ambisonics and wavefield synthesis for holophonic sound imaging,

    J. Daniel, S. Moreau, and R. Nicol, “Further investiga- tions of high-order ambisonics and wavefield synthesis for holophonic sound imaging,” in Audio Engineering Society Convention 114 . Audio Engineering Society, 2003

  16. [21]

    Rumsey, Spatial audio

    F. Rumsey, Spatial audio. Routledge, 2012

  17. [22]

    Headphone simula- JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2025 18 tion of free-field listening. I: Stimulus synthesis,

    F. L. Wightman and D. J. Kistler, “Headphone simula- JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2025 18 tion of free-field listening. I: Stimulus synthesis,” The Journal of the Acoustical Society of America , vol. 85, no. 2, pp. 858–867, 1989

  18. [23]

    Sound localiza- tion by human listeners,

    J. C. Middlebrooks and D. M. Green, “Sound localiza- tion by human listeners,” Annual Review of Psychology, vol. 42, no. 1991, pp. 135–159, 1991

  19. [24]

    Xie, Head-related transfer function and virtual au- ditory display

    B. Xie, Head-related transfer function and virtual au- ditory display. J. Ross Publishing, 2013

  20. [25]

    Fundamentals of binaural technology,

    H. Møller, “Fundamentals of binaural technology,” Ap- plied Acoustics, vol. 36, no. 3-4, pp. 171–218, 1992

  21. [26]

    Individual differences in external- ear transfer functions reduced by scaling in frequency,

    J. C. Middlebrooks, “Individual differences in external- ear transfer functions reduced by scaling in frequency,” The Journal of the Acoustical Society of America , vol. 106, no. 3, pp. 1480–1492, 1999

  22. [27]

    Localization using nonindividualized head- related transfer functions,

    E. M. Wenzel, M. Arruda, D. J. Kistler, and F. L. Wightman, “Localization using nonindividualized head- related transfer functions,” The Journal of the Acousti- cal Society of America , vol. 94, no. 1, pp. 111–123, 1993

  23. [28]

    Binaural technique: Do we need individual recordings?

    H. Møller, M. F. Sørensen, C. B. Jensen, and D. Ham- mershøi, “Binaural technique: Do we need individual recordings?” Journal of the Audio Engineering Society , vol. 44, no. 6, pp. 451–469, 1996

  24. [29]

    Personalized HRTF model- ing based on deep neural network using anthropometric measurements and images of the ear,

    G. W. Lee and H. K. Kim, “Personalized HRTF model- ing based on deep neural network using anthropometric measurements and images of the ear,” Applied Sciences, vol. 8, no. 11, p. 2180, 2018

  25. [30]

    Measurement of head-related transfer functions: A review,

    S. Li and J. Peissig, “Measurement of head-related transfer functions: A review,” Applied Sciences, vol. 10, no. 14, p. 5014, 2020

  26. [31]

    The CIPIC HRTF database,

    V . R. Algazi, R. O. Duda, D. M. Thompson, and C. Avendano, “The CIPIC HRTF database,” inProceed- ings of the 2001 IEEE Workshop on the Applications of Signal Processing to Audio and Acoustics (Cat. No. 01TH8575). IEEE, 2001, pp. 99–102

  27. [32]

    Boundary element method calculation of individual head-related transfer function. I. Rigid model calculation,

    B. F. Katz, “Boundary element method calculation of individual head-related transfer function. I. Rigid model calculation,” The Journal of the Acoustical Society of America, vol. 110, no. 5, pp. 2440–2448, 2001

  28. [33]

    A cross-evaluated database of measured and simulated HRTFs including 3D head meshes, anthropometric features, and headphone im- pulse responses,

    F. Brinkmann, M. Dinakaran, R. Pelzer, P. Grosche, D. V oss, and S. Weinzierl, “A cross-evaluated database of measured and simulated HRTFs including 3D head meshes, anthropometric features, and headphone im- pulse responses,” Journal of the Audio Engineering Society, vol. 67, ...

  29. [34]

    Virtual sound source positioning using vector base amplitude panning,

    V . Pulkki, “Virtual sound source positioning using vector base amplitude panning,” Journal of the Audio Engineering Society, vol. 45, no. 6, pp. 456–466, 1997

  30. [35]

    A new HRTF interpolation approach for nonlinear 3D audio systems,

    V . Bruschi, N. Dourou, A. Carini, and S. Cecchi, “A new HRTF interpolation approach for nonlinear 3D audio systems,” in 2023 Immersive and 3D Audio: from Architecture to Automotive (I3DA) . IEEE, 2023, pp. 1–9

  31. [36]

    A model of head- related transfer functions based on principal compo- nents analysis and minimum-phase reconstruction,

    D. J. Kistler and F. L. Wightman, “A model of head- related transfer functions based on principal compo- nents analysis and minimum-phase reconstruction,” The Journal of the Acoustical Society of America , vol. 91, no. 3, pp. 1637–1647, 1992

  32. [37]

    Implicit HRTF modeling using temporal convolutional networks,

    I. D. Gebru, D. Markovi ´c, A. Richard, S. Krenn, G. A. Butler, F. De la Torre, and Y . Sheikh, “Implicit HRTF modeling using temporal convolutional networks,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, p...

  33. [38]

    Neural synthesis of binaural speech from mono audio,

    A. Richard, D. Markovic, I. D. Gebru, S. Krenn, G. A. Butler, F. Torre, and Y . Sheikh, “Neural synthesis of binaural speech from mono audio,” in International Conference on Learning Representations , 2021

  34. [39]

    A machine learning tutorial for spatial auditory display using head-related transfer functions,

    K. McMullen and Y . Wan, “A machine learning tutorial for spatial auditory display using head-related transfer functions,” The Journal of the Acoustical Society of America, vol. 151, no. 2, pp. 1277–1293, 2022

  35. [40]

    A review on head- related transfer function generation for spatial audio,

    V . Bruschi, L. Grossi, N. A. Dourou, A. Quattrini, A. Vancheri, T. Leidi, and S. Cecchi, “A review on head- related transfer function generation for spatial audio,” Applied Sciences, vol. 14, no. 23, p. 11242, 2024

  36. [41]

    A survey on machine learning techniques for head- related transfer function individualization,

    D. Fantini, M. Geronazzo, F. Avanzini, and S. Ntalampi- ras, “A survey on machine learning techniques for head- related transfer function individualization,” IEEE Open Journal of Signal Processing , 2025

  37. [42]

    An overview of machine learning and other data- based methods for spatial audio capture, processing, and reproduction,

    M. Cobos, J. Ahrens, K. Kowalczyk, and A. Politis, “An overview of machine learning and other data- based methods for spatial audio capture, processing, and reproduction,” EURASIP Journal on Audio, Speech, and Music Processing, vol. 2022, no. 1, p. 10, 2022

  38. [43]

    Perceptually based head- related transfer function database optimization,

    B. F. Katz and G. Parseihian, “Perceptually based head- related transfer function database optimization,” The Journal of the Acoustical Society of America , vol. 131, no. 2, pp. EL99–EL105, 2012

  39. [44]

    Estimation and modeling of pinna-related transfer functions,

    G. Michele, S. Spagnol, A. Federico et al., “Estimation and modeling of pinna-related transfer functions,” in Proceedings of the 13th International Conference on Digital Audio Effects, DAFx 2010 . Institute of Elec- tronic Music and Acoustics (IEM), University of Music and Per...

  40. [45]

    Extracting the frequencies of the pinna spectral notches in measured head related impulse responses,

    V . C. Raykar, R. Duraiswami, and B. Yegnanarayana, “Extracting the frequencies of the pinna spectral notches in measured head related impulse responses,” The Jour- nal of the Acoustical Society of America, vol. 118, no. 1, pp. 364–374, 2005

  41. [46]

    A wide dataset of ear shapes and pinna-related transfer functions generated by random ear drawings,

    C. Guezenoc and R. Seguier, “A wide dataset of ear shapes and pinna-related transfer functions generated by random ear drawings,” The Journal of the Acoustical Society of America , vol. 147, no. 6, pp. 4087–4096, 2020

  42. [47]

    Efficient real spherical harmonic representa- tion of head-related transfer functions,

    G. D. Romigh, D. S. Brungart, R. M. Stern, and B. D. Simpson, “Efficient real spherical harmonic representa- tion of head-related transfer functions,” IEEE Journal of Selected Topics in Signal Processing , vol. 9, no. 5, pp. 921–930, 2015

  43. [48]

    Autoencoding HRTFs for DNN based HRTF personalization using anthropometric features,

    T.-Y . Chen, T.-H. Kuo, and T.-S. Chi, “Autoencoding HRTFs for DNN based HRTF personalization using anthropometric features,” in ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Sig- nal Processing (ICASSP) . IEEE, 2019, pp. 271–275

  44. [49]

    Autoencoders, unsupervised learning, and JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2025 19 deep architectures,

    P. Baldi, “Autoencoders, unsupervised learning, and JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2025 19 deep architectures,” in Proceedings of ICML workshop on unsupervised and transfer learning . JMLR Work- shop and Conference Proceedings, 2012, pp. 37–49

  45. [50]

    HRTF personal- ization based on artificial neural network in individual virtual auditory space,

    H. Hu, L. Zhou, H. Ma, and Z. Wu, “HRTF personal- ization based on artificial neural network in individual virtual auditory space,” Applied Acoustics , vol. 69, no. 2, pp. 163–172, 2008

  46. [51]

    Global HRTF personalization using anthropometric measures,

    Y . Wang, Y . Zhang, Z. Duan, and M. Bocko, “Global HRTF personalization using anthropometric measures,” in Audio Engineering Society Conference: 2020 AES International Conference on Audio for Virtual and Aug- mented Reality. Audio Engineering Society, 2020

  47. [52]

    Modeling of individual HRTFs based on spatial principal com- ponent analysis,

    M. Zhang, Z. Ge, T. Liu, X. Wu, and T. Qu, “Modeling of individual HRTFs based on spatial principal com- ponent analysis,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 28, pp. 785– 797, 2020

  48. [53]

    HRTF individualization using deep learning,

    R. Miccini and S. Spagnol, “HRTF individualization using deep learning,” in 2020 IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops (VRW). IEEE, 2020, pp. 390–395

  49. [54]

    An individualization approach for head-related transfer function in arbitrary directions based on deep learning,

    D. Yao, J. Zhao, L. Cheng, J. Li, X. Li, X. Guo, and Y . Yan, “An individualization approach for head-related transfer function in arbitrary directions based on deep learning,” JASA Express Letters, vol. 2, no. 6, p. 064401, 2022

  50. [56]

    Towards HRTF personalization using denoising diffusion models,

    J. C. A. S ´anchez, L. Comanducci, M. Pezzoli, and F. Antonacci, “Towards HRTF personalization using denoising diffusion models,” in ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025, pp. 1–5

  51. [57]

    A hybrid approach to struc- tural modeling of individualized HRTFs,

    R. Miccini and S. Spagnol, “A hybrid approach to struc- tural modeling of individualized HRTFs,” in 2021 IEEE Conference on Virtual Reality and 3D User Interfaces Abstracts and Workshops (VRW). IEEE, 2021, pp. 80– 85

  52. [58]

    Magnitude modeling of personalized HRTF based on ear images and an- thropometric measurements,

    M. Zhao, Z. Sheng, and Y . Fang, “Magnitude modeling of personalized HRTF based on ear images and an- thropometric measurements,” Applied Sciences, vol. 12, no. 16, p. 8155, 2022

  53. [59]

    PRTFNet: HRTF individualization for accurate spectral cues using a compact prtf,

    B.-Y . Ko, G.-T. Lee, H. Nam, and Y .-H. Park, “PRTFNet: HRTF individualization for accurate spectral cues using a compact prtf,” IEEE Access , vol. 11, pp. 96 119–96 130, 2023

  54. [60]

    A ma- chine learning approach to predicting personalized head related transfer functions and headphone equalization from video capture data,

    N. Javeri, P. B. Dutta, K. Sunder, and K. Jain, “A ma- chine learning approach to predicting personalized head related transfer functions and headphone equalization from video capture data,” in 2023 Immersive and 3D Audio: from Architecture to Automotive (I3DA). IEEE, 2023, pp. 1–9

  55. [61]

    HRTF individualization based on anthropometric mea- surements extracted from 3D head meshes,

    D. Fantini, F. Avanzini, S. Ntalampiras, and G. Presti, “HRTF individualization based on anthropometric mea- surements extracted from 3D head meshes,” in 2021 Im- mersive and 3D Audio: from Architecture to Automotive (I3DA). IEEE, 2021, pp. 1–10

  56. [62]

    On the predictabil- ity of HRTFs from ear shapes using deep networks,

    Y . Zhou, H. Jiang, and V . K. Ithapu, “On the predictabil- ity of HRTFs from ear shapes using deep networks,” in ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2021, pp. 441–445

  57. [63]

    Predicting global head-related transfer functions from scanned head geometry using deep learning and compact rep- resentations,

    Y . Wang, Y . Zhang, Z. Duan, and M. Bocko, “Predicting global head-related transfer functions from scanned head geometry using deep learning and compact rep- resentations,” arXiv preprint arXiv:2207.14352 , 2022

  58. [64]

    Efficient prediction of individual head-related transfer functions based on 3D meshes,

    J. Zhao, D. Yao, J. Gu, and J. Li, “Efficient prediction of individual head-related transfer functions based on 3D meshes,” Applied Acoustics , vol. 219, p. 109938, 2024

  59. [65]

    AudioEar: single-view ear reconstruction for personalized spatial audio,

    X. Huang, Y . Wang, Y . Liu, B. Ni, W. Zhang, J. Liu, and T. Li, “AudioEar: single-view ear reconstruction for personalized spatial audio,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 37, no. 1, 2023, pp. 944–952

  60. [66]

    Denoising of photogrammetric dummy head ear point clouds for individual head-related transfer functions computation,

    F. Di Giusto, F. Llu ´ıs, S. van Ophem, and E. Deckers, “Denoising of photogrammetric dummy head ear point clouds for individual head-related transfer functions computation,” arXiv preprint arXiv:2408.16410 , 2024

  61. [67]

    HRTF estimation in the wild,

    V . Jayaram, I. Kemelmacher-Shlizerman, and S. M. Seitz, “HRTF estimation in the wild,” in Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology, 2023, pp. 1–9

  62. [68]

    HRTF estimation using a score- based prior,

    E. Thuillier, J.-M. Lemercier, E. Moliner, T. Gerkmann, and V . V ¨alim¨aki, “HRTF estimation using a score- based prior,” in ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025, pp. 1–5

  63. [69]

    Recovery of individual head-related transfer functions from a small set of measurements,

    B.-S. Xie, “Recovery of individual head-related transfer functions from a small set of measurements,” The Journal of the Acoustical Society of America , vol. 132, no. 1, pp. 282–294, 2012

  64. [70]

    Head-related impulse response interpolation in virtual sound system,

    L. Chen, H. Hu, and Z. Wu, “Head-related impulse response interpolation in virtual sound system,” in 2008 Fourth International Conference on Natural Computa- tion, vol. 6. IEEE, 2008, pp. 162–166

  65. [71]

    Head-related transfer function interpolation from spa- tially sparse measurements using autoencoder with source position conditioning,

    Y . Ito, T. Nakamura, S. Koyama, and H. Saruwatari, “Head-related transfer function interpolation from spa- tially sparse measurements using autoencoder with source position conditioning,” in 2022 International Workshop on Acoustic Signal Enhancement (IWAENC) , 2022, pp. 1–5

  66. [72]

    Individualizing head-related transfer functions for bin- aural acoustic applications,

    N. H. Zandi, A. M. El-Mohandes, and R. Zheng, “Individualizing head-related transfer functions for bin- aural acoustic applications,” in 2022 21st ACM/IEEE International Conference on Information Processing in Sensor Networks (IPSN) . IEEE, 2022, pp. 105–117

  67. [73]

    Spatial upsampling of sparse head related transfer functions-a VQ-V AE & Transformer based approach,

    D. Zurale and S. Dubnov, “Spatial upsampling of sparse head related transfer functions-a VQ-V AE & Transformer based approach,” in Audio Engineering Society Conference: AES 2023 International Conference on Spatial and Immersive Audio . Audio Engineering JOURNAL OF LATEX CLASS ...

  68. [74]

    Spatial group- ing as a method to improve personalized head-related transfer function prediction,

    K.-W. Chang, Y .-L. Shen, and T.-S. Chi, “Spatial group- ing as a method to improve personalized head-related transfer function prediction,” JASA Express Letters , vol. 5, no. 3, p. 034801, 03 2025

  69. [75]

    Deep HRTF encoding & interpolation: Exploring spatial correlations using convolutional neural networks,

    D. Zurale, S. Yadegari, and S. Dubnov, “Deep HRTF encoding & interpolation: Exploring spatial correlations using convolutional neural networks,” in 19th Sound and Music Computing Conference, SMC 2022 . Sound and Music Computing Network, 2022, pp. 350–357

  70. [77]

    Head-related transfer function inter- polation with a spherical CNN,

    X. Chen, F. Ma, Y . Zhang, A. Bastine, and P. N. Samarasinghe, “Head-related transfer function inter- polation with a spherical CNN,” arXiv preprint arXiv:2309.08290, 2023

  71. [78]

    HRTF inter- polation using a spherical neural process meta-learner,

    E. Thuillier, C. T. Jin, and V . V ¨alim¨aki, “HRTF inter- polation using a spherical neural process meta-learner,” IEEE/ACM Transactions on Audio, Speech, and Lan- guage Processing, vol. 32, pp. 1790–1802, 2024

  72. [79]

    Head-related transfer func- tion upsampling with spatial extrapolation features,

    J. Zhao, D. Yao, and J. Li, “Head-related transfer func- tion upsampling with spatial extrapolation features,” IEEE Transactions on Audio, Speech and Language Processing, vol. 33, pp. 1034–1048, 2025

  73. [80]

    HRTF upsampling with a generative adversarial network using a gnomonic equiangular pro- jection,

    A. O. Hogg, M. Jenkins, H. Liu, I. Squires, S. J. Cooper, and L. Picinali, “HRTF upsampling with a generative adversarial network using a gnomonic equiangular pro- jection,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2024

  74. [81]

    HRTF spatial upsampling in the spherical harmonics domain employ- ing a generative adversarial network,

    X. Hu, L. Picinali, J. Li, A. Hogg et al., “HRTF spatial upsampling in the spherical harmonics domain employ- ing a generative adversarial network,” in Proceedings of the 27th International Conference on Digital Audio Effects, DAFx 2024 , 2024

  75. [82]

    A ma- chine learning approach for denoising and upsampling HRTFs,

    X. Hu, J. Li, L. Picinali, and A. O. Hogg, “A ma- chine learning approach for denoising and upsampling HRTFs,” arXiv preprint arXiv:2504.17586 , 2025

  76. [83]

    Global HRTF in- terpolation via learned affine transformation of hyper- conditioned features,

    J. W. Lee, S. Lee, and K. Lee, “Global HRTF in- terpolation via learned affine transformation of hyper- conditioned features,” in ICASSP 2023 - 2023 IEEE In- ternational Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023, pp. 1–5

  77. [84]

    HRTF Field: unifying measured HRTF magnitude representation with neural fields,

    Y . Zhang, Y . Wang, and Z. Duan, “HRTF Field: unifying measured HRTF magnitude representation with neural fields,” in ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023, pp. 1–5

  78. [85]

    Spatial upsampling of head-related transfer functions using a physics-informed neural network,

    F. Ma, T. D. Abhayapala, P. N. Samarasinghe, and X. Chen, “Spatial upsampling of head-related transfer functions using a physics-informed neural network,” arXiv preprint arXiv:2307.14650 , 2023

  79. [86]

    NIIRF: neural IIR filter field for HRTF upsampling and personalization,

    Y . Masuyama, G. Wichern, F. G. Germain, Z. Pan, S. Khurana, C. Hori, and J. Le Roux, “NIIRF: neural IIR filter field for HRTF upsampling and personalization,” in ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2024, pp. 1016–1020

  80. [87]

    Neural Steerer: novel steering vector synthesis with a causal neural field over frequency and direction,

    D. Di Carlo, A. A. Nugraha, M. Fontaine, Y . Bando, and K. Yoshii, “Neural Steerer: novel steering vector synthesis with a causal neural field over frequency and direction,” in 2024 IEEE International Conference on Acoustics, Speech and Signal Processing Workshops (ICASSPW), 2...

  81. [88]

    Retrieval-augmented neural field for HRTF upsampling and personalization,

    Y . Masuyama, G. Wichern, F. G. Germain, C. Ick, and J. Le Roux, “Retrieval-augmented neural field for HRTF upsampling and personalization,” in ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025, pp. 1–5

  82. [90]

    Neural fields in visual computing and beyond,

    Y . Xie, T. Takikawa, S. Saito, O. Litany, S. Yan, N. Khan, F. Tombari, J. Tompkin, V . sitzmann, and S. Sridhar, “Neural fields in visual computing and beyond,” Computer Graphics Forum, vol. 41, no. 2, pp. 641–676, 2022

  83. [91]

    Implicit neural representations with peri- odic activation functions,

    V . Sitzmann, J. Martel, A. Bergman, D. Lindell, and G. Wetzstein, “Implicit neural representations with peri- odic activation functions,” in Advances in Neural Infor- mation Processing Systems, vol. 33. Curran Associates, Inc., 2020, pp. 7462–7473

  84. [92]

    Fourier features let networks learn high frequency functions in low dimensional domains,

    M. Tancik, P. Srinivasan, B. Mildenhall, S. Fridovich- Keil, N. Raghavan, U. Singhal, R. Ramamoorthi, J. Bar- ron, and R. Ng, “Fourier features let networks learn high frequency functions in low dimensional domains,” in Advances in Neural Information Processing Systems , vol. ...

  85. [93]

    NeRF: Representing scenes as neural radiance fields for view synthesis,

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “NeRF: Representing scenes as neural radiance fields for view synthesis,” Communications of the ACM , vol. 65, no. 1, pp. 99– 106, 2021

  86. [94]

    Learning neural acoustic fields,

    A. Luo, Y . Du, M. Tarr, J. Tenenbaum, A. Torralba, and C. Gan, “Learning neural acoustic fields,” in Advances in Neural Information Processing Systems , vol. 35. Curran Associates, Inc., 2022, pp. 3165–3177

  87. [95]

    Neural acoustic context field: Rendering realistic room impulse response with neural fields,

    S. Liang, C. Huang, Y . Tian, A. Kumar, and C. Xu, “Neural acoustic context field: Rendering realistic room impulse response with neural fields,” arXiv preprint arXiv:2309.15977, 2023

  88. [96]

    E. G. Williams, Fourier acoustics: sound radiation and nearfield acoustical holography . Elsevier, 1999

  89. [97]

    Listen HRTF database,

    O. Warusfel, “Listen HRTF database,” online, IR- CAM and AK, Available: http://recherche. ircam. fr/equipes/salles/listen/index. html, 2003

  90. [98]

    Dataset of head-related transfer functions measured with a circular loudspeaker array,

    K. Watanabe, Y . Iwaya, Y . Suzuki, S. Takane, and S. Sato, “Dataset of head-related transfer functions measured with a circular loudspeaker array,” Acoustical Science and Technology , vol. 35, no. 3, pp. 159–165, 2014

  91. [99]

    Measurement of a head-related transfer function database with high spatial resolution,

    T. Carpentier, H. Bahu, M. Noisternig, and O. Warus- JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2025 21 fel, “Measurement of a head-related transfer function database with high spatial resolution,” in 7th forum acusticum (EAA), 2014

  92. [100]

    Sound localiza- tion in individualized and non-individualized crosstalk cancellation systems,

    P. Majdak, B. Masiero, and J. Fels, “Sound localiza- tion in individualized and non-individualized crosstalk cancellation systems,” The Journal of the Acoustical Society of America , vol. 133, no. 4, pp. 2055–2068, 2013

  93. [101]

    A high-resolution head-related transfer function and three-dimensional ear model database,

    R. Bomhardt, M. de la Fuente Klein, and J. Fels, “A high-resolution head-related transfer function and three-dimensional ear model database,” Proceedings of Meetings on Acoustics , vol. 29, no. 1, 2016

  94. [102]

    A database of head-related transfer function and morphological measurements,

    R. Sridhar, J. G. Tylka, and E. Y . Choueiri, “A database of head-related transfer function and morphological measurements,” in 143rd Audio Engineering Society Convention 2017, 2017, pp. 851 – 855

  95. [103]

    A perceptual evaluation of individual and non- individual HRTFs: A case study of the SADIE II database,

    C. Armstrong, L. Thresh, D. Murphy, and G. Kear- ney, “A perceptual evaluation of individual and non- individual HRTFs: A case study of the SADIE II database,” Applied Sciences , vol. 8, no. 11, p. 2029, 2018

  96. [104]

    Adapting hearing devices to the individual ear acous- tics: Database and target response correction functions for various device styles,

    F. Denk, S. M. Ernst, S. D. Ewert, and B. Kollmeier, “Adapting hearing devices to the individual ear acous- tics: Database and target response correction functions for various device styles,” Trends in Hearing , vol. 22, p. 2331216518779313, 2018

  97. [105]

    Computed HRIRs and ears database for acoustic research,

    S. Ghorbal, X. Bonjour, and R. S ´eguier, “Computed HRIRs and ears database for acoustic research,” in Audio Engineering Society Convention 148 . Audio Engineering Society, 2020

  98. [106]

    The SONI- COM HRTF dataset,

    I. Engel, R. Daugintis, T. Vicente, A. O. Hogg, J. Pauwels, A. J. Tournier, and L. Picinali, “The SONI- COM HRTF dataset,” Journal of the Audio Engineering Society, vol. 71, no. 5, pp. 241–253, 2023

  99. [107]

    Inter-laboratory round robin HRTF measurement com- parison,

    A. Andreopoulou, D. R. Begault, and B. F. Katz, “Inter-laboratory round robin HRTF measurement com- parison,” IEEE Journal of Selected Topics in Signal Processing, vol. 9, no. 5, pp. 895–906, 2015

  100. [108]

    On the relevance of the differences between HRTF measurement setups for machine learning,

    J. Pauwels and L. Picinali, “On the relevance of the differences between HRTF measurement setups for machine learning,” in ICASSP 2023-2023 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023, pp. 1–5

  101. [109]

    Mitigating cross- database differences for learning unified HRTF repre- sentation,

    Y . Wen, Y . Zhang, and Z. Duan, “Mitigating cross- database differences for learning unified HRTF repre- sentation,” in 2023 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) . IEEE, 2023, pp. 1–5

  102. [110]

    Personalization of head-related transfer functions (HRTF) based on automatic photo- anthropometry and inference from a database,

    E. A. Torres-Gallegos, F. Orduna-Bustamante, and F. Ar ´ambula-Cos´ıo, “Personalization of head-related transfer functions (HRTF) based on automatic photo- anthropometry and inference from a database,” Applied Acoustics, vol. 97, pp. 84–95, 2015

  103. [111]

    Towards fast and convenient end-to-end HRTF personalization,

    B. Zhi, D. N. Zotkin, and R. Duraiswami, “Towards fast and convenient end-to-end HRTF personalization,” in ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2022, pp. 441–445

  104. [112]

    HRTF recommendation based on the predicted binaural colouration model,

    N. Marggraf-Turley, M. Lovedee-Turner, and E. De Sena, “HRTF recommendation based on the predicted binaural colouration model,” in ICASSP 2024- 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2024, pp. 1106–1110

  105. [113]

    Virtual autoencoder based recommendation system for individ- ualizing head-related transfer functions,

    Y . Luo, D. N. Zotkin, and R. Duraiswami, “Virtual autoencoder based recommendation system for individ- ualizing head-related transfer functions,” in 2013 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA). IEEE, 2013, pp. 1–4

  106. [114]

    Temporal convolutional neural networks to generate a head-related impulse response from one direction to another,

    T. Kobayashi, Y . Maruyama, I. Nambu, S. Yano, and Y . Wada, “Temporal convolutional neural networks to generate a head-related impulse response from one direction to another,” arXiv preprint arXiv:2310.14018, 2023

  107. [115]

    Head-related transfer functions of human subjects,

    H. Møller, M. F. Sørensen, D. Hammershøi, and C. B. Jensen, “Head-related transfer functions of human subjects,” Journal of the Audio Engineering Society , vol. 43, no. 5, pp. 300–321, 1995

  108. [116]

    On the external- ization of sound images,

    W. M. Hartmann and A. Wittenberg, “On the external- ization of sound images,” The Journal of the Acoustical Society of America, vol. 99, no. 6, pp. 3678–3688, 1996

  109. [117]

    Perception of spatial sound,

    E. M. Wenzel, D. R. Begault, and M. Godfroy-Cooper, “Perception of spatial sound,” in Immersive Sound . Routledge, 2017, pp. 5–39

  110. [118]

    Sound externalization: A review of recent research,

    V . Best, R. Baumgartner, M. Lavandier, P. Maj- dak, and N. Kop ˇco, “Sound externalization: A review of recent research,” Trends in Hearing , vol. 24, p. 2331216520948390, 2020

  111. [119]

    Majdak, C

    P. Majdak, C. Hollomey, and R. Baumgartner, “AMT

  112. [120]

    x: a toolbox for reproducible research in auditory modeling,” Acta Acustica, vol. 6, p. 19, 2022

  113. [121]

    Predicting speech intelligibil- ity in hearing-impaired listeners using a physiologically inspired auditory model,

    J. Zaar and L. H. Carney, “Predicting speech intelligibil- ity in hearing-impaired listeners using a physiologically inspired auditory model,” Hearing Research, vol. 426, p. 108553, 2022

  114. [122]

    Modeling sound-source localization in sagittal planes for human listeners,

    R. Baumgartner, P. Majdak, and B. Laback, “Modeling sound-source localization in sagittal planes for human listeners,” The Journal of the Acoustical Society of America, vol. 136, no. 2, pp. 791–802, 2014

  115. [123]

    An ideal-observer model of human sound localization,

    J. Reijniers, D. Vanderelst, C. Jin, S. Carlile, and H. Peremans, “An ideal-observer model of human sound localization,” Biological Cybernetics, vol. 108, pp. 169– 181, 2014

  116. [124]

    A bayesian model for human directional localization of broadband static sound sources,

    R. Barumerli, P. Majdak, M. Geronazzo, D. Meijer, F. Avanzini, and R. Baumgartner, “A bayesian model for human directional localization of broadband static sound sources,” Acta Acustica, vol. 7, p. 12, 2023

  117. [125]

    Ideal-observer model of human sound local- ization of sources with unknown spectrum,

    J. Reijniers, G. McLachlan, B. Partoens, and H. Pere- mans, “Ideal-observer model of human sound local- ization of sources with unknown spectrum,” Scientific Reports, vol. 15, no. 1, p. 7289, 2025

  118. [126]

    Classifying non-individual head-related transfer functions with a computational auditory model: Cali- bration and metrics,

    R. Daugintis, R. Barumerli, L. Picinali, and M. Geron- azzo, “Classifying non-individual head-related transfer functions with a computational auditory model: Cali- bration and metrics,” in ICASSP 2023 - 2023 IEEE In- ternational Conference on Acoustics, Speech and Signal JOURN...

  119. [127]

    MOSNet: Deep learn- ing based objective assessment for voice conversion,

    C.-C. Lo, S.-W. Fu, W.-C. Huang, X. Wang, J. Yamag- ishi, Y . Tsao, and H.-M. Wang, “MOSNet: Deep learn- ing based objective assessment for voice conversion,” in Proc. Interspeech 2019 , 2019

  120. [128]

    NORESQA: A framework for speech quality assessment using non- matching references,

    P. Manocha, B. Xu, and A. Kumar, “NORESQA: A framework for speech quality assessment using non- matching references,” in Advances in Neural Informa- tion Processing Systems , vol. 34. Curran Associates, Inc., 2021, pp. 22 363–22 378

  121. [129]

    Deep learning-based non-intrusive multi-objective speech assessment model with cross- domain features,

    R. E. Zezario, S.-W. Fu, F. Chen, C.-S. Fuh, H.-M. Wang, and Y . Tsao, “Deep learning-based non-intrusive multi-objective speech assessment model with cross- domain features,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , vol. 31, pp. 54–70, 2022

  122. [130]

    APG-MOS: Auditory perception guided-mos predictor for synthetic speech,

    Z. Lian, L. Wang, and H. Huang, “APG-MOS: Auditory perception guided-mos predictor for synthetic speech,” arXiv preprint arXiv:2504.20447 , 2025

  123. [131]

    DPLM: A deep per- ceptual spatial-audio localization metric,

    P. Manocha, A. Kumar, B. Xu, A. Menon, I. D. Gebru, V . K. Ithapu, and P. Calamia, “DPLM: A deep per- ceptual spatial-audio localization metric,” in 2021 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA). IEEE, 2021, pp. 6–10

  124. [132]

    SAQAM: Spatial audio quality assessment met- ric,

    ——, “SAQAM: Spatial audio quality assessment met- ric,” in Proc. Interspeech 2022 , 2022, pp. 649–653

  125. [133]

    Spatialization quality metric for binaural speech,

    P. Manocha, I. D. Gebru, A. Kumar, D. Markovic, and A. Richard, “Spatialization quality metric for binaural speech,” in Proc. Interspeech 2023 , 2023, pp. 5426– 5430

  126. [134]

    HAPG-SAQAM: Human auditory percep- tion guided spatial audio quality assessment metric,

    Y . Zheng, J. Yao, X. Deng, Y . Yang, R. Liao, W. Tu, and C. Lin, “HAPG-SAQAM: Human auditory percep- tion guided spatial audio quality assessment metric,” in ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2025, pp. 1–5

  127. [135]

    A survey on bias and fairness in machine learning,

    N. Mehrabi, F. Morstatter, N. Saxena, K. Lerman, and A. Galstyan, “A survey on bias and fairness in machine learning,” ACM Computing Surveys (CSUR) , vol. 54, no. 6, pp. 1–35, 2021

  128. [136]

    Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead,

    C. Rudin, “Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead,” Nature Machine Intelligence , vol. 1, no. 5, pp. 206–215, 2019

  129. [137]

    Unbiased look at dataset bias,

    A. Torralba and A. A. Efros, “Unbiased look at dataset bias,” in CVPR 2011. IEEE, 2011, pp. 1521–1528

  130. [138]

    AudioGen: Textually guided audio generation,

    F. Kreuk, G. Synnaeve, A. Polyak, U. Singer, A. D ´efossez, J. Copet, D. Parikh, Y . Taigman, and Y . Adi, “AudioGen: Textually guided audio generation,” in The Eleventh International Conference on Learning Representations, 2023

  131. [139]

    AudioLDM: Text-to- audio generation with latent diffusion models,

    H. Liu, Z. Chen, Y . Yuan, X. Mei, X. Liu, D. Mandic, W. Wang, and M. D. Plumbley, “AudioLDM: Text-to- audio generation with latent diffusion models,” in Pro- ceedings of the International Conference on Machine Learning, 2023, pp. 21 450–21 474

  132. [140]

    Stable audio open,

    Z. Evans, J. D. Parker, C. Carr, Z. Zukowski, J. Taylor, and J. Pons, “Stable audio open,” in ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025, pp. 1–5

  133. [141]

    Visual to sound: Generating natural sound for videos in the wild,

    Y . Zhou, Z. Wang, C. Fang, T. Bui, and T. L. Berg, “Visual to sound: Generating natural sound for videos in the wild,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 3550–3558

  134. [142]

    Diff-Foley: Syn- chronized video-to-audio synthesis with latent diffusion models,

    S. Luo, C. Yan, C. Hu, and H. Zhao, “Diff-Foley: Syn- chronized video-to-audio synthesis with latent diffusion models,” in Advances in Neural Information Processing Systems, vol. 36. Curran Associates, Inc., 2023, pp. 48 855–48 876

  135. [143]

    Tri-Ergon: fine-grained video-to-audio generation with multi-modal conditions and lufs control,

    B. Li, F. Yang, Y . Mao, Q. Ye, H. Chen, and Y . Zhong, “Tri-Ergon: fine-grained video-to-audio generation with multi-modal conditions and lufs control,” in Proceed- ings of the AAAI Conference on Artificial Intelligence , vol. 39, no. 5, 2025, pp. 4616–4624

  136. [144]

    U-Net: Con- volutional networks for biomedical image segmenta- tion,

    O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Con- volutional networks for biomedical image segmenta- tion,” in Medical Image Computing and Computer- assisted Intervention–MICCAI 2015: 18th International Conference. Springer, 2015, pp. 234–241

  137. [145]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems , vol. 30. Curran Associates, Inc., 2017

  138. [146]

    Denoising diffusion probabilistic models,

    J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Advances in Neural Informa- tion Processing Systems , vol. 33. Curran Associates, Inc., 2020, pp. 6840–6851

  139. [147]

    An overview of multi-task learning in deep neural networks,

    S. Ruder, “An overview of multi-task learning in deep neural networks,” arXiv preprint arXiv:1706.05098 , 2017

  140. [148]

    Self-supervised generation of spatial audio for 360°video,

    P. Morgado, N. Nvasconcelos, T. Langlois, and O. Wang, “Self-supervised generation of spatial audio for 360°video,” in Advances in Neural Information Processing Systems, vol. 31. Curran Associates, Inc., 2018

  141. [149]

    2.5 D visual sound,

    R. Gao and K. Grauman, “2.5 D visual sound,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 324–333

  142. [150]

    Sep- stereo: Visually guided stereophonic audio generation by associating source separation,

    H. Zhou, X. Xu, D. Lin, X. Wang, and Z. Liu, “Sep- stereo: Visually guided stereophonic audio generation by associating source separation,” in Computer Vision– ECCV 2020: 16th European Conference . Springer, 2020, pp. 52–69

  143. [151]

    Cyclic learning for bin- aural audio generation and localization,

    Z. Li, B. Zhao, and Y . Yuan, “Cyclic learning for bin- aural audio generation and localization,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 26 669–26 678

  144. [152]

    A V- NeRF: learning neural fields for real-world audio-visual scene synthesis,

    S. Liang, C. Huang, Y . Tian, A. Kumar, and C. Xu, “A V- NeRF: learning neural fields for real-world audio-visual scene synthesis,” in Advances in Neural Information Processing Systems, vol. 36. Curran Associates, Inc., 2023, pp. 37 472–37 490

  145. [153]

    A V-GS: Learning material and geometry aware pri- JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2025 23 ors for novel view acoustic synthesis,

    S. Bhosale, H. Yang, D. Kanojia, J. Deng, and X. Zhu, “A V-GS: Learning material and geometry aware pri- JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2025 23 ors for novel view acoustic synthesis,” in Advances in Neural Information Processing Systems . Curran Associate...

  146. [154]

    SOAF: Scene occlusion-aware neural acoustic field,

    H. Gao, J. Ma, D. Ahmedt-Aristizabal, C. Nguyen, and M. Liu, “SOAF: Scene occlusion-aware neural acoustic field,” arXiv preprint arXiv:2407.02264 , 2024

  147. [155]

    A V-Surf: Surface-enhanced geometry- aware novel-view acoustic synthesis,

    H. Baek, H. Shin, J. Seo, C. Kim, S. Kim, H. Kim, and S. Kim, “A V-Surf: Surface-enhanced geometry- aware novel-view acoustic synthesis,” arXiv preprint arXiv:2503.12806, 2025

  148. [156]

    TAS: Personalized text- guided audio spatialization,

    Z. Li, B. Zhao, and Y . Yuan, “TAS: Personalized text- guided audio spatialization,” in Proceedings of the 32nd ACM International Conference on Multimedia , 2024, pp. 9029–9037

  149. [157]

    DualSpec: Text-to-spatial-audio generation via dual- spectrogram guided diffusion model,

    L. Zhao, S. Chen, L. Feng, X.-L. Zhang, and X. Li, “DualSpec: Text-to-spatial-audio generation via dual- spectrogram guided diffusion model,” arXiv preprint arXiv:2502.18952, 2025

  150. [158]

    AudioSpa: Spatializing sound events with text,

    L. Feng, L. Zhao, B. Zhu, X.-L. Zhang, and X. Li, “AudioSpa: Spatializing sound events with text,” arXiv preprint arXiv:2502.11219, 2025

  151. [159]

    ImmerseDiffusion: A generative spatial audio latent diffusion model,

    M. Heydari, M. Souden, B. Conejo, and J. Atkins, “ImmerseDiffusion: A generative spatial audio latent diffusion model,” in ICASSP 2025-2025 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025, pp. 1–5

  152. [160]

    SEE- 2-SOUND: Zero-shot spatial environment-to-spatial sound,

    R. Dagli, S. Prakash, R. Wu, and H. Khosravani, “SEE- 2-SOUND: Zero-shot spatial environment-to-spatial sound,” arXiv preprint arXiv:2406.06612 , 2024

  153. [161]

    Both ears wide open: Towards language-driven spatial audio generation,

    P. Sun, S. Cheng, X. Li, Z. Ye, H. Liu, H. Zhang, W. Xue, and Y . Guo, “Both ears wide open: Towards language-driven spatial audio generation,” in Interna- tional Conference on Learning Representations , 2024

  154. [162]

    ViSAGe: Video-to-spatial audio generation,

    J. Kim, H. Yun, and G. Kim, “ViSAGe: Video-to-spatial audio generation,” in The Thirteenth International Con- ference on Learning Representations , 2025

  155. [163]

    End-to-end binaural speech synthesis,

    W. C. Huang, D. Markovic, A. Richard, I. D. Gebru, and A. Menon, “End-to-end binaural speech synthesis,” in Proc. Interspeech 2022 , 2022

  156. [164]

    BinauralGrad: A two-stage conditional diffusion probabilistic model for binaural audio synthesis,

    Y . Leng, Z. Chen, J. Guo, H. Liu, J. Chen, X. Tan, D. Mandic, L. He, X. Li, T. Qin, s. zhao, and T.-Y . Liu, “BinauralGrad: A two-stage conditional diffusion probabilistic model for binaural audio synthesis,” in Advances in Neural Information Processing Systems , vol. 35. Cur...

  157. [165]

    DIFFBAS: An advanced binaural audio synthesis model focusing on binaural differences recovery,

    Y . Li, Y . Shen, and D. Wang, “DIFFBAS: An advanced binaural audio synthesis model focusing on binaural differences recovery,” Applied Sciences, vol. 14, no. 8, p. 3385, 2024

  158. [166]

    DopplerBAS: Binaural audio synthesis addressing doppler effect,

    J. Liu, Z. Ye, Q. Chen, S. Zheng, W. Wang, Q. Zhang, and Z. Zhao, “DopplerBAS: Binaural audio synthesis addressing doppler effect,” in Findings of the Associa- tion for Computational Linguistics: ACL 2023. Associ- ation for Computational Linguistics, 2023, pp. 11 905– 11 912

  159. [167]

    Dual position attention time-frequency network for binaural audio synthesis,

    C. He, W. Chen, and M. Wang, “Dual position attention time-frequency network for binaural audio synthesis,” in ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2025, pp. 1–5

  160. [168]

    Two- stage unet with Gated-Conv fusion for binaural audio synthesis,

    W. Zhang, C. He, Y . Cao, S. Xu, and M. Wang, “Two- stage unet with Gated-Conv fusion for binaural audio synthesis,” Sensors, vol. 25, no. 6, p. 1790, 2025

  161. [169]

    Neural fourier shift for binaural speech rendering,

    J. W. Lee and K. Lee, “Neural fourier shift for binaural speech rendering,” in ICASSP 2023-2023 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2023, pp. 1–5

  162. [170]

    Zero- shot mono-to-binaural speech synthesis,

    A. Levkovitch, J. Salazar, S. Mariooryad, R. Skerry- Ryan, N. Bar, B. Kleijn, and E. Nachmani, “Zero- shot mono-to-binaural speech synthesis,” arXiv preprint arXiv:2412.08356, 2024

  163. [171]

    Self-supervised audio spatialization with correspon- dence classifier,

    Y .-D. Lu, H.-Y . Lee, H.-Y . Tseng, and M.-H. Yang, “Self-supervised audio spatialization with correspon- dence classifier,” in 2019 IEEE International Confer- ence on Image Processing (ICIP) . IEEE, 2019, pp. 3347–3351

  164. [172]

    Visually informed binaural audio generation without binaural audios,

    X. Xu, H. Zhou, Z. Liu, B. Dai, X. Wang, and D. Lin, “Visually informed binaural audio generation without binaural audios,” in Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition , 2021, pp. 15 485–15 494

  165. [173]

    Binaural audio gen- eration via multi-task learning,

    S. Li, S. Liu, and D. Manocha, “Binaural audio gen- eration via multi-task learning,” ACM Transactions on Graphics (TOG), vol. 40, no. 6, pp. 1–13, 2021

  166. [174]

    Localize to binauralize: Audio spatialization from visual sound source localization,

    K. K. Rachavarapu, V . Sundaresha, A. Rajagopalan et al. , “Localize to binauralize: Audio spatialization from visual sound source localization,” in Proceedings of the IEEE/CVF International Conference on Com- puter Vision, 2021, pp. 1930–1939

  167. [175]

    Exploiting audio-visual consistency with partial supervision for spatial audio generation,

    Y .-B. Lin and Y .-C. F. Wang, “Exploiting audio-visual consistency with partial supervision for spatial audio generation,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 35, no. 3, 2021, pp. 2056– 2063

  168. [176]

    Multi-attention audio-visual fusion network for audio spatialization,

    W. Zhang and J. Shao, “Multi-attention audio-visual fusion network for audio spatialization,” in Proceedings of the 2021 International Conference on Multimedia Retrieval, 2021, pp. 394–401

  169. [177]

    Beyond mono to binaural: Generating binaural audio from mono audio with depth and cross modal attention,

    K. K. Parida, S. Srivastava, and G. Sharma, “Beyond mono to binaural: Generating binaural audio from mono audio with depth and cross modal attention,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2022, pp. 3347–3356

  170. [178]

    Points2Sound: from mono to binaural audio using 3D point cloud scenes,

    F. Llu ´ıs, V . Chatziioannou, and A. Hofmann, “Points2Sound: from mono to binaural audio using 3D point cloud scenes,” EURASIP Journal on Audio, Speech, and Music Processing , vol. 2022, no. 1, p. 33, 2022

  171. [179]

    Visually-guided au- dio spatialization in video with geometry-aware multi- task learning,

    R. Garg, R. Gao, and K. Grauman, “Visually-guided au- dio spatialization in video with geometry-aware multi- task learning,” International Journal of Computer Vi- sion, vol. 131, no. 10, pp. 2723–2737, 2023

  172. [180]

    Visually guided binaural audio generation with cross-modal con- sistency,

    M. Liu, J. Wang, X. Qian, and X. Xie, “Visually guided binaural audio generation with cross-modal con- sistency,” in ICASSP 2024-2024 IEEE International JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2025 24 Conference on Acoustics, Speech and Signal Processing (ICASSP)....

  173. [181]

    Cross-modal generative model for visual-guided binaural stereo generation,

    Z. Li, B. Zhao, and Y . Yuan, “Cross-modal generative model for visual-guided binaural stereo generation,” Knowledge-Based Systems, vol. 296, p. 111814, 2024

  174. [182]

    CCStereo: Audio-visual contextual and contrastive learning for binaural audio generation,

    Y . Chen, K. Shimada, C. Simon, Y . Ikemiya, T. Shibuya, and Y . Mitsufuji, “CCStereo: Audio-visual contextual and contrastive learning for binaural audio generation,” arXiv preprint arXiv:2501.02786 , 2025

  175. [183]

    OmniAudio: generating spatial audio from 360-degree video,

    H. Liu, T. Luo, K. Luo, Q. Jiang, P. Sun, J. Wang, R. Huang, Q. Chen, W. Wang, X. Li, S. Zhang, Z. Yan, Z. Zhao, and W. Xue, “OmniAudio: generating spatial audio from 360-degree video,” in Forty-second Interna- tional Conference on Machine Learning , 2025

  176. [184]

    NeRAF: 3D scene infused neural radiance and acoustic fields,

    A. Brunetto, S. Hornauer, and F. Moutarde, “NeRAF: 3D scene infused neural radiance and acoustic fields,” in The Thirteenth International Conference on Learning Representations, 2025

  177. [185]

    A V-Cloud: Spatial au- dio rendering through audio-visual cloud splatting,

    M. Chen and E. Shlizerman, “A V-Cloud: Spatial au- dio rendering through audio-visual cloud splatting,” in Advances in Neural Information Processing Systems , vol. 37. Curran Associates, Inc., 2024, pp. 141 021– 141 044

  178. [186]

    SoundVista: Novel- view ambient sound synthesis via visual-acoustic bind- ing,

    M. Chen, I. D. Gebru, I. Ananthabhotla, C. Richardt, D. Markovic, J. Sandakly, S. Krenn, T. Keebler, E. Shlizerman, and A. Richard, “SoundVista: Novel- view ambient sound synthesis via visual-acoustic bind- ing,” in Proceedings of the IEEE/CVF Conference on Computer Vision and...

  179. [187]

    In- the-wild audio spatialization with flexible text-guided localization,

    T. Pan, J. Liu, Z. Huang, J. Tang, and G. Wu, “In- the-wild audio spatialization with flexible text-guided localization,” in Proceedings of the Annual Meeting of the Association for Computational Linguistics , 2025

  180. [188]

    Diff-SAGe: End-to-end spatial audio gener- ation using diffusion models,

    S. S. Kushwaha, J. Ma, M. R. Thomas, Y . Tian, and A. Bruni, “Diff-SAGe: End-to-end spatial audio gener- ation using diffusion models,” in ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2025, pp. 1–5

  181. [189]

    ISDrama: Immersive spatial drama generation through multimodal prompting,

    Y . Zhang, W. Guo, C. Pan, Z. Zhu, T. Jin, and Z. Zhao, “ISDrama: Immersive spatial drama generation through multimodal prompting,” in ACM International Confer- ence on Multimedia (ACM MM) , 2025

  182. [190]

    INRAS: Implicit neural representation for audio scenes,

    K. Su, M. Chen, and E. Shlizerman, “INRAS: Implicit neural representation for audio scenes,” in Advances in Neural Information Processing Systems , vol. 35. Curran Associates, Inc., 2022, pp. 8144–8158

  183. [191]

    MESH2IR: Neural acoustic impulse response genera- tor for complex 3D scenes,

    A. Ratnarajah, Z. Tang, R. Aralikatti, and D. Manocha, “MESH2IR: Neural acoustic impulse response genera- tor for complex 3D scenes,” in Proceedings of the 30th ACM International Conference on Multimedia , 2022, pp. 924–933

  184. [192]

    Listen2Scene: Inter- active material-aware binaural sound propagation for reconstructed 3D scenes,

    A. Ratnarajah and D. Manocha, “Listen2Scene: Inter- active material-aware binaural sound propagation for reconstructed 3D scenes,” in 2024 IEEE Conference Virtual Reality and 3D User Interfaces (VR) . IEEE, 2024, pp. 254–264

  185. [193]

    Can large language models understand spatial audio?

    C. Tang, W. Yu, G. Sun, X. Chen, T. Tan, W. Li, J. Zhang, L. Lu, Z. Ma, Y . Wang et al. , “Can large language models understand spatial audio?” in Proc. Interspeech 2024, 2024

  186. [194]

    Learning spatially-aware language and audio embeddings,

    B. Devnani, S. Seto, Z. Aldeneh, A. Toso, E. Menyaylenko, B.-J. Theobald, J. Sheaffer, and M. Sarabia, “Learning spatially-aware language and audio embeddings,” in Advances in Neural Information Processing Systems, vol. 37. Curran Associates, Inc., 2024, pp. 33 505–33 537

  187. [195]

    BAT: learning to reason about spatial sounds with large language models,

    Z. Zheng, P. Peng, Z. Ma, X. Chen, E. Choi, and D. Harwath, “BAT: learning to reason about spatial sounds with large language models,” in Proceedings of the 41st International Conference on Machine Learning. JMLR.org, 2024

  188. [196]

    SoundSpaces: Audio-visual navigation in 3D environ- ments,

    C. Chen, U. Jain, C. Schissler, S. V . A. Gari, Z. Al- Halah, V . K. Ithapu, P. Robinson, and K. Grauman, “SoundSpaces: Audio-visual navigation in 3D environ- ments,” in Computer Vision–ECCV 2020: 16th Euro- pean Conference. Springer, 2020, pp. 17–36

  189. [197]

    SoundSpaces 2.0: a simulation platform for visual- acoustic learning,

    C. Chen, C. Schissler, S. Garg, P. Kobernik, A. Clegg, P. Calamia, D. Batra, P. Robinson, and K. Grauman, “SoundSpaces 2.0: a simulation platform for visual- acoustic learning,” in Advances in Neural Information Processing Systems, vol. 35. Curran Associates, Inc., 2022, pp. 8896–8911

  190. [198]

    GW A: a large high-quality acoustic dataset for audio processing,

    Z. Tang, R. Aralikatti, A. J. Ratnarajah, and D. Manocha, “GW A: a large high-quality acoustic dataset for audio processing,” in ACM SIGGRAPH 2022 Conference Proceedings, ser. SIGGRAPH ’22. New York, NY , USA: Association for Computing Machinery, 2022

  191. [199]

    Replay: Multi-modal multi-view acted videos for casual holography,

    R. Shapovalov, Y . Kleiman, I. Rocco, D. Novotny, A. Vedaldi, C. Chen, F. Kokkinos, B. Graham, and N. Neverova, “Replay: Multi-modal multi-view acted videos for casual holography,” in Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, 2023, pp. 20 338–20 348

  192. [200]

    SoundCam: A dataset for finding humans using room acoustics,

    M. Wang, S. Clarke, J.-H. Wang, R. Gao, and J. Wu, “SoundCam: A dataset for finding humans using room acoustics,” in Advances in Neural Information Process- ing Systems , vol. 36. Curran Associates, Inc., 2023, pp. 52 238–52 264

  193. [201]

    Real Acoustic Fields: an audio-visual room acoustics dataset and benchmark,

    Z. Chen, I. D. Gebru, C. Richardt, A. Kumar, W. Laney, A. Owens, and A. Richard, “Real Acoustic Fields: an audio-visual room acoustics dataset and benchmark,” in Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, 2024, pp. 21 886– 21 896

  194. [202]

    RealMAN: a real-recorded and annotated microphone array dataset for dynamic speech enhancement and localization,

    B. Yang, C. Quan, Y . Wang, P. Wang, Y . Yang, Y . Fang, N. Shao, H. Bu, X. Xu, and X. Li, “RealMAN: a real-recorded and annotated microphone array dataset for dynamic speech enhancement and localization,” in Advances in Neural Information Processing Systems , vol. 37. Curran ...

  195. [203]

    SonicSim: a customizable simulation platform for speech processing in moving sound source sce- JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2025 25 narios,

    K. Li, W. Sang, C. Zeng, R. Yang, G. Chen, and X. Hu, “SonicSim: a customizable simulation platform for speech processing in moving sound source sce- JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2025 25 narios,” in The Thirteenth International Conference on Learning Re...

  196. [204]

    Nord: Non-matching reference based relative depth estimation from binaural speech,

    P. Manocha, I. D. Gebru, A. Kumar, D. Markovic, and A. Richard, “Nord: Non-matching reference based relative depth estimation from binaural speech,” in ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5

  197. [205]

    Matterport3D: learning from RGB-D data in indoor environments,

    A. Chang, A. Dai, T. Funkhouser, M. Halber, M. Nieb- ner, M. Savva, S. Song, A. Zeng, and Y . Zhang, “Matterport3D: learning from RGB-D data in indoor environments,” in 2017 International Conference on 3D Vision (3DV), 2017, pp. 667–676

  198. [206]

    Novel-view acoustic synthesis,

    C. Chen, A. Richard, R. Shapovalov, V . K. Ithapu, N. Neverova, K. Grauman, and A. Vedaldi, “Novel-view acoustic synthesis,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2023, pp. 6409–6419

  199. [207]

    End-to-end paired ambisonic- binaural audio rendering,

    Y . Zhu, Q. Kong, J. Shi, S. Liu, X. Ye, J.-C. Wang, H. Shan, and J. Zhang, “End-to-end paired ambisonic- binaural audio rendering,” IEEE/CAA Journal of Auto- matica Sinica, vol. 11, no. 2, pp. 502–513, 2024

  200. [208]

    Fr´echet audio distance: A metric for evaluating music enhancement algorithms,

    K. Kilgour, M. Zuluaga, D. Roblek, and M. Sharifi, “Fr´echet audio distance: A metric for evaluating music enhancement algorithms,” in Proc. Interspeech 2019 , 2019

  201. [209]

    The effect of audio on the experience in virtual reality: a scoping review,

    I. d. V . Bosman, O. O. Buruk, K. Jørgensen, and J. Hamari, “The effect of audio on the experience in virtual reality: a scoping review,” Behaviour & Infor- mation Technology, vol. 43, no. 1, pp. 165–199, 2024

  202. [210]

    Spatial audio in virtual reality: A systematic review,

    G. Corr ˆea De Almeida, V . Costa de Souza, L. G. Da Silveira J´unior, and M. R. Veronez, “Spatial audio in virtual reality: A systematic review,” in Proceedings of the 25th Symposium on Virtual and Augmented Reality , 2023, pp. 264–268

  203. [211]

    Place illusion and plausibility can lead to realistic behaviour in immersive virtual environments,

    M. Slater, “Place illusion and plausibility can lead to realistic behaviour in immersive virtual environments,” Philosophical Transactions of the Royal Society B: Biological Sciences, vol. 364, no. 1535, pp. 3549–3557, 2009

  204. [212]

    Beyond reality,

    A. Eames, “Beyond reality,” in Proceedings of the 17th ACM SIGGRAPH International Conference on Virtual- Reality Continuum and its Applications in Industry , 2019, pp. 1–2

  205. [213]

    Brain activa- tion in virtual reality for attention guidance,

    P. Ulsamer, K. Pfeffel, and N. H. M ¨uller, “Brain activa- tion in virtual reality for attention guidance,” in Inter- national Conference on Human-Computer Interaction . Springer, 2020, pp. 190–200

  206. [214]

    Augmented/mixed reality audio for hear- ables: Sensing, control, and rendering,

    R. Gupta, J. He, R. Ranjan, W.-S. Gan, F. Klein, C. Schneiderwind, A. Neidhardt, K. Brandenburg, and V . V¨alim¨aki, “Augmented/mixed reality audio for hear- ables: Sensing, control, and rendering,” IEEE Signal Processing Magazine, vol. 39, no. 3, pp. 63–89, 2022

  207. [215]

    The benefit of binaural hearing in a cocktail party: Effect of location and type of interferer,

    M. L. Hawley, R. Y . Litovsky, and J. F. Culling, “The benefit of binaural hearing in a cocktail party: Effect of location and type of interferer,” The Journal of the Acoustical Society of America, vol. 115, no. 2, pp. 833– 843, 2004

  208. [216]

    Sixty years of frequency-domain monaural speech enhancement: From traditional to deep learning methods,

    C. Zheng, H. Zhang, W. Liu, X. Luo, A. Li, X. Li, and B. C. Moore, “Sixty years of frequency-domain monaural speech enhancement: From traditional to deep learning methods,” Trends in Hearing , vol. 27, p. 23312165231209913, 2023

  209. [217]

    The virtual reality lab: Realization and application of virtual sound environments,

    V . Hohmann, R. Paluch, M. Krueger, M. Meis, and G. Grimm, “The virtual reality lab: Realization and application of virtual sound environments,” Ear and Hearing, vol. 41, pp. 31S–38S, 2020

  210. [218]

    Virtual-reality-based research in hearing science: a platforming approach,

    R. L. Pedersen, L. Picinali, N. Kajs, and F. Patou, “Virtual-reality-based research in hearing science: a platforming approach,” Journal of the Audio Engineer- ing Society, vol. 71, no. 6, pp. 374–389, 2023

  211. [219]

    AI-driven innovations in hearing health: A review of artificial intelligence applications in audiology and hearing technologies,

    S. Chitra Thara, K. Vidhya Lekshmi, and N. Venkateswaramurthy, “AI-driven innovations in hearing health: A review of artificial intelligence applications in audiology and hearing technologies,” Current Aging Science , 2025

  212. [220]

    Pay self- attention to audio-visual navigation,

    Y . Yu, L. Cao, F. Sun, X. Liu, and L. Wang, “Pay self- attention to audio-visual navigation,” in British Machine Vision Conference (BMVC) . British Machine Vision Association, 2022

  213. [221]

    Catch me if you hear me: Audio-visual naviga- tion in complex unmapped environments with moving sounds,

    A. Younes, D. Honerkamp, T. Welschehold, and A. Val- ada, “Catch me if you hear me: Audio-visual naviga- tion in complex unmapped environments with moving sounds,” IEEE Robotics and Automation Letters , vol. 8, no. 2, pp. 928–935, 2023

  214. [222]

    Knowledge-driven scene priors for semantic audio-visual embodied navigation,

    G. Tatiya, J. Francis, L. Bondi, I. Navarro, E. Nyberg, J. Sinapov, and J. Oh, “Knowledge-driven scene priors for semantic audio-visual embodied navigation,” arXiv preprint arXiv:2212.11345, 2022

  215. [223]

    Seman- tic audio-visual navigation,

    C. Chen, Z. Al-Halah, and K. Grauman, “Seman- tic audio-visual navigation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 15 516–15 525

  216. [224]

    Sound adversarial audio-visual navigation,

    Y . Yu, W. Huang, F. Sun, C. Chen, Y . Wang, and X. Liu, “Sound adversarial audio-visual navigation,” in The Tenth International Conference on Learning Representations, 2022

  217. [225]

    Sim2Real transfer for audio-visual navigation with frequency-adaptive acoustic field prediction,

    C. Chen, J. Ramos, A. Tomar, and K. Grau- man, “Sim2Real transfer for audio-visual navigation with frequency-adaptive acoustic field prediction,” in 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2024, pp. 8595– 8602

  218. [226]

    Learning to set waypoints for audio-visual navigation,

    C. Chen, S. Majumder, A.-H. Ziad, R. Gao, S. Ku- mar Ramakrishnan, and K. Grauman, “Learning to set waypoints for audio-visual navigation,” in The Ninth International Conference on Learning Representations , 2021

  219. [227]

    Stereo depth estimation with echoes,

    C. Zhang, K. Tian, B. Ni, G. Meng, B. Fan, Z. Zhang, and C. Pan, “Stereo depth estimation with echoes,” in European Conference on Computer Vision . Springer, 2022, pp. 496–513

  220. [228]

    Beyond visual field of view: Perceiving 3D environment with echoes and vision,

    L. Zhu, E. Rahtu, and H. Zhao, “Beyond visual field of view: Perceiving 3D environment with echoes and vision,” arXiv preprint arXiv:2207.01136 , 2022

  221. [229]

    Be- JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2025 26 yond image to depth: Improving depth prediction using echoes,

    K. K. Parida, S. Srivastava, and G. Sharma, “Be- JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2025 26 yond image to depth: Improving depth prediction using echoes,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 8268–8277

  222. [230]

    Dense 2D-3D indoor prediction with sound via aligned cross-modal distil- lation,

    H. Yun, J. Na, and G. Kim, “Dense 2D-3D indoor prediction with sound via aligned cross-modal distil- lation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 7863–7872

  223. [231]

    Semantic object prediction and spatial sound super-resolution with binaural sounds,

    A. B. Vasudevan, D. Dai, and L. Van Gool, “Semantic object prediction and spatial sound super-resolution with binaural sounds,” in European Conference on Computer Vision. Springer, 2020, pp. 638–655

  224. [232]

    3D audio-visual segmentation,

    A. Sokolov, S. Bhosale, and X. Zhu, “3D audio-visual segmentation,” in NeurIPS 2024 Workshop on Audio Imagination, 2024

  225. [233]

    VisualEchoes: Spatial image repre- sentation learning through echolocation,

    R. Gao, C. Chen, Z. Al-Halah, C. Schissler, and K. Grauman, “VisualEchoes: Spatial image repre- sentation learning through echolocation,” in Com- puter Vision–ECCV 2020: 16th European Conference . Springer, 2020, pp. 658–676

  226. [234]

    Generating diverse audio-visual 360 soundscapes for sound event localization and detection,

    A. S. Roman, A. Chang, G. Meza, and I. R. Roman, “Generating diverse audio-visual 360 soundscapes for sound event localization and detection,” arXiv preprint arXiv:2504.02988, 2025

  227. [235]

    Sounding Bodies: modeling 3D spa- tial sound of humans using body pose and audio,

    X. XU, D. Markovic, J. Sandakly, T. Keebler, S. Krenn, and A. Richard, “Sounding Bodies: modeling 3D spa- tial sound of humans using body pose and audio,” in Advances in Neural Information Processing Systems , vol. 36. Curran Associates, Inc., 2023, pp. 44 740– 44 752

  228. [236]

    Fed- erated learning: Challenges, methods, and future direc- tions,

    T. Li, A. K. Sahu, A. Talwalkar, and V . Smith, “Fed- erated learning: Challenges, methods, and future direc- tions,” IEEE Signal Processing Magazine, vol. 37, no. 3, pp. 50–60, 2020

  229. [237]

    Fedaudio: A feder- ated learning benchmark for audio tasks,

    T. Zhang, T. Feng, S. Alam, S. Lee, M. Zhang, S. S. Narayanan, and S. Avestimehr, “Fedaudio: A feder- ated learning benchmark for audio tasks,” in ICASSP 2023-2023 IEEE International Conference on Acous- tics, Speech and Signal Processing (ICASSP) . IEEE, 2023, pp. 1–5

  230. [238]

    Evaluating HRTF similarity through subjective assessments: Factors that can affect judgment,

    A. Andreopoulou and A. Roginska, “Evaluating HRTF similarity through subjective assessments: Factors that can affect judgment,” in Proceedings - 40th Interna- tional Computer Music Conference, ICMC 2014 and 11th Sound and Music Computing Conference, SMC 2014 - Music Technology...

  231. [239]

    Investigation into consistency of subjective and objective perceptual selec- tion of non-individual head-related transfer functions,

    C. Kim, V . Lim, and L. Picinali, “Investigation into consistency of subjective and objective perceptual selec- tion of non-individual head-related transfer functions,” Journal of the Audio Engineering Society , vol. 68, no. 11, pp. 819–831, 2020

  232. [240]

    Compar- ing subjective similarity ratings and quantitative er- rors for the evaluation of free-field binaural panning techniques,

    Z. T. Rusk, M. Neal, and M. C. Vigeant, “Compar- ing subjective similarity ratings and quantitative er- rors for the evaluation of free-field binaural panning techniques,” The Journal of the Acoustical Society of America, vol. 155, no. 3 Supplement, pp. A215–A215, 2024

  233. [241]

    3D Tune-In Toolkit: An open- source library for real-time binaural spatialisation,

    M. Cuevas-Rodr ´ıguez, L. Picinali, D. Gonz ´alez-Toledo, C. Garre, E. de la Rubia-Cuestas, L. Molina-Tanco, and A. Reyes-Lecuona, “3D Tune-In Toolkit: An open- source library for real-time binaural spatialisation,”PloS one, vol. 14, no. 3, p. e0211899, 2019

  234. [242]

    A data-driven explo- ration of elevation cues in HRTFs: An explainable AI perspective across multiple datasets,

    J. A. De Rus, M. Montagud, J. Lopez-Ballester, F. J. Ferri, and M. Cobos, “A data-driven explo- ration of elevation cues in HRTFs: An explainable AI perspective across multiple datasets,” arXiv preprint arXiv:2503.11312, 2025

  235. [243]

    Domain-adversarial training of neural networks,

    Y . Ganin, E. Ustinova, H. Ajakan, P. Germain, H. Larochelle, F. Laviolette, M. March, and V . Lempit- sky, “Domain-adversarial training of neural networks,” Journal of Machine Learning Research, vol. 17, no. 59, pp. 1–35, 2016

  236. [244]

    A survey of unsupervised deep domain adaptation,

    G. Wilson and D. J. Cook, “A survey of unsupervised deep domain adaptation,” ACM Transactions on Intelli- gent Systems and Technology (TIST), vol. 11, no. 5, pp. 1–46, 2020

  237. [245]

    Deep convolutional neu- ral networks and data augmentation for environmental sound classification,

    J. Salamon and J. P. Bello, “Deep convolutional neu- ral networks and data augmentation for environmental sound classification,” IEEE Signal Processing Letters , vol. 24, no. 3, pp. 279–283, 2017

  238. [246]

    Generative adversarial nets,

    I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y . Bengio, “Generative adversarial nets,” in Advances in Neural Information Processing Systems , vol. 27. Curran Associates, Inc., 2014

  239. [247]

    Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations,

    M. Raissi, P. Perdikaris, and G. E. Karniadakis, “Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations,” Jour- nal of Computational Physics , vol. 378, pp. 686–707, 2019

  240. [248]

    Physics-informed machine learning,

    G. E. Karniadakis, I. G. Kevrekidis, L. Lu, P. Perdikaris, S. Wang, and L. Yang, “Physics-informed machine learning,” Nature Reviews Physics , vol. 3, no. 6, pp. 422–440, 2021

  241. [249]

    Physics and geometry informed neural operator net- work with application to acoustic scattering,

    S. Nair, T. F. Walsh, G. Pickrell, and F. Semperlotti, “Physics and geometry informed neural operator net- work with application to acoustic scattering,” arXiv preprint arXiv:2406.03407, 2024

  242. [250]

    Physics- informed neural network for volumetric sound field reconstruction of speech signals,

    M. Olivieri, X. Karakonstantis, M. Pezzoli, F. An- tonacci, A. Sarti, and E. Fernandez-Grande, “Physics- informed neural network for volumetric sound field reconstruction of speech signals,” EURASIP Journal on Audio, Speech, and Music Processing , vol. 2024, no. 1, p. 42, 2024

  243. [251]

    Implicit neural representation with physics-informed neural networks for the reconstruction of the early part of room impulse responses,

    M. Pezzoli, F. Antonacci, and A. Sarti, “Implicit neural representation with physics-informed neural networks for the reconstruction of the early part of room impulse responses,” in 10th Convention of the European Acous- tics Association, 2023, pp. 2127–2184

  244. [252]

    Physics-informed machine learning for sound field estimation: Fundamentals, state of the art, and challenges,

    S. Koyama, J. G. C. Ribeiro, T. Nakamura, N. Ueno, and M. Pezzoli, “Physics-informed machine learning for sound field estimation: Fundamentals, state of the art, and challenges,” IEEE Signal Processing Magazine, JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2025 27 vol....

  245. [253]

    Spatial interpolation of head- related transfer functions using a physics-informed au- toencoder,

    W. Chen and X. Wei, “Spatial interpolation of head- related transfer functions using a physics-informed au- toencoder,” Multimedia Systems, vol. 31, no. 3, p. 247, 2025

  246. [254]

    Acoustic field reconstruction in tubes via physics-informed neural networks,

    X. Luan, K. Yokota, and G. Scavone, “Acoustic field reconstruction in tubes via physics-informed neural networks,” arXiv preprint arXiv:2505.12557 , 2025

  247. [255]

    Explain- able artificial intelligence: Understanding, visualizing and interpreting deep learning models,

    W. Samek, T. Wiegand, and K.-R. M ¨uller, “Explain- able artificial intelligence: Understanding, visualizing and interpreting deep learning models,” arXiv preprint arXiv:1708.08296, 2017

  248. [256]

    beta- V AE: learning basic visual concepts with a constrained variational framework,

    I. Higgins, L. Matthey, A. Pal, C. Burgess, X. Glorot, M. Botvinick, S. Mohamed, and A. Lerchner, “beta- V AE: learning basic visual concepts with a constrained variational framework,” in International Conference on Learning Representations, 2017

  249. [257]

    Distilling the knowledge in a neural network,

    G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” arXiv preprint arXiv:1503.02531, 2015

  250. [258]

    Learning both weights and connections for efficient neural network,

    S. Han, J. Pool, J. Tran, and W. Dally, “Learning both weights and connections for efficient neural network,” in Advances in Neural Information Processing Systems , vol. 28. Curran Associates, Inc., 2015

  251. [259]

    Progressive distillation for fast sampling of diffusion models,

    T. Salimans and J. Ho, “Progressive distillation for fast sampling of diffusion models,” in International Conference on Learning Representations , 2022

  252. [260]

    Denoising diffusion implicit models,

    J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” in International Conference on Learn- ing Representations, 2021

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.