REVIEW 4 major objections 7 minor 1 cited by
Binaural Sound Event Localization and Detection based on HRTF Cues for Humanoid Robots
T0 review · 4 major / 7 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read HRTF cue maps let two microphones localize sounds to 4.4 degrees
desk verdict The feature design and ablation logic are solid, but the headline 4.4° localization error is measured on the same 48-direction HRTF grid used for training, so it may reflect grid-node recognition rather than continuous spatial hearing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Binaural Time-Frequency Feature (BTFF), an eight-channel input map built from a binaural pair: left and right mel-spectrograms; left and right velocity maps (time-differences of the magnitude spectrogram); an ITD-map, computed as $\frac{1}{\omega}\mathrm{Im}[\ln(P_R/P_L)]$ for bins below 1.5 kHz and projected to the mel scale; an ILD-map, $10\log_{10}|P_R/P_L|^2$ for bins above 5 kHz, also mel-projected; and left and right SC-maps, mel spectra restricted to bands above 5 kHz. These channels are the mechanism because they parse the HRTF's spatial information into complementary, frequency-segregated cues before the network sees the data: ITD carries low-frequency azimuth, ILD carries high-frequency azimuth plus front-back asymmetry, and SC carries elevation. The network itself, BiSELDnet, is a CRNN with depthwise separable convolutions, bidirectional GRUs, and a fully connected head that outputs one 3D direction vector per event class per frame, using the activity-coupled Cartesian DOA training target.
What would settle it
Render binaural test clips from HRIRs at directions not seen in training (for example, 15 degree rather than 30 degree azimuth steps, or elevations at 45 degrees), or record real binaural sounds with a manikin head, and rerun BiSELDnet; if localization error jumps well above 4.4 degrees or detection F-score drops, the central generalization claim fails.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that each class of HRTF-derived cue maps to a specific sub-problem of binaural SELD, and encoding them explicitly makes the whole task learnable from two channels. The velocity map improves detection by marking onsets and transients; the ITD-map, derived from the imaginary part of the log spectral ratio below 1.5 kHz without phase unwrapping, and the ILD-map, computed as log-power difference above 5 kHz, jointly resolve azimuth, with ILD's front-back asymmetry disambiguating ITD's symmetry; and the SC-map, mel bands above 5 kHz, supplies elevation-dependent pinna notch cues. With these eight channels, BiSELDnet detects all 12 event classes and outputs a 3D direction vector per class per frame, reaching a 4.4 degree localization error and 92.1% localization recall on the Binaural Set test split. The authors take this as evidence that HRTF-based spatial cues integrated as input features are a viable path to binaural 3D SELD in humanoid robots.
Load-bearing premise
The load-bearing premise is that performance measured on synthetic mixtures made from a fixed 48-direction HRTF grid transfers to real binaural hearing; if the grid is not representative of continuous, real-world directions, the reported accuracy is a property of the benchmark rather than of the robot's ears.
Editorial extensions
If this is right
- A two-channel binaural front end is enough for joint 3D detection and localization in the synthetic setting, so robots can avoid the size, calibration, and data costs of four-channel arrays.
- Explicitly separating ITD, ILD, and spectral cues makes each cue's contribution inspectable: V-map for detection, ITD/ILD for azimuth, and SC for elevation.
- The reported 4.4 degree localization error and 92.1% recall imply the model can pick out which of 12 classes is active and point to it with near-grid resolution on the Binaural Set.
- The Binaural Set itself, with clean and noisy conditions and separate horizontal and vertical test subsets, gives the community a controlled benchmark for comparing binaural SELD methods.
Reading between the lines
- Because training and test directions come from the same 48-point HRTF grid, the 4.4 degree error may partly measure interpolation inside that grid; the paper does not establish accuracy on directions between grid points or on real recorded binaural scenes.
- A natural next test is to evaluate the same BTFF on continuous azimuths and on real binaural recordings from a manikin head; if accuracy degrades, the cues remain valid but the current benchmark overstates deployable performance.
- The feature design suggests a hearing-aid or telepresence variant could work with generic HRTFs per user, provided individual pinna cues are preserved; the SC-map channel is the load-bearing part for elevation and the most likely to need per-listener adaptation.
- One could ablate the hand-crafted maps against learned spatial features, for example a network given only the two mel-spectrograms, to quantify how much of the 4.4 degree result comes from the engineered cues versus the CRNN itself.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes BiSELD, a binaural (two-channel) variant of sound event localization and detection, together with a synthetic benchmark dataset (the Binaural Set) built by convolving isolated event recordings from NIGENS and DCASE2016 Task 2 with measured HRIRs from the authors' KAIST HRTF database at 12 azimuths and 4 elevations (48 directions). The proposed input representation BTFF concatenates eight channels—left/right mel-spectrograms, left/right velocity maps, an ITD map (phase-derived delay below 1.5 kHz), an ILD map (above 5 kHz), and left/right spectral-cue maps (mel bands above 5 kHz)—and feeds a 763K-parameter CRNN, BiSELDnet, with ACCDOA-style per-class Cartesian DOA outputs. Ablation experiments show consistent gains from each feature group: V-map improves detection, ITD/ILD maps reduce horizontal-plane localization error (LE from 17.3° to 4.2°), and the SC map reduces median-plane localization error (25.2° to 12.2°). The full system is reported to achieve a SELD error of 0.110, F-score of 87.1%, LE_CD of 4.4°, and LR_CD of 92.1% on the test set.
Significance. If the results hold, the work provides a compact, lightweight two-channel SELD pipeline whose feature design is directly grounded in psychoacoustic cues (ITD below 1.5 kHz, ILD above 5 kHz, pinna-related spectral notches), which is a plausible and falsifiable design hypothesis of genuine interest to the humanoid robotics and binaural-audio communities. The internal ablation logic is clean and consistently executed: each feature is added to a fixed mel-spectrogram baseline and evaluated on the sub-task it is designed for, with ten independent training runs and median reporting in Tables IV–VI. The authors are also transparent about the synthetic nature of the benchmark and about their use of the measured KAIST HRTF database. The main weaknesses are the scope of the evaluation—training and testing on the same 48-direction grid, with no off-grid, reverberant, or real binaural signals—the best-of-ten reporting in Table VII, the absence of any external SELD baseline, and several dataset-description inconsistencies. These issues are local and fixable, but they currently limit the practical force of the headline localization claims.
major comments (4)
- [III-C and Table VII] The entire evaluation is confined to the same 48-direction KAIST HRTF grid used for training, with 30° azimuth and 30° elevation spacing; no test sample is rendered from an off-grid direction, a reverberant scene, a second head or ear geometry, or a real binaural recording. The headline LE_CD of 4.4° in Table VII is an order of magnitude finer than the grid spacing and is consistent with near-perfect selection of the correct grid node rather than with interpolation or continuous spatial inference. Because the abstract and title make claims about 'human-like auditory perception' for humanoid robots, the authors should add an off-grid evaluation (for example, directions at 15° offsets in azimuth and elevation from the training grid) and, ideally, a different HRTF set or real binaural data; until then, the localization claims should be explicitly re-scoped to the grid-matched synthetic setting.
- [V-B4, Table VII versus Tables IV–VI] Tables IV–VI report median values over ten training runs, but Table VII and the abstract report the 'best' performance of BiSELDnet with no measure of variance; the headline SELD error of 0.110, F-score of 87.1%, and LE_CD of 4.4° are therefore the upper envelope of the runs and are not statistically comparable with the median-based ablation tables. The authors should report the median and spread (standard deviation or the full range) for the final configuration in Table VII and ensure that the abstract uses the same statistic.
- [Section V and Table I] The evaluation contains no comparison with any external SELD or binaural-localization baseline, even though Table I identifies Wilkins et al. [64] as the state of the art for binaural SELD; without at least one matched baseline trained and evaluated under the same protocol on the Binaural Set, the absolute performance figures have no external calibration and the 'effectiveness' claim rests entirely on self-ablations. Please add such a comparison or explicitly limit the claims to the feature-ablation findings.
- [III-C, Tables II and III] The dataset description is internally inconsistent: the text states that 'for each sound class, 20 samples were prepared and split into training, validation, and test sets in a 14:3:3 ratio,' but Table III reports 672/144/144 mixtures whose construction requires far more than 20 event clips per class, and the 14:3:3 ratio does not correspond to 672/144/144 (the actual per-class split implied by Table III is 56/12/12). Additionally, NIGENS is described as having 14 classes but Table II lists 15 class names, the 12 classes actually used in BiSELD are never enumerated, and the exact direction sets underlying Test-H and Test-V are not specified; these details must be corrected and documented for the Binaural Set to serve as a reproducible benchmark.
minor comments (7)
- [Conclusion versus Section IV] The Conclusion states that BiSELDnet is built with depthwise separable convolutions, but Section IV and Fig. 4 describe only generic convolution, normalization, activation, and pooling modules; the architecture description and the Conclusion should be reconciled.
- [Section IV] The detection rule 'exceeds 0.5v' appears to refer to a magnitude threshold of 0.5 on the output DOA vector, but the notation is undefined; it should be stated explicitly.
- [Section III-C] Test-H and Test-V are defined only by their sample counts (36 and 12); the exact azimuth and elevation grid points in each subset should be stated so that Tables V and VI are interpretable from the text alone.
- [Section III-A3] The spectral-notch analysis quotes elevations from −40° to 90°, whereas the Binaural Set covers only −30° to +60°; the two ranges should be harmonized or the discrepancy explained.
- [Section III-C] The paper should state explicitly that polyphony in the Binaural Set is restricted to across-class overlap, since each mixture contains exactly one instance of each class, and should note the corresponding limitation relative to the same-class overlap scenarios addressed by multi-ACCDOA.
- [Abstract and Table VI] The abstract's '4.4° localization error' should be qualified: Table VI shows that the median-plane localization error remains 12.2° even with the SC map, so the 4.4° LE_CD is dominated by horizontal-plane performance rather than uniform 3D accuracy.
- [Section IV] Section IV reports only the total parameter count (763,020); layer-wise specifications such as convolution filter counts, GRU units, and pooling sizes are needed to reproduce BiSELDnet from the text.
Circularity Check
No load-bearing circularity: each BTFF component is validated by in-paper ablation on a held-out split, and the self-citations supply data and design precedents rather than the predicted result.
full rationale
The claimed derivation chain is: psychoacoustic analysis of HRIRs (Sec. III-A) → construction of the eight-channel BTFF (Eqs. 3–9) → supervised training of BiSELDnet on the synthetic Binaural Set → measured SELD error of 0.110 and LE_CD of 4.4° on a held-out split (Table VII). Every link is empirical rather than definitional. The ITD-, ILD-, and SC-maps are computed from raw binaural audio (Eqs. 5–9), while the azimuth/elevation labels come from the synthesis procedure (Sec. III-C, 14:3:3 split, Table III), so the input features are not defined in terms of the labels, and no fitted parameter is renamed as a prediction. The sub-feature claims are supported by in-paper ablations on held-out data (V-map: SED error 0.321→0.286; +ITD+ILD: LE 17.3°→4.2° on Test-H; +SC: LE 25.2°→12.2° on Test-V), so they are measured, not assumed. The self-citations ([65] KAIST HRTF database, [70] velocity features, [71][72] ITD estimation, [79] preliminary ICA version, [16] Ph.D. thesis) provide a measured data asset and design precedents, but none is invoked as the proof of the target result; the load-bearing evidence for each claim is in-paper and reproducible from the described pipeline. Two caveats are real but fall outside circularity. First, Section III-C builds the entire dataset from 12 azimuths × 4 elevations of the same KAIST HRTF grid, and all test splits draw from this same grid, so the 4.4° LE_CD (far below the 30° spacing) is demonstrated only on that discrete grid, with no off-grid, reverberant, or real-binaural evaluation; this limits external validity and belongs under correctness risk rather than circularity. Second, Tables IV–VI report median values over ten runs while Table VII reports the 'BEST PERFORMANCES', so the headline number is the upper envelope of the runs; this is a reporting caveat, not a reduction of prediction to input. Nothing in the construction forces the network to achieve these numbers, so the derivation is not circular.
Assumptions & free parameters
free parameters (4)
- ITD low-frequency cutoff =
1.5 kHz
- ILD/SC high-frequency cutoff =
5 kHz
- Event detection threshold =
0.5 (vector magnitude)
- Spatial direction grid =
12 azimuth x 4 elevation = 48 directions
assumptions (5)
- domain assumption Measured KAIST HRTF database accurately captures binaural localization cues for a humanoid robot with human-like ears.
- domain assumption Synthetic binaural mixtures created by convolution with anechoic HRIRs and additive background noise are representative of real-world binaural audio.
- domain assumption The NIGENS and DCASE2016 Task 2 sound event databases provide sufficient coverage of the target 12 sound classes.
- ad hoc to paper The ITD, ILD, and spectral notch cues are the dominant cues for binaural localization and can be captured by the proposed maps.
- domain assumption Evaluation on the synthetic test set is a valid measure of sound event localization and detection performance.
Cite this review
Pith. "Pith review of Binaural Sound Event Localization and Detection based on HRTF Cues for Humanoid Robots." pith.science (2026). https://pith.science/paper/7PAFJBMM
@misc{pith2026250720530,
author = {Pith},
title = {Pith review of: Binaural Sound Event Localization and Detection based on HRTF Cues for Humanoid Robots},
year = {2026},
howpublished = {\url{https://pith.science/paper/7PAFJBMM}},
note = {Machine review of arXiv:2507.20530}
}
read the original abstract
This paper introduces Binaural Sound Event Localization and Detection (BiSELD), a task that aims to jointly detect and localize multiple sound events using binaural audio, inspired by the spatial hearing mechanism of humans. To support this task, we present a synthetic benchmark dataset, called the Binaural Set, which simulates realistic auditory scenes using measured head-related transfer functions (HRTFs) and diverse sound events. To effectively address the BiSELD task, we propose a new input feature representation called the Binaural Time-Frequency Feature (BTFF), which encodes interaural time difference (ITD), interaural level difference (ILD), and high-frequency spectral cues (SC) from binaural signals. BTFF is composed of eight channels, including left and right mel-spectrograms, velocity-maps, SC-maps, and ITD-/ILD-maps, designed to cover different spatial cues across frequency bands and spatial axes. A CRNN-based model, BiSELDnet, is then developed to learn both spectro-temporal patterns and HRTF-based localization cues from BTFF. Experiments on the Binaural Set show that each BTFF sub-feature enhances task performance: V-map improves detection, ITD-/ILD-maps enable accurate horizontal localization, and SC-map captures vertical spatial cues. The final system achieves a SELD error of 0.110 with 87.1% F-score and 4.4{\deg} localization error, demonstrating the effectiveness of the proposed framework in mimicking human-like auditory perception.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 1 Pith paper
-
Auditory Intelligence: Understanding the World Through Sound
A position paper proposing four cognitively inspired audio tasks (ASPIRE, SODA, AUX, AUGMENT) to push machine hearing beyond recognition toward explanation, reasoning, and interaction.
Reference graph
Works this paper leans on
-
[64]
Two vs. four-channel sound event localization and detection,
J. Wilkins, M. Fuentes, L. Bondi, S. Ghaffarzadegan, A. Abavi sani, and J. P. Bello, “Two vs. four-channel sound event localization and detection,” in Proc. Detection Classification Acoust. Scenes Events Workshop (DCASE Workshop), Tampere, Finland, Sep. 2023, pp. 1–5
work page 2023
-
[1]
A comprehensive survey on humanoid robot development,
S. Saeedvand, M. Jafari, H. S. Aghdasi, and J. Baltes, “A comprehensive survey on humanoid robot development,” Knowl. Eng. Rev., vol. 34, pp. 1– 18, Dec. 2019, doi: 10.1017/S02698889190001582
-
[2]
Overview of the torqu e-controlled humanoid robot TORO,
J. Englsberger et al., “Overview of the torqu e-controlled humanoid robot TORO,” in Proc. 14th IEEE-RAS Int. Conf. Humanoid Robots, Madrid, Spain, Nov. 2014, pp. 916–923
work page 2014
-
[3]
Development of a biped walking robot compensating for three -axis moment by trunk motio n,
J.-I. Yamaguchi , A. Takanishi , and I. Kato , “ Development of a biped walking robot compensating for three -axis moment by trunk motio n,” in Proc. IEEE/RSJ Int. Conf. Intelligent Robots and Systems (IROS) , Yokohama, Japan, Jul. 1993, pp. 561–566
work page 1993
-
[4]
The development of Honda humanoid robot,
K. Hirai, M. Hirose, Y. Haikawa, and T. Takenaka, “The development of Honda humanoid robot,” in Proc. IEEE Int. Conf. Robotics and Automation (ICRA), Leuven, Belgium, May 1998, pp. 1321–1326
work page 1998
-
[5]
Towards the design of a biped jogging robot,
M. Gienger, K. Löffler, and F. Pfeiffer, “Towards the design of a biped jogging robot,” in Proc. IEEE Int. Conf. Robotics and Automation (ICRA), Seoul, Korea, May 2001, pp. 4140–4145
work page 2001
-
[6]
Design of prototype humanoid robotics platform for HRP,
K. Kaneko et al. , “ Design of prototype humanoid robotics platform for HRP,” in Proc. IEEE/RSJ Int. Conf. Intelligent Robots and Systems (IROS), Lausanne, Switzerland, Dec. 2002, pp. 2431–2436
work page 2002
-
[7]
The intelligent ASIMO: System overview and integration ,
Y. Sakagami, R. Watanabe, C. Aoyama, S. Matsunaga, N. Higaki, and K. Fujimura, “The intelligent ASIMO: System overview and integration ,” in Proc. IEEE/RSJ Int. Conf. Intelligent Robots and System s (IROS), Lausanne, Switzerland, Dec. 2002, pp. 2478–2483
work page 2002
Show all 80 references
-
[8]
Mechanical design of humanoid robot platform KHR-3 (KAIST humanoid robot - 3: HUBO),
I.-W. Park, J .-Y. Kim, J . Lee, and J .-H. Oh, “ Mechanical design of humanoid robot platform KHR-3 (KAIST humanoid robot - 3: HUBO),” in Proc. 5th IEEE-RAS Int. Conf. Humanoid Robots (ICHR), Tsukuba, Japan, Dec. 2005, pp. 321–326
2005
-
[9]
Design of android type humanoid robot Albert HUBO,
J.-H. Oh, D. Hanson, W.-S. Kim, I.-Y. Han, J.-Y. Kim, and I.-W. Park, “Design of android type humanoid robot Albert HUBO,” in Proc. IEEE/RSJ Int. Conf. Intelligent Robots and Systems (IROS), Beijing, China, Oct. 2006, pp. 1428–1433
2006
-
[10]
A literature review of sensor heads for humanoid robots,
J. A. Rojas-Quintero and M. C. Rodrí guez-Liñán, “A literature review of sensor heads for humanoid robots,” Robot. Auton. Syst., vol. 143, pp. 1– 21, Sep. 2021, doi: 10.1016/j.robot.2021.103834
2021
-
[11]
The Harvard binocular head ,
N. J. Ferrier , “ The Harvard binocular head ,” in Proc. SPIE 1708, Applications of Artificial Intelligenc e X: Machine Vision and Robotic s, Orlando, FL, USA, Mar. 1992, pp. 1–13. TABLE VI MEDIAN VALUES FOR BISELDNET PERFORMANCE WITH MS AND MS + SC-MAP AS INPUT FEATURES ON TE...
1992
-
[12]
The robot musician ‘wabot-2’ (waseda robot-2),
I. Kato et al., “The robot musician ‘wabot-2’ (waseda robot-2),” Robotics, vol. 3, no. 2, pp. 143–155, Jun. 1987, doi: 10.1016/0167-8493(87)90002- 7
1987 doi
-
[13]
Odor and airflow: Complementary senses for a humanoid robot ,
R. A. Russell and A. H. Purnamadjaja , “ Odor and airflow: Complementary senses for a humanoid robot ,” in Proc. IEEE Int. Conf. Robotics and Automation (ICRA), Washington, DC, USA, May 2002, pp. 1842–1847
2002
-
[14]
Electronic nose technology and application s,
M. Bonnefille , “Electronic nose technology and application s,” in Practical Analysis of Flavor and Fragrance Materials , K. Goodner and R. Rouseff, Eds. Chichester, UK: Wiley, 2011, pp. 111–154
2011
-
[15]
Development of auditory-evoked reflexes: Visuo -acoustic cues integration in a bino cular head ,
L. Natale, G. Metta, and G. Sandini, “Development of auditory-evoked reflexes: Visuo -acoustic cues integration in a bino cular head ,” Robot. Auton. Syst., vol. 39, no. 2, pp. 87–106, May 2002, doi: 10.1016/S0921- 8890(02)00174-4
2002 doi
-
[16]
Binaural sound event localization and detection neural network based on HRTF localization cues for humanoid robots,
G.-T. Lee, “Binaural sound event localization and detection neural network based on HRTF localization cues for humanoid robots,” Ph.D. Dissertation, KAIST, 2024
2024
-
[17]
A. S. Bregman, Auditory Scene Analysis. The Perceptual Organization of Sound. Cambridge, MA, USA: MIT Press, 1990, pp. 3–9
1990
-
[18]
Fundamentals of computational auditory scene analysis,
D. Wang and G. J. Brown , “Fundamentals of computational auditory scene analysis,” in Computational Auditory Scene Analysis: Principl es, Algorithms, and Applications, D. Wang and G. J. Brown, Eds. Piscataway, NJ, USA: Wiley-IEEE Press, 2006, pp. 11–14
2006
-
[19]
Computational models of auditory scene analysis: A review,
B. T. Szabó, S. L. Denham, and I. Winkler , “Computational models of auditory scene analysis: A review,” Front. Neurosci., vol. 10, pp. 1–16, Nov. 2016, doi: 10.3389/fnins.2016.00524
2016
-
[20]
Auditory streaming as an online classification process with evidence accumulation,
D. Barniv and I. Nelken, “Auditory streaming as an online classification process with evidence accumulation,” PLoS One, vol. 10, no. 12, pp. 1– 20, Dec. 2015, doi: 10.1371/journal.pone.0144788
2015 doi
-
[21]
Combined estimation of spectral envelopes and sound source direction of concurrent voices by multidimensional statistical filtering,
J. Nix and V. Hohmann, “Combined estimation of spectral envelopes and sound source direction of concurrent voices by multidimensional statistical filtering,” IEEE Trans. Audio Speech Lang. Process. , vol. 15, no. 3, pp. 995–1008, Mar. 2007, doi: 10.1109/TASL.2006.889788
2007
-
[22]
An oscillatory correlation model of auditory streaming,
D. Wang and P. Chang , “An oscillatory correlation model of auditory streaming,” Cogn. Neurodynamics, vol. 2, no. 1, pp. 7–19, Mar. 2008, doi: 10.1007/s11571-007-9035-8
2008 doi
-
[23]
Monophonic sound source separation with an unsupervised network of spiking neurons,
R. Pichevar and J. Rouat, “Monophonic sound source separation with an unsupervised network of spiking neurons,” Neurocomputing, vol. 71, no. 1-3, pp. 109–120, Dec. 2007, doi: 10.1016/j.neucom.2007.08.001
2007 doi
-
[24]
Modelling the emergence and dynamics of perceptual organisation in auditory streaming,
R. W. Mill, T. M. Böhm, A. Bendixen, I. Winkler, and S. L. Denham , “Modelling the emergence and dynamics of perceptual organisation in auditory streaming,” PLoS Comput. Biol. , vol. 9, no. 3, pp. 1–21, Mar. 2013, doi: 10.1371/journal.pcbi.1002925
2013 doi
-
[25]
Neuromechanistic model of auditory bistability,
J. Rankin, E. Sussman, and J. Rinzel , “Neuromechanistic model of auditory bistability,” PLoS Comput. Biol., vol. 11, no. 11, pp. 1–34, Nov. 2015, doi: 10.1371/journal.pcbi.1004555
2015 doi
-
[26]
Segregating complex sound sources through temporal coherence ,
L. Krishnan, M. Elhilali, and S. Shamma , “Segregating complex sound sources through temporal coherence ,” PLoS Comput. Biol. , vol. 10, no. 12, pp. 1–10, Dec. 2014, doi: 10.1371/journal.pcbi.1003985
2014 doi
-
[27]
A cocktail party with a cortical twist: How cortical mechanisms contribute to sound segregation,
M. Elhilali and S. A. Shamm a, “A cocktail party with a cortical twist: How cortical mechanisms contribute to sound segregation,” J. Acoust. Soc. Am., vol. 124, no. 6, pp. 3751–3771, Dec. 2008, doi: 10.1121/1.3001672
2008 doi
-
[28]
Sound event localization and detection of overlapping sources using convolutional recurrent neural networks,
S. Adavanne, A. Politis, J. Nikunen, and T. Virtanen , “Sound event localization and detection of overlapping sources using convolutional recurrent neural networks,” IEEE J. Sel. Top. Signal Process., vol. 13, no. 1, pp. 34–48, Mar. 2019, doi: 10.1109/JSTSP.2018.2885636
2019
-
[29]
Sound source localization based on deep neural networks with directional activate function exploiting phase information,
R. Takeda and K. Komatani , “Sound source localization based on deep neural networks with directional activate function exploiting phase information,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP), Shanghai, China, Mar. 2016, pp. 405–409
2016
-
[30]
Sound source localiz ation using deep learning models,
N. Yalta, K. Nakadai, and T. Ogata , “Sound source localiz ation using deep learning models,” J. Robot. Mechatron., vol. 29, no. 1, pp. 37–48, Feb. 2017, doi: 10.20965/jrm.2017.p0037
2017 doi
-
[31]
Deep neural networks for multiple speaker detection and locali zation,
W. He, P. Motlicek, and J. -M. Odobez , “ Deep neural networks for multiple speaker detection and locali zation,” in Proc. IEEE Int . Conf. Robotics and Automation (ICRA), Brisbane, Australia, May 2018, pp. 74– 79
2018
-
[32]
Audio surveillance: A systematic review,
M. Crocco, M. Cristani, A. Trucco, and V. Murino, “Audio surveillance: A systematic review,” ACM Comput. Surv., vol. 48, no. 4, pp. 1–46, Feb. 2016, doi: 10.1145/2871183
2016 doi
-
[33]
Scream and gunshot detection and localization for audio -surveillance systems,
G. Valenzise, L. Gerosa, M. Tagliasacchi, F. Antonacci, and A. Sarti , “Scream and gunshot detection and localization for audio -surveillance systems,” in Proc. IEEE Conf . Advanced Video and Signal Based Surveillance (AVSS), London, UK, Sep. 2007, pp. 21–26
2007
-
[34]
Environmental sound recognition wi th time –frequency audio feature s,
S. Chu, S. Narayanan, and C. -C. J. Kuo , “Environmental sound recognition wi th time –frequency audio feature s,” IEEE Trans. Audio Speech Lang. Process ., vol. 17, no. 6, pp. 1142–1158, Aug. 2009, doi: 10.1109/TASL.2009.2017438
2009
-
[35]
Emergency vehicles audio detection and localization in autonomous driving ,
H. Sun, X. Liu, K. Xu, J. Miao, and Q. Luo , “Emergency vehicles audio detection and localization in autonomous driving ,” 2021 , arXiv:2109.14797
2021 arXiv
-
[36]
Overview and evaluation of sound event localization and detection in DCASE 2019,
A. Politis, A. Mesaros, S. Adavanne, T. Heittola, and T. Virtanen , “Overview and evaluation of sound event localization and detection in DCASE 2019,” IEEE-ACM Trans. Audio Speech Lang., vol. 29, pp. 684– 698, Dec. 2020, doi: 10.1109/TASLP.2020.3047233
2019
-
[37]
Sound event detection and localization based on CNN and LSTM,
Z. Lu, “Sound event detection and localization based on CNN and LSTM,” Detection Classification Acoust. Scenes Events (DCASE) Challenge , Tech. Rep., pp. 1–3, 2019
2019
-
[38]
Online direction of arrival estimation based on deep learning ,
Q. Li, X. Zhang, and H. Li, “Online direction of arrival estimation based on deep learning ,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP), Calgary, AB, Canada, Apr. 2018, pp. 2616–2620
2018
-
[39]
Exploiting temporal context in CNN based multisource DOA estimation,
A. Bohlender, A. Spriet, W. Tirry, and N. Madhu, “Exploiting temporal context in CNN based multisource DOA estimation,” IEEE/ACM Trans. Audio Speech Lang. Process., vol. 29, no. 1, pp. 1594–1608, Mar. 2021, doi: 10.1109/TASLP.2021.3067113
2021
-
[40]
Sound event localization and detection based on CRNN using rectangular filters and channel rotation data augmentation ,
F. Ronchini, D. Arteaga, and A. P érez-López, “Sound event localization and detection based on CRNN using rectangular filters and channel rotation data augmentation ,” in Proc. Detection Classification Acoust. Scenes Events Workshop (DCASE Workshop), Tokyo, Japan, Nov. 2020, p...
2020
-
[41]
GCC-PHAT cross -correlation audio features for simultaneous sound event localization and detection (SELD) in multiple rooms ,
H. A. C. Maruri, P. L. Meyer, J. H uang, JAdH. Ontiveros, and H. L u, “GCC-PHAT cross -correlation audio features for simultaneous sound event localization and detection (SELD) in multiple rooms ,” Detection Classification Acoust. Scenes Events (DCASE) Challenge, Tech. Rep., p...
2019
-
[42]
Polyphonic sound event detection and localization using a two-stage strategy ,
Y. Cao et al., “Polyphonic sound event detection and localization using a two-stage strategy ,” in Proc. Detecti on Classification Acoust. Scenes Events Workshop (DCASE Workshop), New York, NY, USA, Oct. 2019, pp. 30–34
2019
-
[43]
Two-stage sound event localization and detection using intensity vector and generalized cross -correlation,
Y. Cao et al., “Two-stage sound event localization and detection using intensity vector and generalized cross -correlation,” Detection Classification Acoust. Scenes Events (DCASE) Challenge, Tech. Rep., pp. 1–4, 2019
2019
-
[44]
Sound event detection an d localization using CRNN model s,
A. Sampathkumar and D. Kowerko , “ Sound event detection an d localization using CRNN model s,” Detection Classification Acoust. Scenes Events (DCASE) Challenge, Tech. Rep., pp. 1–3, 2020
2020
-
[45]
An improved event -independent network for polyphonic sound event localization and detection ,
Y. Cao, T. Iqbal, Q. Kong, F. An, W. Wang, and M. D. Plumbley , “An improved event -independent network for polyphonic sound event localization and detection ,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP), Toronto, ON, Canada, Jun. 2021, pp. 885–889
2021
-
[46]
A general network architecture for sound event localization and detection using transfer learning and recurrent neural network,
T. N. T. Nguyen et al., “A general network architecture for sound event localization and detection using transfer learning and recurrent neural network,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP), Toronto, ON, Canada, Jun. 2021, pp. 935–939
2021
-
[47]
ACCDOA: Activity -coupled Cartesian direction of arrival representation for sound event localization and detection,
K. Shimada, Y. Koyama, N. Takahashi, S. Takahashi, and Y. Mitsufuji , “ACCDOA: Activity -coupled Cartesian direction of arrival representation for sound event localization and detection,” in Proc. IEEE Int. Conf. Acoust., Speech, Sign al Process. (ICASSP ), Toronto, ON , Canad...
2021
-
[48]
Multi-ACCDOA: Localizing and detecting overlapping sounds from the same class with auxiliary duplicating per mutation invariant training ,
K. Shimada, Y. Koyama, S. Takahashi, N. Takahashi, E. Tsunoo, and Y. Mitsufuji, “Multi-ACCDOA: Localizing and detecting overlapping sounds from the same class with auxiliary duplicating per mutation invariant training ,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process...
2022
-
[49]
SALSA: Spatial cue-augmented log-spectrogram features for polyphonic sound event localization and detection ,
T. N. T. Nguyen, K. N. Watcharasupat, N. K. Ng uyen, D. L. Jones, and W.-S. Gan, “SALSA: Spatial cue-augmented log-spectrogram features for polyphonic sound event localization and detection ,” IEEE-ACM Trans. Audio Speech Lang ., vol. 30 , pp. 1749–1762, May 2022, doi: 10.1109...
2022
-
[50]
SALSA-Lite: A fast and effective feature for polyphonic sound event localization and detection with microphone arrays,
T. N. T. Nguyen, D. L. Jones, K. N. Watcharasupat, H. Phan, and W. -S. Gan, “SALSA-Lite: A fast and effective feature for polyphonic sound event localization and detection with microphone arrays,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP ), Singapore, ...
2022
-
[51]
A report on sound event detection with different binaural features ,
S. Adavanne and T. Virtanen , “A report on sound event detection with different binaural features ,” in Proc. Detection Classification Acoust. Scenes Events Workshop (DCASE Workshop ), Munich, Germany, Nov. 2017, pp. 1–5. 12 > REPLACE THIS LINE WITH YOUR MANUSCRIPT ID NUMBER (...
2017
-
[52]
Binaural signal representations for joint sound event detection and acoustic scene classification ,
D. A. Krause and A. Mesaros , “Binaural signal representations for joint sound event detection and acoustic scene classification ,” in Proc. 30th Euro. Signal Process. Conf. (EUSIPCO), Belgrade, Serbia, Aug. 2022, pp. 399–403
2022
-
[53]
A learning-based approach to robust binaural sound localization ,
K. Youssef, S. Argentieri, and J.-L. Zarader, “A learning-based approach to robust binaural sound localization ,” in Proc. IEEE/RSJ Int. Conf. Intelligent Robots and Systems (IROS), Tokyo, Japan , Nov. 2013, pp. 2927–2932
2013
-
[54]
On sound source localization of speech signals using deep neural networks,
R. Roden, N. Moritz, S. Gerlach, S. Weinzierl, and S. Goetze, “On sound source localization of speech signals using deep neural networks,” in Proc. Deutsche Jahrestagung Akustik (DAGA ), Nuremberg, Germany , Mar. 2015, pp. 1510–1513
2015
-
[55]
Autonomous sensorimotor learning for sound source localization by a humanoid robot,
Q. V. Nguyen, L. Girin , G. Bailly, F. Elisei, and D. C. Nguyen , “Autonomous sensorimotor learning for sound source localization by a humanoid robot,” in Proc. IEEE/RSJ Int. Conf. Intelligent Robots and Systems (IROS), Madrid, Spain, Oct. 2018, pp. 1–4
2018
-
[56]
Multitask learning of time-frequency CNN for sound source localization ,
C. Pang, H. Liu, and X. L i, “Multitask learning of time-frequency CNN for sound source localization ,” IEEE Access, vol. 7, pp. 40725–40737, Mar. 2019, doi: 10.1109/ACCESS.2019.2905617
2019
-
[57]
Full-sphere binaural sound source localization using multi -task neural network ,
Y. Yang, J. Xi, W. Zhang, and L. Zhang , “ Full-sphere binaural sound source localization using multi -task neural network ,” in Proc. Asia - Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), Auckland, New Zealand, Dec. 2020, pp. 432–436
2020
-
[58]
Deep neural network based audio source separation,
A. Zermini, Y. Yu, Y. Xu, M. D. Plumbley, and W. Wang, “Deep neural network based audio source separation,” in Proc. Int. Conf. Math. Signal Proc. (IMA), Birmingham, UK, May 2016, pp. 1–4
2016
-
[59]
Exploiting deep neural networks and head movements for robust binaural localization of multiple sources in reverberant environments ,
N. Ma, T. Ma y, and G. J. Brown , “Exploiting deep neural networks and head movements for robust binaural localization of multiple sources in reverberant environments ,” IEEE/ACM Trans. Audio Speech Lang. Process., vol. 25 , no. 12 , pp. 2444–2453, Dec. 2017, doi: 10.1109/TASL...
2017
-
[60]
Learning deep direct-path relative transfer function for binaural sound source localization,
B. Yang, H. Liu, and X. Li , “Learning deep direct-path relative transfer function for binaural sound source localization,” IEEE/ACM Trans. Audio Speech Lang. Process ., vol. 29 , pp. 3491–3503, Oct. 2021, doi: 10.1109/TASLP.2021.3120641
2021
-
[61]
Binaural source localization using deep learning and head rotation information,
G. García-Barrios, D. A. Krause, A. Politis, A. Mesaros, J. M. Gutiérrez- Arriola, and R. Fraile, “Binaural source localization using deep learning and head rotation information,” in Proc. 30th Euro. Signal Process. Conf. (EUSIPCO), Belgrade, Serbia, Aug. 2022, pp. 36–40
2022
-
[62]
Binaural source localization in median plane using learning based method for robot audition,
P. Dwivedi, G. Routray, and R. M. Hegde, “Binaural source localization in median plane using learning based method for robot audition,” in Proc. 24th Int. Congr. Acoust. (ICA), Gyeongju, Korea, Oct. 2022, pp. 1–8
2022
-
[63]
Goal-driven, neurobiological- inspired convolutional neural network models of human spatial hearing,
K. van der Heijden and S. M ehrkanoon, “Goal-driven, neurobiological- inspired convolutional neural network models of human spatial hearing,” Neurocomputing, vol. 470 , pp. 432–442, Jan. 2022, doi: 10.1016/j.neucom.2021.05.104
2022 doi
-
[65]
HRTF measurement for accurate sound localization cues,
G.-T. Lee, S.-M. Choi, B.-Y. Ko, and Y. -H. Park, “HRTF measurement for accurate sound localization cues,” 2022, arXiv:2203.03166v2
2022 arXiv
-
[66]
Iida, Head-Related Transfer Function and Acoustic Virtual Reality
K. Iida, Head-Related Transfer Function and Acoustic Virtual Reality . Narashino, Japan: Springer, 2019, pp. 15–55
2019
-
[67]
Xie, Head-Related Transfer Function and Virtual Auditor y Display, 2nd ed., Plantation, FL, USA: J
B. Xie, Head-Related Transfer Function and Virtual Auditor y Display, 2nd ed., Plantation, FL, USA: J. Ross Publishing, 2013, pp. 81–85
2013
-
[68]
Sound pressure generated in an external- ear replica and real human ears by a nearby point source,
E. A. G. Shaw and R. Teranishi, “Sound pressure generated in an external- ear replica and real human ears by a nearby point source,” J. Acoust. Soc. Am., vol. 44, no. 1, pp. 240–249, Jul. 1968, doi: 10.1121/1.1911059
1968 doi
-
[69]
Mechanism for generating pe aks and notches of head -related transf er functions in the median plan e,
H. Takemoto, P. Mokhtari, H. Kato, R. Nishimura, and K. Iida , “Mechanism for generating pe aks and notches of head -related transf er functions in the median plan e,” J. Acoust. Soc. Am ., vol. 132, no. 6, pp. 3832–3841, Dec. 2012, doi: 10.1121/1.4765083
2012 doi
-
[70]
Deep learning based cough detection camera using enhanced features ,
G.-T. Lee, H. Nam, S. -H. Kim, S. -M. Choi, Y. Kim, and Y. -H. Park , “Deep learning based cough detection camera using enhanced features ,” Expert Syst. Appl ., vol. 206, pp. 1–20, Nov. 2022, doi: 10.1016/j.eswa.2022.117811
2022
-
[71]
Estimation of interaural time difference based on cochlear filter bank and ZCPA auditory model ,
G.-T. Lee and Y.-H. Park, “Estimation of interaural time difference based on cochlear filter bank and ZCPA auditory model ,” Trans. Korean Soc. Noise Vib. Eng ., vol. 29, no. 6, pp. 722–734, Dec. 2019, doi: 10.5050/KSNVE.2019.29.6.722
2019 doi
-
[72]
Method for estimating interaural time differences based on cochlea filter bank and EUZ auditory model in noisy environment,
G.-T. Lee and Y.-H. Park , “ Method for estimating interaural time differences based on cochlea filter bank and EUZ auditory model in noisy environment,” in Proc. Int. Conf. Noise Control Eng. (Inter-noise), Seoul, Korea, Aug. 2020, pp. 1–12
2020
-
[73]
Significance of the modified group delay feature in speech recognition,
R. M. Hegde, H. A. Murthy, and V. R. R. Gad de, “Significance of the modified group delay feature in speech recognition,” IEEE Trans. Audio Speech Lang. Process ., vol. 1 5, no. 1, pp. 190–202, Jan. 2007, doi: 10.1109/TASL.2006.876858
2007
-
[74]
Group delay spectrogram of speech signals without phase wrapping,
B. Yegnanarayana, “Group delay spectrogram of speech signals without phase wrapping,” J. Acoust. Soc. Am., vol. 151, no. 3, pp. 2181–2191, Mar. 2022, doi: 10.1121/10.0009922
2022 doi
-
[75]
R. F. Lyon , Human and Machine Hearing: Extracting Meaning from Sound. Cambridge, UK: Cambridge Univ. Press, 2017, pp. 66–67
2017
-
[76]
The NIGENS general sound events database,
I. Trowitzsch, J. Taghia , Y. Kashef, and K. Obermayer , “The NIGENS general sound events database,” 2020, arXiv:1902.08314
2020 arXiv
-
[77]
Detection and classification of acoustic scenes and events: Outcome of the DCASE 2016 challenge ,
A. Mesaros et al., “Detection and classification of acoustic scenes and events: Outcome of the DCASE 2016 challenge ,” IEEE-ACM Trans. Audio Speech Lang ., vol. 26, no. 2 , pp. 379–393, Feb. 2018, doi: 10.1109/TASLP.2017.2778423
2016
-
[78]
A multi -device d ataset for urban acoustic scene classificatio n,
A. Mesaros, T. Heittola, and T. Virtanen , “ A multi -device d ataset for urban acoustic scene classificatio n,” in Proc. Detection Classification Acoust. Scenes Events Workshop (DCASE Workshop), Surrey, UK, Nov. 2018, pp. 1–5
2018
-
[79]
Binaural sound event localization and detection for humanoid robo t,
G.-T. Lee and Y. -H. Park , “ Binaural sound event localization and detection for humanoid robo t,” in Proc. 24th Int. Congr. Acoust. (ICA ), Gyeongju, Korea, Oct. 2022, pp. 1–12
2022
-
[80]
Joint measurement of localization and detection of sound events ,
A. Mesaros, S. Adavanne, A. Politis, T. Heittola, and T. Virtanen, “Joint measurement of localization and detection of sound events ,” in Proc. IEEE Workshop Appl. Signal Process. Audio Acoust. (WASPAA ), New Paltz, NY, USA, Oct. 2019, pp. 333–337
2019
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.