Pith. sign in

REVIEW 4 major objections 7 minor 1 cited by

Binaural Sound Event Localization and Detection based on HRTF Cues for Humanoid Robots

T0 review · 4 major / 7 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read HRTF cue maps let two microphones localize sounds to 4.4 degrees

desk verdict The feature design and ablation logic are solid, but the headline 4.4° localization error is measured on the same 48-direction HRTF grid used for training, so it may reflect grid-node recognition rather than continuous spatial hearing. read the letter →

arxiv 2507.20530 v1 pith:7PAFJBMM submitted 2025-07-28 eess.AS cs.SD

classification eess.AScs.SD
keywords binauralsoundeventlocalizationanddetectionhead-relatedtransferfunctioninterauraltimedifferencelevelspectralcuesdirectionofarrivalestimationhumanoidrobots
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that binaural sound event localization and detection (BiSELD), detecting what is sounding and from where using only two ear-like microphones, can be tackled with hand-designed spatial cues rather than four-channel microphone arrays. To test this, the authors build a synthetic Binaural Set by convolving isolated sound events with measured head-related impulse responses at 48 directions, and propose an eight-channel time-frequency input, the Binaural Time-Frequency Feature (BTFF), that encodes interaural time difference, interaural level difference, and high-frequency spectral cues along with mel-spectrograms and velocity maps. A compact CRNN, BiSELDnet, trained on BTFF reports a SELD error of 0.110, an F-score of 87.1%, and a 4.4 degree localization error on the test set. The paper argues this shows a two-channel, HRTF-informed system can approximate human-like spatial hearing for humanoid robots with much lighter hardware than conventional arrays.

What carries the argument

The central object is the Binaural Time-Frequency Feature (BTFF), an eight-channel input map built from a binaural pair: left and right mel-spectrograms; left and right velocity maps (time-differences of the magnitude spectrogram); an ITD-map, computed as $\frac{1}{\omega}\mathrm{Im}[\ln(P_R/P_L)]$ for bins below 1.5 kHz and projected to the mel scale; an ILD-map, $10\log_{10}|P_R/P_L|^2$ for bins above 5 kHz, also mel-projected; and left and right SC-maps, mel spectra restricted to bands above 5 kHz. These channels are the mechanism because they parse the HRTF's spatial information into complementary, frequency-segregated cues before the network sees the data: ITD carries low-frequency azimuth, ILD carries high-frequency azimuth plus front-back asymmetry, and SC carries elevation. The network itself, BiSELDnet, is a CRNN with depthwise separable convolutions, bidirectional GRUs, and a fully connected head that outputs one 3D direction vector per event class per frame, using the activity-coupled Cartesian DOA training target.

What would settle it

Render binaural test clips from HRIRs at directions not seen in training (for example, 15 degree rather than 30 degree azimuth steps, or elevations at 45 degrees), or record real binaural sounds with a manikin head, and rerun BiSELDnet; if localization error jumps well above 4.4 degrees or detection F-score drops, the central generalization claim fails.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that each class of HRTF-derived cue maps to a specific sub-problem of binaural SELD, and encoding them explicitly makes the whole task learnable from two channels. The velocity map improves detection by marking onsets and transients; the ITD-map, derived from the imaginary part of the log spectral ratio below 1.5 kHz without phase unwrapping, and the ILD-map, computed as log-power difference above 5 kHz, jointly resolve azimuth, with ILD's front-back asymmetry disambiguating ITD's symmetry; and the SC-map, mel bands above 5 kHz, supplies elevation-dependent pinna notch cues. With these eight channels, BiSELDnet detects all 12 event classes and outputs a 3D direction vector per class per frame, reaching a 4.4 degree localization error and 92.1% localization recall on the Binaural Set test split. The authors take this as evidence that HRTF-based spatial cues integrated as input features are a viable path to binaural 3D SELD in humanoid robots.

Load-bearing premise

The load-bearing premise is that performance measured on synthetic mixtures made from a fixed 48-direction HRTF grid transfers to real binaural hearing; if the grid is not representative of continuous, real-world directions, the reported accuracy is a property of the benchmark rather than of the robot's ears.

Editorial extensions

If this is right

  • A two-channel binaural front end is enough for joint 3D detection and localization in the synthetic setting, so robots can avoid the size, calibration, and data costs of four-channel arrays.
  • Explicitly separating ITD, ILD, and spectral cues makes each cue's contribution inspectable: V-map for detection, ITD/ILD for azimuth, and SC for elevation.
  • The reported 4.4 degree localization error and 92.1% recall imply the model can pick out which of 12 classes is active and point to it with near-grid resolution on the Binaural Set.
  • The Binaural Set itself, with clean and noisy conditions and separate horizontal and vertical test subsets, gives the community a controlled benchmark for comparing binaural SELD methods.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because training and test directions come from the same 48-point HRTF grid, the 4.4 degree error may partly measure interpolation inside that grid; the paper does not establish accuracy on directions between grid points or on real recorded binaural scenes.
  • A natural next test is to evaluate the same BTFF on continuous azimuths and on real binaural recordings from a manikin head; if accuracy degrades, the cues remain valid but the current benchmark overstates deployable performance.
  • The feature design suggests a hearing-aid or telepresence variant could work with generic HRTFs per user, provided individual pinna cues are preserved; the SC-map channel is the load-bearing part for elevation and the most likely to need per-listener adaptation.
  • One could ablate the hand-crafted maps against learned spatial features, for example a network given only the two mel-spectrograms, to quantify how much of the 4.4 degree result comes from the engineered cues versus the CRNN itself.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. The paper proposes BiSELD, a binaural (two-channel) variant of sound event localization and detection, together with a synthetic benchmark dataset (the Binaural Set) built by convolving isolated event recordings from NIGENS and DCASE2016 Task 2 with measured HRIRs from the authors' KAIST HRTF database at 12 azimuths and 4 elevations (48 directions). The proposed input representation BTFF concatenates eight channels—left/right mel-spectrograms, left/right velocity maps, an ITD map (phase-derived delay below 1.5 kHz), an ILD map (above 5 kHz), and left/right spectral-cue maps (mel bands above 5 kHz)—and feeds a 763K-parameter CRNN, BiSELDnet, with ACCDOA-style per-class Cartesian DOA outputs. Ablation experiments show consistent gains from each feature group: V-map improves detection, ITD/ILD maps reduce horizontal-plane localization error (LE from 17.3° to 4.2°), and the SC map reduces median-plane localization error (25.2° to 12.2°). The full system is reported to achieve a SELD error of 0.110, F-score of 87.1%, LE_CD of 4.4°, and LR_CD of 92.1% on the test set.

Significance. If the results hold, the work provides a compact, lightweight two-channel SELD pipeline whose feature design is directly grounded in psychoacoustic cues (ITD below 1.5 kHz, ILD above 5 kHz, pinna-related spectral notches), which is a plausible and falsifiable design hypothesis of genuine interest to the humanoid robotics and binaural-audio communities. The internal ablation logic is clean and consistently executed: each feature is added to a fixed mel-spectrogram baseline and evaluated on the sub-task it is designed for, with ten independent training runs and median reporting in Tables IV–VI. The authors are also transparent about the synthetic nature of the benchmark and about their use of the measured KAIST HRTF database. The main weaknesses are the scope of the evaluation—training and testing on the same 48-direction grid, with no off-grid, reverberant, or real binaural signals—the best-of-ten reporting in Table VII, the absence of any external SELD baseline, and several dataset-description inconsistencies. These issues are local and fixable, but they currently limit the practical force of the headline localization claims.

major comments (4)
  1. [III-C and Table VII] The entire evaluation is confined to the same 48-direction KAIST HRTF grid used for training, with 30° azimuth and 30° elevation spacing; no test sample is rendered from an off-grid direction, a reverberant scene, a second head or ear geometry, or a real binaural recording. The headline LE_CD of 4.4° in Table VII is an order of magnitude finer than the grid spacing and is consistent with near-perfect selection of the correct grid node rather than with interpolation or continuous spatial inference. Because the abstract and title make claims about 'human-like auditory perception' for humanoid robots, the authors should add an off-grid evaluation (for example, directions at 15° offsets in azimuth and elevation from the training grid) and, ideally, a different HRTF set or real binaural data; until then, the localization claims should be explicitly re-scoped to the grid-matched synthetic setting.
  2. [V-B4, Table VII versus Tables IV–VI] Tables IV–VI report median values over ten training runs, but Table VII and the abstract report the 'best' performance of BiSELDnet with no measure of variance; the headline SELD error of 0.110, F-score of 87.1%, and LE_CD of 4.4° are therefore the upper envelope of the runs and are not statistically comparable with the median-based ablation tables. The authors should report the median and spread (standard deviation or the full range) for the final configuration in Table VII and ensure that the abstract uses the same statistic.
  3. [Section V and Table I] The evaluation contains no comparison with any external SELD or binaural-localization baseline, even though Table I identifies Wilkins et al. [64] as the state of the art for binaural SELD; without at least one matched baseline trained and evaluated under the same protocol on the Binaural Set, the absolute performance figures have no external calibration and the 'effectiveness' claim rests entirely on self-ablations. Please add such a comparison or explicitly limit the claims to the feature-ablation findings.
  4. [III-C, Tables II and III] The dataset description is internally inconsistent: the text states that 'for each sound class, 20 samples were prepared and split into training, validation, and test sets in a 14:3:3 ratio,' but Table III reports 672/144/144 mixtures whose construction requires far more than 20 event clips per class, and the 14:3:3 ratio does not correspond to 672/144/144 (the actual per-class split implied by Table III is 56/12/12). Additionally, NIGENS is described as having 14 classes but Table II lists 15 class names, the 12 classes actually used in BiSELD are never enumerated, and the exact direction sets underlying Test-H and Test-V are not specified; these details must be corrected and documented for the Binaural Set to serve as a reproducible benchmark.
minor comments (7)
  1. [Conclusion versus Section IV] The Conclusion states that BiSELDnet is built with depthwise separable convolutions, but Section IV and Fig. 4 describe only generic convolution, normalization, activation, and pooling modules; the architecture description and the Conclusion should be reconciled.
  2. [Section IV] The detection rule 'exceeds 0.5v' appears to refer to a magnitude threshold of 0.5 on the output DOA vector, but the notation is undefined; it should be stated explicitly.
  3. [Section III-C] Test-H and Test-V are defined only by their sample counts (36 and 12); the exact azimuth and elevation grid points in each subset should be stated so that Tables V and VI are interpretable from the text alone.
  4. [Section III-A3] The spectral-notch analysis quotes elevations from −40° to 90°, whereas the Binaural Set covers only −30° to +60°; the two ranges should be harmonized or the discrepancy explained.
  5. [Section III-C] The paper should state explicitly that polyphony in the Binaural Set is restricted to across-class overlap, since each mixture contains exactly one instance of each class, and should note the corresponding limitation relative to the same-class overlap scenarios addressed by multi-ACCDOA.
  6. [Abstract and Table VI] The abstract's '4.4° localization error' should be qualified: Table VI shows that the median-plane localization error remains 12.2° even with the SC map, so the 4.4° LE_CD is dominated by horizontal-plane performance rather than uniform 3D accuracy.
  7. [Section IV] Section IV reports only the total parameter count (763,020); layer-wise specifications such as convolution filter counts, GRU units, and pooling sizes are needed to reproduce BiSELDnet from the text.

Circularity Check

0 steps flagged · score 2.0 of 10

No load-bearing circularity: each BTFF component is validated by in-paper ablation on a held-out split, and the self-citations supply data and design precedents rather than the predicted result.

full rationale

The claimed derivation chain is: psychoacoustic analysis of HRIRs (Sec. III-A) → construction of the eight-channel BTFF (Eqs. 3–9) → supervised training of BiSELDnet on the synthetic Binaural Set → measured SELD error of 0.110 and LE_CD of 4.4° on a held-out split (Table VII). Every link is empirical rather than definitional. The ITD-, ILD-, and SC-maps are computed from raw binaural audio (Eqs. 5–9), while the azimuth/elevation labels come from the synthesis procedure (Sec. III-C, 14:3:3 split, Table III), so the input features are not defined in terms of the labels, and no fitted parameter is renamed as a prediction. The sub-feature claims are supported by in-paper ablations on held-out data (V-map: SED error 0.321→0.286; +ITD+ILD: LE 17.3°→4.2° on Test-H; +SC: LE 25.2°→12.2° on Test-V), so they are measured, not assumed. The self-citations ([65] KAIST HRTF database, [70] velocity features, [71][72] ITD estimation, [79] preliminary ICA version, [16] Ph.D. thesis) provide a measured data asset and design precedents, but none is invoked as the proof of the target result; the load-bearing evidence for each claim is in-paper and reproducible from the described pipeline. Two caveats are real but fall outside circularity. First, Section III-C builds the entire dataset from 12 azimuths × 4 elevations of the same KAIST HRTF grid, and all test splits draw from this same grid, so the 4.4° LE_CD (far below the 30° spacing) is demonstrated only on that discrete grid, with no off-grid, reverberant, or real-binaural evaluation; this limits external validity and belongs under correctness risk rather than circularity. Second, Tables IV–VI report median values over ten runs while Table VII reports the 'BEST PERFORMANCES', so the headline number is the upper envelope of the runs; this is a reporting caveat, not a reduction of prediction to input. Nothing in the construction forces the network to achieve these numbers, so the derivation is not circular.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The central empirical claims rest on the synthetic Binaural Set and on hand-selected feature cutoffs. No new physical entities are introduced. The free parameters are design choices that shape the reported numbers.

free parameters (4)
  • ITD low-frequency cutoff = 1.5 kHz
    Used to restrict ITD-map computation to frequencies below 1.5 kHz (Eq. 6); chosen based on psychoacoustic dominance of ITD, but the exact value is a design choice.
  • ILD/SC high-frequency cutoff = 5 kHz
    Used to isolate interaural level and spectral cues above 5 kHz (Eqs. 7 and 9); motivated by head shadow and pinna effects, but the threshold is hand-selected.
  • Event detection threshold = 0.5 (vector magnitude)
    An event is detected when the output DOA vector magnitude exceeds 0.5v (Fig. 4c); this threshold directly affects the reported F-score and error rate and is set without sensitivity analysis.
  • Spatial direction grid = 12 azimuth x 4 elevation = 48 directions
    The Binaural Set generates sounds only at these 48 directions; the grid size is a trade-off between coverage and dataset size and limits evaluation to on-grid directions.
assumptions (5)
  • domain assumption Measured KAIST HRTF database accurately captures binaural localization cues for a humanoid robot with human-like ears.
    The entire Binaural Set is generated by convolving sound events with these HRIRs (Section III-C1).
  • domain assumption Synthetic binaural mixtures created by convolution with anechoic HRIRs and additive background noise are representative of real-world binaural audio.
    The method is not evaluated on any real binaural recordings.
  • domain assumption The NIGENS and DCASE2016 Task 2 sound event databases provide sufficient coverage of the target 12 sound classes.
    The class set is not explicitly enumerated, and the listed NIGENS count is inconsistent (14 stated, 15 listed in Table II).
  • ad hoc to paper The ITD, ILD, and spectral notch cues are the dominant cues for binaural localization and can be captured by the proposed maps.
    The feature maps are engineered to encode these cues, but the model's reliance on them is assumed, not measured.
  • domain assumption Evaluation on the synthetic test set is a valid measure of sound event localization and detection performance.
    The test set shares the same HRTF database and direction grid as the training set.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Binaural Sound Event Localization and Detection based on HRTF Cues for Humanoid Robots." pith.science (2026). https://pith.science/paper/7PAFJBMM

@misc{pith2026250720530,
  author       = {Pith},
  title        = {Pith review of: Binaural Sound Event Localization and Detection based on HRTF Cues for Humanoid Robots},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7PAFJBMM}},
  note         = {Machine review of arXiv:2507.20530}
}
read the original abstract

This paper introduces Binaural Sound Event Localization and Detection (BiSELD), a task that aims to jointly detect and localize multiple sound events using binaural audio, inspired by the spatial hearing mechanism of humans. To support this task, we present a synthetic benchmark dataset, called the Binaural Set, which simulates realistic auditory scenes using measured head-related transfer functions (HRTFs) and diverse sound events. To effectively address the BiSELD task, we propose a new input feature representation called the Binaural Time-Frequency Feature (BTFF), which encodes interaural time difference (ITD), interaural level difference (ILD), and high-frequency spectral cues (SC) from binaural signals. BTFF is composed of eight channels, including left and right mel-spectrograms, velocity-maps, SC-maps, and ITD-/ILD-maps, designed to cover different spatial cues across frequency bands and spatial axes. A CRNN-based model, BiSELDnet, is then developed to learn both spectro-temporal patterns and HRTF-based localization cues from BTFF. Experiments on the Binaural Set show that each BTFF sub-feature enhances task performance: V-map improves detection, ITD-/ILD-maps enable accurate horizontal localization, and SC-map captures vertical spatial cues. The final system achieves a SELD error of 0.110 with 87.1% F-score and 4.4{\deg} localization error, demonstrating the effectiveness of the proposed framework in mimicking human-like auditory perception.

Figures

Figures reproduced from arXiv: 2507.20530 by the authors.

Figure 1
Figure 1. , detecting the cries of a baby trapped under debris requires the robot to interpret binaural signals from its microphones, highlighting the vital role of sound event localization and detection (SELD) capabilities in real-world deployment. The main objective of this research is to develop a two￾channel SELD framework tailored for humanoid robots [16]. Conventional horizontal two-channel input systems struggle with f… view at source ↗
Figure 2
Figure 2. Abstract concept of binaural sound event localization and detection (BiSELD) for a humanoid robot [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 4
Figure 4. CRNN based BiSELD model: (a) BiSELDnet architecture with BTFF, (b) coordinate system conversion of output vectors (Cartesian  spherical), (c) result of sound event detection, and (d) result of sound event localization [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figures from the paper (2 more)
Figure 5
Figure 5. Figure 5: and summarized in Table IV, MS + V achieved a lower error rate (ER) of 0.350 compared to 0.392, and a higher F-score of 77.9% compared to 75.0%, reducing the SED error from 0.321 to 0.286. These results demonstrate that incorporating V￾map effectively enhances detectio…
Figure 7
Figure 7. Figure 7: Evaluation results of BiSELDnet on test set V with MS and MS + SC-map (MS + SC) as input features [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Auditory Intelligence: Understanding the World Through Sound

    eess.AS 2025-08 conditional novelty 4.0 of 10

    A position paper proposing four cognitively inspired audio tasks (ASPIRE, SODA, AUX, AUGMENT) to push machine hearing beyond recognition toward explanation, reasoning, and interaction.

Reference graph

Works this paper leans on

80 extracted references · 67 canonical work pages · cited by 1 Pith paper

  1. [64]

    Two vs. four-channel sound event localization and detection,

    J. Wilkins, M. Fuentes, L. Bondi, S. Ghaffarzadegan, A. Abavi sani, and J. P. Bello, “Two vs. four-channel sound event localization and detection,” in Proc. Detection Classification Acoust. Scenes Events Workshop (DCASE Workshop), Tampere, Finland, Sep. 2023, pp. 1–5

  2. [1]

    A comprehensive survey on humanoid robot development,

    S. Saeedvand, M. Jafari, H. S. Aghdasi, and J. Baltes, “A comprehensive survey on humanoid robot development,” Knowl. Eng. Rev., vol. 34, pp. 1– 18, Dec. 2019, doi: 10.1017/S02698889190001582

  3. [2]

    Overview of the torqu e-controlled humanoid robot TORO,

    J. Englsberger et al., “Overview of the torqu e-controlled humanoid robot TORO,” in Proc. 14th IEEE-RAS Int. Conf. Humanoid Robots, Madrid, Spain, Nov. 2014, pp. 916–923

  4. [3]

    Development of a biped walking robot compensating for three -axis moment by trunk motio n,

    J.-I. Yamaguchi , A. Takanishi , and I. Kato , “ Development of a biped walking robot compensating for three -axis moment by trunk motio n,” in Proc. IEEE/RSJ Int. Conf. Intelligent Robots and Systems (IROS) , Yokohama, Japan, Jul. 1993, pp. 561–566

  5. [4]

    The development of Honda humanoid robot,

    K. Hirai, M. Hirose, Y. Haikawa, and T. Takenaka, “The development of Honda humanoid robot,” in Proc. IEEE Int. Conf. Robotics and Automation (ICRA), Leuven, Belgium, May 1998, pp. 1321–1326

  6. [5]

    Towards the design of a biped jogging robot,

    M. Gienger, K. Löffler, and F. Pfeiffer, “Towards the design of a biped jogging robot,” in Proc. IEEE Int. Conf. Robotics and Automation (ICRA), Seoul, Korea, May 2001, pp. 4140–4145

  7. [6]

    Design of prototype humanoid robotics platform for HRP,

    K. Kaneko et al. , “ Design of prototype humanoid robotics platform for HRP,” in Proc. IEEE/RSJ Int. Conf. Intelligent Robots and Systems (IROS), Lausanne, Switzerland, Dec. 2002, pp. 2431–2436

  8. [7]

    The intelligent ASIMO: System overview and integration ,

    Y. Sakagami, R. Watanabe, C. Aoyama, S. Matsunaga, N. Higaki, and K. Fujimura, “The intelligent ASIMO: System overview and integration ,” in Proc. IEEE/RSJ Int. Conf. Intelligent Robots and System s (IROS), Lausanne, Switzerland, Dec. 2002, pp. 2478–2483

Show all 80 references
  1. [8]

    Mechanical design of humanoid robot platform KHR-3 (KAIST humanoid robot - 3: HUBO),

    I.-W. Park, J .-Y. Kim, J . Lee, and J .-H. Oh, “ Mechanical design of humanoid robot platform KHR-3 (KAIST humanoid robot - 3: HUBO),” in Proc. 5th IEEE-RAS Int. Conf. Humanoid Robots (ICHR), Tsukuba, Japan, Dec. 2005, pp. 321–326

  2. [9]

    Design of android type humanoid robot Albert HUBO,

    J.-H. Oh, D. Hanson, W.-S. Kim, I.-Y. Han, J.-Y. Kim, and I.-W. Park, “Design of android type humanoid robot Albert HUBO,” in Proc. IEEE/RSJ Int. Conf. Intelligent Robots and Systems (IROS), Beijing, China, Oct. 2006, pp. 1428–1433

  3. [10]

    A literature review of sensor heads for humanoid robots,

    J. A. Rojas-Quintero and M. C. Rodrí guez-Liñán, “A literature review of sensor heads for humanoid robots,” Robot. Auton. Syst., vol. 143, pp. 1– 21, Sep. 2021, doi: 10.1016/j.robot.2021.103834

  4. [11]

    The Harvard binocular head ,

    N. J. Ferrier , “ The Harvard binocular head ,” in Proc. SPIE 1708, Applications of Artificial Intelligenc e X: Machine Vision and Robotic s, Orlando, FL, USA, Mar. 1992, pp. 1–13. TABLE VI MEDIAN VALUES FOR BISELDNET PERFORMANCE WITH MS AND MS + SC-MAP AS INPUT FEATURES ON TE...

  5. [12]

    The robot musician ‘wabot-2’ (waseda robot-2),

    I. Kato et al., “The robot musician ‘wabot-2’ (waseda robot-2),” Robotics, vol. 3, no. 2, pp. 143–155, Jun. 1987, doi: 10.1016/0167-8493(87)90002- 7

  6. [13]

    Odor and airflow: Complementary senses for a humanoid robot ,

    R. A. Russell and A. H. Purnamadjaja , “ Odor and airflow: Complementary senses for a humanoid robot ,” in Proc. IEEE Int. Conf. Robotics and Automation (ICRA), Washington, DC, USA, May 2002, pp. 1842–1847

  7. [14]

    Electronic nose technology and application s,

    M. Bonnefille , “Electronic nose technology and application s,” in Practical Analysis of Flavor and Fragrance Materials , K. Goodner and R. Rouseff, Eds. Chichester, UK: Wiley, 2011, pp. 111–154

  8. [15]

    Development of auditory-evoked reflexes: Visuo -acoustic cues integration in a bino cular head ,

    L. Natale, G. Metta, and G. Sandini, “Development of auditory-evoked reflexes: Visuo -acoustic cues integration in a bino cular head ,” Robot. Auton. Syst., vol. 39, no. 2, pp. 87–106, May 2002, doi: 10.1016/S0921- 8890(02)00174-4

  9. [16]

    Binaural sound event localization and detection neural network based on HRTF localization cues for humanoid robots,

    G.-T. Lee, “Binaural sound event localization and detection neural network based on HRTF localization cues for humanoid robots,” Ph.D. Dissertation, KAIST, 2024

  10. [17]

    A. S. Bregman, Auditory Scene Analysis. The Perceptual Organization of Sound. Cambridge, MA, USA: MIT Press, 1990, pp. 3–9

  11. [18]

    Fundamentals of computational auditory scene analysis,

    D. Wang and G. J. Brown , “Fundamentals of computational auditory scene analysis,” in Computational Auditory Scene Analysis: Principl es, Algorithms, and Applications, D. Wang and G. J. Brown, Eds. Piscataway, NJ, USA: Wiley-IEEE Press, 2006, pp. 11–14

  12. [19]

    Computational models of auditory scene analysis: A review,

    B. T. Szabó, S. L. Denham, and I. Winkler , “Computational models of auditory scene analysis: A review,” Front. Neurosci., vol. 10, pp. 1–16, Nov. 2016, doi: 10.3389/fnins.2016.00524

  13. [20]

    Auditory streaming as an online classification process with evidence accumulation,

    D. Barniv and I. Nelken, “Auditory streaming as an online classification process with evidence accumulation,” PLoS One, vol. 10, no. 12, pp. 1– 20, Dec. 2015, doi: 10.1371/journal.pone.0144788

  14. [21]

    Combined estimation of spectral envelopes and sound source direction of concurrent voices by multidimensional statistical filtering,

    J. Nix and V. Hohmann, “Combined estimation of spectral envelopes and sound source direction of concurrent voices by multidimensional statistical filtering,” IEEE Trans. Audio Speech Lang. Process. , vol. 15, no. 3, pp. 995–1008, Mar. 2007, doi: 10.1109/TASL.2006.889788

  15. [22]

    An oscillatory correlation model of auditory streaming,

    D. Wang and P. Chang , “An oscillatory correlation model of auditory streaming,” Cogn. Neurodynamics, vol. 2, no. 1, pp. 7–19, Mar. 2008, doi: 10.1007/s11571-007-9035-8

  16. [23]

    Monophonic sound source separation with an unsupervised network of spiking neurons,

    R. Pichevar and J. Rouat, “Monophonic sound source separation with an unsupervised network of spiking neurons,” Neurocomputing, vol. 71, no. 1-3, pp. 109–120, Dec. 2007, doi: 10.1016/j.neucom.2007.08.001

  17. [24]

    Modelling the emergence and dynamics of perceptual organisation in auditory streaming,

    R. W. Mill, T. M. Böhm, A. Bendixen, I. Winkler, and S. L. Denham , “Modelling the emergence and dynamics of perceptual organisation in auditory streaming,” PLoS Comput. Biol. , vol. 9, no. 3, pp. 1–21, Mar. 2013, doi: 10.1371/journal.pcbi.1002925

  18. [25]

    Neuromechanistic model of auditory bistability,

    J. Rankin, E. Sussman, and J. Rinzel , “Neuromechanistic model of auditory bistability,” PLoS Comput. Biol., vol. 11, no. 11, pp. 1–34, Nov. 2015, doi: 10.1371/journal.pcbi.1004555

  19. [26]

    Segregating complex sound sources through temporal coherence ,

    L. Krishnan, M. Elhilali, and S. Shamma , “Segregating complex sound sources through temporal coherence ,” PLoS Comput. Biol. , vol. 10, no. 12, pp. 1–10, Dec. 2014, doi: 10.1371/journal.pcbi.1003985

  20. [27]

    A cocktail party with a cortical twist: How cortical mechanisms contribute to sound segregation,

    M. Elhilali and S. A. Shamm a, “A cocktail party with a cortical twist: How cortical mechanisms contribute to sound segregation,” J. Acoust. Soc. Am., vol. 124, no. 6, pp. 3751–3771, Dec. 2008, doi: 10.1121/1.3001672

  21. [28]

    Sound event localization and detection of overlapping sources using convolutional recurrent neural networks,

    S. Adavanne, A. Politis, J. Nikunen, and T. Virtanen , “Sound event localization and detection of overlapping sources using convolutional recurrent neural networks,” IEEE J. Sel. Top. Signal Process., vol. 13, no. 1, pp. 34–48, Mar. 2019, doi: 10.1109/JSTSP.2018.2885636

  22. [29]

    Sound source localization based on deep neural networks with directional activate function exploiting phase information,

    R. Takeda and K. Komatani , “Sound source localization based on deep neural networks with directional activate function exploiting phase information,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP), Shanghai, China, Mar. 2016, pp. 405–409

  23. [30]

    Sound source localiz ation using deep learning models,

    N. Yalta, K. Nakadai, and T. Ogata , “Sound source localiz ation using deep learning models,” J. Robot. Mechatron., vol. 29, no. 1, pp. 37–48, Feb. 2017, doi: 10.20965/jrm.2017.p0037

  24. [31]

    Deep neural networks for multiple speaker detection and locali zation,

    W. He, P. Motlicek, and J. -M. Odobez , “ Deep neural networks for multiple speaker detection and locali zation,” in Proc. IEEE Int . Conf. Robotics and Automation (ICRA), Brisbane, Australia, May 2018, pp. 74– 79

  25. [32]

    Audio surveillance: A systematic review,

    M. Crocco, M. Cristani, A. Trucco, and V. Murino, “Audio surveillance: A systematic review,” ACM Comput. Surv., vol. 48, no. 4, pp. 1–46, Feb. 2016, doi: 10.1145/2871183

  26. [33]

    Scream and gunshot detection and localization for audio -surveillance systems,

    G. Valenzise, L. Gerosa, M. Tagliasacchi, F. Antonacci, and A. Sarti , “Scream and gunshot detection and localization for audio -surveillance systems,” in Proc. IEEE Conf . Advanced Video and Signal Based Surveillance (AVSS), London, UK, Sep. 2007, pp. 21–26

  27. [34]

    Environmental sound recognition wi th time –frequency audio feature s,

    S. Chu, S. Narayanan, and C. -C. J. Kuo , “Environmental sound recognition wi th time –frequency audio feature s,” IEEE Trans. Audio Speech Lang. Process ., vol. 17, no. 6, pp. 1142–1158, Aug. 2009, doi: 10.1109/TASL.2009.2017438

  28. [35]

    Emergency vehicles audio detection and localization in autonomous driving ,

    H. Sun, X. Liu, K. Xu, J. Miao, and Q. Luo , “Emergency vehicles audio detection and localization in autonomous driving ,” 2021 , arXiv:2109.14797

  29. [36]

    Overview and evaluation of sound event localization and detection in DCASE 2019,

    A. Politis, A. Mesaros, S. Adavanne, T. Heittola, and T. Virtanen , “Overview and evaluation of sound event localization and detection in DCASE 2019,” IEEE-ACM Trans. Audio Speech Lang., vol. 29, pp. 684– 698, Dec. 2020, doi: 10.1109/TASLP.2020.3047233

  30. [37]

    Sound event detection and localization based on CNN and LSTM,

    Z. Lu, “Sound event detection and localization based on CNN and LSTM,” Detection Classification Acoust. Scenes Events (DCASE) Challenge , Tech. Rep., pp. 1–3, 2019

  31. [38]

    Online direction of arrival estimation based on deep learning ,

    Q. Li, X. Zhang, and H. Li, “Online direction of arrival estimation based on deep learning ,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP), Calgary, AB, Canada, Apr. 2018, pp. 2616–2620

  32. [39]

    Exploiting temporal context in CNN based multisource DOA estimation,

    A. Bohlender, A. Spriet, W. Tirry, and N. Madhu, “Exploiting temporal context in CNN based multisource DOA estimation,” IEEE/ACM Trans. Audio Speech Lang. Process., vol. 29, no. 1, pp. 1594–1608, Mar. 2021, doi: 10.1109/TASLP.2021.3067113

  33. [40]

    Sound event localization and detection based on CRNN using rectangular filters and channel rotation data augmentation ,

    F. Ronchini, D. Arteaga, and A. P érez-López, “Sound event localization and detection based on CRNN using rectangular filters and channel rotation data augmentation ,” in Proc. Detection Classification Acoust. Scenes Events Workshop (DCASE Workshop), Tokyo, Japan, Nov. 2020, p...

  34. [41]

    GCC-PHAT cross -correlation audio features for simultaneous sound event localization and detection (SELD) in multiple rooms ,

    H. A. C. Maruri, P. L. Meyer, J. H uang, JAdH. Ontiveros, and H. L u, “GCC-PHAT cross -correlation audio features for simultaneous sound event localization and detection (SELD) in multiple rooms ,” Detection Classification Acoust. Scenes Events (DCASE) Challenge, Tech. Rep., p...

  35. [42]

    Polyphonic sound event detection and localization using a two-stage strategy ,

    Y. Cao et al., “Polyphonic sound event detection and localization using a two-stage strategy ,” in Proc. Detecti on Classification Acoust. Scenes Events Workshop (DCASE Workshop), New York, NY, USA, Oct. 2019, pp. 30–34

  36. [43]

    Two-stage sound event localization and detection using intensity vector and generalized cross -correlation,

    Y. Cao et al., “Two-stage sound event localization and detection using intensity vector and generalized cross -correlation,” Detection Classification Acoust. Scenes Events (DCASE) Challenge, Tech. Rep., pp. 1–4, 2019

  37. [44]

    Sound event detection an d localization using CRNN model s,

    A. Sampathkumar and D. Kowerko , “ Sound event detection an d localization using CRNN model s,” Detection Classification Acoust. Scenes Events (DCASE) Challenge, Tech. Rep., pp. 1–3, 2020

  38. [45]

    An improved event -independent network for polyphonic sound event localization and detection ,

    Y. Cao, T. Iqbal, Q. Kong, F. An, W. Wang, and M. D. Plumbley , “An improved event -independent network for polyphonic sound event localization and detection ,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP), Toronto, ON, Canada, Jun. 2021, pp. 885–889

  39. [46]

    A general network architecture for sound event localization and detection using transfer learning and recurrent neural network,

    T. N. T. Nguyen et al., “A general network architecture for sound event localization and detection using transfer learning and recurrent neural network,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP), Toronto, ON, Canada, Jun. 2021, pp. 935–939

  40. [47]

    ACCDOA: Activity -coupled Cartesian direction of arrival representation for sound event localization and detection,

    K. Shimada, Y. Koyama, N. Takahashi, S. Takahashi, and Y. Mitsufuji , “ACCDOA: Activity -coupled Cartesian direction of arrival representation for sound event localization and detection,” in Proc. IEEE Int. Conf. Acoust., Speech, Sign al Process. (ICASSP ), Toronto, ON , Canad...

  41. [48]

    Multi-ACCDOA: Localizing and detecting overlapping sounds from the same class with auxiliary duplicating per mutation invariant training ,

    K. Shimada, Y. Koyama, S. Takahashi, N. Takahashi, E. Tsunoo, and Y. Mitsufuji, “Multi-ACCDOA: Localizing and detecting overlapping sounds from the same class with auxiliary duplicating per mutation invariant training ,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process...

  42. [49]

    SALSA: Spatial cue-augmented log-spectrogram features for polyphonic sound event localization and detection ,

    T. N. T. Nguyen, K. N. Watcharasupat, N. K. Ng uyen, D. L. Jones, and W.-S. Gan, “SALSA: Spatial cue-augmented log-spectrogram features for polyphonic sound event localization and detection ,” IEEE-ACM Trans. Audio Speech Lang ., vol. 30 , pp. 1749–1762, May 2022, doi: 10.1109...

  43. [50]

    SALSA-Lite: A fast and effective feature for polyphonic sound event localization and detection with microphone arrays,

    T. N. T. Nguyen, D. L. Jones, K. N. Watcharasupat, H. Phan, and W. -S. Gan, “SALSA-Lite: A fast and effective feature for polyphonic sound event localization and detection with microphone arrays,” in Proc. IEEE Int. Conf. Acoust., Speech, Signal Process. (ICASSP ), Singapore, ...

  44. [51]

    A report on sound event detection with different binaural features ,

    S. Adavanne and T. Virtanen , “A report on sound event detection with different binaural features ,” in Proc. Detection Classification Acoust. Scenes Events Workshop (DCASE Workshop ), Munich, Germany, Nov. 2017, pp. 1–5. 12 > REPLACE THIS LINE WITH YOUR MANUSCRIPT ID NUMBER (...

  45. [52]

    Binaural signal representations for joint sound event detection and acoustic scene classification ,

    D. A. Krause and A. Mesaros , “Binaural signal representations for joint sound event detection and acoustic scene classification ,” in Proc. 30th Euro. Signal Process. Conf. (EUSIPCO), Belgrade, Serbia, Aug. 2022, pp. 399–403

  46. [53]

    A learning-based approach to robust binaural sound localization ,

    K. Youssef, S. Argentieri, and J.-L. Zarader, “A learning-based approach to robust binaural sound localization ,” in Proc. IEEE/RSJ Int. Conf. Intelligent Robots and Systems (IROS), Tokyo, Japan , Nov. 2013, pp. 2927–2932

  47. [54]

    On sound source localization of speech signals using deep neural networks,

    R. Roden, N. Moritz, S. Gerlach, S. Weinzierl, and S. Goetze, “On sound source localization of speech signals using deep neural networks,” in Proc. Deutsche Jahrestagung Akustik (DAGA ), Nuremberg, Germany , Mar. 2015, pp. 1510–1513

  48. [55]

    Autonomous sensorimotor learning for sound source localization by a humanoid robot,

    Q. V. Nguyen, L. Girin , G. Bailly, F. Elisei, and D. C. Nguyen , “Autonomous sensorimotor learning for sound source localization by a humanoid robot,” in Proc. IEEE/RSJ Int. Conf. Intelligent Robots and Systems (IROS), Madrid, Spain, Oct. 2018, pp. 1–4

  49. [56]

    Multitask learning of time-frequency CNN for sound source localization ,

    C. Pang, H. Liu, and X. L i, “Multitask learning of time-frequency CNN for sound source localization ,” IEEE Access, vol. 7, pp. 40725–40737, Mar. 2019, doi: 10.1109/ACCESS.2019.2905617

  50. [57]

    Full-sphere binaural sound source localization using multi -task neural network ,

    Y. Yang, J. Xi, W. Zhang, and L. Zhang , “ Full-sphere binaural sound source localization using multi -task neural network ,” in Proc. Asia - Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), Auckland, New Zealand, Dec. 2020, pp. 432–436

  51. [58]

    Deep neural network based audio source separation,

    A. Zermini, Y. Yu, Y. Xu, M. D. Plumbley, and W. Wang, “Deep neural network based audio source separation,” in Proc. Int. Conf. Math. Signal Proc. (IMA), Birmingham, UK, May 2016, pp. 1–4

  52. [59]

    Exploiting deep neural networks and head movements for robust binaural localization of multiple sources in reverberant environments ,

    N. Ma, T. Ma y, and G. J. Brown , “Exploiting deep neural networks and head movements for robust binaural localization of multiple sources in reverberant environments ,” IEEE/ACM Trans. Audio Speech Lang. Process., vol. 25 , no. 12 , pp. 2444–2453, Dec. 2017, doi: 10.1109/TASL...

  53. [60]

    Learning deep direct-path relative transfer function for binaural sound source localization,

    B. Yang, H. Liu, and X. Li , “Learning deep direct-path relative transfer function for binaural sound source localization,” IEEE/ACM Trans. Audio Speech Lang. Process ., vol. 29 , pp. 3491–3503, Oct. 2021, doi: 10.1109/TASLP.2021.3120641

  54. [61]

    Binaural source localization using deep learning and head rotation information,

    G. García-Barrios, D. A. Krause, A. Politis, A. Mesaros, J. M. Gutiérrez- Arriola, and R. Fraile, “Binaural source localization using deep learning and head rotation information,” in Proc. 30th Euro. Signal Process. Conf. (EUSIPCO), Belgrade, Serbia, Aug. 2022, pp. 36–40

  55. [62]

    Binaural source localization in median plane using learning based method for robot audition,

    P. Dwivedi, G. Routray, and R. M. Hegde, “Binaural source localization in median plane using learning based method for robot audition,” in Proc. 24th Int. Congr. Acoust. (ICA), Gyeongju, Korea, Oct. 2022, pp. 1–8

  56. [63]

    Goal-driven, neurobiological- inspired convolutional neural network models of human spatial hearing,

    K. van der Heijden and S. M ehrkanoon, “Goal-driven, neurobiological- inspired convolutional neural network models of human spatial hearing,” Neurocomputing, vol. 470 , pp. 432–442, Jan. 2022, doi: 10.1016/j.neucom.2021.05.104

  57. [65]

    HRTF measurement for accurate sound localization cues,

    G.-T. Lee, S.-M. Choi, B.-Y. Ko, and Y. -H. Park, “HRTF measurement for accurate sound localization cues,” 2022, arXiv:2203.03166v2

  58. [66]

    Iida, Head-Related Transfer Function and Acoustic Virtual Reality

    K. Iida, Head-Related Transfer Function and Acoustic Virtual Reality . Narashino, Japan: Springer, 2019, pp. 15–55

  59. [67]

    Xie, Head-Related Transfer Function and Virtual Auditor y Display, 2nd ed., Plantation, FL, USA: J

    B. Xie, Head-Related Transfer Function and Virtual Auditor y Display, 2nd ed., Plantation, FL, USA: J. Ross Publishing, 2013, pp. 81–85

  60. [68]

    Sound pressure generated in an external- ear replica and real human ears by a nearby point source,

    E. A. G. Shaw and R. Teranishi, “Sound pressure generated in an external- ear replica and real human ears by a nearby point source,” J. Acoust. Soc. Am., vol. 44, no. 1, pp. 240–249, Jul. 1968, doi: 10.1121/1.1911059

  61. [69]

    Mechanism for generating pe aks and notches of head -related transf er functions in the median plan e,

    H. Takemoto, P. Mokhtari, H. Kato, R. Nishimura, and K. Iida , “Mechanism for generating pe aks and notches of head -related transf er functions in the median plan e,” J. Acoust. Soc. Am ., vol. 132, no. 6, pp. 3832–3841, Dec. 2012, doi: 10.1121/1.4765083

  62. [70]

    Deep learning based cough detection camera using enhanced features ,

    G.-T. Lee, H. Nam, S. -H. Kim, S. -M. Choi, Y. Kim, and Y. -H. Park , “Deep learning based cough detection camera using enhanced features ,” Expert Syst. Appl ., vol. 206, pp. 1–20, Nov. 2022, doi: 10.1016/j.eswa.2022.117811

  63. [71]

    Estimation of interaural time difference based on cochlear filter bank and ZCPA auditory model ,

    G.-T. Lee and Y.-H. Park, “Estimation of interaural time difference based on cochlear filter bank and ZCPA auditory model ,” Trans. Korean Soc. Noise Vib. Eng ., vol. 29, no. 6, pp. 722–734, Dec. 2019, doi: 10.5050/KSNVE.2019.29.6.722

  64. [72]

    Method for estimating interaural time differences based on cochlea filter bank and EUZ auditory model in noisy environment,

    G.-T. Lee and Y.-H. Park , “ Method for estimating interaural time differences based on cochlea filter bank and EUZ auditory model in noisy environment,” in Proc. Int. Conf. Noise Control Eng. (Inter-noise), Seoul, Korea, Aug. 2020, pp. 1–12

  65. [73]

    Significance of the modified group delay feature in speech recognition,

    R. M. Hegde, H. A. Murthy, and V. R. R. Gad de, “Significance of the modified group delay feature in speech recognition,” IEEE Trans. Audio Speech Lang. Process ., vol. 1 5, no. 1, pp. 190–202, Jan. 2007, doi: 10.1109/TASL.2006.876858

  66. [74]

    Group delay spectrogram of speech signals without phase wrapping,

    B. Yegnanarayana, “Group delay spectrogram of speech signals without phase wrapping,” J. Acoust. Soc. Am., vol. 151, no. 3, pp. 2181–2191, Mar. 2022, doi: 10.1121/10.0009922

  67. [75]

    R. F. Lyon , Human and Machine Hearing: Extracting Meaning from Sound. Cambridge, UK: Cambridge Univ. Press, 2017, pp. 66–67

  68. [76]

    The NIGENS general sound events database,

    I. Trowitzsch, J. Taghia , Y. Kashef, and K. Obermayer , “The NIGENS general sound events database,” 2020, arXiv:1902.08314

  69. [77]

    Detection and classification of acoustic scenes and events: Outcome of the DCASE 2016 challenge ,

    A. Mesaros et al., “Detection and classification of acoustic scenes and events: Outcome of the DCASE 2016 challenge ,” IEEE-ACM Trans. Audio Speech Lang ., vol. 26, no. 2 , pp. 379–393, Feb. 2018, doi: 10.1109/TASLP.2017.2778423

  70. [78]

    A multi -device d ataset for urban acoustic scene classificatio n,

    A. Mesaros, T. Heittola, and T. Virtanen , “ A multi -device d ataset for urban acoustic scene classificatio n,” in Proc. Detection Classification Acoust. Scenes Events Workshop (DCASE Workshop), Surrey, UK, Nov. 2018, pp. 1–5

  71. [79]

    Binaural sound event localization and detection for humanoid robo t,

    G.-T. Lee and Y. -H. Park , “ Binaural sound event localization and detection for humanoid robo t,” in Proc. 24th Int. Congr. Acoust. (ICA ), Gyeongju, Korea, Oct. 2022, pp. 1–12

  72. [80]

    Joint measurement of localization and detection of sound events ,

    A. Mesaros, S. Adavanne, A. Politis, T. Heittola, and T. Virtanen, “Joint measurement of localization and detection of sound events ,” in Proc. IEEE Workshop Appl. Signal Process. Audio Acoust. (WASPAA ), New Paltz, NY, USA, Oct. 2019, pp. 333–337

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.