REVIEW 3 major objections 5 minor 238 references
A Review on Sound Source Localization in Robotics: Focusing on Deep Learning Methods
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read No prior survey has focused on deep-learning sound localization for robots; this review supplies that map from 78 studies.
desk verdict Solid, useful survey for robotics researchers entering SSL, but the gap claim against Lv et al. is overstated and the corpus needs to be published. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a four-way taxonomy applied to a curated corpus of 78 papers: method family, microphone-array configuration, robot type, and application domain. The taxonomy does the argumentative work by turning scattered experiments into a map, and the map is the review's product. A second load-bearing mechanism is the paper's pairing of data (synthetic room simulations and robot-recorded corpora) with learning paradigms (supervised, semi-supervised, weakly supervised, transfer), which the review treats as the two cornerstones that decide whether deep models transfer to robots.
What would settle it
Run the same boolean search over the major robotics and audio publication venues listed in Section 2 for 2013–2025; finding a peer-reviewed Transformer-based SSL study tested on a robot, or an earlier robotics-focused deep-learning SSL review, would undercut the gap claim and the roadmap built on it.
Extended reading notes
Core claim
At its center, the paper claims that 'no review has explored SSL in robotic platforms with a particular focus on deep-learning models,' and positions itself as that missing review. It synthesizes 78 peer-reviewed studies and arranges the field along four axes: classical method families (TDOA, beamforming, steered-response power, subspace analysis), deep-learning architectures from MLPs and CNNs through CRNNs to attention models, microphone-array configurations, and robot types and application domains. The trajectory it reports is a shift from single-source analytic methods to data-driven models that handle noise, reverberation, multiple sources, and ego-noise, with Transformer-based networks still untested aboard robots. On the paper's own terms, the central discovery is that robotics imposes its own constraints—platform motion, motor noise, limited compute—and that general-audio SSL surveys have underplayed those constraints.
Load-bearing premise
The load-bearing premise is the Section 1 claim that no earlier review has covered sound source localization on robotic platforms with a deep-learning focus; that premise is strained by the paper's own description of reference [5], which already covers machine-learning and deep-learning SSL for condition-monitoring robots, and it would collapse if the literature search missed relevant studies.
Editorial extensions
If this is right
- For a given robot class, the array geometry is a design constraint, not a detail: binaural setups dominate humanoids, while circular and spherical arrays are favored for 360-degree and aerial platforms.
- The absence of Transformer-based SSL on robots is an open research slot, not a sign of impossibility; attention models are already state of the art in general audio tasks.
- Robotic SSL evaluation should include ego-noise and platform motion, because models trained on clean, static corpora may fail on real robots.
- Data and learning strategy decisions—synthetic room simulation, augmentation, and transfer from pre-trained audio models—carry as much weight as architecture choice for deep SSL on robots.
- The field needs shared robotics-specific benchmarks with accurate spatial labels; without them, array-geometry and method comparisons stay anecdotal.
Reading between the lines
- Beyond the paper: the same evidence implies a lagged adoption pattern—general SSL moved to attention around 2020–2022 while robotic SSL is still mostly CNN/CRNN—so Transformer-based robotic SSL is a testable near-term prediction.
- Beyond the paper: the call for shared benchmarks implies that robotics needs an evaluation protocol with fixed tasks, metrics, and baseline systems; without it, claims about array geometry or ego-noise robustness cannot be compared across robots.
- Beyond the paper: if large language models are added after localization and speech recognition, then localization error—not transcription accuracy—becomes the bottleneck for grounding verbal commands to physical referents.
- Beyond the paper: the gap claim is better read as 'deep-learning SSL across general robotic platforms' than as 'deep-learning SSL in robotics' in absolute terms, because reference [5] already covers condition-monitoring robots.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript is a survey of sound source localization (SSL) in robotics, with a stated emphasis on deep learning. It reviews classical methods (TDOA, beamforming, SRP, subspace analysis), machine/deep learning architectures (MLP, CNN, CRNN, attention-based models), datasets and learning paradigms, and it categorizes a corpus of 78 papers by robot type and application domain. The paper's central positioning is that it fills a gap in the survey literature: no prior review has explored SSL on robotic platforms with a particular focus on deep learning. It closes with a discussion of open challenges and future research directions, including a roadmap centered on robustness, multi-source analysis, efficiency, and explainability.
Significance. If the corpus is representative, the review is a useful synthesis: it organizes a sizeable body of work into a clear taxonomy, provides detailed tabular summaries of methods and robot platforms, and gives a thoughtful treatment of data generation, augmentation, and learning paradigms, including robotics-specific datasets such as DREGON, SSLR, AVQ, and UaVirBASE. The classical-method overview is accurate and well structured, and the application-domain categorization should help practitioners locate relevant work. However, the value of the survey depends on two load-bearing assumptions that are not currently established: that the claimed gap in the review literature is real, and that the 78-paper corpus is representative and auditable. Both need to be addressed before the review's central claims can be accepted.
major comments (3)
- [Section 1, Table 1] Section 1 states that 'no review has explored SSL in robotic platforms with a particular focus on deep-learning models.' This absolute claim is not reconciled with Lv et al. [5], which the paper's own Table 1 lists as a peer-reviewed SSL review for condition-monitoring robots that covers 'traditional and machine learning models.' Condition-monitoring robots are robotic platforms, and the paper later classifies industrial condition monitoring as one of its own application domains (Section 6.2). If Lv et al. is excluded because its focus is narrower or because its machine-learning coverage is not specifically deep learning, the authors need to state that criterion explicitly. As written, the central novelty assertion is not robustly established.
- [Section 2.2, Tables 3 and 4] The methodology states that the search converged on a corpus of 78 papers, but no machine-readable corpus list, search log, or screening decisions are provided. Tables 3 and 4 enumerate only 53 method-classified papers (33 traditional and 20 ML/DL), leaving the remaining 25 papers itemized nowhere in a verifiable form. This limits independent audit of the representative-coverage claim and makes the Section 4.2.4 statement that Transformer-based networks have not been adopted for robotic SSL impossible to verify. The authors should provide the full list of 78 papers with DOIs or arXiv identifiers, and ideally the search strings and inclusion/exclusion outcomes.
- [Section 4.2.4] The sentence 'Transformer-based networks—now state-of-the-art in many audio tasks—have not yet been adopted for robotic SSL' is a concrete absence claim that anchors a future-work direction. Given Section 2.1's own admission that the search prioritized 'breadth of coverage over exhaustive enumeration,' a single robot-mounted Transformer SSL study outside the corpus would invalidate the claim. The claim should be restricted to the reviewed corpus, or supported by a reproducible search protocol that can substantiate the absence.
minor comments (5)
- [Section 3 heading] The heading 'SSL foundementals' contains a typo; it should read 'SSL fundamentals.'
- [Table 2] The table header contains 'T ransaction' with an erroneous space in 'IEEE T ransaction on Instrumentation and Measurement'; the same issue appears in the second IEEE/ACM Transactions row.
- [Section 4.2.4] There is a typo, 'for instnae,' which should read 'for instance.'
- [Section 6.2 / Figure 7] Figure 7 shows a pie-chart distribution of application domains, but the text provides counts for only some categories (social 14, search and rescue 12, service 9, industrial 9) and no count for the 'General' category; adding exact counts would make the figure verifiable.
- [References] Reference formatting is inconsistent: for example, [5] is cited as 'ISA transactions (2024)' without volume or page details, and some entries lack complete venue information; a consistent style would improve usability.
Circularity Check
No significant circularity; the review's synthesis is independent of the authors' prior results, with only minor non-load-bearing self-citations.
full rationale
This is a survey paper, not a derivation chain. It contains no equations from which results are derived and no fitted parameter that is later relabeled as a prediction. The central contribution is a literature synthesis: a corpus of 78 robotics-related SSL papers organized by method, robot type, and application domain, plus a qualitative account of challenges and future directions. The novelty claim in Section 1 ('no review has explored SSL in robotic platforms with a particular focus on deep-learning models') is a literature-coverage assertion, checkable against Table 1. Even if one disputes it by pointing to Lv et al. [5], which reviews SSL in condition-monitoring robots including machine-learning methods, that is a completeness or scoping criticism, not circularity. The paper's own Section 7.1.4 also flags the field's lack of comprehensive benchmarks, which is a limitation statement about the research area rather than a self-referential justification. The only in-house items are the authors' own papers [4,72], which appear as surveyed examples (e.g., Section 4.2.3 for CRNN-based SSL and Section 6.2 for industrial applications) and in Table 4. These self-citations are illustrative, not load-bearing: removing them would not change any methodological conclusion, taxonomy, or future-work roadmap. No imported uniqueness theorem, no ansatz smuggled via citation, and no renaming of a known result are present. The score of 2 reflects the presence of minor self-citations among the reviewed corpus; the central synthesis remains independent of those citations.
Assumptions & free parameters
assumptions (3)
- domain assumption The 78-paper corpus identified by the search strategy is representative of the robotic SSL field.
- domain assumption The inclusion rule that a paper must evaluate SSL on a robot or target a robotic use-case is a sufficient boundary for 'SSL in robotics'.
- domain assumption The authors' summaries of the cited papers accurately reflect the original source results.
Cite this review
Pith. "Pith review of A Review on Sound Source Localization in Robotics: Focusing on Deep Learning Methods." pith.science (2026). https://pith.science/paper/MPYJU4GP
@misc{pith2026250701143,
author = {Pith},
title = {Pith review of: A Review on Sound Source Localization in Robotics: Focusing on Deep Learning Methods},
year = {2026},
howpublished = {\url{https://pith.science/paper/MPYJU4GP}},
note = {Machine review of arXiv:2507.01143}
}
read the original abstract
Sound source localization (SSL) adds a spatial dimension to auditory perception, allowing a system to pinpoint the origin of speech, machinery noise, warning tones, or other acoustic events, capabilities that facilitate robot navigation, human-machine dialogue, and condition monitoring. While existing surveys provide valuable historical context, they typically address general audio applications and do not fully account for robotic constraints or the latest advancements in deep learning. This review addresses these gaps by offering a robotics-focused synthesis, emphasizing recent progress in deep learning methodologies. We start by reviewing classical methods such as Time Difference of Arrival (TDOA), beamforming, Steered-Response Power (SRP), and subspace analysis. Subsequently, we delve into modern machine learning (ML) and deep learning (DL) approaches, discussing traditional ML and neural networks (NNs), convolutional neural networks (CNNs), convolutional recurrent neural networks (CRNNs), and emerging attention-based architectures. The data and training strategy that are the two cornerstones of DL-based SSL are explored. Studies are further categorized by robot types and application domains to facilitate researchers in identifying relevant work for their specific contexts. Finally, we highlight the current challenges in SSL works in general, regarding environmental robustness, sound source multiplicity, and specific implementation constraints in robotics, as well as data and learning strategies in DL-based SSL. Also, we sketch promising directions to offer an actionable roadmap toward robust, adaptable, efficient, and explainable DL-based SSL for next-generation robots.
Reference graph
Works this paper leans on
-
[5]
D.Lv,W.Tang,G.Feng,D.Zhen,F.Gu,A.D.Ball, An overview of sound source localization based condition monitoring robots, ISA transactions (2024)
2024
-
[1]
Rascon, I
C. Rascon, I. Meza, Localization of sound sources in robotics: A review, Robotics and Autonomous Systems 96 (2017) 184–210
2017
-
[2]
Jo, T.-W
H.-M. Jo, T.-W. Kim, K.-C. Kwak, Sound source localization using deep learning for human–robot interaction under intelligent robot environments, Electronics 14 (2025) 1043
2025
-
[3]
Korayem, S
M. Korayem, S. Azargoshasb, A. Korayem, S. Tabibian, Design and implementation of the voice command recognition and the sound source localization system for human–robot interaction, Robotica 39 (2021) 1779–1790
2021
-
[4]
Jalayer, M
R. Jalayer, M. Jalayer, C. Orsenigo, C. Vercellis, A conceptual framework for localization of ac- tive sound sources in manufacturing environment based on artificial intelligence, in: International Conference on Flexible Automation and Intelligent Manufacturing, 2023, pp. 699–707
2023
-
[6]
D. Lv, G. Feng, D. Zhen, X. Liang, G. Sun, F. Gu, Motor bearing fault source localization based on soundandrobotmovementcharacteristics,in: 2024 InternationalConferenceonSensing,Measurement &DataAnalyticsintheeraofArtificialIntelligence (ICSMD), IEEE, 2024, pp. 1–6
2024
-
[7]
Marques, J
I. Marques, J. Sousa, B. Sá, D. Costa, P. Sousa, S.Pereira,A.Santos,C.Lima,N.Hammerschmidt, S. Pinto, et al., Microphone array for speaker lo- calization and identification in shared autonomous vehicles, Electronics 11 (2022) 766
2022
-
[8]
Yamada, K
T. Yamada, K. Itoyama, K. Nishida, K. Nakadai, Sound source tracking by drones with microphone arrays, in: 2020 IEEE/SICE International Sympo- sium on System Integration (SII), IEEE, 2020, pp. 796–801
2020
Show all 238 references
-
[9]
Yamamoto, K
T. Yamamoto, K. Hoshiba, B. Yen, K. Nakadai, Implementation of a robot operation system-based network for sound source localization using mul- tiple drones, in: 2024 Asia Pacific Signal and Information Processing Association Annual Sum- 25 Reza Jalayer et al.- Preprint 1–35 mi...
2024
-
[10]
T.Latif,E.Whitmire,T.Novak,A.Bozkurt, Sound localization sensors for search and rescue biobots, IEEE Sensors Journal 16 (2015) 3444–3453
2015
-
[11]
Zhang, K
B. Zhang, K. Masahide, H. Lim, Sound source localization and interaction based human search- ing robot under disaster environment, in: 2019 SICEInternationalSymposiumonControlSystems (SICE ISCS), IEEE, 2019, pp. 16–20
2019
-
[12]
Yamada, H
N.Mae,Y.Mitsui,S.Makino,D.Kitamura,N.Ono, T. Yamada, H. Saruwatari, Sound source local- ization using binaural difference for hose-shaped rescue robot, in: 2017 Asia-Pacific Signal and In- formation Processing Association Annual Summit and Conference (APSIPA ASC), IEEE, 2017, ...
2017
-
[13]
J.-H. Park, K. B. Sim, A design of mobile robot based on network camera and sound source local- izationforintelligentsurveillancesystem, in: 2008 International Conference on Control, Automation and Systems, IEEE, 2008, pp. 674–678
2008
-
[14]
Z. Han, T. Li, Research on sound source local- ization and real-time facial expression recognition for security robot, in: Journal of Physics: Confer- ence Series, volume 1621, IOP Publishing, 2020, p. 012045
2020
-
[15]
Obeidat, W
H. Obeidat, W. Shuaieb, O. Obeidat, R. Abd- Alhameed, A review of indoor localization tech- niques and wireless technologies, Wireless Per- sonal Communications 119 (2021) 289–327
2021
-
[16]
Tarokh, P
M. Tarokh, P. Merloti, Vision-based robotic per- son following under light variations and difficult walking maneuvers, Journal of Field Robotics 27 (2010) 387–398
2010
-
[17]
D. Hall, B. Talbot, S. R. Bista, H. Zhang, R. Smith, F. Dayoub, N. Sünderhauf, The robotic vision scene understanding challenge, arXiv preprint arXiv:2009.05246 (2020)
2020 arXiv
-
[18]
Belkin, A
I. Belkin, A. Abramenko, D. Yudin, Real-time lidar-based localization of mobile ground robot, Procedia Computer Science 186 (2021) 440–448
2021
-
[19]
Z. Yu, A wifi indoor localization system based on robot data acquisition and deep learning model, in: 2024 6th International Conference on Internet of Things, Automation and Artificial Intelligence (IoTAAI), IEEE, 2024, pp. 367–371
2024
-
[20]
Aun, et al., Indoor positioning system: A re- view, International Journal of Advanced Computer Science and Applications 13 (2022)
N.H.A.Wahab,N.Sunar,S.H.Ariffin,K.Y.Wong, Y. Aun, et al., Indoor positioning system: A re- view, International Journal of Advanced Computer Science and Applications 13 (2022)
2022
-
[21]
I. S. Alfurati, A. T. Rashid, Performance compar- ison of three types of sensor matrices for indoor multi-robot localization, International Journal of Computer Applications 181 (2018) 22–29
2018
-
[22]
A. M. Flynn, R. A. Brooks, W. M. Wells III, D. S. Barrett, Squirt: The prototypical mobile robot for autonomous graduate students (1989)
1989
-
[23]
Grumiaux, S
P.-A. Grumiaux, S. Kitić, L. Girin, A. Guérin, A survey of sound source localization with deep learning methods, The Journal of the Acoustical Society of America 152 (2022) 107–151
2022
-
[24]
M. U. Liaquat, H. S. Munawar, A. Rahman, Z. Qadir, A. Z. Kouzani, M. P. Mahmud, Lo- calization of sound sources: A systematic review, Energies 14 (2021) 3910
2021
-
[25]
D.Desai,N.Mehendale, Areviewonsoundsource localization systems, Archives of Computational Methods in Engineering 29 (2022) 4631–4642
2022
-
[26]
B. J. Zhang, N. T. Fitter, Nonverbal sound in human-robot interaction: A systematic review, ACM Transactions on Human-Robot Interaction 12 (2023) 1–46
2023
-
[27]
G.Jekateryńczuk,Z.Piotrowski, Asurveyofsound sourcelocalizationanddetectionmethodsandtheir applications, Sensors 24 (2023) 68
2023
-
[28]
A. Khan, A. Waqar, B. Kim, D. Park, A review on recent advances in sound source localization techniques, challenges, and applications, Sensors and Actuators Reports (2025) 100313
2025
-
[29]
He, Deep Learning Approaches for Auditory Perception in Robotics, Ph.D
W. He, Deep Learning Approaches for Auditory Perception in Robotics, Ph.D. thesis, EPFL, 2021
2021
-
[30]
2927–2932
K.Youssef,S.Argentieri,J.-L.Zarader, Alearning- based approach to robust binaural sound localiza- tion, in: 2013 IEEE/RSJ International Conference on Intelligent Robots and Systems, IEEE, 2013, pp. 2927–2932
2013
-
[31]
Nakamura, R
K. Nakamura, R. Gomez, K. Nakadai, Real-time super-resolution three-dimensional sound source localization for robots, in: 2013 IEEE/RSJ In- ternational Conference on Intelligent Robots and Systems, IEEE, 2013, pp. 3949–3954
2013
-
[32]
Ohata, K
T. Ohata, K. Nakamura, T. Mizumoto, T. Taiki, K. Nakadai, Improvement in outdoor sound source detection using a quadrotor-embedded microphone array, in: 2014IEEE/RSJInternationalConference on Intelligent Robots and Systems, IEEE, 2014, pp. 1902–1907
2014
-
[33]
Grondin, F
F. Grondin, F. Michaud, Time difference of arrival estimation based on binary frequency mask for sound source localization on mobile robots, in: 2015 IEEE/RSJ International Conference on Intel- ligentRobotsandSystems(IROS),IEEE,2015, pp. 6149–6154
2015
-
[34]
Nakamura, L
K. Nakamura, L. Sinapayen, K. Nakadai, Interac- tive sound source localization using robot audition for tablet devices, in: 2015 IEEE/RSJ Interna- tional Conference on Intelligent Robots and Sys- tems (IROS), IEEE, 2015, pp. 6137–6142
2015
-
[35]
X. Li, L. Girin, F. Badeig, R. Horaud, Reverberant 26 Reza Jalayer et al.- Preprint 1–35 soundlocalizationwitharobotheadbasedondirect- path relative transfer function, in: 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE, 2016, pp. 2819–2826
2016
-
[36]
Nakadai, M
K. Nakadai, M. Kumon, H. G. Okuno, K. Hoshiba, M. Wakabayashi, K. Washizaki, T. Ishiki, D. Gabriel, Y. Bando, T. Morito, et al., Devel- opment of microphone-array-embedded uav for search and rescue task, in: 2017 IEEE/RSJ In- ternational Conference on Intelligent Robots and Sy...
2017
-
[37]
Strauss, P
M. Strauss, P. Mordel, V. Miguet, A. Deleforge, Dregon: Dataset and methods for uav-embedded sound source localization, in: 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE, 2018, pp. 1–8
2018
-
[38]
5320–5325
L.Wang,R.Sanchez-Matilla,A.Cavallaro, Audio- visualsensingfromaquadcopter: datasetandbase- lines for source localization and sound enhance- ment, in: 2019IEEE/RSJInternationalConference on Intelligent Robots and Systems (IROS), IEEE, 2019, pp. 5320–5325
2019
-
[39]
Michaud, S
S. Michaud, S. Faucher, F. Grondin, J.-S. Lauzon, M. Labbé, D. Létourneau, F. Ferland, F. Michaud, 3d localization of a sound source using mobile microphone arrays referenced by slam, in: 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE, 20...
2020
-
[40]
Sewtz, T
M. Sewtz, T. Bodenmüller, R. Triebel, Robust music-based sound source localization in reverber- ant and echoic environments, in: 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), IEEE, 2020, pp. 2474–2480
2020
-
[41]
Tourbabin, H
V. Tourbabin, H. Barfuss, B. Rafaely, W. Keller- mann, Enhanced robot audition by dynamic acous- tic sensing in moving humanoids, in: 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2015, pp. 5625–5629
2015
-
[42]
R.Takeda,K.Komatani, Soundsourcelocalization based on deep neural networks with directional activate function exploiting phase information, in: 2016 IEEE international conference on acoustics, speech and signal processing (ICASSP), IEEE, 2016, pp. 405–409
2016
-
[43]
Takeda, K
R. Takeda, K. Komatani, Unsupervised adaptation of deep neural networks for sound source localiza- tion using entropy minimization, in: 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2017, pp. 2217–2221
2017
-
[44]
E. L. Ferguson, S. B. Williams, C. T. Jin, Sound source localization in a multipath environment us- ing convolutional neural networks, in: 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2018, pp. 2386–2390
2018
-
[45]
4530–4535
F.Grondin,F.Michaud, Noisemaskfortdoasound source localization of speech on mobile robots in noisy environments, in: 2016 IEEE International Conference on Robotics and Automation (ICRA), IEEE, 2016, pp. 4530–4535
2016
-
[46]
W. He, P. Motlicek, J.-M. Odobez, Deep neural networks for multiple speaker detection and local- ization, in: 2018 IEEE International Conference on Robotics and Automation (ICRA), IEEE, 2018, pp. 74–79
2018
-
[47]
I.An,M.Son,D.Manocha,S.-E.Yoon, Reflection- aware sound source localization, in: 2018 IEEE international conference on robotics and automa- tion (ICRA), IEEE, 2018, pp. 66–73
2018
-
[48]
L. Song, H. Wang, P. Chen, Automatic patrol and inspection method for machinery diagnosis robot—sound signal-based fuzzy search approach, IEEE Sensors Journal 20 (2020) 8276–8286
2020
-
[49]
M.Clayton, L.Wang, A.McPherson, A.Cavallaro, An embedded multichannel sound acquisition sys- tem for drone audition, IEEE Sensors Journal 23 (2023) 13377–13386
2023
-
[50]
F.Grondin,F.Michaud, Lightweightandoptimized sound source localization and tracking methods for open and closed microphone array configurations, Robotics and Autonomous Systems 113 (2019) 63–80
2019
-
[51]
Y.-J.Go,J.-S.Choi, Anacousticsourcelocalization methodusingadrone-mountedphasedmicrophone array, Drones 5 (2021) 75
2021
-
[52]
Yamada, K
T. Yamada, K. Itoyama, K. Nishida, K. Nakadai, Placement planning for sound source tracking in active drone audition, Drones 7 (2023) 405
2023
-
[53]
Skoczylas, P
A. Skoczylas, P. Stefaniak, S. Anufriiev, B. Jach- nik, Belt conveyors rollers diagnostics based on acoustic signal collected using autonomous legged inspectionrobot, AppliedSciences11(2021)2299
2021
-
[54]
Z. Shi, L. Zhang, D. Wang, Audio–visual sound source localization and tracking based on mobile robot for the cocktail party problem, Applied Sciences 13 (2023) 6056
2023
-
[55]
Keyrouz, Advanced binaural sound localization in 3-d for humanoid robots, IEEE Transactions on Instrumentation and Measurement 63 (2014) 2098–2107
F. Keyrouz, Advanced binaural sound localization in 3-d for humanoid robots, IEEE Transactions on Instrumentation and Measurement 63 (2014) 2098–2107
2014
-
[56]
Z. Wang, W. Zou, H. Su, Y. Guo, D. Li, Multiple sound source localization exploiting robot motion and approaching control, IEEE Transactions on Instrumentation and Measurement 72 (2023) 1–16
2023
-
[57]
Tourbabin, B
V. Tourbabin, B. Rafaely, Theoretical framework for the optimization of microphone array config- uration for humanoid robot audition, IEEE/ACM 27 Reza Jalayer et al.- Preprint 1–35 Transactions on Audio, Speech, and Language Pro- cessing 22 (2014) 1803–1814
2014
-
[58]
Manamperi, T
W. Manamperi, T. D. Abhayapala, J. Zhang, P. N. Samarasinghe, Droneaudition: Soundsourcelocal- ization using on-board microphones, IEEE/ACM Transactions on Audio, Speech, and Language Pro- cessing 30 (2022) 508–519
2022
-
[59]
Gonzalez-Billandon, G
J. Gonzalez-Billandon, G. Belgiovine, M. Tata, A. Sciutti, G. Sandini, F. Rea, Self-supervised learning framework for speaker localisation with a humanoid robot, in: 2021 IEEE International Conference on Development and Learning (ICDL), IEEE, 2021, pp. 1–7
2021
-
[60]
J. J. Gamboa-Montero, M. Basiri, J. C. Castillo, S. Marques-Villarroya, M. A. Salichs, Real-time acoustic touch localization in human-robot inter- action based on steered response power, in: 2022 IEEE International Conference on Development and Learning (ICDL), IEEE, 2022, pp. 101–106
2022
-
[61]
Yalta, K
N. Yalta, K. Nakadai, T. Ogata, Sound source localization using deep learning models, Journal of Robotics and Mechatronics 29 (2017) 37–48
2017
-
[62]
L. Chen, G. Chen, L. Huang, Y.-S. Choy, W. Sun, Multiplesoundsourcelocalization,separation,and reconstruction by microphone array: A dnn-based approach, Applied Sciences 12 (2022) 3428
2022
-
[63]
Tian, Multiple CRNN for SELD, Parameters 488211 (2020) 490326
C. Tian, Multiple CRNN for SELD, Parameters 488211 (2020) 490326
2020
-
[64]
Bohlender, A
A. Bohlender, A. Spriet, W. Tirry, N. Madhu, Ex- ploiting temporal context in CNN based multi- source DOA estimation, IEEE/ACM Transactions on Audio, Speech, and Language Processing 29 (2021) 1594–1608
2021
-
[65]
X. Li, H. Liu, X. Yang, Sound source localization for mobile robot based on time difference feature and space grid matching, in: 2011 IEEE/RSJ International Conference on Intelligent Robots and Systems, IEEE, 2011, pp. 2879–2886
2011
-
[66]
A. H. El Zooghby, C. G. Christodoulou, M. Geor- giopoulos, A neural network-based smart antenna formultiplesourcetracking, IEEETransactionson Antennas and Propagation 48 (2000) 768–776
2000
-
[67]
A.Ishfaque,B.Kim, Real-timesoundsourcelocal- ization in robots using fly ormia ochracea inspired mems directional microphone, IEEE Sensors Let- ters 7 (2022) 1–4
2022
-
[68]
Athanasopoulos, T
G. Athanasopoulos, T. Dekens, H. Brouckxon, W. Verhelst, The effect of speech denoising algo- rithms on sound source localization for humanoid robots, in: 2012 11th International Conference on Information Science, Signal Processing and their Applications (ISSPA), IEEE, 2012, p...
2012
-
[69]
A. S. Subramanian, C. Weng, S. Watanabe, M. Yu, D. Yu, Deep learning based multi-source localiza- tion with source splitting and its effectiveness in multi-talker speech recognition, Computer Speech & Language 75 (2022) 101360
2022
-
[70]
P. Goli, S. van de Par, Deep learning-based speech specific source localization by using bin- aural and monaural microphone arrays in hearing aids, IEEE/ACM Transactions on Audio, Speech, and Language Processing 31 (2023) 1652–1666
2023
-
[71]
S. Wu, Y. Zheng, K. Ye, H. Cao, X. Zhang, H. Sun, Sound source localization for unmanned aerial vehicles in low signal-to-noise ratio environments, Remote Sensing 16 (2024) 1847
2024
-
[72]
R.Jalayer, M.Jalayer, A.Mor, C.Orsenigo, C.Ver- cellis, Convlstm-based sound source localization in a manufacturing workplace, Computers & In- dustrial Engineering 192 (2024) 110213
2024
-
[73]
Risoud, J.-N
M. Risoud, J.-N. Hanson, F. Gauvrit, C. Renard, P.-E. Lemesre, N.-X. Bonne, C. Vincent, Sound source localization, European annals of otorhino- laryngology, head and neck diseases 135 (2018) 259–264
2018
-
[74]
T.Hirvonen, Classificationofspatialaudiolocation and content using convolutional neural networks, in: Audio Engineering Society Convention 138, Audio Engineering Society, 2015
2015
-
[75]
X. Xiao, S. Zhao, X. Zhong, D. L. Jones, E. S. Chng, H. Li, A learning-based approach to direc- tion of arrival estimation in noisy and reverberant environments, in: 2015 IEEE international con- ference on acoustics, speech and signal processing (ICASSP), IEEE, 2015, pp. 2814–2818
2015
-
[76]
Y. Geng, J. Jung, D. Seol, Sound-source localiza- tion system based on neural network for mobile robots, in: 2008 IEEE International Joint Confer- ence on Neural Networks (IEEE World Congress on Computational Intelligence), IEEE, 2008, pp. 3126–3130
2008
-
[77]
G. Liu, S. Yuan, J. Wu, R. Zhang, A sound source localization method based on microphone array for mobile robot, in: 2018 Chinese Automation Congress (CAC), IEEE, 2018, pp. 1621–1625
2018
-
[78]
S. Y. Lee, J. Chang, S. Lee, Deep learning-based methodformultiplesoundsourcelocalizationwith high resolution and accuracy, Mechanical Systems and Signal Processing 161 (2021) 107959
2021
-
[79]
Gaona, M
O.Acosta,L.Hermida,M.Herrera,C.Montenegro, E. Gaona, M. Bejarano, K. Gordillo, I. Pavón, C. Asensio, Remote binaural system (rbs) for noise acoustic monitoring, Journal of Sensor and Actuator Networks 12 (2023) 63
2023
-
[80]
Deleforge, F
A. Deleforge, F. Forbes, R. Horaud, Acoustic space learning for sound-source separation and localization on binaural manifolds, International journal of neural systems 25 (2015) 1440003
2015
-
[81]
D. Gala, N. Lindsay, L. Sun, Realtime active sound source localization for unmanned ground 28 Reza Jalayer et al.- Preprint 1–35 robots using a self-rotational bi-microphone array, JournalofIntelligent&RoboticSystems95(2019) 935–954
2019
-
[82]
M. D. Baxendale, M. Nibouche, E. L. Secco, A. G. Pipe, M. J. Pearson, Feed-forward selection of cerebellar models for calibration of robot sound source localization, in: Conference on Biomimetic and Biohybrid Systems, Springer, 2019, pp. 3–14
2019
-
[83]
D.Gala, L.Sun, Movingsoundsourcelocalization andtrackingforanautonomousrobotequippedwith a self-rotating bi-microphone array, The Journal of the Acoustical Society of America 154 (2023) 1261–1273
2023
-
[84]
E.Mumolo,M.Nolich,G.Vercelli, Algorithmsfor acoustic localization based on microphone array in service robotics, Robotics and Autonomous systems 42 (2003) 69–88
2003
-
[85]
Q. V. Nguyen, F. Colas, E. Vincent, F. Charpillet, Long-term robot motion planning for active sound source localization with monte carlo tree search, in: 2017 Hands-free Speech Communications and Microphone Arrays (HSCMA), IEEE, 2017, pp. 61–65
2017
-
[86]
Tamai, S
Y. Tamai, S. Kagami, Y. Amemiya, Y. Sasaki, H. Mizoguchi, T. Takano, Circular microphone array for robot’s audition, in: SENSORS, 2004 IEEE, IEEE, 2004, pp. 565–570
2004
-
[87]
C. Choi, D. Kong, J. Kim, S. Bang, Speech en- hancement and recognition using circular micro- phone array for service robots, in: Proceedings 2003 IEEE/RSJ International Conference on Intel- ligent Robots and Systems (IROS 2003)(Cat. No. 03CH37453), volume 4, IEEE, 2003, pp. 3...
2003
-
[88]
Sasaki, M
Y. Sasaki, M. Kabasawa, S. Thompson, S. Kagami, K. Oro, Spherical microphone array for spatial sound localization for a mobile robot, in: 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, IEEE, 2012, pp. 713–718
2012
-
[89]
L. Jin, J. Yan, X. Du, X. Xiao, D. Fu, Rnn for solv- ingtime-variantgeneralizedsylvesterequationwith applications to robots and acoustic source localiza- tion, IEEE Transactions on Industrial Informatics 16 (2020) 6359–6369
2020
-
[90]
Bando, T
Y. Bando, T. Mizumoto, K. Itoyama, K. Nakadai, H. G. Okuno, Posture estimation of hose-shaped robotusingmicrophonearraylocalization, in: 2013 IEEE/RSJ International Conference on Intelligent Robots and Systems, IEEE, 2013, pp. 3446–3451
2013
-
[91]
Kim, Improvement of sound source local- ization for a binaural robot of spherical head with pinnae (2013)
U.-H. Kim, Improvement of sound source local- ization for a binaural robot of spherical head with pinnae (2013)
2013
-
[92]
Kumon, Y
M. Kumon, Y. Noda, Active soft pinnae for robots, in: 2011 IEEE/RSJ International Conference on Intelligent Robots and Systems, IEEE, 2011, pp. 112–117
2011
-
[93]
J. C. Murray, H. R. Erwin, A neural network classifier for notch filter classification of sound- source elevation in a mobile robot, in: The 2011 InternationalJointConferenceonNeuralNetworks, IEEE, 2011, pp. 763–769
2011
-
[94]
Zhang, J
Y. Zhang, J. Weng, Grounded auditory develop- ment by a developmental robot, in: IJCNN’01. InternationalJointConferenceonNeuralNetworks. Proceedings (Cat. No. 01CH37222), volume 2, IEEE, 2001, pp. 1059–1064
2001
-
[95]
B. Xu, G. Sun, R. Yu, Z. Yang, High-accuracy tdoa-based localization without time synchroniza- tion, IEEETransactionsonParallelandDistributed Systems 24 (2012) 1567–1576
2012
-
[96]
Knapp, G
C. Knapp, G. Carter, The generalized correlation method for estimation of time delay, IEEE transac- tions on acoustics, speech, and signal processing 24 (2003) 320–327
2003
-
[97]
Park, K.-D
B.-C. Park, K.-D. Ban, K.-C. Kwak, H.-S. Yoon, Performance analysis of gcc-phat-based sound source localization for intelligent robots, The Jour- nal of Korea Robotics Society 2 (2007) 270–274
2007
-
[98]
J. Wang, X. Qian, Z. Pan, M. Zhang, H. Li, Gcc- phat with speech-oriented attention for robotic sound source localization, in: 2021 IEEE Inter- national Conference on Robotics and Automation (ICRA), IEEE, 2021, pp. 5876–5883
2021
-
[99]
A.Lombard,Y.Zheng,H.Buchner,W.Kellermann, Tdoaestimationformultiplesoundsourcesinnoisy and reverberant environments using broadband in- dependent component analysis, IEEE Transactions on Audio, Speech, and Language Processing 19 (2010) 1490–1503
2010
-
[100]
Huang, L
B. Huang, L. Xie, Z. Yang, Tdoa-based source localization with distance-dependent noises, IEEE Transactions on Wireless Communications 14 (2014) 468–480
2014
-
[101]
Scheuing, B
J. Scheuing, B. Yang, Disambiguation of tdoa estimation for multiple sources in reverberant en- vironments, IEEE transactions on audio, speech, and language processing 16 (2008) 1479–1489
2008
-
[102]
U.-H. Kim, K. Nakadai, H. G. Okuno, Improved sound source localization and front-back disam- biguation for humanoid robots with two ears, in: Recent Trends in Applied Artificial Intelligence: 26th International Conference on Industrial, Engi- neering and Other Applications of ...
2013
-
[103]
U.-H. Kim, K. Nakadai, H. G. Okuno, Improved sound source localization in horizontal plane for binaural robot audition, Applied Intelligence 42 (2015) 63–74. 29 Reza Jalayer et al.- Preprint 1–35
2015
-
[104]
G.Chen,Y.Xu, Asoundsourcelocalizationdevice based on rectangular pyramid structure for mobile robot, Journal of Sensors 2019 (2019) 4639850
2019
-
[105]
H. Chen, C. Liu, Q. Chen, Efficient and robust approaches for three-dimensional sound source recognitionandlocalizationusinghumanoidrobots sensor arrays, International Journal of Advanced Robotic Systems 17 (2020) 1729881420941357
2020
-
[106]
Q. Xu, P. Yang, Sound source localization strategy based on mobile robot, in: Proceedings of 2013 Chinese Intelligent Automation Conference: Intel- ligent Automation & Intelligent Technology and Systems, Springer, 2013, pp. 469–478
2013
-
[107]
Alameda-Pineda, R
X. Alameda-Pineda, R. Horaud, A geometric ap- proach to sound source localization from time- delay estimates, IEEE/ACM Transactions on Au- dio, Speech, and Language Processing 22 (2014) 1082–1095
2014
-
[108]
Valin, F
J.-M. Valin, F. Michaud, J. Rouat, Robust localiza- tion and tracking of simultaneous moving sound sources using beamforming and particle filtering, RoboticsandAutonomousSystems55(2007)216– 228
2007
-
[109]
Luzanto, N
A. Luzanto, N. Bohmer, R. Mahu, E. Alvarado, R. M. Stern, N. Becerra Yoma, Effective acoustic model-based beamforming training for static and dynamic hri applications, Sensors 24 (2024) 6644
2024
-
[110]
3689–3692
S.Kagami,S.Thompson,Y.Sasaki,H.Mizoguchi, T.Enomoto, 2dsoundsourcemappingfrommobile robot using beamforming and particle filtering, in: 2009 IEEE International Conference on Acoustics, Speech and Signal Processing, IEEE, 2009, pp. 3689–3692
2009
-
[111]
Capon, High-resolution frequency-wavenumber spectrum analysis, Proceedings of the IEEE 57 (2005) 1408–1418
J. Capon, High-resolution frequency-wavenumber spectrum analysis, Proceedings of the IEEE 57 (2005) 1408–1418
2005
-
[112]
Dmochowski, J
J. Dmochowski, J. Benesty, S. Affes, Linearly constrained minimum variance source localization and spectral estimation, IEEE transactions on audio, speech, and language processing 16 (2008) 1490–1502
2008
-
[113]
C. Yang, Y. Wang, Y. Wang, D. Hu, H. Guo, An improved functional beamforming algorithm for far-field multi-sound source localization based on hilbert curve, Applied Acoustics 192 (2022) 108729
2022
-
[114]
M. Liu, S. Qu, X. Zhao, Minimum variance distor- tionlessresponse—hanburybrownandtwisssound source localization, Applied Sciences 13 (2023) 6013
2023
-
[115]
Zhang, R
C. Zhang, R. Wang, L. Yu, Y. Xiao, Q. Guo, H. Ji, Localizationofcyclostationaryacousticsourcesvia cyclostationary beamforming and its high spatial resolution implementation, Mechanical Systems and Signal Processing 204 (2023) 110718
2023
-
[116]
M. M. Faraji, S. B. Shouraki, E. Iranmehr, B. Linares-Barranco, Sound source localization in wide-range outdoor environment using distributed sensor network, IEEE Sensors Journal 20 (2019) 2234–2246
2019
-
[117]
J. H. DiBiase, H. F. Silverman, M. S. Brandstein, Robust localization in reverberant rooms, in: Mi- crophone arrays: signal processing techniques and applications, Springer, 2001, pp. 157–180
2001
-
[118]
D. Yook, T. Lee, Y. Cho, Fast sound source lo- calization using two-level search space clustering, IEEE transactions on cybernetics 46 (2015) 20–26
2015
-
[119]
2027–2032
C.T.Ishi,O.Chatot,H.Ishiguro,N.Hagita, Evalu- ationofamusic-basedreal-timesoundlocalization of multiple sound sources in real noisy environ- ments, in: 2009 IEEE/RSJ International Confer- ence on Intelligent Robots and Systems, IEEE, 2009, pp. 2027–2032
2009
-
[120]
Zhang, D
X. Zhang, D. Feng, An efficient music algorithm enhanced by iteratively estimating signal subspace anditsapplicationsinspatialcolorednoise, Remote Sensing 14 (2022) 4260
2022
-
[121]
Wang, Doa estimation of indoor sound sources based on spherical harmonic domain beam-space music, Symmetry 15 (2023) 187
L.Weng,X.Song,Z.Liu,X.Liu,H.Zhou,H.Qiu, M. Wang, Doa estimation of indoor sound sources based on spherical harmonic domain beam-space music, Symmetry 15 (2023) 187
2023
-
[122]
Suzuki, T
R. Suzuki, T. Takahashi, H. G. Okuno, Develop- ment of a robotic pet using sound source localiza- tion with the hark robot audition system, Journal of Robotics and Mechatronics 29 (2017) 146–153
2017
-
[123]
L. Chen, W. Sun, L. Huang, L. Yu, Broadband soundsourcelocalisationvianon-synchronousmea- surements for service robots: A tensor completion approach, IEEE Robotics and Automation Letters 7 (2022) 12193–12200
2022
-
[124]
L.Chen,L.Huang,G.Chen,W.Sun, Alargescale 3dsoundsourcelocalisationapproachachievedvia small size microphone array for service robots, in: 2022 5th International Conference on Information Communication and Signal Processing (ICICSP), IEEE, 2022, pp. 589–594
2022
-
[125]
Hoshiba, K
K. Hoshiba, K. Washizaki, M. Wakabayashi, T. Ishiki, M. Kumon, Y. Bando, D. Gabriel, K. Nakadai, H. G. Okuno, Design of UAV- embedded microphone array system for sound source localization in outdoor environments, Sen- sors 17 (2017) 2535
2017
-
[126]
Azrad, A
S. Azrad, A. Salman, S. A. R. Al-Haddad, Perfor- mance of doa estimation algorithms for acoustic localization of indoor flying drones using artificial soundsource, JournalofAeronautics,Astronautics and Aviation 56 (2024) 469–476
2024
-
[127]
Nakamura, K
K. Nakamura, K. Nakadai, F. Asano, Y. Hasegawa, H. Tsujino, Intelligent sound source localization for dynamic environments, in: 2009 IEEE/RSJ 30 Reza Jalayer et al.- Preprint 1–35 international conference on Intelligent Robots and Systems, IEEE, 2009, pp. 664–669
2009
-
[128]
Narang, K
G. Narang, K. Nakamura, K. Nakadai, Auditory- aware navigation for mobile robots based on reflection-robust sound source localization and vi- sualslam, in: 2014IEEEInternationalConference on Systems, Man, and Cybernetics (SMC), IEEE, 2014, pp. 4021–4026
2014
-
[129]
Asano, M
F. Asano, M. Morisawa, K. Kaneko, K. Yokoi, Sound source localization using a single-point stereomicrophoneforrobots, in: 2015Asia-Pacific SignalandInformationProcessingAssociationAn- nual Summit and Conference (APSIPA), IEEE, 2015, pp. 76–85
2015
-
[130]
H. D. Tran, H. Li, Sound event recognition with probabilistic distance svms, IEEE transactions on audio, speech, and language processing 19 (2010) 1556–1568
2010
-
[131]
Wang, J.-F
J.-C. Wang, J.-F. Wang, K. W. He, C.-S. Hsu, Environmental sound classification using hybrid svm/knn classifier and mpeg-7 audio low-level descriptor, in: The 2006 IEEE international joint conference on neural network proceedings, IEEE, 2006, pp. 1731–1735
2006
-
[132]
Yussif, H
A.-M. Yussif, H. Sadeghi, T. Zayed, Application of machine learning for leak localization in water supply networks, Buildings 13 (2023) 849
2023
-
[133]
H. Chen, W. Ser, Sound source doa estimation and localization in noisy reverberant environments us- ing least-squares support vector machines, Journal of Signal Processing Systems 63 (2011) 287–300
2011
-
[134]
Salvati, C
D. Salvati, C. Drioli, G. L. Foresti, A weighted mvdrbeamformerbasedonsvmlearningforsound source localization, Pattern Recognition Letters 84 (2016) 15–21
2016
-
[135]
Salvati, C
D. Salvati, C. Drioli, G. L. Foresti, On the use of machine learning in microphone array beamform- ingforfar-fieldsoundsourcelocalization, in: 2016 IEEE 26th International Workshop on Machine Learning for Signal Processing (MLSP), IEEE, 2016, pp. 1–6
2016
-
[136]
C. M. Gadre, R. K. Patole, S. P. Metkar, Compar- ative analysis of knn and cnn for localization of single sound source, in: 2023 International Con- ference on Network, Multimedia and Information Technology (NMITCON), IEEE, 2023, pp. 1–6
2023
-
[137]
G.Putrada, M
P.Nando, A. G.Putrada, M. Abdurohman, Increas- ing the precision of noise source detection system using knn method, Kinetik: Game Technology, In- formationSystem,ComputerNetwork,Computing, Electronics, and Control (2019) 157–168
2019
-
[138]
Y.Sun,J.Chen,C.Yuen,S.Rahardja, Indoorsound source localization with probabilistic neural net- work, IEEE Transactions on Industrial Electronics 65 (2017) 6403–6413
2017
-
[139]
LeCun, Y
Y. LeCun, Y. Bengio, G. Hinton, Deep learning, nature 521 (2015) 436–444
2015
-
[140]
T. Fu, Z. Zhang, Y. Liu, J. Leng, Development of an artificial neural network for source localization using a fiber optic acoustic emission sensor array, Structural Health Monitoring 14 (2015) 168–177
2015
-
[141]
C. Jin, M. Schenkel, S. Carlile, Neural system identification model of human sound localization, The Journal of the Acoustical Society of America 108 (2000) 1215–1235
2000
-
[142]
C.-J. Pu, J. G. Harris, J. C. Principe, A neuromor- phic microphone for sound localization, in: 1997 IEEE International Conference on Systems, Man, and Cybernetics. Computational Cybernetics and Simulation, volume 2, IEEE, 1997, pp. 1469–1474
1997
-
[143]
Y. Kim, H. Ling, Direction of arrival estimation of humans with a small sensor array using an artifi- cial neural network, Progress In Electromagnetics Research B 27 (2011) 127–149
2011
-
[144]
Davila-Chacon, J
J. Davila-Chacon, J. Twiefel, J. Liu, S. Wermter, Improvinghumanoidrobotspeechrecognitionwith sound source localisation, in: Artificial Neural Networks and Machine Learning–ICANN 2014: 24th International Conference on Artificial Neural Networks, Hamburg, Germany, September 15-...
2014
-
[145]
Dávila-Chacón, J
J. Dávila-Chacón, J. Liu, S. Wermter, Enhanced robotspeechrecognitionusingbiomimeticbinaural sound source localization, IEEE transactions on neural networks and learning systems 30 (2018) 138–150
2018
-
[146]
Takeda, K
R. Takeda, K. Komatani, Discriminative multiple sound source localization based on deep neural networks using independent location model, in: 2016 IEEE Spoken Language Technology Work- shop (SLT), IEEE, 2016, pp. 603–609
2016
-
[147]
M. J. Bianco, P. Gerstoft, J. Traer, E. Ozanich, M. A. Roch, S. Gannot, C.-A. Deledalle, Machine learninginacoustics: Theoryandapplications, The Journal of the Acoustical Society of America 146 (2019) 3590–3628
2019
-
[148]
A.Krizhevsky,I.Sutskever,G.E.Hinton, Imagenet classification with deep convolutional neural net- works, Advances in neural information processing systems 25 (2012)
2012
-
[149]
J.M.Vera-Diaz,D.Pizarro,J.Macias-Guarasa, To- wards end-to-end acoustic localization using deep learning: From audio signals to source position coordinates, Sensors 18 (2018) 3418
2018
-
[150]
Suvorov, G
D. Suvorov, G. Dong, R. Zhukov, Deep residual network for sound source localization in the time domain, arXiv preprint arXiv:1808.06429 (2018)
2018 arXiv
-
[151]
Huang, R
D. Huang, R. F. Perez, Sseldnet: a fully end- to-end sample-level framework for sound event localization and detection, DCASE (2021). 31 Reza Jalayer et al.- Preprint 1–35
2021
-
[152]
Vincent, T
E. Vincent, T. Virtanen, S. Gannot, Audio source separation and speech enhancement, John Wiley & Sons, 2018
2018
-
[153]
Chakrabarty, E
S. Chakrabarty, E. A. Habets, Multi-speaker DOAestimationusingdeepconvolutionalnetworks trainedwithnoisesignals, IEEEJournalofSelected Topics in Signal Processing 13 (2019) 8–21
2019
-
[154]
Y. Wu, R. Ayyalasomayajula, M. J. Bianco, D.Bharadia,P.Gerstoft, Soundsourcelocalization based on multi-task learning and image translation network, The Journal of the Acoustical Society of America 150 (2021) 3374–3386
2021
-
[155]
S. S. Butt, M. Fatima, A. Asghar, W. Muhammad, Active binaural auditory perceptual system for a socially interactive humanoid robot, Engineering Proceedings 12 (2022) 83
2022
-
[156]
Krause, A
D. Krause, A. Politis, K. Kowalczyk, Comparison ofconvolutiontypesincnn-basedfeatureextraction for sound source localization, in: 2020 28th Eu- ropean Signal Processing Conference (EUSIPCO), IEEE, 2021, pp. 820–824
2020
-
[157]
Diaz-Guerra, A
D. Diaz-Guerra, A. Miguel, J. R. Beltran, Ro- bust sound source tracking using SRP-PHAT and 3d convolutional neural networks, IEEE/ACM Transactions on Audio, Speech, and Language Pro- cessing 29 (2020) 300–311
2020
-
[158]
Bologni, R
G. Bologni, R. Heusdens, J. Martinez, Acoustic reflectors localization from stereo recordings using neural networks, in: ICASSP 2021-2021 IEEE InternationalConferenceonAcoustics,Speechand Signal Processing (ICASSP), IEEE, 2021, pp. 1–5
2021
-
[159]
Nguyen, L
Q. Nguyen, L. Girin, G. Bailly, F. Elisei, D.-C. Nguyen, Autonomous sensorimotor learning for sound source localization by a humanoid robot, in: IROS 2018-Workshop on Crossmodal Learning for Intelligent Robotics in conjunction with IEEE/RSJ IROS, 2018
2018
-
[160]
Boztas, Sound source localization for auditory perception of a humanoid robot using deep neural networks, Neural Computing and Applications 35 (2023) 6801–6811
G. Boztas, Sound source localization for auditory perception of a humanoid robot using deep neural networks, Neural Computing and Applications 35 (2023) 6801–6811
2023
-
[161]
C. Pang, H. Liu, X. Li, Multitask learning of time- frequency cnn for sound source localization, IEEE Access 7 (2019) 40725–40737
2019
-
[162]
J. Ko, H. Kim, J. Kim, Real-time sound source localization for low-power iot devices based on multi-stream cnn, Sensors 22 (2022) 4650
2022
-
[163]
De Groot, S
A.Y.Mjaid,V.Prasad,M.Jonker,C.VanDerHorst, L. De Groot, S. Narayana, Ai-based simultaneous audio localization and communication for robots, in: Proceedings of the 8th ACM/IEEE Conference on Internet of Things Design and Implementation, 2023, pp. 172–183
2023
-
[164]
K. Cho, B. Van Merriënboer, C. Gulcehre, D. Bah- danau, F. Bougares, H. Schwenk, Y. Bengio, Learning phrase representations using rnn encoder- decoder for statistical machine translation, arXiv preprint arXiv:1406.1078 (2014)
2014 arXiv
-
[165]
Hochreiter, J
S. Hochreiter, J. Schmidhuber, Long short-term memory, Neural computation 9 (1997) 1735–1780
1997
-
[166]
T. N. T. Nguyen, N. K. Nguyen, H. Phan, L. Pham, K. Ooi, D. L. Jones, W.-S. Gan, A general network architecture for sound event localization and detec- tion using transfer learning and recurrent neural network, in: ICASSP 2021-2021 IEEE Interna- tional Conference on Acoustics,...
2021
-
[167]
Z.-Q. Wang, X. Zhang, D. Wang, Robust speaker localization guided by deep learning-based time- frequency masking, IEEE/ACM Transactions on Audio,Speech,andLanguageProcessing27(2018) 178–188
2018
-
[168]
M. B. Andra, T. Usagawa, Portable keyword spot- ting and sound source detection system design on mobilerobotwithminimicrophonearray, in: 2020 6thinternationalconferenceoncontrol,automation and robotics (ICCAR), IEEE, 2020, pp. 170–174
2020
-
[169]
Adavanne, A
S. Adavanne, A. Politis, J. Nikunen, T. Virtanen, Sound event localization and detection of overlap- ping sources using convolutional recurrent neural networks, IEEE Journal of Selected Topics in Signal Processing 13 (2018) 34–48
2018
-
[170]
Scenes Events Challenge, Tech
Z.Lu, Soundeventdetectionandlocalizationbased on cnn and lstm, Detection Classification Acoust. Scenes Events Challenge, Tech. Rep (2019)
2019
-
[171]
Perotin, R
L. Perotin, R. Serizel, E. Vincent, A. Guérin, CRNN-based multiple doa estimation using acous- tic intensity features for ambisonics recordings, IEEE Journal of Selected Topics in Signal Process- ing 13 (2019) 22–33
2019
-
[172]
Grumiaux, S
P.-A. Grumiaux, S. Kitić, L. Girin, A. Guérin, Improvedfeatureextractionforcrnn-basedmultiple sound source localization, in: 2021 29th European Signal Processing Conference (EUSIPCO), IEEE, 2021, pp. 231–235
2021
-
[173]
J.-H. Kim, J. Choi, J. Son, G.-S. Kim, J. Park, J.-H. Chang, Mimo noise suppression preserving spatialcuesforsoundsourcelocalizationinmobile robot, in: 2021 IEEE International Symposium on Circuits and Systems (ISCAS), IEEE, 2021, pp. 1–5
2021
-
[174]
C. Han, Y. Luo, N. Mesgarani, Real-time binau- ral speech separation with preserved spatial cues, in: ICASSP 2020-2020 IEEE International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2020, pp. 6404–6408
2020
-
[175]
Ronneberger, P
O. Ronneberger, P. Fischer, T. Brox, U-net: Convo- lutional networks for biomedical image segmenta- tion, in: Medical image computing and computer- assisted intervention–MICCAI 2015: 18th interna- 32 Reza Jalayer et al.- Preprint 1–35 tional conference, Munich, Germany, Octobe...
2015
-
[176]
W. Mack, U. Bharadwaj, S. Chakrabarty, E. A. Habets, Signal-aware broadband doa estimation using attention mechanisms, in: ICASSP 2020- 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2020, pp. 4930–4934
2020
-
[177]
Bazhikov, D
A.Altayeva,N.Omarov,S.Tileubay,A.Zhaksylyk, K. Bazhikov, D. Kambarov, Convolutional lstm network for real-time impulsive sound detection and classification in urban environments., Interna- tional Journal of Advanced Computer Science & Applications 14 (2023)
2023
-
[178]
Akter, M
R. Akter, M. R. Islam, S. K. Debnath, P. K. Sarker, M. K. Uddin, A hybrid cnn-lstm model for envi- ronmental sound classification: Leveraging feature engineering and transfer learning, Digital Signal Processing 163 (2025) 105234
2025
-
[179]
L. S. S. Varnita, K. Subramanyam, M. Ananya, P. Mathilakath, M. Krishnan, S. Tiwari, R. T. Shankarappa, et al., Precision in audio: Cnn+ lstm-based 3d sound event localization and detec- tion in real-world environments, in: 2024 2nd International Conference on Networking and C...
2024
-
[180]
Bahdanau, K
D. Bahdanau, K. Cho, Y. Bengio, Neural machine translationbyjointlylearningtoalignandtranslate, arXiv preprint arXiv:1409.0473 (2014)
2014 arXiv
-
[181]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, I. Polosukhin, Attention is all you need, Advances in neural information processing systems 30 (2017)
2017
-
[182]
H. Phan, L. Pham, P. Koch, N. Q. Duong, I. McLoughlin, A. Mertins, On multitask loss function for audio event detection and localization, arXiv preprint arXiv:2009.05527 (2020)
2020 arXiv
-
[183]
Nakatani, S
C.Schymura,T.Ochiai,M.Delcroix,K.Kinoshita, T. Nakatani, S. Araki, D. Kolossa, Exploiting attention-basedsequence-to-sequencearchitectures for sound event localization, in: 2020 28th Euro- pean Signal Processing Conference (EUSIPCO), IEEE, 2021, pp. 231–235
2020
-
[184]
Emmanuel, N
P. Emmanuel, N. Parrish, M. Horton, Multi-scale network for sound event localization and detection, Tech. report of DCASE Challenge, 2021 (2021)
2021
-
[185]
Yalta, Y
N. Yalta, Y. Sumiyoshi, Y. Kawaguchi, The hitachi dcase 2021 task 3 system: Handling directive in- terference with self attention layers, DCASE Chal- lenge, Music Technol. Group, Universitat Pompeu Fabra, Barcelona, Spain, Tech. Rep 29 (2021)
2021
-
[186]
Zhang, L
G. Zhang, L. Geng, F. Xie, C.-D. He, A dynamic convolution-transformer neural network for multi- ple sound source localization based on functional beamforming, Mechanical Systems and Signal Processing 211 (2024) 111272
2024
-
[187]
Zhang, X
R. Zhang, X. Shen, A novel sound source local- ization method using transformer, in: 2024 9th InternationalConferenceonIntelligentInformatics and Biomedical Sciences (ICIIBMS), volume 9, IEEE, 2024, pp. 361–366
2024
-
[188]
X. Chen, L. Zhao, J. Cui, H. Li, X. Wang, Hybrid convolutional neural network-transformer model for end-to-end binaural sound source localization in reverberant environments, IEEE Access (2025)
2025
-
[189]
Zhang, J
D. Zhang, J. Chen, J. Bai, M. Wang, M. S. Ayub, Q.Yan, D.Shi, W.-S.Gan, Multiplesoundsources localization using sub-band spatial features and attentionmechanism, Circuits,Systems,andSignal Processing 44 (2025) 2592–2620
2025
-
[190]
T. Lin, Y. Wang, X. Liu, X. Qiu, A survey of transformers, AI open 3 (2022) 111–132
2022
-
[191]
H. Mu, W. Xia, W. Che, Improving domain gen- eralization for sound classification with sparse frequency-regularized transformer, in: 2023 IEEE International Conference on Multimedia and Expo (ICME), IEEE, 2023, pp. 1104–1108
2023
-
[192]
Z.Liu,Y.Wang,K.Han,W.Zhang,S.Ma,W.Gao, Post-training quantization for vision transformer, Advances in Neural Information Processing Sys- tems 34 (2021) 28092–28103
2021
-
[193]
J. B. Allen, D. A. Berkley, Image method for efficiently simulating small-room acoustics, The Journal of the Acoustical Society of America 65 (1979) 943–950
1979
-
[194]
Diaz-Guerra, A
D. Diaz-Guerra, A. Miguel, J. R. Beltran, gpurir: A python library for room impulse response sim- ulation with gpu acceleration, Multimedia Tools and Applications 80 (2021) 5653–5671
2021
-
[195]
E. A. Habets, Room impulse response generator, Technische Universiteit Eindhoven, Tech. Rep 2 (2006) 1
2006
-
[196]
Scheibler, E
R. Scheibler, E. Bezzam, I. Dokmanić, Pyrooma- coustics: A python package for audio room simu- lation and array processing algorithms, in: 2018 IEEEinternationalconferenceonacoustics,speech and signal processing (ICASSP), IEEE, 2018, pp. 351–355
2018
-
[197]
shoebox
D. Campbell, K. Palomaki, G. Brown, A matlab simulation of" shoebox" room acoustics for use in researchandteaching, ComputingandInformation Systems 9 (2005) 48
2005
-
[198]
S.M.Schimmel,M.F.Muller,N.Dillier, Afastand accurate “shoebox” room acoustics simulator, in: 2009 IEEE International Conference on Acoustics, Speech and Signal Processing, IEEE, 2009, pp. 241–244
2009
-
[199]
Varanasi, H
V. Varanasi, H. Gupta, R. M. Hegde, A deep learningframeworkforrobustdoaestimationusing 33 Reza Jalayer et al.- Preprint 1–35 spherical harmonic decomposition, IEEE/ACM Transactions on Audio, Speech, and Language Processing 28 (2020) 1248–1259
2020
-
[200]
Gaultier, S
C. Gaultier, S. Kataria, A. Deleforge, Vast: The virtual acoustic space traveler dataset, in: Latent Variable Analysis and Signal Separation: 13th In- ternational Conference, LVA/ICA 2017, Grenoble, France, February 21-23, 2017, Proceedings 13, Springer, 2017, pp. 68–79
2017
-
[201]
R.Cheng,C.Bao,Z.Cui, Mass: Microphonearray speech simulator in room acoustic environment for multi-channel speech coding and enhancement, Applied sciences 10 (2020) 1484
2020
-
[202]
D.Krause,A.Politis,K.Kowalczyk, Datadiversity forimprovingdnn-basedlocalizationofconcurrent sound events, in: 2021 29th european signal pro- cessing conference (EUSIPCO), IEEE, 2021, pp. 236–240
2021
-
[203]
Evers, A
C. Evers, A. H. Moore, P. A. Naylor, Localization of moving microphone arrays from moving sound sourcesforrobotaudition, in: 201624thEuropean Signal Processing Conference (EUSIPCO), IEEE, 2016, pp. 1008–1012
2016
-
[204]
Mesaros, R
A. Mesaros, R. Serizel, T. Heittola, T. Virtanen, M. D. Plumbley, A decade of dcase: Achieve- ments, practices, evaluations and future challenges, in: ICASSP 2025-2025 IEEE International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2025, pp. 1–5
2025
-
[205]
S.Adavanne,A.Politis,T.Virtanen, Amulti-room reverberant dataset for sound event localization and detection, arXiv preprint arXiv:1905.08546 (2019)
2019 arXiv
-
[206]
Politis, A
A. Politis, A. Mesaros, S. Adavanne, T. Heittola, T. Virtanen, Overview and evaluation of sound event localization and detection in DCASE 2019, IEEE/ACM Transactions on Audio, Speech, and Language Processing 29 (2020) 684–698
2020
-
[207]
Politis, S
A. Politis, S. Adavanne, D. Krause, A. Deleforge, P. Srivastava, T. Virtanen, A dataset of dynamic re- verberant sound scenes with directional interferers for sound event localization and detection, arXiv preprint arXiv:2106.06999 (2021)
2021 arXiv
-
[208]
Politis, K
A. Politis, K. Shimada, P. Sudarsanam, S. Ada- vanne, D. Krause, Y. Koyama, N. Takahashi, S. Takahashi, Y. Mitsufuji, T. Virtanen, Starss22: A dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events, arXiv preprint arXiv:2206.01948 (2022)
2022 arXiv
-
[209]
Shimada, A
K. Shimada, A. Politis, P. Sudarsanam, D. A. Krause, K. Uchida, S. Adavanne, A. Hakala, Y. Koyama, N. Takahashi, S. Takahashi, et al., Starss23: An audio-visual dataset of spatial record- ingsofrealsceneswithspatiotemporalannotations of sound events, Advances in neural informa...
2023
-
[210]
D. A. Krause, A. Politis, A. Mesaros, Sound event detection and localization with distance estima- tion, in: 2024 32nd European Signal Processing Conference(EUSIPCO),IEEE,2024, pp.286–290
2024
-
[211]
J. W. Yeow, E.-L. Tan, J. Bai, S. Peksi, W.-S. Gan, Squeeze-and-excite resnet-conformers for sound event localization, detection, and distance estimationfordcase2024challenge, arXivpreprint arXiv:2407.09021 (2024)
2024 arXiv
-
[212]
Shimada, N
K. Shimada, N. Takahashi, S. Takahashi, Y. Mitsu- fuji, Sound event localization and detection using activity-coupled cartesian doa vector and rd3net, arXiv preprint arXiv:2006.12014 (2020)
2020 arXiv
-
[213]
Evers, H
C. Evers, H. W. Löllmann, H. Mellmann, A. Schmidt, H. Barfuss, P. A. Naylor, W. Keller- mann, The locata challenge: Acoustic source localizationandtracking, IEEE/ACMTransactions on Audio, Speech, and Language Processing 28 (2020) 1620–1643
2020
-
[214]
Jekateryńczuk, R
G. Jekateryńczuk, R. Szadkowski, Z. Piotrowski, Uavirbase: A public-access unmanned aerial ve- hicle sound source localization dataset, Applied Sciences 15 (2025) 5378
2025
-
[215]
Zhang, W
J. Zhang, W. Ding, L. He, Data augmentation and prior knowledge-based regularization for sound event localization and detection, DCASE 2019 Detection and Classification of Acoustic Scenes and Events 2019 Challenge (2019)
2019
-
[216]
D.S.Park,W.Chan,Y.Zhang,C.-C.Chiu,B.Zoph, E.D.Cubuk,Q.V.Le,Specaugment: Asimpledata augmentation method for automatic speech recog- nition, arXiv preprint arXiv:1904.08779 (2019)
2019 arXiv
-
[217]
Mazzon, Y
L. Mazzon, Y. Koizumi, M. Yasuda, N. Harada, Firstorderambisonicsdomainspatialaugmentation fordnn-baseddirectionofarrivalestimation, arXiv preprint arXiv:1910.04388 (2019)
2019 arXiv
-
[218]
Pratik, W
P. Pratik, W. J. Jee, S. Nagisetty, R. Mars, C. Lim, Sound event localization and detection using crnn architecture with mixup for model generalization (2019)
2019
-
[219]
K. Noh, C. Jeong-Hwan, J. Dongyeop, C. Joon- Hyuk, Three-stage approach for sound event lo- calization and detection, Detection Classifica- tion Acoust. Scenes Events Challenge, Tech. Rep (2019)
2019
-
[220]
Q. Wang, J. Du, H.-X. Wu, J. Pan, F. Ma, C.-H. Lee, A four-stage data augmentation approach to resnet-conformer based acoustic modeling for soundeventlocalizationanddetection, IEEE/ACM Transactions on Audio, Speech, and Language Processing 31 (2023) 1251–1264
2023
-
[221]
Takahashi, M
N. Takahashi, M. Gygli, B. Pfister, L. Van Gool, Deep convolutional neural networks and data aug- mentation for acoustic event detection, arXiv 34 Reza Jalayer et al.- Preprint 1–35 preprint arXiv:1604.07160 (2016)
2016 arXiv
-
[222]
W. He, P. Motlicek, J.-M. Odobez, Neural network adaptationanddataaugmentationformulti-speaker direction-of-arrival estimation, IEEE/ACM Trans- actions on Audio, Speech, and Language Process- ing 29 (2021) 1303–1317
2021
-
[223]
T.Jenrungrot,V.Jayaram,S.Seitz,I.Kemelmacher- Shlizerman,Theconeofsilence: Speechseparation by localization, Advances in Neural Information Processing Systems 33 (2020) 20925–20938
2020
-
[224]
Y.Hu,P.N.Samarasinghe,S.Gannot,T.D.Abhaya- pala, Semi-supervised multiple source localization using relative harmonic coefficients under noisy andreverberantenvironments, IEEE/ACMTransac- tions on Audio, Speech, and Language Processing 28 (2020) 3108–3123
2020
-
[225]
Takeda, Y
R. Takeda, Y. Kudo, K. Takashima, Y. Kitamura, K. Komatani, Unsupervised adaptation of neural networks for discriminative sound source localiza- tion with eliminative constraint, in: 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, ...
2018
-
[226]
M. J. Bianco, S. Gannot, P. Gerstoft, Semi- supervisedsourcelocalizationwithdeepgenerative modeling, in: 2020 IEEE 30th International Work- shop on Machine Learning for Signal Processing (MLSP), IEEE, 2020, pp. 1–6
2020
-
[227]
Le Moing, P
G. Le Moing, P. Vinayavekhin, D. J. Agra- vante, T. Inoue, J. Vongkulbhisal, A. Munawar, R. Tachibana, Data-efficient framework for real- world multiple sound source 2d localization, in: ICASSP 2021-2021 IEEE International Confer- ence on Acoustics, Speech and Signal Processin...
2021
-
[228]
Z.Ren,S.Wang,Y.Zhang, Weaklysupervisedma- chine learning, CAAI Transactions on Intelligence Technology 8 (2023) 549–580
2023
-
[229]
W. He, P. Motlicek, J.-M. Odobez, Adaptation of multiplesoundsourcelocalizationneuralnetworks with weak supervision and domain-adversarial training, in: ICASSP 2019-2019 IEEE Interna- tional Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2019, pp. 770–774
2019
-
[230]
Opochinsky, B
R. Opochinsky, B. Laufer-Goldshtein, S. Gannot, G. Chechik, Deep ranking-based sound source localization, in: 2019 IEEE Workshop on Applica- tions of Signal Processing to Audio and Acoustics (WASPAA), IEEE, 2019, pp. 283–287
2019
-
[231]
S.Niu,Y.Liu,J.Wang,H.Song,Adecadesurveyof transfer learning (2010–2020), IEEE Transactions on Artificial Intelligence 1 (2021) 151–166
2021
-
[232]
S. Park, Y. Jeong, T. Lee, Many-to-many audio spectrogram tansformer: Transformer for sound event localization and detection., in: DCASE, 2021, pp. 105–109
2021
-
[233]
Lawrence, R
J.F.Gemmeke,D.P.Ellis,D.Freedman,A.Jansen, W. Lawrence, R. C. Moore, M. Plakal, M. Ritter, Audio set: An ontology and human-labeleddataset for audio events, in: 2017 IEEE international con- ference on acoustics, speech and signal processing (ICASSP), IEEE, 2017, pp. 776–780
2017
-
[234]
Zhang, S
M. Zhang, S. Yu, Z. Hu, K. Xia, J. Wang, H. Zhu, Sound source localization with sparse bayesian- based feature matching via deep transfer learning in shallow sea, Measurement (2025) 117873
2025
-
[235]
Y.Yu,K.V.Horoshenkov,G.Sailor,S.Tait, Sparse representation for artefact/defect localization with anacousticarrayonamobilepipeinspectionrobot, Applied Acoustics 231 (2025) 110545
2025
-
[236]
Kothig, M
A. Kothig, M. Ilievski, L. Grasse, F. Rea, M. Tata, A bayesian system for noise-robust binaural sound localisation for humanoid robots, in: 2019 IEEE International Symposium on Robotic and Sensors Environments (ROSE), IEEE, 2019, pp. 1–7
2019
-
[237]
Shiri, J
H. Shiri, J. Wodecki, B. Ziętek, R. Zimroz, In- spection robotic ugv platform and the procedure for an acoustic signal-based fault detection in belt conveyor idler, Energies 14 (2021) 7646
2021
-
[238]
D.Yang,J.Zhao, Acousticwake-uptechnologyfor microsystems: a review, Micromachines 14 (2023) 129. 35
2023
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.