REVIEW 4 major objections 5 minor 69 references
SpeechCompass: Enhancing Mobile Captioning with Diarization and Directional Guidance via Multi-Microphone Localization
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Sound direction lets mobile captions separate speakers in group conversations.
desk verdict Solid prototype with a real gap: diarization only works for turn-taking speech, and the paper never tests overlap. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the time-difference-of-arrival (TDOA) angle estimate: for each of six microphone pairs, GCC-PHAT (generalized cross-correlation with phase transform) finds the delay that maximizes the normalized cross-correlation; with known microphone geometry and the far-field approximation, the delay is converted into an azimuth angle, and a kernel density estimate over the most recent 600 samples selects the most likely source angle. That angle is bound to automatic speech recognition output so each transcript segment inherits a speaker direction, which the interface renders as text color, a directional glyph, or selective suppression of entire directions.
What would settle it
An experiment that feeds the system a two-speaker conversation with a known fraction of overlapping speech (e.g., 20 to 30 percent of speaking time) and computes diarization error rate would show whether direction alone can keep speakers separate when turns overlap; if the error rate rises to near-chance levels under overlap, the central benefit would fail in exactly the group conversations where it is needed.
Extended reading notes
Core claim
The paper's central claim is that the direction of arrival of speech, measured with a small microphone array, is sufficient to diarize a group conversation into separate speakers and to guide the user's attention with directional visual cues. The authors demonstrate this by implementing a complete system: four synchronized microphones on a phone case feed a GCC-PHAT time-delay estimator running on a low-power microcontroller, which fuses six inter-microphone delays through kernel density estimation into a 360-degree azimuth angle with mean error 11 to 22 degrees at conversational loudness, comparable to human localization error. The angle drives a captioning interface that colors transcript lines, shows arrows or moving indicators, and lets the user hide speech from chosen directions. In a lab study with eight deaf or hard-of-hearing frequent users of captioning, all participants rated directional feedback as valuable and said they would recommend it, with colored text and arrows the most preferred visualization styles.
Load-bearing premise
The diarization evaluation only tests turn-taking speech with no overlap, so the central benefit—separating speakers by direction—is unproven for overlapping conversations, where two people speak at once.
Editorial extensions
If this is right
- If direction-based diarization works, mobile captioning apps can separate speakers in turn-taking group conversations without requiring speakers to install apps, wear microphones, or enroll voice samples.
- Users can know where to look when the speaker changes, and can filter out speech from a given direction, directly addressing the two top challenges reported by frequent users: background noise and merged text.
- The localization runs on a low-power microcontroller at under 20 milliseconds per frame, and the whole system draws little enough current to run for about 18 hours on a 500 milliamp-hour battery.
- Diarization error rate improved by roughly 32 percent relative with four microphones versus three, suggesting that modest hardware additions materially improve speaker separation.
- A two-microphone phone can already provide limited 180-degree directional guidance in the app without the extra phone-case hardware.
Reading between the lines
- The same direction signal could be combined with speaker-embedding diarization to handle overlapping speech, where pure angle assignment fails; this is a natural next step the paper leaves open.
- If phones gain more microphones with wider spacing, the approach could become a pure software layer on existing devices, making directional captioning widely available without the phone-case hardware.
- Because the localization is language-agnostic and works for non-speech sounds, the compass interface could be extended to alert users to environmental sounds such as alarms with spatial cues.
- The system's privacy profile—no voice enrollment and no audio leaving the device—could make it attractive beyond accessibility, for example in meeting transcription or language learning at a table.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. SpeechCompass is a mobile captioning system that augments automatic speech recognition with real-time azimuth localization from a custom four-microphone phone case and a low-power microcontroller. The estimated direction of speech is used to diarize transcripts by color-coding or visually separating speakers, to display directional indicators (arrows, edge dot, minimap), and to suppress speech from selected directions. The paper contributes the embedded hardware and GCC-PHAT-based localization pipeline with kernel density estimation, an Android application, technical evaluations of localization accuracy, latency, power consumption, and diarization error, as well as a foundational survey (n=263), an online interface survey (n=494), and an in-person lab study with eight deaf or hard-of-hearing frequent users of captioning technology.
Significance. If the technical claims hold, this is a valuable accessibility contribution: it shows that low-power, privacy-preserving, calibration-free microphone-array processing can add spatial speaker cues to mobile ASR without requiring conversation partners to enroll or wear devices. The paper's strengths include the released code and design files, the concrete low-latency implementation (2.9 ms per microphone pair, ~263 ms onset-to-estimate for speech), localization errors of 11-22 degrees at conversational loudness that are comparable to human azimuth accuracy, and the user evaluation with the target deaf and hard-of-hearing population. The main weakness is that the diarization evaluation, which is central to the group-conversation claim, is limited to turn-taking speech without overlap and still reports diarization error rates of 30-37% for the four-microphone configuration; the benefit for realistic overlapping group conversations is therefore not yet demonstrated.
major comments (4)
- [Sec. 5.6] The diarization evaluation uses only synthesized turn-taking conversations with no overlap in the speech content, as explicitly stated. For the four-microphone configuration, the reported DER is 0.30-0.37 even in this simplest scenario. Because the localization pipeline selects a single KDE peak per frame (Sec. 4.3) and GCC-PHAT picks up the loudest sound (Sec. 5.3), overlapping speech from different directions cannot be represented or diarized. Since the motivating scenario in Sec. 1 is a dinner-table group conversation, where overlap is common, the central claim that localization 'allows diarization' for group conversations is only supported for the least challenging condition. I recommend either evaluating with overlapping speech and reporting per-speaker coverage, or explicitly limiting the claim to turn-taking conversations.
- [Sec. 6.3] The lab study has only eight participants, and while a Kruskal-Wallis test shows a significant overall effect on visualization preferences, the post-hoc pairwise comparisons are not significant after Bonferroni correction. The user-facing claim that 'the value of diarization and visualizing localization was consistent across participants' rests entirely on subjective Likert ratings from n=8, with no objective measure of comprehension, speaker-attribution accuracy, or reading speed. The manuscript should either temper this claim to 'initial preference evidence' or add a task-based outcome measure to support the accessibility benefit.
- [Secs. 5.1, 5.3, 5.6] The technical evaluations use a single stationary source at 1.5 m with the device rotated for localization, and four fixed sources at cardinal angles for diarization. Real group conversations involve moving speakers, multiple simultaneous sources, and reverberation, which are exactly the conditions in the motivating scenario. The paper does not characterize performance under these conditions, and the discussion in Sec. 5.3 notes that the algorithm tracks the loudest sound. The authors should either add experiments with moving or concurrent sources, or explicitly state these as scope limitations in Sec. 7.
- [Secs. 4.3 and 5.6] The KDE bandwidth (25), the KDE buffer size (600 samples), and the diarization running histogram window (522 ms) are fixed values with no reported sensitivity analysis or independent justification. If these parameters were tuned on the evaluation data, the reported DER and localization errors may be optimistic. A sensitivity sweep over these parameters, or a statement that they were chosen a priori, is needed to support the robustness of the reported numbers.
minor comments (5)
- [Sec. 5.6] The text says 'as can be seen in Table 12', but the DER results are presented in a figure, not a table; this should be 'Figure 12'.
- [Sec. 5.5] The text reports 'power consumption of the whole system was 28 mAh'; mAh is a charge unit, so this should likely be '28 mA' or 'mW'.
- [References] References [31] and [32] are duplicate entries for the same paper by Jain et al. (2015) on head-mounted display visualizations; one should be removed and the in-text citations updated.
- [Sec. 6.2] The paragraph describing the lab study mentions the 'SoundCompass UI', but the system is consistently named SpeechCompass elsewhere; this should be corrected.
- [Sec. 8] The conclusion states 'All the participants found the diarization, localization, and visualization features to be useful'; given n=8 and the lack of significant pairwise results, the wording 'all eight participants in our study' would be more precise than the unqualified 'all participants'.
Circularity Check
No significant circularity: localization is measured against external ground truth, and diarization is evaluated with independent speaker-turn labels.
full rationale
The derivation chain in SpeechCompass is self-contained. The localization pipeline (GCC-PHAT, TDOA, KDE) is characterized against servo-controlled ground-truth azimuth angles in Section 5.1–5.3, not against its own outputs; the KDE bandwidth and PHAT normalization exponent are fixed algorithm parameters, not fitted to the evaluation data. Diarization accuracy in Section 5.6 is scored with PyAnnotate against ground-truth speaker turns in synthesized turn-taking conversations, so the labels are independent of the algorithm's outputs; the no-overlap condition limits ecological validity but is not circularity. The user studies measure self-reported preferences after using the system, and the n=263 and n=494 surveys motivate or inform the design rather than serve as tests of the implemented claims. No load-bearing conclusion rests on a self-citation: reference [50] (which includes an author of this paper) is cited only for prior captioning and head-worn display work, not to justify SpeechCompass's central capabilities. The noted limitations—small sample size, quiet lab setting, and no-overlap diarization test—are external-validity and generalizability concerns, not reductions of the results to their own inputs.
Assumptions & free parameters
free parameters (4)
- KDE Gaussian kernel bandwidth =
25
- GCC-PHAT partial normalization exponent =
-0.3
- KDE sample buffer size =
600 latest samples
- Diarization running histogram window =
522 ms
assumptions (3)
- domain assumption Far-field plane-wave approximation for sound propagation
- domain assumption One active speaker per time frame (no overlapping speech)
- domain assumption GCC-PHAT tracks the loudest sound source
Cite this review
Pith. "Pith review of SpeechCompass: Enhancing Mobile Captioning with Diarization and Directional Guidance via Multi-Microphone Localization." pith.science (2026). https://pith.science/paper/FVQ4XRVF
@misc{pith2026250208848,
author = {Pith},
title = {Pith review of: SpeechCompass: Enhancing Mobile Captioning with Diarization and Directional Guidance via Multi-Microphone Localization},
year = {2026},
howpublished = {\url{https://pith.science/paper/FVQ4XRVF}},
note = {Machine review of arXiv:2502.08848}
}
read the original abstract
Speech-to-text capabilities on mobile devices have proven helpful for hearing and speech accessibility, language translation, note-taking, and meeting transcripts. However, our foundational large-scale survey (n=263) shows that the inability to distinguish and indicate speaker direction makes them challenging in group conversations. SpeechCompass addresses this limitation through real-time, multi-microphone speech localization, where the direction of speech allows visual separation and guidance (e.g., arrows) in the user interface. We introduce efficient real-time audio localization algorithms and custom sound perception hardware running on a low-power microcontroller and four integrated microphones, which we characterize in technical evaluations. Informed by a large-scale survey (n=494), we conducted an in-person study of group conversations with eight frequent users of mobile speech-to-text, who provided feedback on five visualization styles. The value of diarization and visualizing localization was consistent across participants, with everyone agreeing on the value and potential of directional guidance for group conversations.
Figures
Figures from the paper (13 more)
Reference graph
Works this paper leans on
-
[1]
Chandykunju Alex, Kevin Martin Jose, and Arun Joseph. 2013. Sound localization and Visualization device. In2013 IEEE Global Humanitarian Technology Conference (GHTC). 262–264. doi:10.1109/GHTC.2013.6713692
arXiv 2013
-
[2]
Android. 2022. Introducing Live Transcribe. https://www.android.com/ accessibility/live-transcribe/. Accessed 2022-03-26
work page 2022
-
[3]
Android. 2022. SpeechRecognizer API Documentation). https://developer.android. com/reference/android/speech/SpeechRecognizer. Accessed 2022-10-25
work page 2022
-
[4]
Xavier Anguera, Chuck Wooters, and Javier Hernando. 2007. Acoustic beam- forming for speaker diarization of meetings. IEEE Transactions on Audio, Speech, and Language Processing 15, 7 (2007), 2011–2022
work page 2007
-
[5]
ARM. 2022. CMSIS DSP Software Library. https://www.keil.com/pack/doc/ CMSIS/DSP/html/index.html. Accessed 2022-05-12
work page 2022
-
[6]
Ava. 2022. Ava Captioning Solution. https://www.ava.me/. Accessed 2024-12-10
work page 2022
-
[7]
Jacob Benesty, Jingdong Chen, and Yiteng Huang. 2008. Conventional beamform- ing techniques. Microphone array signal processing (2008), 39–65
work page 2008
-
[8]
Larwan Berke, Khaled Albusays, Matthew Seita, and Matt Huenerfauth. 2019. Preferred Appearance of Captions Generated by Automatic Speech Recognition for Deaf and Hard-of-Hearing Viewers. In Extended Abstracts of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland Uk) (CHI EA ’19). Association for Computing Machinery, New York,...
arXiv 2019
Show all 69 references
-
[9]
Larwan Berke, Christopher Caulfield, and Matt Huenerfauth. 2017. Deaf and Hard-of-Hearing Perspectives on Imperfect Automatic Speech Recognition for Captioning One-on-One Meetings. In Proceedings of the 19th International ACM SIGACCESS Conference on Computers and Accessibility...
2017
-
[10]
Rachel Boll, Shruti Mahajan, Jeanne Reis, and Erin T. Solovey. 2020. Creating Questionnaires That Align with ASL Linguistic Principles and Cultural Practices within the Deaf Community. , Article 61 (2020), 4 pages. doi:10.1145/3373625. 3418071
2020 doi
-
[11]
Danielle Bragg, Nicholas Huynh, and Richard E. Ladner. 2016. A Personalizable Mobile Sound Detector App Design for Deaf and Hard-of-Hearing Users. In Proceedings of the 18th International ACM SIGACCESS Conference on Computers and Accessibility (Reno, Nevada, USA) (ASSETS ’16)....
2016
-
[12]
Hervé Bredin, Ruiqing Yin, Juan Manuel Coria, Gregory Gelly, Pavel Korshunov, Marvin Lavechin, Diego Fustes, Hadrien Titeux, Wassim Bouaziz, and Marie- Philippe Gill. 2020. pyannote.audio: neural building blocks for speaker diarization. In ICASSP 2020, IEEE International Confe...
2020
-
[13]
Janine Butler, Brian Trager, and Byron Behm. 2019. Exploration of Automatic Speech Recognition for Deaf and Hard of Hearing Students in Higher Education Classes. In Proceedings of the 21st International ACM SIGACCESS Conference on Computers and Accessibility (Pittsburgh, PA, U...
2019
-
[14]
Chao Cai, Henglin Pu, Peng Wang, Zhe Chen, and Jun Luo. 2021. We Hear Your PACE: Passive Acoustic Localization of Multiple Walking Persons. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 5, 2, Article 55 (jun 2021), 24 pages. doi:10.1145/3463510
2021 doi
-
[15]
Caluã de Lacerda Pataca, Saad Hassan, Nathan Tinker, Roshan Lalintha Peiris, and Matt Huenerfauth. 2024. Caption Royale: Exploring the Design Space of Affective Captions from the Perspective of Deaf and Hard-of-Hearing Individuals. In Proceedings of the 2024 CHI Conference on ...
2024
-
[16]
Artem Dementyev, Pascal Getreuer, Dimitri Kanevsky, Malcolm Slaney, and Richard F Lyon. 2021. VHP: Vibrotactile Haptics Platform for On-Body Ap- plications. In The 34th Annual ACM Symposium on User Interface Software and Technology (Virtual Event, USA) (UIST ’21). Association ...
2021
-
[17]
Hoang Do, Harvey F Silverman, and Ying Yu. 2007. A real-time SRP-PHAT source location implementation using stochastic region contraction (SRC) on a large- aperture microphone array. In 2007 IEEE International Conference on Acoustics, Speech and Signal Processing-ICASSP’07 , Vo...
2007
-
[18]
Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T Freeman, and Michael Rubinstein. 2018. Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation. arXiv preprint arXiv:1804.03619 (2018)
2018 arXiv
-
[19]
Hakan Erdogan, John R Hershey, Shinji Watanabe, and Jonathan Le Roux. 2015. Phase-sensitive and recognition-boosted speech separation using deep recurrent neural networks. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 708–712
2015
-
[20]
Leah Findlater, Bonnie Chinh, Dhruv Jain, Jon Froehlich, Raja Kushalnagar, and Angela Carey Lin. 2019. Deaf and Hard-of-Hearing Individuals’ Preferences for Wearable and Mobile Sound Awareness Technologies. InProceedings of the 2019 CHI Conference on Human Factors in Computing...
2019
-
[21]
Israel D Gebru, Sileye Ba, Xiaofei Li, and Radu Horaud. 2017. Audio-visual speaker diarization based on spatiotemporal bayesian fusion. IEEE transactions on pattern analysis and machine intelligence 40, 5 (2017), 1086–1099
2017
-
[22]
Abraham Glasser, Kesavan Kushalnagar, and Raja Kushalnagar. 2017. Deaf, Hard of Hearing, and Hearing Perspectives on Using Automatic Speech Recog- nition in Conversation. In Proceedings of the 19th International ACM SIGAC- CESS Conference on Computers and Accessibility (Baltim...
2017
-
[23]
Goodman, Ping Liu, Dhruv Jain, Emma J
Steven M. Goodman, Ping Liu, Dhruv Jain, Emma J. McDonnell, Jon E. Froehlich, and Leah Findlater. 2021. Toward User-Driven Sound Recognizer Personalization with People Who Are d/Deaf or Hard of Hearing. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 5, 2, Article 63 (ju...
2021
-
[24]
Google. 2018. Google Surveys Methodology. http://services.google.com/fh/files/ misc/white_paper_how_google_surveys_works.pdf. Accessed 2022-03-22
2018
-
[25]
Beth G Greene, David B Pisoni, and Thomas D Carrell. 1984. Recognition of speech spectrograms. The Journal of the Acoustical Society of America 76, 1 (1984), 32–43
1984
-
[26]
François Grondin and François Michaud. 2019. Lightweight and optimized sound source localization and tracking methods for open and closed microphone array configurations. Robotics and Autonomous Systems 113 (2019), 63–80
2019
-
[27]
Ru Guo, Yiru Yang, Johnson Kuang, Xue Bin, Dhruv Jain, Steven Goodman, Leah Findlater, and Jon Froehlich. 2020. HoloSound: Combining Speech and Sound Identification for Deaf or Hard of Hearing Users on a Head-Mounted Display. In Proceedings of the 22nd International ACM SIGACC...
2020
-
[28]
Jahn Heymann, Lukas Drude, and Reinhold Haeb-Umbach. 2016. Neural network based spectral mask estimation for acoustic beamforming. In 2016 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 196–200
2016
-
[29]
Pulsar Instruments. 2022. Decibel chart – decibel levels of common sounds. https://pulsarinstruments.com/news/decibel-chart-noise-level. Accessed 2022- 07-27
2022
-
[30]
Yusuf Isik, Jonathan Le Roux, Zhuo Chen, Shinji Watanabe, and John R Hershey
-
[32]
Froehlich
Dhruv Jain, Leah Findlater, Jamie Gilkeson, Benjamin Holland, Ramani Du- raiswami, Dmitry Zotkin, Christian Vogler, and Jon E. Froehlich. 2015. Head- Mounted Display Visualizations to Support Sound Awareness for the Deaf and Hard of Hearing. In Proceedings of the 33rd Annual A...
2015 doi
-
[33]
Froehlich
Dhruv Jain, Kelly Mack, Akli Amrous, Matt Wright, Steven Goodman, Leah Findlater, and Jon E. Froehlich. 2020. HomeSound: An Iterative Field Deployment of an In-Home Sound Awareness System for Deaf or Hard of Hearing Users. In Proceedings of the 2020 CHI Conference on Human Fac...
2020
-
[34]
Vahid Ahmadi Kalkhorani, Anurag Kumar, Ke Tan, Buye Xu, and DeLiang Wang
-
[35]
Yoshihiro Kaneko, Inho Chung, and Kenji Suzuki. 2013. Light-emitting device for supporting auditory awareness of hearing-impaired people during group conver- sations. In 2013 IEEE International Conference on Systems, Man, and Cybernetics . IEEE, 3567–3572
2013
-
[36]
Ellington Kirby, Seoyoon Park, Yan Wang, and Yingying Chen. 2016. HearHere: Smartphone Based Audio Localization Using Time Difference of Arrival: Demo. In Proceedings of the 22nd Annual International Conference on Mobile Computing and Networking (New York City, New York)(MobiC...
2016
-
[37]
Charles Knapp and Glifford Carter. 1976. The generalized correlation method for estimation of time delay. IEEE transactions on acoustics, speech, and signal processing 24, 4 (1976), 320–327
1976
-
[38]
Kushalnagar, Gary W
Raja S. Kushalnagar, Gary W. Behm, Aaron W. Kelstone, and Shareef Ali
-
[39]
Ahmet Köse, Aleksei Tepljakov, and Sergei Astapov. 2017. Real-time localization and visualization of a sound source for virtual reality applications. In 2017 25th International Conference on Software, Telecommunications and Computer Networks (SoftCOM). 1–6. doi:10.23919/SOFTCO...
2017
-
[40]
Bowon Lee, Amir Said, Ton Kalker, and Ronald W Schafer. 2008. Maximum likelihood time delay estimation with phase domain analysis in the generalized cross correlation framework. In 2008 Hands-Free Speech Communication and Microphone Arrays. IEEE, 89–92
2008
-
[41]
Bo Li, Tara N Sainath, Ron J Weiss, Kevin W Wilson, and Michiel Bacchiani
-
[42]
LibriVox. 2022. Alice’s Adventures in Wonderland by Lewis Carroll (Version 2). https://librivox.org/alices-adventures-in-wonderland-by-lewis-carroll-4/. Accessed 2022-07-12
2022
-
[43]
Hong Liu and Miao Shen. 2010. Continuous sound source localization based on microphone array for mobile robots. In 2010 IEEE/RSJ International Conference on Intelligent Robots and Systems . IEEE, 4332–4339
2010
-
[44]
Richard F Lyon. 2017. Human and machine hearing: extracting meaning from sound. Cambridge University Press
2017
-
[45]
Neural network adaptive beamforming for robust multichannel speech recognition. (2016)
2016
-
[46]
Microsoft. 2022. Translator. https://translator.microsoft.com/. Accessed 2022-03- 26
2022
-
[47]
MiniDSP. 2022. USB Mic array. https://www.minidsp.com/products/usb-audio- interface/uma-8-16-usb-mic-array. Accessed 2022-05-15
2022
-
[48]
Kai Morich. 2022. usb-serial-for-android). https://github.com/mik3y/usb-serial- for-android. Accessed 2022-11-09
2022
-
[49]
James C Makous and John C Middlebrooks. 1990. Two-dimensional sound local- ization by human listeners. The journal of the Acoustical Society of America 87, 5 (1990), 2188–2200
1990
-
[50]
Alex Olwal, Kevin Balke, Dmitrii Votintcev, Thad Starner, Paula Conn, Bonnie Chinh, and Benoit Corda. 2020. Wearable Subtitles: Augmenting Spoken Com- munication with Lightweight Eyewear for All-Day Captioning . Association for Computing Machinery, New York, NY, USA, 1108–1120...
2020
-
[51]
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015. Librispeech: an asr corpus based on public domain audio books. In 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 5206–5210
2015
-
[52]
Tae Jin Park, Naoyuki Kanda, Dimitrios Dimitriadis, Kyu J Han, Shinji Watanabe, and Shrikanth Narayanan. 2022. A review of speaker diarization: Recent advances with deep learning. Computer Speech & Language 72 (2022), 101317
2022
-
[53]
Pius Kavuma Basajjabaka Mugagga and Simon Winberg. 2015. Sound source localisation on Android smartphones: A first step to using smartphones as au- ditory sensors for training A.I systems with Big Data. In AFRICON 2015. 1–5. doi:10.1109/AFRCON.2015.7331970
2015
-
[54]
Ashutosh Saxena and Andrew Y Ng. 2009. Learning sound location from a single microphone. In 2009 IEEE International Conference on Robotics and Automation . IEEE, 1737–1742
2009
-
[55]
Matthew Seita. 2020. Designing Automatic Speech Recognition Technologies to Improve Accessibility for Deaf and Hard-of-Hearing People in Small Group Meetings. In Extended Abstracts of the 2020 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI EA ’20...
2020
-
[56]
Deep Sleep. 2022. Rain Sound. https://www.youtube.com/watch?v= 13EL6Mgeocc&t=3448s. Accessed 2022-07-12
2022
-
[57]
Priyantha, Anit Chakraborty, and Hari Balakrishnan
Nissanka B. Priyantha, Anit Chakraborty, and Hari Balakrishnan. 2000. The Cricket Location-Support System. In Proceedings of the 6th Annual International Conference on Mobile Computing and Networking (Boston, Massachusetts, USA) (MobiCom ’00). Association for Computing Machine...
2000
-
[58]
Speaksee. 2022. Speaksee microphone kit. https://speak-see.com/. Accessed 2022-10-25
2022
-
[59]
Hassan Taherian, Ashutosh Pandey, Daniel Wong, Buye Xu, and DeLiang Wang
-
[60]
Hassan Taherian and DeLiang Wang. 2024. Multi-Channel Conversational Speaker Separation via Neural Diarization. IEEE/ACM Transactions on Audio, Speech, and Language Processing 32 (2024), 2467–2476. doi:10.1109/TASLP.2024. 3393726
2024 doi
-
[61]
David Snyder, Daniel Garcia-Romero, Gregory Sell, Alan McCree, Daniel Povey, and Sanjeev Khudanpur. 2019. Speaker recognition for multi-speaker conver- sations using x-vectors. In ICASSP 2019-2019 IEEE International conference on acoustics, speech and signal processing (ICASSP...
2019
-
[62]
Quan Wang, Carlton Downey, Li Wan, Philip Andrew Mansfield, and Ignacio Lopz Moreno. 2018. Speaker diarization with LSTM. In 2018 IEEE International confer- ence on acoustics, speech and signal processing (ICASSP) . IEEE, 5239–5243
2018
-
[63]
Lucinda Wilder. 1975. Articulatory and acoustic characteristics of speech sounds. In Understanding Language. Elsevier, 31–76
1975
-
[64]
In ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
Leveraging Sound Localization to Improve Continuous Speaker Separation. In ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 621–625. doi:10.1109/ICASSP48485.2024.10446934
2024
-
[65]
Yang Yang, George Sung, Shao-Fu Shih, Hakan Erdogan, Chehung Lee, and Matthias Grundmann. 2024. Binaural Angular Separation Network. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1201–1205
2024
-
[66]
Giuseppe Valenzise, Luigi Gerosa, Marco Tagliasacchi, Fabio Antonacci, and Augusto Sarti. 2007. Scream and gunshot detection and localization for audio- surveillance systems. In 2007 IEEE Conference on Advanced Video and Signal Based Surveillance. IEEE, 21–26
2007
-
[69]
Yuteng Xiao, Jihang Yin, Honggang Qi, Hongsheng Yin, and Gang Hua. 2017. MVDR algorithm based on estimated diagonal loading for beamforming. Mathe- matical Problems in Engineering 2017 (2017)
2017
-
[71]
Jiahui Yu, Chung-Cheng Chiu, Bo Li, Shuo-yiin Chang, Tara N Sainath, Yanzhang He, Arun Narayanan, Wei Han, Anmol Gulati, Yonghui Wu, et al. 2021. Fastemit: Low-latency streaming asr with sequence-level emission regularization. In ICASSP 2021-2021 IEEE International Conference ...
2021
-
[2015]
InProceedings of the 17th International ACM SIGACCESS Conference on Computers & Accessibility (Lisbon, Portugal) (AS- SETS ’15)
Tracked Speech-To-Text Display: Enhancing Accessibility and Readabil- ity of Real-Time Speech-To-Text. InProceedings of the 17th International ACM SIGACCESS Conference on Computers & Accessibility (Lisbon, Portugal) (AS- SETS ’15). Association for Computing Machinery, New York...
-
[2016]
arXiv preprint arXiv:1607.02173 (2016)
Single-channel multi-speaker separation using deep clustering. arXiv preprint arXiv:1607.02173 (2016)
2016 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.