Pith. sign in

REVIEW 4 major objections 5 minor 69 references

SpeechCompass: Enhancing Mobile Captioning with Diarization and Directional Guidance via Multi-Microphone Localization

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Sound direction lets mobile captions separate speakers in group conversations.

desk verdict Solid prototype with a real gap: diarization only works for turn-taking speech, and the paper never tests overlap. read the letter →

arxiv 2502.08848 v2 pith:FVQ4XRVF submitted 2025-02-12 cs.HC cs.SD

classification cs.HCcs.SD
keywords assistivetechnologyhearingaccessibilitysoundlocalizationspeakerdiarizationmicrophonearraymobilecaptioningreal-timespeechrecognition
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that adding multi-microphone localization to mobile speech-to-text can fix a problem single-microphone captioning cannot: in group conversations, the transcript loses who said what and from where. The authors build SpeechCompass, a phone case with four microphones and a low-power microcontroller that estimates the 360-degree direction of speech in real time, and a captioning app that uses that direction to color text, place arrows, and let users suppress unwanted directions. They report that a survey of 263 frequent captioning users identifies speaker separation as a top challenge, and that eight frequent users in a lab study all agreed that directional guidance is valuable. If correct, mobile captioning becomes useful in group conversations without requiring conversation partners to wear or enroll devices, at lower computational cost and with less privacy risk than speaker-embedding approaches.

What carries the argument

The load-bearing mechanism is the time-difference-of-arrival (TDOA) angle estimate: for each of six microphone pairs, GCC-PHAT (generalized cross-correlation with phase transform) finds the delay that maximizes the normalized cross-correlation; with known microphone geometry and the far-field approximation, the delay is converted into an azimuth angle, and a kernel density estimate over the most recent 600 samples selects the most likely source angle. That angle is bound to automatic speech recognition output so each transcript segment inherits a speaker direction, which the interface renders as text color, a directional glyph, or selective suppression of entire directions.

What would settle it

An experiment that feeds the system a two-speaker conversation with a known fraction of overlapping speech (e.g., 20 to 30 percent of speaking time) and computes diarization error rate would show whether direction alone can keep speakers separate when turns overlap; if the error rate rises to near-chance levels under overlap, the central benefit would fail in exactly the group conversations where it is needed.

Watch

Extended reading notes

Core claim

The paper's central claim is that the direction of arrival of speech, measured with a small microphone array, is sufficient to diarize a group conversation into separate speakers and to guide the user's attention with directional visual cues. The authors demonstrate this by implementing a complete system: four synchronized microphones on a phone case feed a GCC-PHAT time-delay estimator running on a low-power microcontroller, which fuses six inter-microphone delays through kernel density estimation into a 360-degree azimuth angle with mean error 11 to 22 degrees at conversational loudness, comparable to human localization error. The angle drives a captioning interface that colors transcript lines, shows arrows or moving indicators, and lets the user hide speech from chosen directions. In a lab study with eight deaf or hard-of-hearing frequent users of captioning, all participants rated directional feedback as valuable and said they would recommend it, with colored text and arrows the most preferred visualization styles.

Load-bearing premise

The diarization evaluation only tests turn-taking speech with no overlap, so the central benefit—separating speakers by direction—is unproven for overlapping conversations, where two people speak at once.

Editorial extensions

If this is right

  • If direction-based diarization works, mobile captioning apps can separate speakers in turn-taking group conversations without requiring speakers to install apps, wear microphones, or enroll voice samples.
  • Users can know where to look when the speaker changes, and can filter out speech from a given direction, directly addressing the two top challenges reported by frequent users: background noise and merged text.
  • The localization runs on a low-power microcontroller at under 20 milliseconds per frame, and the whole system draws little enough current to run for about 18 hours on a 500 milliamp-hour battery.
  • Diarization error rate improved by roughly 32 percent relative with four microphones versus three, suggesting that modest hardware additions materially improve speaker separation.
  • A two-microphone phone can already provide limited 180-degree directional guidance in the app without the extra phone-case hardware.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same direction signal could be combined with speaker-embedding diarization to handle overlapping speech, where pure angle assignment fails; this is a natural next step the paper leaves open.
  • If phones gain more microphones with wider spacing, the approach could become a pure software layer on existing devices, making directional captioning widely available without the phone-case hardware.
  • Because the localization is language-agnostic and works for non-speech sounds, the compass interface could be extended to alert users to environmental sounds such as alarms with spatial cues.
  • The system's privacy profile—no voice enrollment and no audio leaving the device—could make it attractive beyond accessibility, for example in meeting transcription or language learning at a table.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. SpeechCompass is a mobile captioning system that augments automatic speech recognition with real-time azimuth localization from a custom four-microphone phone case and a low-power microcontroller. The estimated direction of speech is used to diarize transcripts by color-coding or visually separating speakers, to display directional indicators (arrows, edge dot, minimap), and to suppress speech from selected directions. The paper contributes the embedded hardware and GCC-PHAT-based localization pipeline with kernel density estimation, an Android application, technical evaluations of localization accuracy, latency, power consumption, and diarization error, as well as a foundational survey (n=263), an online interface survey (n=494), and an in-person lab study with eight deaf or hard-of-hearing frequent users of captioning technology.

Significance. If the technical claims hold, this is a valuable accessibility contribution: it shows that low-power, privacy-preserving, calibration-free microphone-array processing can add spatial speaker cues to mobile ASR without requiring conversation partners to enroll or wear devices. The paper's strengths include the released code and design files, the concrete low-latency implementation (2.9 ms per microphone pair, ~263 ms onset-to-estimate for speech), localization errors of 11-22 degrees at conversational loudness that are comparable to human azimuth accuracy, and the user evaluation with the target deaf and hard-of-hearing population. The main weakness is that the diarization evaluation, which is central to the group-conversation claim, is limited to turn-taking speech without overlap and still reports diarization error rates of 30-37% for the four-microphone configuration; the benefit for realistic overlapping group conversations is therefore not yet demonstrated.

major comments (4)
  1. [Sec. 5.6] The diarization evaluation uses only synthesized turn-taking conversations with no overlap in the speech content, as explicitly stated. For the four-microphone configuration, the reported DER is 0.30-0.37 even in this simplest scenario. Because the localization pipeline selects a single KDE peak per frame (Sec. 4.3) and GCC-PHAT picks up the loudest sound (Sec. 5.3), overlapping speech from different directions cannot be represented or diarized. Since the motivating scenario in Sec. 1 is a dinner-table group conversation, where overlap is common, the central claim that localization 'allows diarization' for group conversations is only supported for the least challenging condition. I recommend either evaluating with overlapping speech and reporting per-speaker coverage, or explicitly limiting the claim to turn-taking conversations.
  2. [Sec. 6.3] The lab study has only eight participants, and while a Kruskal-Wallis test shows a significant overall effect on visualization preferences, the post-hoc pairwise comparisons are not significant after Bonferroni correction. The user-facing claim that 'the value of diarization and visualizing localization was consistent across participants' rests entirely on subjective Likert ratings from n=8, with no objective measure of comprehension, speaker-attribution accuracy, or reading speed. The manuscript should either temper this claim to 'initial preference evidence' or add a task-based outcome measure to support the accessibility benefit.
  3. [Secs. 5.1, 5.3, 5.6] The technical evaluations use a single stationary source at 1.5 m with the device rotated for localization, and four fixed sources at cardinal angles for diarization. Real group conversations involve moving speakers, multiple simultaneous sources, and reverberation, which are exactly the conditions in the motivating scenario. The paper does not characterize performance under these conditions, and the discussion in Sec. 5.3 notes that the algorithm tracks the loudest sound. The authors should either add experiments with moving or concurrent sources, or explicitly state these as scope limitations in Sec. 7.
  4. [Secs. 4.3 and 5.6] The KDE bandwidth (25), the KDE buffer size (600 samples), and the diarization running histogram window (522 ms) are fixed values with no reported sensitivity analysis or independent justification. If these parameters were tuned on the evaluation data, the reported DER and localization errors may be optimistic. A sensitivity sweep over these parameters, or a statement that they were chosen a priori, is needed to support the robustness of the reported numbers.
minor comments (5)
  1. [Sec. 5.6] The text says 'as can be seen in Table 12', but the DER results are presented in a figure, not a table; this should be 'Figure 12'.
  2. [Sec. 5.5] The text reports 'power consumption of the whole system was 28 mAh'; mAh is a charge unit, so this should likely be '28 mA' or 'mW'.
  3. [References] References [31] and [32] are duplicate entries for the same paper by Jain et al. (2015) on head-mounted display visualizations; one should be removed and the in-text citations updated.
  4. [Sec. 6.2] The paragraph describing the lab study mentions the 'SoundCompass UI', but the system is consistently named SpeechCompass elsewhere; this should be corrected.
  5. [Sec. 8] The conclusion states 'All the participants found the diarization, localization, and visualization features to be useful'; given n=8 and the lack of significant pairwise results, the wording 'all eight participants in our study' would be more precise than the unqualified 'all participants'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: localization is measured against external ground truth, and diarization is evaluated with independent speaker-turn labels.

full rationale

The derivation chain in SpeechCompass is self-contained. The localization pipeline (GCC-PHAT, TDOA, KDE) is characterized against servo-controlled ground-truth azimuth angles in Section 5.1–5.3, not against its own outputs; the KDE bandwidth and PHAT normalization exponent are fixed algorithm parameters, not fitted to the evaluation data. Diarization accuracy in Section 5.6 is scored with PyAnnotate against ground-truth speaker turns in synthesized turn-taking conversations, so the labels are independent of the algorithm's outputs; the no-overlap condition limits ecological validity but is not circularity. The user studies measure self-reported preferences after using the system, and the n=263 and n=494 surveys motivate or inform the design rather than serve as tests of the implemented claims. No load-bearing conclusion rests on a self-citation: reference [50] (which includes an author of this paper) is cited only for prior captioning and head-worn display work, not to justify SpeechCompass's central capabilities. The noted limitations—small sample size, quiet lab setting, and no-overlap diarization test—are external-validity and generalizability concerns, not reductions of the results to their own inputs.

Assumptions & free parameters 4 free parameters · 3 assumptions · 0 invented entities

The system has no invented physical entities. Its technical parameters are hand-tuned constants rather than fitted scientific constants; they affect localization smoothing and latency but not the conceptual claim. The main load-bearing assumptions are geometric (far field), acoustic (loudest source dominates), and conversational (no overlap), each acknowledged in the paper.

free parameters (4)
  • KDE Gaussian kernel bandwidth = 25
    Chosen for the kernel density estimate over angle samples; not derived from theory. Affects localization smoothing and convergence latency.
  • GCC-PHAT partial normalization exponent = -0.3
    Introduced in Appendix A.1 to give less weight to end-fire delays; a hand-tuned value that deviates from standard full normalization.
  • KDE sample buffer size = 600 latest samples
    Chosen as the window for angle density estimation; trades latency against localization stability.
  • Diarization running histogram window = 522 ms
    Applied to TDOA frames in Section 5.6 to obtain speaker labels; hand-selected smoothing interval.
assumptions (3)
  • domain assumption Far-field plane-wave approximation for sound propagation
    Used in Appendix A.1 Eq. 4 to convert time delay to azimuth angle. It is valid only when the source is roughly a meter or more from the array, but the phone-case prototype can be placed on a table with closer speakers.
  • domain assumption One active speaker per time frame (no overlapping speech)
    Section 5.6 explicitly excludes overlap: "no overlap in the speech content." Direction-based diarization cannot assign simultaneous speakers from different angles to separate transcript lines.
  • domain assumption GCC-PHAT tracks the loudest sound source
    Section 5.3 states: "The GCC-PHAT picks up the loudest sound, thus making it difficult to estimate TDOA if the sound source is under or at the environmental noise level." This bounds localization reliability in noisy everyday settings.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SpeechCompass: Enhancing Mobile Captioning with Diarization and Directional Guidance via Multi-Microphone Localization." pith.science (2026). https://pith.science/paper/FVQ4XRVF

@misc{pith2026250208848,
  author       = {Pith},
  title        = {Pith review of: SpeechCompass: Enhancing Mobile Captioning with Diarization and Directional Guidance via Multi-Microphone Localization},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FVQ4XRVF}},
  note         = {Machine review of arXiv:2502.08848}
}
read the original abstract

Speech-to-text capabilities on mobile devices have proven helpful for hearing and speech accessibility, language translation, note-taking, and meeting transcripts. However, our foundational large-scale survey (n=263) shows that the inability to distinguish and indicate speaker direction makes them challenging in group conversations. SpeechCompass addresses this limitation through real-time, multi-microphone speech localization, where the direction of speech allows visual separation and guidance (e.g., arrows) in the user interface. We introduce efficient real-time audio localization algorithms and custom sound perception hardware running on a low-power microcontroller and four integrated microphones, which we characterize in technical evaluations. Informed by a large-scale survey (n=494), we conducted an in-person study of group conversations with eight frequent users of mobile speech-to-text, who provided feedback on five visualization styles. The value of diarization and visualizing localization was consistent across participants, with everyone agreeing on the value and potential of directional guidance for group conversations.

Figures

Figures reproduced from arXiv: 2502.08848 by the authors.

Figure 1
Figure 1. SpeechCompass creates user-friendly speech transcripts for group conversations with multiple speakers. Left: Current solutions concatenate and mix the transcribed speech when multiple people participate in a conversation, which makes it challenging to read and understand the transcript. Right: SpeechCompass addresses this limitation through real-time, multi-microphone speech localization, where the direction of spee… view at source ↗
Figure 2
Figure 2. Overview of the SpeechCompass phone case prototype. A) A mobile phone application interface with a mounted [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Participant responses to the question What are the biggest challenges with your current captioning or transcription device/technology? (select all that apply)? 0%5% 10%15% 20%25%30%35%40%45%50%55%60%65%70%75%80%85%90%95%100% Rarely/never A few times/month Multiple times/week Daily Multiple times/day 69 (26%) 43 (16%) 68 (26%) 52 (20%) 31 (12%) 61 (23%) 56 (21%) 65 (25%) 53 (20%) 28 (11%) 70 (27%) 34 (13%) 76 (30%) 4… view at source ↗
Figures from the paper (13 more)
Figure 4
Figure 4. Figure 4: Survey results of how often participants encountered challenging scenarios with today’s transcription technology. [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: SpeechCompass system diagram. The phone case contains four microphones connected to a microcontroller. The audio [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Visualization of localization methods with 2 and 3 microphone configurations. A) Localization with two microphones. [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: The mobile phone application with different direction visualization options. A) Directional glyphs are arrows next to [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: The top and bottom microphones are on the side edge of [PITH_FULL_IMAGE:figures/full_fig_p009_8.png]
Figure 8
Figure 8. Figure 8: Localization on a mobile phone. A) Microphone positioning and distance. B) Examples of the raw microphone data [PITH_FULL_IMAGE:figures/full_fig_p010_8.png]
Figure 9
Figure 9. Figure 9: Effect of source elevation angle on the azimuth [PITH_FULL_IMAGE:figures/full_fig_p010_9.png]
Figure 10
Figure 10. Figure 10: A) SpeechCompass azimuth angle measurement errors for speech and noise sound at different loudness levels. A [PITH_FULL_IMAGE:figures/full_fig_p011_10.png]
Figure 11
Figure 11. Figure 11: Diarization experiment. A) Data collection setup for diarization testing. The SpeechCompass phone case is placed on [PITH_FULL_IMAGE:figures/full_fig_p012_11.png]
Figure 12
Figure 12. Figure 12: Diarization error rate (DER) for three and four [PITH_FULL_IMAGE:figures/full_fig_p012_12.png]
Figure 13
Figure 13. Figure 13: Participants’ preferences for different visualization techniques in the online survey. A) Results indicating how [PITH_FULL_IMAGE:figures/full_fig_p013_13.png]
Figure 14
Figure 14. Figure 14: Examples of seven visualization scenarios that participants experienced in the in-person study. [PITH_FULL_IMAGE:figures/full_fig_p013_14.png]
Figure 15
Figure 15. Figure 15: Boxplots of results of the in-person study. A) Participants’ preferences for different visualization techniques. B) [PITH_FULL_IMAGE:figures/full_fig_p013_15.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

69 extracted references · 50 canonical work pages

  1. [1]

    Chandykunju Alex, Kevin Martin Jose, and Arun Joseph. 2013. Sound localization and Visualization device. In2013 IEEE Global Humanitarian Technology Conference (GHTC). 262–264. doi:10.1109/GHTC.2013.6713692

  2. [2]

    Android. 2022. Introducing Live Transcribe. https://www.android.com/ accessibility/live-transcribe/. Accessed 2022-03-26

  3. [3]

    Android. 2022. SpeechRecognizer API Documentation). https://developer.android. com/reference/android/speech/SpeechRecognizer. Accessed 2022-10-25

  4. [4]

    Xavier Anguera, Chuck Wooters, and Javier Hernando. 2007. Acoustic beam- forming for speaker diarization of meetings. IEEE Transactions on Audio, Speech, and Language Processing 15, 7 (2007), 2011–2022

  5. [5]

    ARM. 2022. CMSIS DSP Software Library. https://www.keil.com/pack/doc/ CMSIS/DSP/html/index.html. Accessed 2022-05-12

  6. [6]

    Ava. 2022. Ava Captioning Solution. https://www.ava.me/. Accessed 2024-12-10

  7. [7]

    Jacob Benesty, Jingdong Chen, and Yiteng Huang. 2008. Conventional beamform- ing techniques. Microphone array signal processing (2008), 39–65

  8. [8]

    Larwan Berke, Khaled Albusays, Matthew Seita, and Matt Huenerfauth. 2019. Preferred Appearance of Captions Generated by Automatic Speech Recognition for Deaf and Hard-of-Hearing Viewers. In Extended Abstracts of the 2019 CHI Conference on Human Factors in Computing Systems (Glasgow, Scotland Uk) (CHI EA ’19). Association for Computing Machinery, New York,...

Show all 69 references
  1. [9]

    Larwan Berke, Christopher Caulfield, and Matt Huenerfauth. 2017. Deaf and Hard-of-Hearing Perspectives on Imperfect Automatic Speech Recognition for Captioning One-on-One Meetings. In Proceedings of the 19th International ACM SIGACCESS Conference on Computers and Accessibility...

  2. [10]

    Rachel Boll, Shruti Mahajan, Jeanne Reis, and Erin T. Solovey. 2020. Creating Questionnaires That Align with ASL Linguistic Principles and Cultural Practices within the Deaf Community. , Article 61 (2020), 4 pages. doi:10.1145/3373625. 3418071

  3. [11]

    Danielle Bragg, Nicholas Huynh, and Richard E. Ladner. 2016. A Personalizable Mobile Sound Detector App Design for Deaf and Hard-of-Hearing Users. In Proceedings of the 18th International ACM SIGACCESS Conference on Computers and Accessibility (Reno, Nevada, USA) (ASSETS ’16)....

  4. [12]

    Hervé Bredin, Ruiqing Yin, Juan Manuel Coria, Gregory Gelly, Pavel Korshunov, Marvin Lavechin, Diego Fustes, Hadrien Titeux, Wassim Bouaziz, and Marie- Philippe Gill. 2020. pyannote.audio: neural building blocks for speaker diarization. In ICASSP 2020, IEEE International Confe...

  5. [13]

    Janine Butler, Brian Trager, and Byron Behm. 2019. Exploration of Automatic Speech Recognition for Deaf and Hard of Hearing Students in Higher Education Classes. In Proceedings of the 21st International ACM SIGACCESS Conference on Computers and Accessibility (Pittsburgh, PA, U...

  6. [14]

    Chao Cai, Henglin Pu, Peng Wang, Zhe Chen, and Jun Luo. 2021. We Hear Your PACE: Passive Acoustic Localization of Multiple Walking Persons. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 5, 2, Article 55 (jun 2021), 24 pages. doi:10.1145/3463510

  7. [15]

    Caluã de Lacerda Pataca, Saad Hassan, Nathan Tinker, Roshan Lalintha Peiris, and Matt Huenerfauth. 2024. Caption Royale: Exploring the Design Space of Affective Captions from the Perspective of Deaf and Hard-of-Hearing Individuals. In Proceedings of the 2024 CHI Conference on ...

  8. [16]

    Artem Dementyev, Pascal Getreuer, Dimitri Kanevsky, Malcolm Slaney, and Richard F Lyon. 2021. VHP: Vibrotactile Haptics Platform for On-Body Ap- plications. In The 34th Annual ACM Symposium on User Interface Software and Technology (Virtual Event, USA) (UIST ’21). Association ...

  9. [17]

    Hoang Do, Harvey F Silverman, and Ying Yu. 2007. A real-time SRP-PHAT source location implementation using stochastic region contraction (SRC) on a large- aperture microphone array. In 2007 IEEE International Conference on Acoustics, Speech and Signal Processing-ICASSP’07 , Vo...

  10. [18]

    Ariel Ephrat, Inbar Mosseri, Oran Lang, Tali Dekel, Kevin Wilson, Avinatan Hassidim, William T Freeman, and Michael Rubinstein. 2018. Looking to listen at the cocktail party: A speaker-independent audio-visual model for speech separation. arXiv preprint arXiv:1804.03619 (2018)

  11. [19]

    Hakan Erdogan, John R Hershey, Shinji Watanabe, and Jonathan Le Roux. 2015. Phase-sensitive and recognition-boosted speech separation using deep recurrent neural networks. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 708–712

  12. [20]

    Leah Findlater, Bonnie Chinh, Dhruv Jain, Jon Froehlich, Raja Kushalnagar, and Angela Carey Lin. 2019. Deaf and Hard-of-Hearing Individuals’ Preferences for Wearable and Mobile Sound Awareness Technologies. InProceedings of the 2019 CHI Conference on Human Factors in Computing...

  13. [21]

    Israel D Gebru, Sileye Ba, Xiaofei Li, and Radu Horaud. 2017. Audio-visual speaker diarization based on spatiotemporal bayesian fusion. IEEE transactions on pattern analysis and machine intelligence 40, 5 (2017), 1086–1099

  14. [22]

    Abraham Glasser, Kesavan Kushalnagar, and Raja Kushalnagar. 2017. Deaf, Hard of Hearing, and Hearing Perspectives on Using Automatic Speech Recog- nition in Conversation. In Proceedings of the 19th International ACM SIGAC- CESS Conference on Computers and Accessibility (Baltim...

  15. [23]

    Goodman, Ping Liu, Dhruv Jain, Emma J

    Steven M. Goodman, Ping Liu, Dhruv Jain, Emma J. McDonnell, Jon E. Froehlich, and Leah Findlater. 2021. Toward User-Driven Sound Recognizer Personalization with People Who Are d/Deaf or Hard of Hearing. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 5, 2, Article 63 (ju...

  16. [24]

    Google. 2018. Google Surveys Methodology. http://services.google.com/fh/files/ misc/white_paper_how_google_surveys_works.pdf. Accessed 2022-03-22

  17. [25]

    Beth G Greene, David B Pisoni, and Thomas D Carrell. 1984. Recognition of speech spectrograms. The Journal of the Acoustical Society of America 76, 1 (1984), 32–43

  18. [26]

    François Grondin and François Michaud. 2019. Lightweight and optimized sound source localization and tracking methods for open and closed microphone array configurations. Robotics and Autonomous Systems 113 (2019), 63–80

  19. [27]

    Ru Guo, Yiru Yang, Johnson Kuang, Xue Bin, Dhruv Jain, Steven Goodman, Leah Findlater, and Jon Froehlich. 2020. HoloSound: Combining Speech and Sound Identification for Deaf or Hard of Hearing Users on a Head-Mounted Display. In Proceedings of the 22nd International ACM SIGACC...

  20. [28]

    Jahn Heymann, Lukas Drude, and Reinhold Haeb-Umbach. 2016. Neural network based spectral mask estimation for acoustic beamforming. In 2016 IEEE Inter- national Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 196–200

  21. [29]

    Pulsar Instruments. 2022. Decibel chart – decibel levels of common sounds. https://pulsarinstruments.com/news/decibel-chart-noise-level. Accessed 2022- 07-27

  22. [30]

    Yusuf Isik, Jonathan Le Roux, Zhuo Chen, Shinji Watanabe, and John R Hershey

  23. [32]

    Froehlich

    Dhruv Jain, Leah Findlater, Jamie Gilkeson, Benjamin Holland, Ramani Du- raiswami, Dmitry Zotkin, Christian Vogler, and Jon E. Froehlich. 2015. Head- Mounted Display Visualizations to Support Sound Awareness for the Deaf and Hard of Hearing. In Proceedings of the 33rd Annual A...

  24. [33]

    Froehlich

    Dhruv Jain, Kelly Mack, Akli Amrous, Matt Wright, Steven Goodman, Leah Findlater, and Jon E. Froehlich. 2020. HomeSound: An Iterative Field Deployment of an In-Home Sound Awareness System for Deaf or Hard of Hearing Users. In Proceedings of the 2020 CHI Conference on Human Fac...

  25. [34]

    Vahid Ahmadi Kalkhorani, Anurag Kumar, Ke Tan, Buye Xu, and DeLiang Wang

  26. [35]

    Yoshihiro Kaneko, Inho Chung, and Kenji Suzuki. 2013. Light-emitting device for supporting auditory awareness of hearing-impaired people during group conver- sations. In 2013 IEEE International Conference on Systems, Man, and Cybernetics . IEEE, 3567–3572

  27. [36]

    Ellington Kirby, Seoyoon Park, Yan Wang, and Yingying Chen. 2016. HearHere: Smartphone Based Audio Localization Using Time Difference of Arrival: Demo. In Proceedings of the 22nd Annual International Conference on Mobile Computing and Networking (New York City, New York)(MobiC...

  28. [37]

    Charles Knapp and Glifford Carter. 1976. The generalized correlation method for estimation of time delay. IEEE transactions on acoustics, speech, and signal processing 24, 4 (1976), 320–327

  29. [38]

    Kushalnagar, Gary W

    Raja S. Kushalnagar, Gary W. Behm, Aaron W. Kelstone, and Shareef Ali

  30. [39]

    Ahmet Köse, Aleksei Tepljakov, and Sergei Astapov. 2017. Real-time localization and visualization of a sound source for virtual reality applications. In 2017 25th International Conference on Software, Telecommunications and Computer Networks (SoftCOM). 1–6. doi:10.23919/SOFTCO...

  31. [40]

    Bowon Lee, Amir Said, Ton Kalker, and Ronald W Schafer. 2008. Maximum likelihood time delay estimation with phase domain analysis in the generalized cross correlation framework. In 2008 Hands-Free Speech Communication and Microphone Arrays. IEEE, 89–92

  32. [41]

    Bo Li, Tara N Sainath, Ron J Weiss, Kevin W Wilson, and Michiel Bacchiani

  33. [42]

    LibriVox. 2022. Alice’s Adventures in Wonderland by Lewis Carroll (Version 2). https://librivox.org/alices-adventures-in-wonderland-by-lewis-carroll-4/. Accessed 2022-07-12

  34. [43]

    Hong Liu and Miao Shen. 2010. Continuous sound source localization based on microphone array for mobile robots. In 2010 IEEE/RSJ International Conference on Intelligent Robots and Systems . IEEE, 4332–4339

  35. [44]

    Richard F Lyon. 2017. Human and machine hearing: extracting meaning from sound. Cambridge University Press

  36. [45]

    Neural network adaptive beamforming for robust multichannel speech recognition. (2016)

  37. [46]

    Microsoft. 2022. Translator. https://translator.microsoft.com/. Accessed 2022-03- 26

  38. [47]

    MiniDSP. 2022. USB Mic array. https://www.minidsp.com/products/usb-audio- interface/uma-8-16-usb-mic-array. Accessed 2022-05-15

  39. [48]

    Kai Morich. 2022. usb-serial-for-android). https://github.com/mik3y/usb-serial- for-android. Accessed 2022-11-09

  40. [49]

    James C Makous and John C Middlebrooks. 1990. Two-dimensional sound local- ization by human listeners. The journal of the Acoustical Society of America 87, 5 (1990), 2188–2200

  41. [50]

    Alex Olwal, Kevin Balke, Dmitrii Votintcev, Thad Starner, Paula Conn, Bonnie Chinh, and Benoit Corda. 2020. Wearable Subtitles: Augmenting Spoken Com- munication with Lightweight Eyewear for All-Day Captioning . Association for Computing Machinery, New York, NY, USA, 1108–1120...

  42. [51]

    Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015. Librispeech: an asr corpus based on public domain audio books. In 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 5206–5210

  43. [52]

    Tae Jin Park, Naoyuki Kanda, Dimitrios Dimitriadis, Kyu J Han, Shinji Watanabe, and Shrikanth Narayanan. 2022. A review of speaker diarization: Recent advances with deep learning. Computer Speech & Language 72 (2022), 101317

  44. [53]

    Pius Kavuma Basajjabaka Mugagga and Simon Winberg. 2015. Sound source localisation on Android smartphones: A first step to using smartphones as au- ditory sensors for training A.I systems with Big Data. In AFRICON 2015. 1–5. doi:10.1109/AFRCON.2015.7331970

  45. [54]

    Ashutosh Saxena and Andrew Y Ng. 2009. Learning sound location from a single microphone. In 2009 IEEE International Conference on Robotics and Automation . IEEE, 1737–1742

  46. [55]

    Matthew Seita. 2020. Designing Automatic Speech Recognition Technologies to Improve Accessibility for Deaf and Hard-of-Hearing People in Small Group Meetings. In Extended Abstracts of the 2020 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI EA ’20...

  47. [56]

    Deep Sleep. 2022. Rain Sound. https://www.youtube.com/watch?v= 13EL6Mgeocc&t=3448s. Accessed 2022-07-12

  48. [57]

    Priyantha, Anit Chakraborty, and Hari Balakrishnan

    Nissanka B. Priyantha, Anit Chakraborty, and Hari Balakrishnan. 2000. The Cricket Location-Support System. In Proceedings of the 6th Annual International Conference on Mobile Computing and Networking (Boston, Massachusetts, USA) (MobiCom ’00). Association for Computing Machine...

  49. [58]

    Speaksee. 2022. Speaksee microphone kit. https://speak-see.com/. Accessed 2022-10-25

  50. [59]

    Hassan Taherian, Ashutosh Pandey, Daniel Wong, Buye Xu, and DeLiang Wang

  51. [60]

    Hassan Taherian and DeLiang Wang. 2024. Multi-Channel Conversational Speaker Separation via Neural Diarization. IEEE/ACM Transactions on Audio, Speech, and Language Processing 32 (2024), 2467–2476. doi:10.1109/TASLP.2024. 3393726

  52. [61]

    David Snyder, Daniel Garcia-Romero, Gregory Sell, Alan McCree, Daniel Povey, and Sanjeev Khudanpur. 2019. Speaker recognition for multi-speaker conver- sations using x-vectors. In ICASSP 2019-2019 IEEE International conference on acoustics, speech and signal processing (ICASSP...

  53. [62]

    Quan Wang, Carlton Downey, Li Wan, Philip Andrew Mansfield, and Ignacio Lopz Moreno. 2018. Speaker diarization with LSTM. In 2018 IEEE International confer- ence on acoustics, speech and signal processing (ICASSP) . IEEE, 5239–5243

  54. [63]

    Lucinda Wilder. 1975. Articulatory and acoustic characteristics of speech sounds. In Understanding Language. Elsevier, 31–76

  55. [64]

    In ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

    Leveraging Sound Localization to Improve Continuous Speaker Separation. In ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). 621–625. doi:10.1109/ICASSP48485.2024.10446934

  56. [65]

    Yang Yang, George Sung, Shao-Fu Shih, Hakan Erdogan, Chehung Lee, and Matthias Grundmann. 2024. Binaural Angular Separation Network. In ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 1201–1205

  57. [66]

    Giuseppe Valenzise, Luigi Gerosa, Marco Tagliasacchi, Fabio Antonacci, and Augusto Sarti. 2007. Scream and gunshot detection and localization for audio- surveillance systems. In 2007 IEEE Conference on Advanced Video and Signal Based Surveillance. IEEE, 21–26

  58. [69]

    Yuteng Xiao, Jihang Yin, Honggang Qi, Hongsheng Yin, and Gang Hua. 2017. MVDR algorithm based on estimated diagonal loading for beamforming. Mathe- matical Problems in Engineering 2017 (2017)

  59. [71]

    Jiahui Yu, Chung-Cheng Chiu, Bo Li, Shuo-yiin Chang, Tara N Sainath, Yanzhang He, Arun Narayanan, Wei Han, Anmol Gulati, Yonghui Wu, et al. 2021. Fastemit: Low-latency streaming asr with sequence-level emission regularization. In ICASSP 2021-2021 IEEE International Conference ...

  60. [2015]

    InProceedings of the 17th International ACM SIGACCESS Conference on Computers & Accessibility (Lisbon, Portugal) (AS- SETS ’15)

    Tracked Speech-To-Text Display: Enhancing Accessibility and Readabil- ity of Real-Time Speech-To-Text. InProceedings of the 17th International ACM SIGACCESS Conference on Computers & Accessibility (Lisbon, Portugal) (AS- SETS ’15). Association for Computing Machinery, New York...

  61. [2016]

    arXiv preprint arXiv:1607.02173 (2016)

    Single-channel multi-speaker separation using deep clustering. arXiv preprint arXiv:1607.02173 (2016)

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.