REVIEW 3 major objections 5 minor 87 references
Spatial Speech Translation: Translating Across Space With Binaural Hearables
T0 review · 3 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read A binaural hearable pipeline translates multiple concurrent speakers in real time while preserving each speaker's direction and voice characteristics, and it generalizes from synthetic training to unseen real-world environments.
desk verdict A solid proof-of-concept system paper that integrates separation, streaming expressive translation, and binaural rendering for hearables; the real-world evaluation is genuine evidence, but the synthetic-to-real BRIR transfer is the main open risk. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing component is a search-based joint localization and separation network. The 360-degree space is divided into 36 angular regions; for each region, the binaural input is time-shifted by the interaural time difference corresponding to that angle and fed to a streaming TF-GridNet that is trained to output the separated source if a speaker is present at that angle and silence otherwise. Interaural phase and level differences are concatenated with the spectrogram as features, and false duplicates from multipath are removed by clustering separation outputs by segment-wise similarity. Around this core, the translation module is a simultaneous speech-to-text model (StreamSpeech-style with Conformer encoder and CTC-guided READ/WRITE policy), followed by a text-to-unit model and an expressivity-preserving vocoder conditioned on an expressive embedding extracted from the source speech; the whole translation model is fine-tuned on the separation model's imperfect outputs to become robust to residual interference. Finally, binaural rendering convolves translated monaural speech with a generic HRTF at the estimated angle for ITD and applies an ILD compensation scale computed from the separated source.
What would settle it
Collect binaural recordings of two concurrent French speakers in a room with a wearer whose head size differs substantially from the 18 cm average assumed in training, or place speakers closer than 0.75 m, then run the released pipeline and measure localization precision and recall plus ASR-BLEU; if precision or recall collapses well below the reported indoor values (97% and 98%) or ASR-BLEU drops to the no-separation baseline, the claimed synthetic-to-real generalization is falsified.
Extended reading notes
Core claim
The paper introduces spatial speech translation, a concept and system that takes a binaural mixture from microphones at the two ears, identifies how many speakers are present and from which angles, separates each voice, translates them simultaneously into the wearer's language while preserving prosody and vocal identity, and renders the translated speech binaurally so it appears to come from the original speaker's direction. The central empirical claim is that this works in real time on Apple M2 silicon and generalizes to unseen real-world environments and wearers: in six indoor and four outdoor venues, the joint localization and separation step reaches 97% precision and recall indoors and 92% and 94% outdoors with a median angle error of 6.8 degrees; fine-tuning the translation model on separation outputs raises ASR-BLEU from 18.06 to 22.07, and the full expressive system raises perceived speaker similarity from 1.81 to 3.45 while keeping median perceived direction error at 16.7 degrees versus 15.0 degrees for the original speech. The paper also reports that generic-HRTF rendering with ILD compensation brings interaural time difference error to 72.3 microseconds and interaural level difference error to 0.16 dB.
Load-bearing premise
The whole system rests on the claim that a separation model trained only on synthetic binaural mixtures, built from 77 room and head configurations, will work on real heads, real rooms, and real distances without any recordings made with the actual hardware; if that synthetic-to-real bridge fails, the spatial translation benefit disappears.
Editorial extensions
If this is right
- Multi-speaker environments become translatable in real time; existing speech translators that assume a single speaker fail under interference.
- The synthetic training recipe removes the need to collect data with each hearable device, room, language, or wearer; new language pairs can be added by generating new synthetic mixtures.
- Listeners can follow who is speaking in a conversation because direction, prosody, and voice identity survive translation.
- Because the pipeline runs on Apple M2 silicon with a real-time factor below one, it is deployable on commodity AR and wearable hardware today.
- Latency can be traded against accuracy via chunk size (one to four seconds), giving tunable behavior for casual versus high-stakes use.
Reading between the lines
- If the synthetic-to-real bridge is as robust as reported, the same separation-trained-on-BRIR recipe should extend to other wearable geometries (for example, earbuds with smaller microphone spacing) and to more languages without any new real-world data, but only if the time-difference-of-arrival shift model is recalibrated to the new inter-microphone distance.
- The fine-tune-on-upstream-distortion trick is a general design principle: cascaded speech systems (speech-to-text plus machine translation, diarization plus translation) should train their downstream model on the actual output distribution of their front end, not only on clean corpora.
- The rendering delay-compensation scheme (apply current spatial cues to delayed translated audio) suggests a broader principle for streaming augmented-reality audio: spatial metadata can be decoupled from audio content latency; this could be tested with moving speakers by measuring whether direction perception remains accurate while a speaker walks during the translation delay.
- Because the system outputs text as an intermediate, a natural extension is spatially anchored transcripts on an AR display; the paper mentions this direction but does not evaluate it.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces 'spatial speech translation', a hearable pipeline that localizes and separates concurrent speakers from binaural microphone input, translates each separated stream in real time with expressive speech-to-speech translation, and renders the translated speech binaurally at each speaker's original direction. The separation/localization model is trained on synthetic mixtures built from CoVoST2 speech, WHAM! noise, and 77 BRIR configurations from four public datasets, and the translation model is fine-tuned on the imperfect separation outputs. Evaluation includes synthetic benchmarks, real-world indoor/outdoor recordings with a Sony WH-1000XM4 prototype, listening studies with 29 participants, latency and noise-cancellation preference studies, and objective metrics such as ASR-BLEU, VSim, localization precision/recall, and delta-ITD/delta-ILD. The central claim is that this is the first real-time binaural hearable system that translates multiple speakers while preserving spatial cues and speaker voice characteristics, and that the synthetic training recipe generalizes to unseen wearers and environments without hardware-specific training data.
Significance. If the claims hold, this is a significant contribution to hearable and speech-translation research. The paper is the first to integrate spatial perception into end-to-end speech translation, and it provides a decomposable, reproducible pipeline with code and dataset release. The strongest evidence is the real-world user study (10 environments, 29 participants), the localization precision/recall results, the delta-ITD/delta-ILD improvements, and the demonstration that fine-tuning on separation outputs improves ASR-BLEU on both synthetic and real-world data. The synthetic-to-real transfer strategy is an appealing practical contribution, and the runtime analysis shows the pipeline is plausible for on-device use. However, the generalization claim rests on a narrow real-world validation, and several headline numbers are reported without measures of variability, which tempers the strength of the conclusions.
major comments (3)
- [§3.1.3, §4, §5] The separation and localization model is trained exclusively on synthetic mixtures generated by convolving CoVoST2 monaural speech with 77 BRIR configurations from CIPIC, RRBRIR, ASH, and CATTRIR, plus WHAM! noise, but the deployed hardware places SP15C microphones on the outside of the Sony WH-1000XM4 earcups. The real-world evaluation uses only loudspeaker reproductions of CoVoST2 test clips, not human talkers, and distances of 0.75–2.5 m. This is the load-bearing bridge for the paper's claim that the system generalizes to unseen wearers and environments without hardware-specific training data, yet §7 does not flag the microphone-placement mismatch, near-field effects, or loudspeaker-only speech sources as open risks. Because separation errors propagate directly to translation and rendering, I would like either additional experiments with human talkers and varied mic placements or an explicit, detailed scope discussion in the limitations section.
- [§5.1, Table 2, Fig. 6] No confidence intervals, standard deviations, or significance tests are reported for any of the headline subjective or objective numbers, including semantic consistency (3.35 vs. 1.15), speaker similarity (1.81 vs. 3.45), ASR-BLEU differences (18.06 vs. 22.07), localization precision/recall, or perceived angular error medians. The claim that the rendered English speech has 'similar' localization error to the original French speech is based on a median comparison without any measure of spread or a paired test. The manuscript should report per-participant/per-sample variability and appropriate statistical tests for the human ratings and for the metric comparisons that support the main claims.
- [§5.2.1, Fig. 11] The localization results are reported as means over participants in indoor and outdoor groups, but the precision/recall values appear to be per-participant binary decisions; no details are given on how partial detections (e.g., two speakers but one false positive) are counted, or how the 90th-percentile AoA error is computed across mixtures with different numbers of sources. Since the clustering false-positive elimination is a key algorithmic component, a precise definition of the evaluation protocol and error aggregation would strengthen the reproducibility of the localization claims.
minor comments (5)
- [§3.1.3] There are typos: 'diferent' and 'binural' should be 'different' and 'binaural'.
- [§6.2] The list of four model configurations labels both item (3) and item (4) as 'Finetuned S2T with Expressive T2S'; one of them should be 'Finetuned S2T with non-Expressive T2S' to match Table 4.
- [Fig. 12 caption] The caption says 'we compute the ΔITD and ΔITD between each input binaural French speech chunk and the rendered English speech chunk'; the second metric should be ΔILD.
- [Abstract, §1, Table 2] The abstract reports BLEU 'up to 22.01' while the introduction and Table 2 report 22.07; the numbers should be made consistent and the model configuration (fine-tuned non-expressive vs. fine-tuned expressive) should be stated in the abstract.
- [§5.1.1 and §5.1.2] The listening survey and spatial perception study report only mean values; adding the number of ratings per condition and error bars in Figs. 6 and 7 would make the results much easier to interpret.
Circularity Check
Rendering ILD metric is by construction (output ILD is set equal to input ILD); central translation and user-study claims are independent.
-
self definitional
[Section 3.3 (ILD compensation equation); Section 6.3 / Table 6.]
"we first compute its ILD from the separated binaural source speech as ILD_i = ||y_i^r||_1 / ||y_i^l||_1. We then scale the translated signal based on the binaural HRTF response and ILD_i as: [o_i^l, o_i^r] = [o_i * h^l(theta=theta_i, phi=0), (ILD_i / (||h^l||_1 / ||h^r||_1)) o_i * h^r(theta=theta_i, phi=0)]"
The right channel is scaled by ILD_i normalized by the HRTF channel norms, so the output ILD is forced to equal ILD_i by the formula itself. Table 6 then reports a ΔILD of 0.16 dB for this method and credits it with preserving spatial cues. That is not an independent prediction: it is the same quantity that was inserted into the output being measured on the output. The nontrivial residual is only how accurately the separated ILD_i matches the original source ILD, plus the ITD contribution from the generic HRTF. The central BLEU and user-study results do not depend on this tautology.
full rationale
The joint separation/localization and translation chain is self-contained and evaluated against external data. Separation is trained on synthetic mixtures from CoVoST2, WHAM!, and four public BRIR datasets and tested on held-out real-world recordings; localization precision/recall and AoA errors are not fitted to the reported numbers. The S2T module uses external StreamSpeech and SeamlessExpressive components, with fine-tuning on separation outputs evaluated on real-world recordings and standard test sets. Existing self-citations to search-based separation and TF-GridNet are architectural borrowings, not load-bearing uniqueness claims. The only reduction-by-construction element is the ILD compensation benchmark: the rendering equation explicitly transfers the input ILD to the output, so the reported ΔILD improvement over the generic-HRTF baseline is a definitional identity rather than an empirical discovery. This is a secondary validation metric; the paper's central spatial-translation claim is independently supported by human localization tests and BLEU results.
Assumptions & free parameters
free parameters (3)
- N=36 angular search regions =
36
- Power threshold for source detection =
1e-2
- Multi-resolution loss weight lambda =
0.1
assumptions (4)
- domain assumption Far-field sources and fixed microphone separation d = 18 cm for TDoA alignment
- domain assumption Synthetic BRIR mixtures from CIPIC, RRBRIR, ASH, and CATTRIR are representative of real-world reverberation and HRTF variability
- domain assumption Generic CIPIC HRTF is adequate for ITD rendering
- domain assumption ASR-BLEU on the rendered speech is a valid proxy for translation quality
Cite this review
Pith. "Pith review of Spatial Speech Translation: Translating Across Space With Binaural Hearables." pith.science (2026). https://pith.science/paper/BEXIYJ42
@misc{pith2026250418715,
author = {Pith},
title = {Pith review of: Spatial Speech Translation: Translating Across Space With Binaural Hearables},
year = {2026},
howpublished = {\url{https://pith.science/paper/BEXIYJ42}},
note = {Machine review of arXiv:2504.18715}
}
read the original abstract
Imagine being in a crowded space where people speak a different language and having hearables that transform the auditory space into your native language, while preserving the spatial cues for all speakers. We introduce spatial speech translation, a novel concept for hearables that translate speakers in the wearer's environment, while maintaining the direction and unique voice characteristics of each speaker in the binaural output. To achieve this, we tackle several technical challenges spanning blind source separation, localization, real-time expressive translation, and binaural rendering to preserve the speaker directions in the translated audio, while achieving real-time inference on the Apple M2 silicon. Our proof-of-concept evaluation with a prototype binaural headset shows that, unlike existing models, which fail in the presence of interference, we achieve a BLEU score of up to 22.01 when translating between languages, despite strong interference from other speakers in the environment. User studies further confirm the system's effectiveness in spatially rendering the translated speech in previously unseen real-world reverberant environments. Taking a step back, this work marks the first step towards integrating spatial perception into speech translation.
Figures
Figures from the paper (9 more)
Reference graph
Works this paper leans on
-
[1]
Alex Agranovich, Eliya Nachmani, Oleg Rybakov, Yifan Ding, Ye Jia, Nadav Bar, Heiga Zen, and Michelle Tadmor Ramanovich. 2024. SimulTron: On-Device Simultaneous Speech to Speech Translation. arXiv:2406.02133 [eess.AS] https: //arxiv.org/abs/2406.02133
arXiv 2024
-
[2]
Dmitry Alexandrovsky, Susanne Putze, Michael Bonfert, Sebastian Höffner, Pitt Michelmann, Dirk Wenig, Rainer Malaka, and Jan David Smeddinck. 2020. Unmet Needs and Opportunities for Mobile Translation AI. CHI (2020)
2020
-
[3]
V.R. Algazi, R.O. Duda, D.M. Thompson, and C. Avendano. 2001. The CIPIC HRTF database. , 99-102 pages. https://doi.org/10.1109/ASPAA.2001.969552
-
[4]
Apple. 2024. Listen with Personalized Spatial Audio for AirPods and Beats. https://support.apple.com/en-us/102596
2024
-
[5]
Shoko Araki, Hiroshi Sawada, and Shoji Makino. 2007. Blind Speech Separation in a Meeting Situation with Maximum SNR Beamformers. In ICASSP
2007
-
[6]
Naveen Arivazhagan, Colin Cherry, Wolfgang Macherey, Chung-Cheng Chiu, Semih Yavuz, Ruoming Pang, Wei Li, and Colin Raffel. 2019. Monotonic Infinite Lookback Attention for Simultaneous Machine Translation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , Anna Korhonen, David Traum, and Lluís Màrquez (Eds.)
2019
-
[7]
Arun Babu, Changhan Wang, Andros Tjandra, Kushal Lakhotia, Qiantong Xu, Naman Goyal, Kritika Singh, Patrick Von Platen, Yatharth Saraf, Juan Pino, et al
-
[8]
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014. Neural Machine Translation by Jointly Learning to Align and Translate. CoRR abs/1409.0473 (2014). https://api.semanticscholar.org/CorpusID:11212020
arXiv 2014
Show all 87 references
-
[9]
Loïc Barrault, Yu-An Chung, Mariano Meglioli, David Dale, Ning Dong, Paul- Ambroise Duquenne, Hady Elsahar, Hongyu Gong, Kevin Heffernan, John Hoff- man, Christopher Klaiber, Pengwei Li, Daniel Licht, Jean Maillard, Alice Rakotoari- son, Kaushik Sadagopan, Guillaume Wenzek, Et...
2025 doi
-
[10]
Google Blog. 2024. Translate with Google Pixel Buds. https://support.google. com/googlepixelbuds/answer/7573100?hl=en
2024
-
[11]
Zalán Borsos, Raphaël Marinier, Damien Vincent, Eugene Kharitonov, Olivier Pietquin, Matt Sharifi, Dominik Roblek, Olivier Teboul, David Grangier, Marco Tagliasacchi, and Neil Zeghidour. 2023. AudioLM: a Language Modeling Approach to Audio Generation. arXiv:2209.03143 [cs.SD]
2023 arXiv
-
[12]
Alessandro Carlini, Camille Bordeau, and Maxime Ambard. 2024. Auditory localization: a comprehensive practical review. Frontiers Psycholo. (2024)
2024
-
[13]
Ishan Chatterjee, Maruchi Kim, Vivek Jayaram, Shyamnath Gollakota, Ira Kemel- macher, Shwetak Patel, and Steven M. Seitz. 2022. ClearBuds: wireless binaural earbuds for learning-based speech enhancement. InProceedings of the 20th Annual International Conference on Mobile Syste...
2022
-
[14]
Sanyuan Chen, Chengyi Wang, Zhengyang Chen, Yu Wu, Shujie Liu, Zhuo Chen, Jinyu Li, Naoyuki Kanda, Takuya Yoshioka, Xiong Xiao, Jian Wu, Long Zhou, Shuo Ren, Yanmin Qian, Yao Qian, Jian Wu, Michael Zeng, Xiangzhan Yu, and Furu Wei. 2022. WavLM: Large-Scale Self-Supervised Pre-...
2022
-
[15]
Tuochao Chen, Malek Itani, Sefik Eskimez, Takuya Yoshioka, and Shyamnath Gollakota. 2024. Hearable devices with sound bubbles. Nature Electronics 7 (11 2024), 1047–1058. https://doi.org/10.1038/s41928-024-01276-z
2024 doi
-
[16]
Tuochao Chen, Qirui Wang, Bohan Wu, Malek Itani, Sefik Emre Eskimez, Takuya Yoshioka, and Shyamnath Gollakota. 2024. Target conversation extraction: Source separation using turn-taking dynamics. In InterSpeech
2024
-
[17]
Shanbo Cheng, Zhichao Huang, Tom Ko, Hang Li, Ningxin Peng, Lu Xu, and Qini Zhang. 2024. Towards Achieving Human Parity on End-to-end Simultaneous Speech Translation via LLM Agent. arXiv:2407.21646 [cs.CL] https://arxiv.org/ abs/2407.21646
2024 arXiv
-
[18]
Kyunghyun Cho and Masha Esipova. 2016. Can neural machine translation do simultaneous translation? arXiv:1606.02012 [cs.CL]
2016 arXiv
-
[19]
Seamless Communication, Loïc Barrault, Yu-An Chung, Mariano Coria Megli- oli, David Dale, Ning Dong, Mark Duppenthaler, Paul-Ambroise Duquenne, Brian Ellis, Hady Elsahar, Justin Haaheim, John Hoffman, Min-Jae Hwang, Hiro- fumi Inaguma, Christopher Klaiber, Ilia Kulikov, Pengwe...
2023 arXiv
-
[20]
Samuele Cornell, Zhong-Qiu Wang, Yoshiki Masuyama, Shinji Watanabe, Manuel Pariente, and Nobutaka Ono. 2023. Multi-Channel Target Speaker Extraction CHI ’25, April 26-May 1, 2025, Yokohama, Japan Chen, Wang, He and Gollakota with Refinement: The WavLab Submission to the Second...
2023 arXiv
-
[21]
Fahim Dalvi, Nadir Durrani, Hassan Sajjad, and Stephan Vogel. 2018. Incremental Decoding and Training Methods for Simultaneous Translation in Neural Machine Translation. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Li...
2018
-
[22]
Qianqian Dong, Zhiying Huang, Qiao Tian, Chen Xu, Tom Ko, Yunlong Zhao, Siyuan Feng, Tang Li, Kexin Wang, Xuxin Cheng, Fengpeng Yue, Ye Bai, Xi Chen, Lu Lu, Zejun Ma, Yuping Wang, Mingxuan Wang, and Yuxuan Wang. 2023. PolyVoice: Language Models for Speech to Speech Translation...
2023 arXiv
-
[23]
Paul-Ambroise Duquenne, Hongyu Gong, Ning Dong, Jingfei Du, Ann Lee, Vedanuj Goswami, Changhan Wang, Juan Pino, Benoît Sagot, and Holger Schwenk. 2023. SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech Translations. In Proceedings of the 61st Annual Mee...
2023
-
[24]
Maha Elbayad, Laurent Besacier, and Jakob Verbeek. 2020. Efficient Wait-k Models for Simultaneous Machine Translation. In InterSpeech
2020
-
[25]
Sefik Emre Eskimez, Takuya Yoshioka, Huaming Wang, Xiaofei Wang, Zhuo Chen, and Xuedong Huang. 2022. Personalized speech enhancement: new models and Comprehensive evaluation. In ICASSP
2022
-
[26]
Qingkai Fang, Zhengrui Ma, Yan Zhou, Min Zhang, and Yang Feng. 2024. CTC- based Non-autoregressive Textless Speech-to-Speech Translation. In Findings of the Association for Computational Linguistics: ACL 2024
2024
-
[27]
Ritwik Giri, Shrikant Venkataramani, Jean-Marc Valin, Umut Isik, and Arvindh Krishnaswamy. 2021. Personalized PercepNet: Real-time, Low-complexity Target Voice Separation and Enhancement. In InterSpeech
2021
-
[28]
Jiatao Gu, James Bradbury, Caiming Xiong, Victor OK Li, and Richard Socher
-
[29]
Anmol Gulati, James Qin, Chung-Cheng Chiu, Niki Parmar, Yu Zhang, Jiahui Yu, Wei Han, Shibo Wang, Zhengdong Zhang, Yonghui Wu, and Ruoming Pang
-
[30]
Cong Han, Yi Luo, and Nima Mesgarani. 2020. Real-time binaural speech separa- tion with preserved spatial cues. In ICASSP). IEEE
2020
-
[31]
IoSR-Surrey. 2016. IoSR-surrey/realroombrirs: Binaural impulse responses cap- tured in real rooms. https://github.com/IoSR-Surrey/RealRoomBRIRs
2016
-
[32]
IoSR-Surrey. 2023. Simulated Room Impulse Responses. https://iosr.uk/software/ index.php
2023
-
[33]
Malek Itani, Tuochao Chen, Takuya Yoshioka, and Shyamnath Gollakota. 2023. Creating speech zones with self-distributing acoustic swarms. Nature Communi- cations 14 (09 2023). https://doi.org/10.1038/s41467-023-40869-8
2023 doi
-
[34]
Vivek Jayaram, Ira Kemelmacher-Shlizerman, and Steven M. Seitz. 2023. HRTF Estimation in the Wild. In UIST. ACM
2023
-
[35]
Teerapat Jenrungrot, Vivek Jayaram, Steve Seitz, and Ira Kemelmacher- Shlizerman. 2020. The Cone of Silence: Speech Separation by Localization. In Advances in Neural Information Processing Systems
2020
-
[36]
Chunyu Kit and Tak Ming Wong. 2008. Comparative Evaluation of Online Machine Translation Systems with Legal Texts. Law Library Journal 100 (2008), 299–321. https://api.semanticscholar.org/CorpusID:13481458
2008
-
[37]
Jungil Kong, Jaehyeon Kim, and Jaekyoung Bae. 2020. Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis. Advances in neural information processing systems 33 (2020), 17022–17033
2020
-
[38]
Matthew Le, Apoorv Vyas, Bowen Shi, Brian Karrer, Leda Sari, Rashel Moritz, Mary Williamson, Vimal Manohar, Yossi Adi, Jay Mahadeokar, and Wei-Ning Hsu. 2024. Voicebox: text-guided multilingual universal speech generation at scale. In Proceedings of the 37th International Conf...
2024
-
[39]
Marie Lebert. 2022. A short history of translation through the ages. https://www.iapti.org/iaptiarticle/a-short-history-of-translation-through-the- ages-marie-lebert-2/
2022
-
[40]
Ann Lee, Peng-Jen Chen, Changhan Wang, Jiatao Gu, Sravya Popuri, Xutai Ma, Adam Polyak, Yossi Adi, Qing He, Yun Tang, Juan Pino, and Wei-Ning Hsu. 2022. Direct Speech-to-Speech Translation With Discrete Units. In Proceedings of the 60th Annual Meeting of the Association for Co...
2022
-
[41]
Tong Lei, Zhongshu Hou, Yuxiang Hu, Wanyu Yang, Tianchi Sun, Xiaobin Rong, Dahan Wang, Kai Chen, and Jing Lu. 2023. A Low-Latency Hybrid Multi-Channel Speech Enhancement System For Hearing Aids. In ICASSP. 1–2
2023
-
[42]
Levinson
Stephen C. Levinson. 2016. Turn-taking in Human Communication – Origins and Implications for Language Processing. Trends in Cog. Sci. (2016)
2016
-
[43]
Liebling, Katherine Heller, Samantha Robertson, and Wesley Hanwen Deng
Daniel J. Liebling, Katherine Heller, Samantha Robertson, and Wesley Hanwen Deng. 2022. Opportunities for Human-centered Evaluation of Machine Transla- tion Systems. In NAACL-HLT
2022
-
[44]
Mingbo Ma, Liang Huang, Hao Xiong, Renjie Zheng, Kaibo Liu, Baigong Zheng, Chuanqiang Zhang, Zhongjun He, Hairong Liu, Xing Li, Hua Wu, and Haifeng Wang. 2019. STACL: Simultaneous Translation with Implicit Anticipation and Controllable Latency using Prefix-to-Prefix Framework....
2019
-
[45]
Xutai Ma, Juan Pino, James Cross, Liezl Puzon, and Jiatao Gu. 2020. Monotonic Multihead Attention. In ICLR
2020
-
[46]
Evgeny Matusov, Stephan Kanthak, and Hermann Ney. 2005. On the Integration of Speech Recognition and Statistical Machine Translation. 3177–3180. https: //doi.org/10.21437/Interspeech.2005-726
2005 doi
-
[47]
Tobias May, Steven Van De Par, and Armin Kohlrausch. 2010. A probabilistic model for robust localization based on a binaural auditory front-end. IEEE Transactions on audio, speech, and language processing 19, 1 (2010), 1–13
2010
-
[48]
Mymanu. 2024. Mymanu Click S. https://mymanu.com/products/mymanu-clik-s
2024
-
[49]
H. Ney. 1999. Speech translation: coupling of recognition and translation. In ICASSP99, Vol. 1
1999
-
[50]
NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia-Gonzalez, Prang...
2022
-
[51]
NPR. 2023. Finding your place in the galaxy with the help of Star Trek. https: //www.npr.org/2023/10/14/1205714903/star-trek
2023
-
[52]
All Things Considered NPR. 1998. Babelfish, a Translator Inspired by ’The Hitchhiker’s Guide’. https://www.npr.org/1998/02/12/1036190/babelfish-a- translator-inspired-by-the-hitchhikers-guide
1998
-
[53]
Lucas Nunes Vieira, Minako O’Hagan, and Carol O’Sullivan. 2020. Understanding the societal impacts of machine translation: a critical review of the literature on medical and legal use cases. Information Communication and Society (06 2020). https://doi.org/10.1080/1369118X.2020.1776370
2020
-
[54]
Sravya Popuri, Peng-Jen Chen, Changhan Wang, Juan Pino, Yossi Adi, Jiatao Gu, Wei-Ning Hsu, and Ann Lee. 2022. Enhanced direct speech-to-speech translation using self-supervised pre-training and data augmentation. In InterSpeech
2022
-
[55]
Santhosh Kumar, Adam Lopez, Damianos G
Matt Post, G. Santhosh Kumar, Adam Lopez, Damianos G. Karakos, Chris Callison- Burch, and Sanjeev Khudanpur. 2013. Improved speech-to-text translation with the Fisher and Callhome Spanish-English speech translation corpus. In Interna- tional Workshop on Spoken Language Translation
2013
-
[56]
Liu, Ron J
Colin Raffel, Minh-Thang Luong, Peter J. Liu, Ron J. Weiss, and Douglas Eck
-
[57]
Yi Ren, Jinglin Liu, Xu Tan, Chen Zhang, Tao Qin, Zhou Zhao, and Tie-Yan Liu
-
[58]
Dario Rethage, Jordi Pons, and Xavier Serra. 2018. A Wavenet for Speech De- noising. In ICASSP
2018
-
[59]
Liebling, Michal Lahav, Katherine Heller, Mark Díaz, Samy Bengio, and Niloufar Salehi (Eds.)
Samantha Robertson, Wesley Deng, Timnit Gebru, Margaret Mitchell, Daniel J. Liebling, Michal Lahav, Katherine Heller, Mark Díaz, Samy Bengio, and Niloufar Salehi (Eds.). 2021. Three Directions for the Design of Human-Centered Machine Translation
2021
-
[60]
In Proceedings of the 34th International Conference on Machine Learning - Volume 70 (Sydney, NSW, Australia)(ICML’17)
Online and linear-time attention by enforcing monotonic alignments. In Proceedings of the 34th International Conference on Machine Learning - Volume 70 (Sydney, NSW, Australia)(ICML’17). JMLR.org, 2837–2846
-
[61]
SDK. 2023. Steam Audio. https://valvesoftware.github.io/steam-audio/
2023
-
[62]
In Annual Meeting of the Association for Computational Linguistics
SimulSpeech: End-to-End Simultaneous Speech to Text Translation. In Annual Meeting of the Association for Computational Linguistics
-
[63]
Kai Shen, Zeqian Ju, Xu Tan, Yanqing Liu, Yichong Leng, Lei He, Tao Qin, Sheng Zhao, and Jiang Bian. 2023. NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers. arXiv:2304.09116 [eess.AS] https://arxiv.org/abs/2304.09116
2023 arXiv
-
[64]
Matthias Sperber and Matthias Paulik. 2020. Speech Translation and the End-to- End Promise: Taking Stock of Where We Are. InAnnual Meeting of the Association for Computational Linguistics
2020
-
[65]
Paul K. Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna, Zalán Borsos, Félix de Chaumont Quitry, Peter Chen, Dalia El Badawy, Wei Han, Eugene Kharitonov, Hannah Muckenhirn, Dirk Padfield, James Qin, Danny Rozenberg, Tara Sainath, Johan Schalkwyk, Matt Sharif...
2023
-
[66]
Timekettle. [n. d.]. Timekettle WT2 Edge/W3 Real-time Translator Earbuds, 2-way simultaneous interpretation. https://www.timekettle.co/products/wt2- edge-online-voice-language-translator-earbuds Spatial Speech Translation: Translating Across Space With Binaural Hearables CHI ’...
2025
-
[67]
ShanonPearce. 2022. Shanonpearce/ash-listening-set: A dataset of filters for head- phone correction and binaural synthesis of spatial audio systems on headphones. https://github.com/ShanonPearce/ASH-Listening-Set/tree/main
2022
-
[68]
Bandhav Veluri, Justin Chan, Malek Itani, Tuochao Chen, Takuya Yoshioka, and Shyamnath Gollakota. 2023. Real-Time Target Sound Extraction. In ICASSP. 1–5
2023
-
[69]
Bandhav Veluri, Malek Itani, Justin Chan, Takuya Yoshioka, and Shyamnath Gollakota. 2023. Semantic Hearing: Programming Acoustic Scenes with Binaural Hearables. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (San Francisco, CA, USA) (...
2023
-
[70]
Cem Subakan, Mirco Ravanelli, Samuele Cornell, Mirko Bronzi, and Jianyuan Zhong. 2021. Attention Is All You Need In Speech Separation. In ICASSP 2021
2021
-
[71]
Anran Wang, Maruchi Kim, Hao Zhang, and Shyamnath Gollakota. 2022. Hybrid Neural Networks for On-device Directional Hearing. AAAI (2022)
2022
-
[72]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems , Vol. 30
2017
-
[73]
Peidong Wang, Eric Sun, Jian Xue, Yu Wu, Long Zhou, Yashesh Gaur, Shujie Liu, and Jinyu Li. 2022. LAMASSU: A Streaming Language-Agnostic Multilin- gual Speech Recognition and Translation Model Using Neural Transducers. In Interspeech. https://api.semanticscholar.org/CorpusID:258968116
2022
-
[74]
Waverly. 2024. Waverly labs Earbuds. https://www.waverlylabs.com/
2024
-
[75]
Bandhav Veluri, Malek Itani, Tuochao Chen, Takuya Yoshioka, and Shyamnath Gollakota. 2024. Look Once to Hear: Target Speech Hearing with Noisy Examples. In CHI (Honolulu, HI, USA) (CHI ’24). ACM, Article 37, 16 pages
2024
-
[76]
Zhongweiyang Xu and Romit Roy Choudhury. 2022. Learning to Separate Voices by Spatial Regions. ICML (2022)
2022
-
[77]
Chengyi Wang, Sanyuan Chen, Yu Wu, Ziqiang Zhang, Long Zhou, Shujie Liu, Zhuo Chen, Yanqing Liu, Huaming Wang, Jinyu Li, Lei He, Sheng Zhao, and Furu Wei. 2023. Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers. arXiv:2301.02111 [cs.CL]
2023 arXiv
-
[78]
Changhan Wang Jiatao Gu Juan Pino Xutai Ma, Mohammad Javad Dousti. 2020. Simuleval: An evaluation toolkit for simultaneous translation. In Proceedings of the EMNLP
2020
-
[79]
Mu Yang, Naoyuki Kanda, Xiaofei Wang, Junkun Chen, Peidong Wang, Jian Xue, Jinyu Li, and Takuya Yoshioka. 2024. Diarist: Streaming Speech Translation with Speaker Diarization. In ICASSP. 10866–10870
2024
-
[80]
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Łukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith St...
2016 arXiv
-
[81]
Shaolei Zhang and Yang Feng. 2021. Universal Simultaneous Machine Translation with Mixture-of-Experts Wait-k Policy. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics
2021
-
[82]
Jian Xue, Peidong Wang, Jinyu Li, Matt Post, and Yashesh Gaur. 2022. Large- Scale Streaming End-to-End Speech Translation with Neural Transducers. In Interspeech
2022
-
[85]
Shaolei Zhang, Qingkai Fang, Shoutao Guo, Zhengrui Ma, Min Zhang, and Yang Feng. 2024. StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
2024
-
[87]
Kateřina Žmolíková, Marc Delcroix, Keisuke Kinoshita, Takuya Higuchi, Atsunori Ogawa, and Tomohiro Nakatani. 2017. Speaker-Aware Neural Network Based Beamformer for Speaker Extraction in Speech Mixtures. InProc. Interspeech 2017
2017
-
[2017]
Non-autoregressive neural machine translation. (2017)
2017
-
[2020]
In InterSpeech
Conformer: Convolution-augmented Transformer for Speech Recognition. In InterSpeech
-
[2022]
In InterSpeech
XLS-R: Self-supervised cross-lingual speech representation learning at scale. In InterSpeech
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.