Pith. sign in

REVIEW 3 major objections 5 minor 3 cited by

DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition

T0 review · 3 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read The paper shows that conditioning Whisper on frame-level diarization labels—silence, target, non-target, overlap—rather than speaker embeddings, produces a target-speaker ASR system that matches or beats previous methods on AMI…

desk verdict Diarization-conditioned Whisper for TS-ASR is a genuine contribution, but the paper overstates real-world robustness and cherry-picks per-dataset variants. read the letter →

arxiv 2501.00114 v1 pith:UORIPP7D submitted 2024-12-30 eess.AS cs.SD

classification eess.AScs.SD
keywords Diarization-ConditionedWhisperTarget-SpeakerASRSpeakerDiarizationSTNOmasksAdaptationFrame-LevelDiarization-DependentTransformationsQuery-KeyBiasingMulti-Speaker
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

DiCoW tries to make a pre-trained single-speaker ASR model do target-speaker transcription by conditioning on who is speaking when, instead of on what the target speaker sounds like. For each target speaker, diarization outputs are collapsed into a four-way per-frame label — silence, target-only, non-target, and overlap — and injected into Whisper through attention biasing (QKb) and per-layer frame-wise affine transformations (FDDT). The paper reports that this matches or beats prior speaker-embedding and enrollment-based systems on AMI, NOTSOFAR-1, Libri2Mix, and LibriCSS, while leaving single-speaker performance close to the original Whisper. If this holds, meeting transcription becomes simpler and more robust to unseen speakers, because no component has to learn a mapping from speaker-embedding space to ASR space.

What carries the argument

The load-bearing object is the STNO mask: for a target speaker $s_k$, each frame $t$ is assigned a probability vector $M_t = (p^S_t,\ p^T_t,\ p^N_t,\ p^O_t)^\top$ whose entries are the probabilities of silence, target-only, non-target, and overlap, computed from the diarization matrix $D\in[0,1]^{S\times T}$. Two mechanisms carry the conditioning: QKb extends the attention query and key with a constant that initially subtracts a bias $c$ from non-target frames' attention scores, and FDDT replaces each encoder layer's input frame $z^l_t$ with a convex combination of four trainable affine maps $W^l_c z^l_t + b^l_c$ weighted by the STNO probabilities. The suppressive initialization (zeroing $W^l_S$ and $W^l_N$, identity for target and overlap) keeps the pre-trained model intact at the start of fine-tuning. A Co-Attention module lets the per-speaker decoding channels exchange information and resolve which instance decodes overlapping speech.

What would settle it

Take the released code, run inference on the reported test sets while randomly corrupting oracle diarization labels at rates from 5% to 30% (missed speech and speaker confusions), and check whether tcpWER rises in proportion to the corruption; the paper's assumption predicts a sharp rise with no recovery of missed segments.

Watch

Extended reading notes

Core claim

The paper's central claim is that target-speaker ASR can be driven by diarization activity labels instead of speaker embeddings. For each target speaker, DiCoW converts diarization probabilities into a fixed-size STNO mask—per-frame probabilities of silence, target speaker only, non-target speaker(s), and overlap—and feeds this mask into Whisper's encoder via frame-level diarization-dependent transformations (FDDT) and query-key biasing (QKb). With ground-truth diarization, the authors report 17.2 cpWER on AMI-sdm, 19.7 tcpWER on NOTSOFAR-1 eval-small, 4.4 cpWER on Libri2Mix test-clean, and 5.6 cpWER on LibriCSS test; with automatic diarization from the DiariZen system, the corresponding numbers are 23.6, 33.5, 6.0, and 8.5. They state that DiCoW achieves the best results to date on Libri2Mix and LibriCSS for both real and oracle diarization, and that the same FDDT mechanism transfers to a non-Whisper Branchformer model.

Load-bearing premise

The model is trained only on perfect, hard diarization labels, so everything depends on the diarizer at test time producing labels close enough to the oracle ones; when it misses or confuses speech, DiCoW has no learned way to recover.

Editorial extensions

If this is right

  • Meeting transcription can be organized as one diarization pass followed by parallel Whisper decoders per speaker, with no enrollment speech and no speaker-embedding conditioning.
  • Systems trained this way inherit Whisper's single-speaker robustness: on LibriSpeech, TED-LIUM, and VoxPopuli the fine-tuned model stays within about one absolute WER point of unmodified Whisper.
  • Because diarization activity is a generic frame-level signal, the same STNO injection should apply to any pre-trained encoder-decoder ASR; the paper's Branchformer experiment supports this.
  • Adding the CTC head and hybrid CTC/attention decoding improves target-speaker accuracy on NOTSOFAR-1 beyond what the Whisper decoder alone achieves, suggesting the conditioning and alignment benefits are additive.
  • The gap between oracle and automatic diarization (e.g. 19.7 to 33.5 tcpWER on NOTSOFAR-1) becomes the main remaining cost, making diarization quality the bottleneck rather than ASR.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension the paper leaves implicit is to train on soft STNO probabilities or to corrupt oracle labels with miss and confusion noise during training; the 13.8-point tcpWER drop on NOTSOFAR-1 suggests that hard-label-only training is the main recoverable loss, and we would expect randomized label corruption to close most of it.
  • Because FDDT is essentially a per-frame, per-class affine mixture, it could be reused as a generic class-conditional conditioning layer for other frame-labeled speech tasks (e.g., language ID, emotion, or source-type conditioning), not just speaker activity.
  • The Co-Attention module's gains appeared mainly on Libri2Mix, where speech is fully overlapped; this suggests that future work on high-overlap meetings should scale up speaker-interaction layers, while for typical meeting overlap rates the per-speaker independent decoders are already near their ceiling.
  • If the oracle-label results hold across more languages and acoustic conditions, the method could make speaker-attributed transcription a post-processing step on top of any off-the-shelf ASR, since the only external requirement is a diarizer that outputs per-frame activity.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The manuscript proposes DiCoW, an adaptation of Whisper for target-speaker ASR in which frame-level speaker diarization probabilities are converted into Silence/Target/Non-target/Overlap (STNO) masks. Two conditioning mechanisms are presented: query-key biasing (QKb), which biases attention scores away from non-target frames via an extended query/key formulation, and frame-level diarization-dependent transformations (FDDT), which apply class-specific affine transforms to encoder hidden states. A CTC head with joint CTC/attention decoding and a Co-Attention module for multi-speaker interaction are also studied. Experiments fine-tune Whisper-large-v3-turbo on AMI, NOTSOFAR-1, and Libri2Mix and evaluate on these plus LibriCSS, comparing oracle and automatic diarization, with a Branchformer experiment for generality. The paper reports state-of-the-art or competitive target-speaker WERs under oracle diarization, along with a substantial degradation when automatic diarization is used, and concludes that the method maintains single-speaker performance.

Significance. If the results are reproducible, the idea of conditioning on diarization outputs rather than speaker embeddings is a useful and economical direction for TS-ASR, with several praiseworthy elements: the code is released; the evaluation spans real (AMI, NOTSOFAR-1) and synthetic corpora; the comparison between oracle and system diarization is presented explicitly; and the Branchformer experiment supports transferability beyond Whisper. The paper is also unusually honest about limitations. However, several headline claims go beyond what the experiments establish, particularly regarding real-diarization robustness, the 'best results to date' statement, and the preservation of single-speaker accuracy.

major comments (3)
  1. [§6.3, Tables 6–7] The central claim of real-world applicability is not established for the automatic-diarization setting. The model is trained only on oracle hard STNO labels (Section 5.3) and decoded with hard DiariZen labels (Section 5.4 and Section 6.3), so it has never seen miss, false-alarm, or confusion errors; Section 6.3 explicitly states that the system 'has not encountered such cases during training' and that the FDDT suppressive initialization makes missed speech unrecoverable. Table 6 shows tcpWER on NOTSOFAR-1 eval-small rising from 19.7% with oracle labels to 33.5% with DiariZen labels, with tcORC-WER rising from 19.1% to 22.6%, and LibriCSS worsening from 8.8% to 11.0% tcpWER. Because the abstract and introduction claim 'more reliable transcription in real-world multi-speaker recordings' and 'even when automatic diarization is used' (Section 1), the authors should either add a training scheme with simulated diarization errors or soft labels, or substantially temper these claims to oracle-conditioned performance.
  2. [§6.1, Table 3] The 'best results to date' claims are not supported by the comparison as presented. Table 3 selects a different system variant for different datasets (MD FDDT for AMI/NOTSOFAR-1/LibriCSS and SD+Co-Attention for Libri2Mix), combines oracle and system-diarization rows, uses different metrics per dataset (cpWER for AMI, tcpWER/tcORC-WER for NOTSOFAR-1, cpWER for Libri2Mix/LibriCSS), includes baseline numbers marked with a dagger that are not directly comparable because they use utterance-group scoring, and reports some ORC-WER values as collar-inflated approximations marked with a star. Under these conditions, the statement that 'The proposed system also achieves the best results to date on the Libri2Mix and LibriCSS datasets for both real- and ground-truth diarization' overreaches. Please provide a consistent protocol, or explicitly restrict the claim to the configurations and metrics that are directly comparable.
  3. [§6.5, Table 9] The abstract's claim that DiCoW maintains Whisper's accuracy and robustness on single-speaker data is stronger than the evidence. With the proposed model and lambda=0.2, TED-LIUM WER is 7.8% versus 4.3% for Whisper with beam size 5, and VoxPopuli is 11.2% versus 10.0%; even without CTC rescoring (lambda=0.0) the model gives 5.0% versus 4.3% and 11.0% versus 10.0%. The conclusion's phrase 'does not substantially degrade' is defensible, but the abstract should be qualified, for example by saying that accuracy is largely preserved on LibriSpeech while some degradation occurs on out-of-domain single-speaker sets, or the authors should provide a model variant that better preserves generalization.
minor comments (5)
  1. [§6.4, Table 8] The description of lambda=1.0 as autoregressive decoding with the top 1000 tokens rescored by CTC contradicts Eq. (3), where lambda=1.0 would mean pure CTC decoding; please clarify the actual decoding objective or rename the parameter.
  2. [§4.3, Eq. (10)] The bias constant c is introduced as a nonnegative value but the non-target score is -c; please define the sign convention explicitly and state whether the same c is used for all layers and heads and how it was selected.
  3. [§4.4, Eq. (13)] The superscript and subscript notation for layer indices is inconsistent (z_l^t versus hat-z_t^l), which makes the FDDT equations harder to follow; a single consistent indexing convention would improve readability.
  4. [§5.3] The spelling 'Librispeech' appears alongside 'LibriSpeech' in the same section, and the description of the CTC preheat stage could state more precisely that monitoring is done on LibriSpeech dev-clean and dev-other while training uses the 960h training set.
  5. [§6.1, Table 3] For the starred ORC-WER approximations obtained by increasing the time collar, the table should state the collar value used for each entry, since a very large collar can make the approximation too loose for meaningful comparison with published ORC-WER numbers.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: DiCoW's headline results rest on held-out test-set evaluations and public benchmark comparisons, not on fitted targets or a load-bearing self-citation chain.

full rationale

The derivation chain is self-contained. DiCoW's central results (Tables 3, 4, 6) are held-out test-set evaluations on AMI, NOTSOFAR-1, Libri2Mix, and LibriCSS against published baselines; no prediction is obtained by construction from the conditioning labels. The STNO masks (Eqs. 4-7) are inputs to the model, not outputs of the ASR system; FDDT (Eq. 13) is a learned transformation with parameters optimized on training folds and evaluated on disjoint test partitions; QK biasing (Eqs. 10-12) is an architectural modification with an empirical constant, not a fitted target quantity. The self-citations in the paper ([35] for the suppressive initialization, [42] for the DiariZen diarizer, [53] for the CHiME-8 system description) are component/design references rather than unverified premises used to force the central result. The paper's own limitation analysis (Sections 6.3 and 7) explicitly acknowledges that oracle-to-automatic diarization transfer is fragile, with tcpWER on NOTSOFAR-1 eval-small rising from 19.7% to 33.5%; this is a robustness or generalization concern about the system's real-world assumptions, not circularity. The selection of 'best-performing variants' in Section 6.1 is a statistical reporting caveat about multiple configurations, not a reduction of a prediction to its input.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

The central claim rests on three domain assumptions: the validity of the STNO decomposition, the transfer from oracle to automatic diarization, and the preservation of Whisper pretraining. Two free hyperparameters are fitted empirically. No new entities are postulated.

free parameters (2)
  • QK biasing constant c = 50
    Empirically chosen initial bias for non-target frames in QK biasing (Section 5.3); affects hallucination behavior and the approximation of a hard attention mask.
  • CTC/attention decoding weight lambda = 0.2 (decoding), 0.3 (training loss)
    Interpolation weight in Eq. 3 for joint CTC/attention decoding; lambda=0.2 selected for best tcpWER on NOTSOFAR-1 (Table 8), and lambda=0.3 used in training as in [30].
assumptions (3)
  • domain assumption Per-speaker diarization posteriors d(s,t) are independent, so STNO probabilities in Eqs. (4)-(7) are valid.
    The four class probabilities are computed as products and marginals over speakers; if posteriors are correlated (e.g., during overlap), the decomposition is a modeling approximation. The paper uses hard labels in practice, which avoids the issue but leaves the soft formulation untested.
  • domain assumption Ground-truth hard diarization labels used in training are representative of the hard labels produced by the automatic DiariZen system at inference.
    Section 5.3 trains on oracle labels while Section 6.3 evaluates with DiariZen; Table 6 shows a substantial tcpWER drop, indicating this assumption is only partially satisfied and is the weakest link.
  • domain assumption Adding affine transforms and bias-extension modules to Whisper and fine-tuning preserves the pre-trained model's linguistic knowledge.
    The method relies on Whisper's pretraining for single-speaker robustness; Table 9 shows partial degradation on out-of-domain sets (TED-LIUM, VoxPopuli), so the preservation is imperfect.

how reviews work

0 comments
Cite this review

Pith. "Pith review of DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition." pith.science (2026). https://pith.science/paper/UORIPP7D

@misc{pith2026250100114,
  author       = {Pith},
  title        = {Pith review of: DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UORIPP7D}},
  note         = {Machine review of arXiv:2501.00114}
}
read the original abstract

Speaker-attributed automatic speech recognition (ASR) in multi-speaker environments remains a significant challenge, particularly when systems conditioned on speaker embeddings fail to generalize to unseen speakers. In this work, we propose Diarization-Conditioned Whisper (DiCoW), a novel approach to target-speaker ASR that leverages speaker diarization outputs as conditioning information. DiCoW extends the pre-trained Whisper model by integrating diarization labels directly, eliminating reliance on speaker embeddings and reducing the need for extensive speaker-specific training data. Our method introduces frame-level diarization-dependent transformations (FDDT) and query-key biasing (QKb) techniques to refine the model's focus on target speakers while effectively handling overlapping speech. By leveraging diarization outputs as conditioning signals, DiCoW simplifies the workflow for multi-speaker ASR, improves generalization to unseen speakers and enables more reliable transcription in real-world multi-speaker recordings. Additionally, we explore the integration of a connectionist temporal classification (CTC) head to Whisper and demonstrate its ability to improve transcription efficiency through hybrid decoding. Notably, we show that our approach is not limited to Whisper; it also provides similar benefits when applied to the Branchformer model. We validate DiCoW on real-world datasets, including AMI and NOTSOFAR-1 from CHiME-8 challenge, as well as synthetic benchmarks such as Libri2Mix and LibriCSS, enabling direct comparisons with previous methods. Results demonstrate that DiCoW enhances the model's target-speaker ASR capabilities while maintaining Whisper's accuracy and robustness on single-speaker data.

Figures

Figures reproduced from arXiv: 2501.00114 by the authors.

Figure 1
Figure 1. Proposed Diarization-Conditioned Whisper model with STNO mask example. [PITH_FULL_IMAGE:figures/full_fig_p012_1.png] view at source ↗
Figure 2
Figure 2. Scheme of Co-Attention module. Dotted lines depict affine transformation [PITH_FULL_IMAGE:figures/full_fig_p014_2.png] view at source ↗
Figure 3
Figure 3. Test TCP-WER as a function of training steps for the NOTSOFAR-1 model [PITH_FULL_IMAGE:figures/full_fig_p024_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Pretraining Multi-Speaker Identification for Neural Speaker Diarization

    eess.AS 2025-05 conditional novelty 6.0 of 10

    Pretraining an encoder to identify multiple speakers from fully overlapped mixtures yields accurate local diarization without simulated conversational data.

  2. MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses

    eess.AS 2025-07 reject novelty 5.0 of 10

    MMW combines a Mamba-based Mix Block, a Frame Diarization Mamba layer, and multi-scale GRPO to reduce side-talk interference in Whisper ASR, reporting WER as low as 3.71%.

  3. SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition

    cs.SD 2025-06 conditional novelty 5.0 of 10

    Conditioning an SOT multi-talker ASR decoder on EEND-EDA speaker embeddings and activity information lowers WER on Libri2Mix and Libri3Mix, provided the diarization branch is accurate.

Reference graph

Works this paper leans on

55 extracted references · 34 canonical work pages · cited by 3 Pith papers

  1. [1]

    Li, et al., Recent advances in end-to-end automatic speech recogni- tion, APSIPA Transactions on Signal and Information Processing 11 (1) (2022)

    J. Li, et al., Recent advances in end-to-end automatic speech recogni- tion, APSIPA Transactions on Signal and Information Processing 11 (1) (2022)

  2. [2]

    Watanabe, M

    S. Watanabe, M. Mandel, J. Barker, E. Vincent, A. Arora, X. Chang, S. Khudanpur, V. Manohar, D. Povey, D. Raj, et al., CHiME-6 chal- lenge: Tackling multispeaker speech recognition for unsegmented record- ings, arXiv preprint arXiv:2004.09249 (2020). 31

  3. [4]

    Cornell, M

    S. Cornell, M. S. Wiesner, S. Watanabe, D. Raj, X. Chang, P. Garcia, Y. Masuyam, Z.-Q. Wang, S. Squartini, S. Khudanpur, The CHiME- 7 DASR Challenge: Distant Meeting Transcription with Multiple De- vices in Diverse Scenarios, in: 7th International Workshop on Speech Processing in Everyday Environments (CHiME 2023), 2023, pp. 1–6. doi:10.21437/CHiME.2023-1

  4. [5]

    Cornell, T

    S. Cornell, T. J. Park, H. Huang, C. Boeddeker, X. Chang, M. Maciejew- ski, M. S. Wiesner, P. Garcia, S. Watanabe, The CHiME-8 DASR Chal- lenge for Generalizable and Array Agnostic Distant Automatic Speech Recognition and Diarization, in: 8th International Workshop on Speech Processing in Everyday Environments (CHiME 2024), 2024, pp. 1–6. doi:10.21437/CHi...

  5. [6]

    Bhandari, D

    N. Bhandari, D. Chen, M. A. del R ´ ıo Fern´ andez, N. Delworth, J. D. Fox, M. Jett´ e, Q. McNamara, C. Miller, O. Novotn´ y, J. Profant, N. Qin, M. Ratajczak, J.-P. Robichaud, Reverb: Open-Source ASR and Diariza- tion from Rev, arXiv preprint arXiv:2410.03930 (2024)

  6. [7]

    Yoshioka, I

    T. Yoshioka, I. Abramovski, C. Aksoylar, Z. Chen, M. David, D. Dim- itriadis, Y. Gong, I. Gurvich, X. Huang, Y. Huang, A. Hurvitz, L. Jiang, S. Koubi, E. Krupka, I. Leichter, C. Liu, P. Parthasarathy, A. Vinnikov, L. Wu, X. Xiao, W. Xiong, H. Wang, Z. Wang, J. Zhang, Y. Zhao, T. Zhou, Advances in Online Audio-Visual Meeting Transcription, in: 2019 IEEE Au...

  7. [8]

    D. Raj, P. Denisov, Z. Chen, H. Erdogan, Z. Huang, M. He, S. Watan- abe, J. Du, T. Yoshioka, Y. Luo, N. Kanda, J. Li, S. Wisdom, J. R. Hershey, Integration of Speech Separation, Diarization, and Recogni- tion for Multi-Speaker Meetings: System Description, Comparison, and 32 Analysis, in: 2021 IEEE Spoken Language Technology Workshop (SLT), 2021, pp. 897–...

  8. [9]

    Kanda, S

    N. Kanda, S. Horiguchi, R. Takashima, Y. Fujita, K. Nagamatsu, S. Watanabe, Auxiliary interference speaker loss for target-speaker speech recognition, Proceedings of the Annual Conference of the In- ternational Speech Communication Association, INTERSPEECH 2019- September (2019) 236–240. doi:10.21437/Interspeech.2019-1126

Show all 55 references
  1. [10]

    Karafi´ at, L

    M. Karafi´ at, L. Burget, P. Matˇ ejka, O. Glembek, J.ˇCernock` y, iVector- based discriminative adaptation for automatic speech recognition, in: 2011 IEEE Workshop on Automatic Speech Recognition & Understand- ing, IEEE, 2011, pp. 152–157

  2. [12]

    Dehak, P

    N. Dehak, P. J. Kenny, R. Dehak, P. Dumouchel, P. Ouellet, Front-end factor analysis for speaker verification, IEEE Transactions on Audio, Speech, and Language Processing 19 (4) (2010) 788–798

  3. [13]

    Snyder, D

    D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, S. Khudanpur, X- vectors: Robust dnn embeddings for speaker recognition, in: 2018 IEEE international conference on acoustics, speech and signal processing (ICASSP), IEEE, 2018, pp. 5329–5333

  4. [14]

    H. Wang, C. Liang, S. Wang, Z. Chen, B. Zhang, X. Xiang, Y. Deng, Y. Qian, Wespeaker: A research and production oriented speaker embed- ding learning toolkit, in: IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2023, pp. 1–5

  5. [15]

    Radford, J

    A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, I. Sutskever, Robust speech recognition via large-scale weak supervision, in: Interna- tional conference on machine learning, PMLR, 2023, pp. 28492–28518

  6. [16]

    Vinnikov, A

    A. Vinnikov, A. Ivry, A. Hurvitz, I. Abramovski, S. Koubi, I. Gurvich, S. Peer, X. Xiao, B. M. Elizalde, N. Kanda, X. Wang, S. Shaer, S. Yagev, 33 Y. Asher, S. Sivasankaran, Y. Gong, M. Tang, H. Wang, E. Krupka, NOTSOF AR-1 Challenge: New Datasets, Baseline, and Tasks for Dis-...

  7. [17]

    Mccowan, J

    I. Mccowan, J. Carletta, W. Kraaij, S. Ashby, S. Bourban, M. Flynn, M. Guillemot, T. Hain, J. Kadlec, V. Karaiskos, M. Kronenthal, G. Lathoud, M. Lincoln, A. Lisowska Masson, W. Post, D. Reidsma, P. Wellner, The ami meeting corpus, Int’l. Conf. on Methods and Tech- niques in B...

  8. [18]

    Cosentino, M

    J. Cosentino, M. Pariente, S. Cornell, A. Deleforge, E. Vincent, Lib- riMix: An Open-Source Dataset for Generalizable Speech Separation, arXiv: Audio and Speech Processing (2020). URL https://api.semanticscholar.org/CorpusID:218862876

  9. [19]

    Kinoshita, M

    K. Kinoshita, M. Delcroix, N. Tawara, Integrating end-to-end neural and clustering-based diarization: Getting the best of both worlds, in: Proc. ICASSP, IEEE, 2021, pp. 7198–7202

  10. [20]

    Bredin, pyannote

    H. Bredin, pyannote. audio 2.1 speaker diarization pipeline: principle, benchmark, and recipe, in: Proc. Interspeech 2023, 2023, pp. 1983–1987

  11. [21]

    Z. Chen, T. Yoshioka, L. Lu, T. Zhou, Z. Meng, Y. Luo, J. Wu, X. Xiao, J. Li, Continuous speech separation: Dataset and analysis, in: ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Sig- nal Processing (ICASSP), IEEE, 2020, pp. 7284–7288

  12. [22]

    L. E. Shafey, H. Soltau, I. Shafran, Joint Speech Recognition and Speaker Diarization via Sequence Transduction, in: Interspeech 2019, 2019, pp. 396–400. doi:10.21437/Interspeech.2019-1943

  13. [23]

    Kanda, X

    N. Kanda, X. Xiao, Y. Gaur, X. Wang, Z. Meng, Z. Chen, T. Yoshioka, Transcribe-to-diarize: Neural speaker diarization for unlimited number of speakers using end-to-end speaker-attributed ASR, in: ICASSP 2022- 2022 IEEE International Conference on Acoustics, Speech and Signal P...

  14. [24]

    Cornell, J.-w

    S. Cornell, J.-w. Jung, S. Watanabe, S. Squartini, One Model to Rule Them All? Towards End-to-End Joint Speaker Diarization and Speech 34 Recognition, in: ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2024, pp. 11856–11860

  15. [25]

    H. Ma, Z. Peng, M. Shao, J. Li, J. Liu, Extending Whisper with prompt tuning to target-speaker ASR, in: ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2024, pp. 12516–12520

  16. [26]

    L. Meng, J. Kang, Y. Wang, Z. Jin, X. Wu, X. Liu, H. Meng, Em- powering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System, in: Interspeech 2024, 2024, pp. 4653–4657. doi:10.21437/Interspeech.2024-971

  17. [27]

    P. Guo, X. Chang, H. Lv, S. Watanabe, L. Xie, SQ-Whisper: Speaker- Querying based Whisper Model for Target-Speaker ASR, arXiv preprint arXiv:2412.05589 (2024)

  18. [28]

    Vaswani, Attention is all you need, Advances in Neural Information Processing Systems (2017)

    A. Vaswani, Attention is all you need, Advances in Neural Information Processing Systems (2017)

  19. [29]

    Gandhi, P

    S. Gandhi, P. von Platen, A. M. Rush, Distil-whisper: Robust knowledge distillation via large-scale pseudo labelling, arXiv preprint arXiv:2311.00430 (2023)

  20. [30]

    Watanabe, T

    S. Watanabe, T. Hori, S. Kim, J. R. Hershey, T. Hayashi, Hybrid CTC/attention architecture for end-to-end speech recognition, IEEE Journal of Selected Topics in Signal Processing 11 (8) (2017) 1240–1253

  21. [31]

    Graves, S

    A. Graves, S. Fern´ andez, F. Gomez, J. Schmidhuber, Connectionist tem- poral classification: labelling unsegmented sequence data with recurrent neural networks, in: Proceedings of the 23rd International Conference on Machine Learning, ICML ’06, Association for Computing Machi...

  22. [32]

    T. Hori, S. Watanabe, J. Hershey, Joint CTC/attention decoding for end-to-end speech recognition, in: R. Barzilay, M.-Y. Kan (Eds.), Proceedings of the 55th Annual Meeting of the Association for Com- putational Linguistics (Volume 1: Long Papers), Association for 35 Computatio...

  23. [33]

    Leviathan, M

    Y. Leviathan, M. Kalman, Y. Matias, Fast Inference from Transformers via Speculative Decoding, in: A. Krause, E. Brunskill, K. Cho, B. En- gelhardt, S. Sabato, J. Scarlett (Eds.), Proceedings of the 40th Inter- national Conference on Machine Learning, Vol. 202 of Proceedings o...

  24. [34]

    Watanabe, T

    S. Watanabe, T. Hori, S. Karita, T. Hayashi, J. Nishitoba, Y. Unno, N. Enrique Yalta Soplin, J. Heymann, M. Wiesner, N. Chen, A. Renduchintala, T. Ochiai, ESPnet: End-to-end speech process- ing toolkit, in: Proceedings of Interspeech, 2018, pp. 2207–2211. doi:10.21437/Interspe...

  25. [35]

    Polok, D

    A. Polok, D. Klement, M. Wiesner, S. Khudanpur, J. ˇCernock´ y, L. Bur- get, Target Speaker ASR with Whisper (2024). arXiv:2409.09543. URL https://arxiv.org/abs/2409.09543

  26. [36]

    Horiguchi, Y

    S. Horiguchi, Y. Takashima, P. Garc ´ ıa, S. Watanabe, Y. Kawaguchi, Multi-Channel End-To-End Neural Diarization with Distributed Micro- phones, in: ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2022, pp. 7332–

  27. [37]

    Moˇ sner, R

    L. Moˇ sner, R. Serizel, L. Burget, O. Plchot, E. Vincent, J. Peng, J. ˇCernock´ y, Multi-channel extension of pre-trained models for speaker verification, in: Proceedings of Interspeech 2024, Vol. 2024, Inter- national Speech Communication Association, 2024, pp. 2135–2139. do...

  28. [38]

    T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cis- tac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. L. Scao, S. Gugger, M. Drame, Q. Lhoest, A. M. Rush, Transformers: State-of-the-art nat- ura...

  29. [39]

    T. v. Neumann, C. B. Boeddeker, M. Delcroix, R. Haeb-Umbach, MeetEval: A Toolkit for Computation of Word Error Rates for Meeting Transcription Systems, in: Proceedings of the 7th International Work- shop on Speech Processing in Everyday Environments (CHiME 2023), 2023, pp. 27–...

  30. [40]

    Panayotov, G

    V. Panayotov, G. Chen, D. Povey, S. Khudanpur, Librispeech: An ASR corpus based on public domain audio books, in: 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2015, pp. 5206–5210. doi:10.1109/ICASSP.2015.7178964

  31. [41]

    Loshchilov, F

    I. Loshchilov, F. Hutter, Decoupled weight decay regularization, in: In- ternational Conference on Learning Representations, 2019, pp. 1–18. URL https://openreview.net/forum?id=Bkg6RiCqY7

  32. [42]

    J. Han, F. Landini, J. Rohdin, A. Silnova, M. Diez, L. Burget, Lever- aging Self-Supervised Learning for Speaker Diarization, arXiv preprint arXiv:2409.09408 (2024)

  33. [43]

    Kinoshita, M

    K. Kinoshita, M. Delcroix, N. Tawara, Advances in integration of end- to-end neural and clustering-based diarization for real conversational speech, in: Proc. Interspeech, 2021, pp. 3565–3569

  34. [44]

    Plaquet, H

    A. Plaquet, H. Bredin, Powerset multi-class cross entropy loss for neural speaker diarization, in: INTERSPEECH 2023, 2023, pp. 3222–3226. doi:10.21437/Interspeech.2023-205

  35. [45]

    S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao, et al., Wavlm: Large-scale self-supervised pre- training for full stack speech processing, IEEE Journal of Selected Topics in Signal Processing 16 (6) (2022) 1505–1518

  36. [46]

    Gulati, J

    A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, et al., Conformer: Convolution-augmented transformer for speech recognition, in: Proc. Interspeech, 2020, pp. 5036–5040. 37

  37. [47]

    T. J. Park, K. J. Han, M. Kumar, S. Narayanan, Auto-tuning spectral clustering for speaker diarization using normalized maximum eigengap, IEEE Signal Processing Letters 27 (2019) 381–385

  38. [48]

    Kanda, G

    N. Kanda, G. Ye, Y. Wu, Y. Gaur, X. Wang, Z. Meng, Z. Chen, T. Yoshioka, Large-Scale Pre-Training of End-to-End Multi- Talker ASR for Meeting Transcription with Single Distant Micro- phone, in: Proceedings of Interspeech 2021, 2021, pp. 3430–3434. doi:10.21437/Interspeech.2021-102

  39. [49]

    D. Raj, D. Povey, S. Khudanpur, SURT 2.0: Advances in Transducer-Based Multi-Talker Speech Recognition, IEEE/ACM Trans. Audio, Speech and Lang. Proc. 31 (2023) 3800–3813. doi:10.1109/TASLP.2023.3318398. URL https://doi.org/10.1109/TASLP.2023.3318398

  40. [50]

    Fazel-Zarandi, W.-N

    M. Fazel-Zarandi, W.-N. Hsu, Cocktail Hubert: Generalized Self- Supervised Pre-Training for Mixture and Single-Source Speech, in: ICASSP 2023 - 2023 IEEE International Conference on Acous- tics, Speech and Signal Processing (ICASSP), 2023, pp. 1–5. doi:10.1109/ICASSP49357.2023...

  41. [51]

    S. Niu, R. Wang, J. Du, G. Yang, Y. Tu, S. Wu, S. Qian, H. Wu, H. Xu, X. Zhang, G. Zhong, X. Yu, J. Chen, M. Wang, D. Cai, T. Gao, G. Wan, F. Ma, J. Pan, J. Gao, The USTC-NERCSLIP Systems for the CHiME- 8 NOTSOF AR-1 Challenge, in: 8th International Workshop on Speech Processi...

  42. [52]

    Zhang, Y

    W. Zhang, Y. Qian, Weakly-Supervised Speech Pre-training: A Case Study on Target Speech Recognition, in: Proc. INTERSPEECH 2023, 2023, pp. 3517–3521. doi:10.21437/Interspeech.2023-1280

  43. [53]

    Polok, D

    A. Polok, D. Klement, J. Han, S. Sedl´ aˇ cek, B. Yusuf, M. Maciejewski, M. Wiesner, L. Burget, BUT/JHU System Description for CHiME-8 NOTSOF AR-1 Challenge, in: 8th International Workshop on Speech Processing in Everyday Environments (CHiME 2024), 2024, pp. 18–22. doi:10.2143...

  44. [54]

    Rousseau, P

    A. Rousseau, P. Del´ eglise, Y. Est` eve, TED-LIUM: an automatic speech recognition dedicated corpus, in: N. Calzolari, K. Choukri, T. Declerck, M. U. Do˘ gan, B. Maegaard, J. Mariani, A. Moreno, J. Odijk, S. Piperidis (Eds.), Proceedings of the Eighth International Conference...

  45. [55]

    C. Wang, M. Riviere, A. Lee, A. Wu, C. Talnikar, D. Haziza, M. Williamson, J. Pino, E. Dupoux, VoxPopuli: A large-scale mul- tilingual speech corpus for representation learning, semi-supervised learning and interpretation, in: C. Zong, F. Xia, W. Li, R. Navigli (Eds.), Proceed...

  46. [56]

    Y. Peng, S. Dalmia, I. Lane, S. Watanabe, Branchformer: Parallel MLP- Attention Architectures to Capture Local and Global Context for Speech Recognition and Understanding, in: K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvari, G. Niu, S. Sabato (Eds.), Proceedings of the 39th I...

  47. [7336]

    doi:10.1109/ICASSP43922.2022.9746749

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.