Prompting Whisper with speaker labels and then fine-tuning it produces more consistent speaker IDs in long audio but error propagation and poor overlap handling limit speaker diarization accuracy.
Word Error Rate Definitions and Algorithms for Long - Form Multi -Talker Speech Recognition
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
eess.AS 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Prompting Whisper for Joint Speech Transcription and Diarization
Prompting Whisper with speaker labels and then fine-tuning it produces more consistent speaker IDs in long audio but error propagation and poor overlap handling limit speaker diarization accuracy.