REVIEW 4 cited by
Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Sortformer is an encoder-based speaker diarization model designed for supervising speaker tagging in speech-to-text models. Instead of relying solely on permutation invariant loss (PIL), Sortformer introduces Sort Loss to resolve the permutation problem, either independently or in tandem with PIL. In addition, we propose a streamlined multi-speaker speech-to-text architecture that leverages Sortformer for speaker supervision, embedding speaker labels into the encoder using sinusoidal kernel functions. This design addresses the speaker permutation problem through sorted objectives, effectively bridging timestamps and tokens to supervise speaker labels in the output transcriptions. Experiments demonstrate that Sort Loss can boost speaker diarization performance, and incorporating the speaker supervision from Sortformer improves multi-speaker transcription accuracy. We anticipate that the proposed Sortformer and multi-speaker architecture will enable the seamless integration of speaker tagging capabilities into foundational speech-to-text systems and multimodal large language models (LLMs), offering an easily adoptable and user-friendly mechanism to enhance their versatility and performance in speaker-aware tasks. The code and trained models are made publicly available through the NVIDIA NeMo Framework.
Forward citations
Cited by 4 Pith papers
-
Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering
A streaming Sortformer with an arrival-ordered speaker cache achieves lower diarization error than prior online systems on DIHARD III and CALLHOME, even at 0.32 second latency.
-
Speaker Targeting via Self-Speaker Adaptation for Multi-talker ASR
A speaker-activity mask injected into an ASR encoder lets one model instance transcribe each talker in overlapped speech without speaker embeddings.
-
A Unified Speech LLM for Diarization and Speech Recognition in Multilingual Conversations
A local-window speech LLM with prompt-based speaker context and an external diarization alignment module achieves 27.25 tcpWER/tcpCER on the MLC-SLM Task II evaluation set, a 54.87% relative improvement over the offic...
-
Joint ASR and Speaker Role Tagging with Serialized Output Training
Fine-tuning Whisper with serialized output training and role-specific tokens produces role-aware transcripts in one pass, cutting multi-talker word error rate by 10 to 40 percent versus a WavLM CTC baseline.
Discussion (0). Continue with ORCID to comment.