Pith. sign in

REVIEW 2 cited by

LSTM based Similarity Measurement with Spectral Clustering for Speaker Diarization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1907.10393 v1 pith:NCJQITVC submitted 2019-07-23 eess.AS cs.LGcs.SDstat.ML

classification eess.AScs.LGcs.SDstat.ML
keywords similarityspeakerclusteringdiarizationmatrixsegmentsspectralsystem
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

More and more neural network approaches have achieved considerable improvement upon submodules of speaker diarization system, including speaker change detection and segment-wise speaker embedding extraction. Still, in the clustering stage, traditional algorithms like probabilistic linear discriminant analysis (PLDA) are widely used for scoring the similarity between two speech segments. In this paper, we propose a supervised method to measure the similarity matrix between all segments of an audio recording with sequential bidirectional long short-term memory networks (Bi-LSTM). Spectral clustering is applied on top of the similarity matrix to further improve the performance. Experimental results show that our system significantly outperforms the state-of-the-art methods and achieves a diarization error rate of 6.63% on the NIST SRE 2000 CALLHOME database.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Revisiting Active Speaker Detection: An In-the-Wild Benchmark for Generalization and Robustness

    cs.CV 2025-05 conditional novelty 6.0 of 10

    UniTalk is a larger and more diverse active speaker detection benchmark on which state-of-the-art models underperform, and it improves cross-dataset generalization when used as a training source.

  2. Whisper Turns Stronger: Augmenting Wav2Vec 2.0 for Superior ASR in Low-Resource Languages

    cs.CL 2024-12 reject novelty 3.0 of 10

    Fine-tuning Wav2Vec2-xlsr-53 on Common Voice audio augmented with pitch shift, Gaussian noise, and band-stop filtering lowers WER and CER in Arabic, Russian, and Portuguese, though the Whisper comparison is overstated.

Pith tools