Pith. sign in

REVIEW

Multi-Scale Speaker Diarization With Neural Affinity Score Fusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.10527 v1 pith:QW2N77MK submitted 2020-11-20 eess.AS

classification eess.AS
keywords speakerrepresentationssegmentsaffinitydiarizationfusionmulti-scaleneural
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Identifying the identity of the speaker of short segments in human dialogue has been considered one of the most challenging problems in speech signal processing. Speaker representations of short speech segments tend to be unreliable, resulting in poor fidelity of speaker representations in tasks requiring speaker recognition. In this paper, we propose an unconventional method that tackles the trade-off between temporal resolution and the quality of the speaker representations. To find a set of weights that balance the scores from multiple temporal scales of segments, a neural affinity score fusion model is presented. Using the CALLHOME dataset, we show that our proposed multi-scale segmentation and integration approach can achieve a state-of-the-art diarization performance.

Discussion (0). Continue with ORCID to comment.

Pith tools