Pith. sign in

REVIEW

Speaker Diarization with Region Proposal Network

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2002.06220 v1 pith:UVN2GX3R submitted 2020-02-14 eess.AS cs.SD

classification eess.AScs.SD
keywords diarizationspeakerspeechnetworkoverlappedrpnsdmethodproposal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Speaker diarization is an important pre-processing step for many speech applications, and it aims to solve the "who spoke when" problem. Although the standard diarization systems can achieve satisfactory results in various scenarios, they are composed of several independently-optimized modules and cannot deal with the overlapped speech. In this paper, we propose a novel speaker diarization method: Region Proposal Network based Speaker Diarization (RPNSD). In this method, a neural network generates overlapped speech segment proposals, and compute their speaker embeddings at the same time. Compared with standard diarization systems, RPNSD has a shorter pipeline and can handle the overlapped speech. Experimental results on three diarization datasets reveal that RPNSD achieves remarkable improvements over the state-of-the-art x-vector baseline.

Discussion (0). Continue with ORCID to comment.

Pith tools