Pith. sign in

REVIEW 1 cited by

Spot the conversation: speaker diarisation in the wild

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2007.01216 v3 pith:V56YWO7Y submitted 2020-07-02 cs.SD cs.CVeess.ASeess.IV

Spot the conversation: speaker diarisation in the wild

classification cs.SD cs.CVeess.ASeess.IV
keywords speakerdiarisationvideosdatasetmethodwildaudio-visualcollected
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The goal of this paper is speaker diarisation of videos collected 'in the wild'. We make three key contributions. First, we propose an automatic audio-visual diarisation method for YouTube videos. Our method consists of active speaker detection using audio-visual methods and speaker verification using self-enrolled speaker models. Second, we integrate our method into a semi-automatic dataset creation pipeline which significantly reduces the number of hours required to annotate videos with diarisation labels. Finally, we use this pipeline to create a large-scale diarisation dataset called VoxConverse, collected from 'in the wild' videos, which we will release publicly to the research community. Our dataset consists of overlapping speech, a large and diverse speaker pool, and challenging background conditions.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. The tttAI System for the TSA-ASR Task of the SmartGlasses Challenge 2026

    eess.AS 2026-07 conditional novelty 4.0

    A cascaded smart-glasses TSA-ASR system with a dominant-speaker overlap fallback achieved 7.10% tcpCER on two-person dialogues and 34.04% on multi-party meetings, ranking second on the meeting track.