Pith. sign in

REVIEW 2 cited by

AVA-Speech: A Densely Labeled Dataset of Speech Activity in Movies

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1808.00606 v2 pith:UYKHNRTN submitted 2018-08-02 cs.SD eess.AS

classification cs.SDeess.AS
keywords speechactivitydatasetapplicationsapproachesava-speechavailableco-occurring
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Speech activity detection (or endpointing) is an important processing step for applications such as speech recognition, language identification and speaker diarization. Both audio- and vision-based approaches have been used for this task in various settings, often tailored toward end applications. However, much of the prior work reports results in synthetic settings, on task-specific datasets, or on datasets that are not openly available. This makes it difficult to compare approaches and understand their strengths and weaknesses. In this paper, we describe a new dataset which we will release publicly containing densely labeled speech activity in YouTube videos, with the goal of creating a shared, available dataset for this task. The labels in the dataset annotate three different speech activity conditions: clean speech, speech co-occurring with music, and speech co-occurring with noise, which enable analysis of model performance in more challenging conditions based on the presence of overlapping noise. We report benchmark performance numbers on AVA-Speech using off-the-shelf, state-of-the-art audio and vision models that serve as a baseline to facilitate future research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 4 citations worldwide. Full citation record

  1. R^3-VQA: "Read the Room" by Video Social Reasoning

    cs.CV 2025-05 conditional novelty 7.0 of 10

    R3-VQA is a new real-world video benchmark on which the best tested model, GPT-4o, scores 83% on generated questions but only 54% on human-written ones, while humans score 91% and 80%.

  2. Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study

    eess.AS 2025-02 conditional novelty 6.0 of 10

    SSL convolutional audio models pre-trained on speech, non-speech, or both perform nearly equally well across speech and non-speech downstream tasks, while domain-specific baselines struggle outside their domains.

Pith tools