Pith. sign in

REVIEW 1 cited by

Recurrent Neural Network Transducer for Audio-Visual Speech Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.04890 v1 pith:6YNX3Y5D submitted 2019-11-08 eess.AS cs.CLcs.CVcs.LGcs.SD

classification eess.AScs.CLcs.CVcs.LGcs.SD
keywords audio-visualsystemspeechlrs3-tednetworkneuralperformancepublic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This work presents a large-scale audio-visual speech recognition system based on a recurrent neural network transducer (RNN-T) architecture. To support the development of such a system, we built a large audio-visual (A/V) dataset of segmented utterances extracted from YouTube public videos, leading to 31k hours of audio-visual training content. The performance of an audio-only, visual-only, and audio-visual system are compared on two large-vocabulary test sets: a set of utterance segments from public YouTube videos called YTDEV18 and the publicly available LRS3-TED set. To highlight the contribution of the visual modality, we also evaluated the performance of our system on the YTDEV18 set artificially corrupted with background noise and overlapping speech. To the best of our knowledge, our system significantly improves the state-of-the-art on the LRS3-TED set.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Negative to Positive Co-learning with Aggressive Modality Dropout

    cs.CL 2025-01 conditional novelty 5.0 of 10

    Applying 80% modality dropout to audio and video features during training reversed negative co-learning on IEMOCAP, lifting unimodal test accuracy from 27% to 47% and beating the unimodal baseline.

Pith tools