Pith. sign in

REVIEW 8 cited by

Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2201.02184 v2 pith:PYGQY5BB submitted 2022-01-05 eess.AS cs.CVcs.SD

classification eess.AScs.CVcs.SD
keywords speechaudio-visualrepresentationhoursav-hubertdatalearninglip-reading
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Video recordings of speech contain correlated audio and visual information, providing a strong signal for speech representation learning from the speaker's lip movements and the produced sound. We introduce Audio-Visual Hidden Unit BERT (AV-HuBERT), a self-supervised representation learning framework for audio-visual speech, which masks multi-stream video input and predicts automatically discovered and iteratively refined multimodal hidden units. AV-HuBERT learns powerful audio-visual speech representation benefiting both lip-reading and automatic speech recognition. On the largest public lip-reading benchmark LRS3 (433 hours), AV-HuBERT achieves 32.5% WER with only 30 hours of labeled data, outperforming the former state-of-the-art approach (33.6%) trained with a thousand times more transcribed video data (31K hours). The lip-reading WER is further reduced to 26.9% when using all 433 hours of labeled data from LRS3 and combined with self-training. Using our audio-visual representation on the same benchmark for audio-only speech recognition leads to a 40% relative WER reduction over the state-of-the-art performance (1.3% vs 2.3%). Our code and models are available at https://github.com/facebookresearch/av_hubert

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE

    eess.AS 2026-08 conditional novelty 6.0 of 10

    A real-scene Mandarin benchmark and a 766-hour curated audio-lip corpus for identity-faithful audio-visual target speaker extraction, with a baseline scoring 0.2261 CER and 82.22% strict identity correctness.

  2. Less is More: Modality-Decoupling for General AIGC Audio-Video Detection

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A decoupled audio-video AIGC detector that fuses independent audio and visual predictions at decision level ranks first in the DDL 2.0 general AIGC detection challenge with a final score of 0.8460.

  3. Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    The paper introduces semantic mismatch between authentic audio and video as a new DeepFake detection challenge via the RARV-SMM class and demonstrates that a semantic reinforcement strategy with ImageBind embeddings i...

  4. Dynamic Uncertainty-aware Multimodal Fusion for Outdoor Health Monitoring

    cs.NI 2025-08 unverdicted novelty 6.0 of 10

    DUAL-Health is an uncertainty-aware multimodal fusion framework that quantifies sensor noise, customizes fusion weights accordingly, and aligns modality distributions to improve outdoor health monitoring.

  5. MuteSwap: Visual-informed Silent Video Identity Conversion

    cs.SD 2025-07 conditional novelty 6.0 of 10

    A single-stage model performs zero-shot voice conversion from silent lip video and target face images, with no acoustic input at inference.

  6. The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models

    eess.AS 2025-06 conditional novelty 6.0 of 10

    Phonetic information becomes decodable in AV-HuBERT only about 20 ms before audio-only HuBERT, not the 100 to 300 ms visual lead in human speech, indicating AV-HuBERT's temporal dynamics are dominated by audio.

  7. Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation

    cs.CL 2025-08 conditional novelty 5.0 of 10

    An audio-visual language model that adds full-face visual features to a pre-trained expressive speech model improves emotion recognition and expressive speech generation by a few F1 points over speech-only on syntheti...

  8. Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition

    cs.LG 2025-05 conditional novelty 5.0 of 10

    On a single-speaker MRI speech corpus, adding vocal-tract video to audio does not improve phoneme recognition, but attention analysis shows articulatory cues can lead acoustic cues in time.

Pith tools