REVIEW 8 cited by
Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Video recordings of speech contain correlated audio and visual information, providing a strong signal for speech representation learning from the speaker's lip movements and the produced sound. We introduce Audio-Visual Hidden Unit BERT (AV-HuBERT), a self-supervised representation learning framework for audio-visual speech, which masks multi-stream video input and predicts automatically discovered and iteratively refined multimodal hidden units. AV-HuBERT learns powerful audio-visual speech representation benefiting both lip-reading and automatic speech recognition. On the largest public lip-reading benchmark LRS3 (433 hours), AV-HuBERT achieves 32.5% WER with only 30 hours of labeled data, outperforming the former state-of-the-art approach (33.6%) trained with a thousand times more transcribed video data (31K hours). The lip-reading WER is further reduced to 26.9% when using all 433 hours of labeled data from LRS3 and combined with self-training. Using our audio-visual representation on the same benchmark for audio-only speech recognition leads to a 40% relative WER reduction over the state-of-the-art performance (1.3% vs 2.3%). Our code and models are available at https://github.com/facebookresearch/av_hubert
Forward citations
Cited by 8 Pith papers
-
Identity-Faithful Audio-Visual Target Speaker Extraction with REAL-2MIX and VOXBLINK2-AVSE
A real-scene Mandarin benchmark and a 766-hour curated audio-lip corpus for identity-faithful audio-visual target speaker extraction, with a baseline scoring 0.2261 CER and 82.22% strict identity correctness.
-
Less is More: Modality-Decoupling for General AIGC Audio-Video Detection
A decoupled audio-video AIGC detector that fuses independent audio and visual predictions at decision level ranks first in the DDL 2.0 general AIGC detection challenge with a final score of 0.8460.
-
Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge
The paper introduces semantic mismatch between authentic audio and video as a new DeepFake detection challenge via the RARV-SMM class and demonstrates that a semantic reinforcement strategy with ImageBind embeddings i...
-
Dynamic Uncertainty-aware Multimodal Fusion for Outdoor Health Monitoring
DUAL-Health is an uncertainty-aware multimodal fusion framework that quantifies sensor noise, customizes fusion weights accordingly, and aligns modality distributions to improve outdoor health monitoring.
-
MuteSwap: Visual-informed Silent Video Identity Conversion
A single-stage model performs zero-shot voice conversion from silent lip video and target face images, with no acoustic input at inference.
-
The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models
Phonetic information becomes decodable in AV-HuBERT only about 20 ms before audio-only HuBERT, not the 100 to 300 ms visual lead in human speech, indicating AV-HuBERT's temporal dynamics are dominated by audio.
-
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
An audio-visual language model that adds full-face visual features to a pre-trained expressive speech model improves emotion recognition and expressive speech generation by a few F1 points over speech-only on syntheti...
-
Towards disentangling the contributions of articulation and acoustics in multimodal phoneme recognition
On a single-speaker MRI speech corpus, adding vocal-tract video to audio does not improve phoneme recognition, but attention analysis shows articulatory cues can lead acoustic cues in time.
Discussion (0). Sign in to comment.