REVIEW 18 cited by
HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Self-supervised approaches for speech representation learning are challenged by three unique problems: (1) there are multiple sound units in each input utterance, (2) there is no lexicon of input sound units during the pre-training phase, and (3) sound units have variable lengths with no explicit segmentation. To deal with these three problems, we propose the Hidden-Unit BERT (HuBERT) approach for self-supervised speech representation learning, which utilizes an offline clustering step to provide aligned target labels for a BERT-like prediction loss. A key ingredient of our approach is applying the prediction loss over the masked regions only, which forces the model to learn a combined acoustic and language model over the continuous inputs. HuBERT relies primarily on the consistency of the unsupervised clustering step rather than the intrinsic quality of the assigned cluster labels. Starting with a simple k-means teacher of 100 clusters, and using two iterations of clustering, the HuBERT model either matches or improves upon the state-of-the-art wav2vec 2.0 performance on the Librispeech (960h) and Libri-light (60,000h) benchmarks with 10min, 1h, 10h, 100h, and 960h fine-tuning subsets. Using a 1B parameter model, HuBERT shows up to 19% and 13% relative WER reduction on the more challenging dev-other and test-other evaluation subsets.
Forward citations
Cited by 18 Pith papers
-
Video2Reaction: Mapping Video to Audience Reaction Distribution in the Wild
A new dataset and benchmark maps movie clips to distributions of audience emotional reactions derived from YouTube comments, showing that finetuned vision-language models can predict these distributions from video alone.
-
Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study
Across six pooling heads and six frozen SSL backbones on English and Mandarin depression speech, a third of configurations collapse to single-class prediction, so backbone- and seed-robustness should be first-class ev...
-
findsylls: A Language-Agnostic Toolkit for Syllable-Level Speech Tokenization and Embedding
A modular toolkit standardizes classical and SSL syllabifiers, enables component recombination, and benchmarks speed–accuracy trade-offs on English, Spanish, and newly annotated Kono.
-
Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models
Backdoors propagate through SLM components with persistence or erasure depending on the targeted part, and poisoned samples are not directly separable from benign ones in shared multitask embeddings.
-
Entropy-based Coarse and Compressed Semantic Speech Representation Learning
Predictive entropy from a token-level speech language model finds merge boundaries, producing compressed semantic tokens that keep ASR and translation accuracy at 15 Hz while lowering latency.
-
Self-supervised learning of speech representations with Dutch archival data
A Dutch-only wav2vec 2.0 large model, pre-trained on 55.7k hours of WhisperX-cleaned TV audio, outperformed reported multi-lingual and Whisper baselines on the Dutch N-Best benchmark.
-
Scaling Self-Supervised Representation Learning for Symbolic Piano Performance
Self-supervised pretraining on 60,000 hours of symbolic piano music produces a generative model and contrastive embeddings that beat leading baselines on continuation quality and several MIR classification benchmarks.
-
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
XY-Tokenizer is a 1 kbps dual-channel speech codec that reports simultaneously strong text alignment and high speaker similarity, comparable to specialized codecs at similar bitrates.
-
Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM
A zero-shot AudioLLM can classify speech as cognitively normal or impaired with accuracy close to supervised baselines, although prompt selection and small margins temper the claim.
-
Multimodal Zero-Shot Framework for Deepfake Hate Speech Detection in Low-Resource Languages
A contrastive audio-text framework detects hate speech in synthesized speech across six languages and outperforms baselines, with a new 127k-sample dataset.
-
Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment
Supervised contrastive alignment of frozen WavLM layers modestly lifts Mandarin depression F1 under LOSO while quantifying that speaker leakage inflated prior Mandarin F1 by ~0.23.
-
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
An audio-visual language model that adds full-face visual features to a pre-trained expressive speech model improves emotion recognition and expressive speech generation by a few F1 points over speech-only on syntheti...
-
A Concept-based approach to Voice Disorder Detection
Concept bottleneck and concept embedding models, trained on clinical concepts extracted from patient notes by a large language model, detect voice pathology from audio almost as accurately as an end-to-end transformer.
-
Benchmarking Foundation Speech and Language Models for Alzheimer's Disease and Related Dementia Detection from Spontaneous Speech
On PREPARE spontaneous speech, Whisper-medium audio embeddings achieved the best three-way ADRD classification (0.731 accuracy, 0.802 AUC), outperforming text-based and traditional acoustic pipelines.
-
DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages
Open w2v-BERT ASR base models for 27 African languages, with a two-step annealing recipe and prefix-frame language conditioning.
-
Leveraging Context for Multimodal Fallacy Classification in Political Debates
A shared-task system that shows adding previous-sentence context helps text-based fallacy classification but not audio, and that late fusion of the two modalities adds little.
-
EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion
EZ-VC combines discrete units from a multilingual self-supervised encoder (Xeus) with an F5-TTS flow-matching decoder to achieve zero-shot any-to-any voice conversion, without text labels or multiple disentangling encoders.
-
Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning
Transferring I-JEPA's masked latent prediction to mel-spectrograms yields competitive audio representations on music and environmental sound tasks with a small fraction of the training data.
Discussion (0). Continue with ORCID to comment.