Pith. sign in

REVIEW 18 cited by

HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.07447 v1 pith:GANGF35R submitted 2021-06-14 cs.CL cs.AIcs.LGeess.AS

classification cs.CLcs.AIcs.LGeess.AS
keywords hubertmodelunitsclusteringlearningpredictionrepresentationself-supervised
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Self-supervised approaches for speech representation learning are challenged by three unique problems: (1) there are multiple sound units in each input utterance, (2) there is no lexicon of input sound units during the pre-training phase, and (3) sound units have variable lengths with no explicit segmentation. To deal with these three problems, we propose the Hidden-Unit BERT (HuBERT) approach for self-supervised speech representation learning, which utilizes an offline clustering step to provide aligned target labels for a BERT-like prediction loss. A key ingredient of our approach is applying the prediction loss over the masked regions only, which forces the model to learn a combined acoustic and language model over the continuous inputs. HuBERT relies primarily on the consistency of the unsupervised clustering step rather than the intrinsic quality of the assigned cluster labels. Starting with a simple k-means teacher of 100 clusters, and using two iterations of clustering, the HuBERT model either matches or improves upon the state-of-the-art wav2vec 2.0 performance on the Librispeech (960h) and Libri-light (60,000h) benchmarks with 10min, 1h, 10h, 100h, and 960h fine-tuning subsets. Using a 1B parameter model, HuBERT shows up to 19% and 13% relative WER reduction on the more challenging dev-other and test-other evaluation subsets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 25 citations worldwide. Full citation record

  1. Video2Reaction: Mapping Video to Audience Reaction Distribution in the Wild

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A new dataset and benchmark maps movie clips to distributions of audience emotional reactions derived from YouTube comments, showing that finetuned vision-language models can predict these distributions from video alone.

  2. Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study

    eess.AS 2026-07 conditional novelty 6.0 of 10

    Across six pooling heads and six frozen SSL backbones on English and Mandarin depression speech, a third of configurations collapse to single-class prediction, so backbone- and seed-robustness should be first-class ev...

  3. findsylls: A Language-Agnostic Toolkit for Syllable-Level Speech Tokenization and Embedding

    cs.CL 2026-03 conditional novelty 6.0 of 10

    A modular toolkit standardizes classical and SSL syllabifiers, enables component recombination, and benchmarks speed–accuracy trade-offs on English, Spanish, and newly annotated Kono.

  4. Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models

    cs.CL 2025-10 unverdicted novelty 6.0 of 10

    Backdoors propagate through SLM components with persistence or erasure depending on the targeted part, and poisoned samples are not directly separable from benign ones in shared multitask embeddings.

  5. Entropy-based Coarse and Compressed Semantic Speech Representation Learning

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Predictive entropy from a token-level speech language model finds merge boundaries, producing compressed semantic tokens that keep ASR and translation accuracy at 15 Hz while lowering latency.

  6. Self-supervised learning of speech representations with Dutch archival data

    cs.SD 2025-07 conditional novelty 6.0 of 10

    A Dutch-only wav2vec 2.0 large model, pre-trained on 55.7k hours of WhisperX-cleaned TV audio, outperformed reported multi-lingual and Whisper baselines on the Dutch N-Best benchmark.

  7. Scaling Self-Supervised Representation Learning for Symbolic Piano Performance

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Self-supervised pretraining on 60,000 hours of symbolic piano music produces a generative model and contrastive embeddings that beat leading baselines on continuation quality and several MIR classification benchmarks.

  8. XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs

    cs.SD 2025-06 conditional novelty 6.0 of 10

    XY-Tokenizer is a 1 kbps dual-channel speech codec that reports simultaneously strong text alignment and high speaker similarity, comparable to specialized codecs at similar bitrates.

  9. Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM

    cs.SD 2025-06 conditional novelty 6.0 of 10

    A zero-shot AudioLLM can classify speech as cognitively normal or impaired with accuracy close to supervised baselines, although prompt selection and small margins temper the claim.

  10. Multimodal Zero-Shot Framework for Deepfake Hate Speech Detection in Low-Resource Languages

    cs.SD 2025-06 conditional novelty 6.0 of 10

    A contrastive audio-text framework detects hate speech in synthesized speech across six languages and outperforms baselines, with a new 127k-sample dataset.

  11. Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment

    eess.AS 2026-07 conditional novelty 5.0 of 10

    Supervised contrastive alignment of frozen WavLM layers modestly lifts Mandarin depression F1 under LOSO while quantifying that speaker leakage inflated prior Mandarin F1 by ~0.23.

  12. Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation

    cs.CL 2025-08 conditional novelty 5.0 of 10

    An audio-visual language model that adds full-face visual features to a pre-trained expressive speech model improves emotion recognition and expressive speech generation by a few F1 points over speech-only on syntheti...

  13. A Concept-based approach to Voice Disorder Detection

    eess.AS 2025-07 conditional novelty 5.0 of 10

    Concept bottleneck and concept embedding models, trained on clinical concepts extracted from patient notes by a large language model, detect voice pathology from audio almost as accurately as an end-to-end transformer.

  14. Benchmarking Foundation Speech and Language Models for Alzheimer's Disease and Related Dementia Detection from Spontaneous Speech

    cs.CL 2025-06 conditional novelty 5.0 of 10

    On PREPARE spontaneous speech, Whisper-medium audio embeddings achieved the best three-way ADRD classification (0.731 accuracy, 0.802 AUC), outperforming text-based and traditional acoustic pipelines.

  15. DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages

    cs.CL 2026-07 reject novelty 4.0 of 10

    Open w2v-BERT ASR base models for 27 African languages, with a two-step annealing recipe and prefix-frame language conditioning.

  16. Leveraging Context for Multimodal Fallacy Classification in Political Debates

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A shared-task system that shows adding previous-sentence context helps text-based fallacy classification but not audio, and that late fusion of the two modalities adds little.

  17. EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion

    cs.SD 2025-05 reject novelty 4.0 of 10

    EZ-VC combines discrete units from a multilingual self-supervised encoder (Xeus) with an F5-TTS flow-matching decoder to achieve zero-shot any-to-any voice conversion, without text labels or multiple disentangling encoders.

  18. Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

    cs.SD 2025-06 conditional novelty 3.0 of 10

    Transferring I-JEPA's masked latent prediction to mel-spectrograms yields competitive audio representations on music and environmental sound tasks with a small fraction of the training data.

Pith tools