Pith. sign in

REVIEW 4 cited by

W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.06209 v2 pith:ZY7R44U3 submitted 2021-08-07 cs.LG cs.SDeess.AS

classification cs.LGcs.SDeess.AS
keywords speechw2v-bertcontrastivelanguagelearningmaskedmodelmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Motivated by the success of masked language modeling~(MLM) in pre-training natural language processing models, we propose w2v-BERT that explores MLM for self-supervised speech representation learning. w2v-BERT is a framework that combines contrastive learning and MLM, where the former trains the model to discretize input continuous speech signals into a finite set of discriminative speech tokens, and the latter trains the model to learn contextualized speech representations via solving a masked prediction task consuming the discretized tokens. In contrast to existing MLM-based speech pre-training frameworks such as HuBERT, which relies on an iterative re-clustering and re-training process, or vq-wav2vec, which concatenates two separately trained modules, w2v-BERT can be optimized in an end-to-end fashion by solving the two self-supervised tasks~(the contrastive task and MLM) simultaneously. Our experiments show that w2v-BERT achieves competitive results compared to current state-of-the-art pre-trained models on the LibriSpeech benchmarks when using the Libri-Light~60k corpus as the unsupervised data. In particular, when compared to published models such as conformer-based wav2vec~2.0 and HuBERT, our model shows~5\% to~10\% relative WER reduction on the test-clean and test-other subsets. When applied to the Google's Voice Search traffic dataset, w2v-BERT outperforms our internal conformer-based wav2vec~2.0 by more than~30\% relatively.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DSA-Tokenizer: Disentangled Semantic-Acoustic Tokenization via Flow Matching-based Hierarchical Fusion

    cs.SD 2026-01 conditional novelty 6.0 of 10

    A dual-stream tokenizer learns separate semantic and acoustic tokens and uses a flow-matching decoder to reconstruct and recombine speech with better attribute control.

  2. Representing Speech Through Autoregressive Prediction of Cochlear Tokens

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Autoregressive prediction over discrete cochlear tokens yields a speech representation that beats prior self-supervised models on lexical-semantic similarity and is competitive on SUPERB tasks.

  3. Different Speech Translation Models Encode and Translate Speaker Gender Differently

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Traditional encoder-decoder speech translation models encode speaker gender in hidden states, while newer adapter-based models largely do not; lower gender encoding tracks with masculine-default translation bias.

  4. DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages

    cs.CL 2026-07 reject novelty 4.0 of 10

    Open w2v-BERT ASR base models for 27 African languages, with a two-step annealing recipe and prefix-frame language conditioning.

Pith tools