Pith. sign in

REVIEW 26 cited by

wav2vec: Unsupervised Pre-training for Speech Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1904.05862 v4 pith:M3GDESGW submitted 2019-04-11 cs.CL

classification cs.CL
keywords dataspeechaudiocharacter-basedpre-trainingrecognitionrepresentationstraining
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We explore unsupervised pre-training for speech recognition by learning representations of raw audio. wav2vec is trained on large amounts of unlabeled audio data and the resulting representations are then used to improve acoustic model training. We pre-train a simple multi-layer convolutional neural network optimized via a noise contrastive binary classification task. Our experiments on WSJ reduce WER of a strong character-based log-mel filterbank baseline by up to 36% when only a few hours of transcribed data is available. Our approach achieves 2.43% WER on the nov92 test set. This outperforms Deep Speech 2, the best reported character-based system in the literature while using two orders of magnitude less labeled training data.

Discussion (0). Sign in to comment.

Forward citations

Cited by 26 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A diffusion-based talking-face generator uses 3D blendshape coefficients to continuously control the emotion intensity of generated facial expressions.

  2. Should Missing Modalities Always Be Necessary to Repair for Multi-modal Sentiment Analysis?

    cs.CL 2026-07 conditional novelty 6.0 of 10

    SIEVE learns a per-sample 'should we repair?' decision from the loss gap between direct and repair branches, improving three missing-modality MSA backbones on CMU-MOSI and IEMOCAP.

  3. STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits

    cs.CV 2025-12 conditional novelty 6.0 of 10

    STARCaster is a 2D spatio-temporal video diffusion model that unifies identity-conditioned audio-driven portrait animation and novel-view synthesis without explicit 3D reconstruction.

  4. Entropy-based Coarse and Compressed Semantic Speech Representation Learning

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Predictive entropy from a token-level speech language model finds merge boundaries, producing compressed semantic tokens that keep ASR and translation accuracy at 15 Hz while lowering latency.

  5. Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning

    eess.AS 2025-08 conditional novelty 6.0 of 10

    The authors release a rule-based Quran Phonetic Script, an 890-hour expert recitation dataset, and a multi-head CTC model that achieves 0.16% average phoneme error rate on held-out reciters.

  6. Scaling and Distilling Transformer Models for sEMG

    eess.AS 2025-07 accept novelty 6.0 of 10

    Vanilla transformers on the emg2qwerty dataset improve cross-user typing accuracy up to 109M parameters, and simple logit distillation recovers most of the gain in a 2.2M-parameter student.

  7. OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder

    cs.SD 2025-07 conditional novelty 6.0 of 10

    OpenBEATs releases the BEATs audio pretraining pipeline, trains 300M-parameter models on 20k hours of multi-domain audio, and reports strong results across 25 datasets, including bioacoustics and reasoning tasks.

  8. Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A two-stage text-guidance framework, Think-Before-Draw, uses chain-of-thought prompting to convert emotion labels into facial muscle descriptions and then progressively guides a diffusion model from coarse emotion to ...

  9. A Dataset for Automatic Assessment of TTS Quality in Spanish

    cs.SD 2025-07 conditional novelty 6.0 of 10

    A new Spanish-language dataset of 4,326 MOS-rated TTS audio clips enables automated naturalness prediction with a mean absolute error around 0.8 on a five-point scale.

  10. Prosodic Structure Beyond Lexical Content: A Study of Self-Supervised Learning

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A masked self-supervised model trained only on pitch, energy, and voice activity captures prosodic structure at multiple timescales, with random masking yielding the most generalizable representations.

  11. StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion

    cs.MM 2025-06 conditional novelty 6.0 of 10

    StarVC is an autoregressive voice conversion model that generates text tokens before acoustic tokens, improving linguistic fidelity while retaining speaker similarity.

  12. Model as Loss: A Self-Consistent Training Paradigm

    cs.SD 2025-05 conditional novelty 6.0 of 10

    Using the model's own encoder as a feature loss improves perceptual quality and iterative stability of a speech enhancement model compared with a WavLM-based loss.

  13. InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Sparse-frame dubbing with adjacent-chunk keyframe sampling lets a streaming audio-video model produce full-body motion synchronized to new audio while preserving identity and camera motion.

  14. Seeing Voices: Generating A-Roll Video from Audio with Mirage

    cs.CV 2025-06 reject novelty 5.0 of 10

    Mirage generates photorealistic A-roll videos of people speaking directly from audio, using only joint self-attention over audio, text, and video tokens.

  15. Seamless Dysfluent Speech Text Alignment for Disordered Speech Analysis

    eess.AS 2025-06 conditional novelty 5.0 of 10

    Neural LCS uses learned phoneme and word similarity instead of exact matches to align dysfluent speech to intended text, and it outperforms DTW and Hard LCS on simulated benchmarks.

  16. Automatic classification of stop realisation with wav2vec2.0

    cs.CL 2025-05 conditional novelty 5.0 of 10

    wav2vec2.0 models classify stop burst presence with 88 to 94 percent accuracy, and a Japanese-pretrained model reaches near-peak accuracy on English with only 500 annotated tokens.

  17. Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes

    cs.SD 2025-05 conditional novelty 5.0 of 10

    ADD-GP, a Gaussian Process classifier with XLS-R speech embeddings, adapts to unseen TTS models with as few as 5 samples and achieves state-of-the-art low error rates on the new LibriFake benchmark.

  18. Contextualized Token Discrimination for Speech Search Query Correction

    cs.SD 2025-09 reject novelty 4.0 of 10

    CTD uses BERT token representations plus a composition layer to correct Chinese spelling errors in ASR queries, but the reported gains lack matched baselines and released data.

  19. EmoSLLM: Parameter-Efficient Adaptation of LLMs for Speech Emotion Recognition

    eess.AS 2025-08 reject novelty 4.0 of 10

    EmoSLLM, a LoRA-fine-tuned 3B LLM with an audio mapper, beats several 7B speech-text LLMs on MSP-Podcast emotion recognition, but only when given the ground-truth transcript at inference.

  20. Weak Supervision Techniques towards Enhanced ASR Models in Industry-level CRM Systems

    cs.SD 2025-07 conditional novelty 4.0 of 10

    Using LLM- and TTS-generated synthetic speech to fine-tune Whisper models cuts character error rates by up to 63% in a luxury retail CRM transcription task, though the evaluation has several methodological weaknesses.

  21. LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models

    cs.CV 2025-06 conditional novelty 4.0 of 10

    Using consistency distillation, INT8 quantization, and pipeline parallelism, the LLIA system generates portrait video from audio at 78 FPS, with 140 ms initial latency.

  22. Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Multimodal LLMs can be organized by fusion mechanism, fusion level, representation paradigm, and training paradigm, with 125 models classified accordingly.

  23. SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction

    cs.SD 2025-06 reject novelty 4.0 of 10

    SALF-MOS, a compact U-Net-style model using frozen wav2vec features, claims state-of-the-art MOS prediction on four benchmarks with only 1,574 parameters.

  24. DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec

    cs.SD 2025-05 conditional novelty 4.0 of 10

    DS-Codec improves low-bitrate speech codec quality by first training a mirrored codec and then switching to a non-mirrored decoder, while using product quantization to form one large codebook.

  25. Amplifying Emotional Signals: Data-Efficient Deep Learning for Robust Speech Emotion Recognition

    eess.AS 2025-08 reject novelty 3.0 of 10

    A pretrained ResNet34 with augmentation reaches 66.7% accuracy on a combined RAVDESS/SAVEE emotion set, but only on a validation split, so the claimed new benchmark is unverified.

  26. DeepEmoNet: Building Machine Learning Models for Automatic Emotion Recognition in Human Speeches

    eess.AS 2025-08 conditional novelty 2.0 of 10

    A ResNet34 pretrained on ImageNet and fine-tuned on log-mel spectrograms, with data augmentation, classifies eight speech emotions at 66.7% accuracy and F1 0.631 on the pooled RAVDESS/SAVEE validation set.

Pith tools