REVIEW 26 cited by
wav2vec: Unsupervised Pre-training for Speech Recognition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We explore unsupervised pre-training for speech recognition by learning representations of raw audio. wav2vec is trained on large amounts of unlabeled audio data and the resulting representations are then used to improve acoustic model training. We pre-train a simple multi-layer convolutional neural network optimized via a noise contrastive binary classification task. Our experiments on WSJ reduce WER of a strong character-based log-mel filterbank baseline by up to 36% when only a few hours of transcribed data is available. Our approach achieves 2.43% WER on the nov92 test set. This outperforms Deep Speech 2, the best reported character-based system in the literature while using two orders of magnitude less labeled training data.
Forward citations
Cited by 26 Pith papers
-
Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation
A diffusion-based talking-face generator uses 3D blendshape coefficients to continuously control the emotion intensity of generated facial expressions.
-
Should Missing Modalities Always Be Necessary to Repair for Multi-modal Sentiment Analysis?
SIEVE learns a per-sample 'should we repair?' decision from the loss gap between direct and repair branches, improving three missing-modality MSA backbones on CMU-MOSI and IEMOCAP.
-
STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits
STARCaster is a 2D spatio-temporal video diffusion model that unifies identity-conditioned audio-driven portrait animation and novel-view synthesis without explicit 3D reconstruction.
-
Entropy-based Coarse and Compressed Semantic Speech Representation Learning
Predictive entropy from a token-level speech language model finds merge boundaries, producing compressed semantic tokens that keep ASR and translation accuracy at 15 Hz while lowering latency.
-
Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning
The authors release a rule-based Quran Phonetic Script, an 890-hour expert recitation dataset, and a multi-head CTC model that achieves 0.16% average phoneme error rate on held-out reciters.
-
Scaling and Distilling Transformer Models for sEMG
Vanilla transformers on the emg2qwerty dataset improve cross-user typing accuracy up to 109M parameters, and simple logit distillation recovers most of the gain in a 2.2M-parameter student.
-
OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
OpenBEATs releases the BEATs audio pretraining pipeline, trains 300M-parameter models on 20k hours of multi-domain audio, and reports strong results across 25 datasets, including bioacoustics and reasoning tasks.
-
Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation
A two-stage text-guidance framework, Think-Before-Draw, uses chain-of-thought prompting to convert emotion labels into facial muscle descriptions and then progressively guides a diffusion model from coarse emotion to ...
-
A Dataset for Automatic Assessment of TTS Quality in Spanish
A new Spanish-language dataset of 4,326 MOS-rated TTS audio clips enables automated naturalness prediction with a mean absolute error around 0.8 on a five-point scale.
-
Prosodic Structure Beyond Lexical Content: A Study of Self-Supervised Learning
A masked self-supervised model trained only on pitch, energy, and voice activity captures prosodic structure at multiple timescales, with random masking yielding the most generalizable representations.
-
StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion
StarVC is an autoregressive voice conversion model that generates text tokens before acoustic tokens, improving linguistic fidelity while retaining speaker similarity.
-
Model as Loss: A Self-Consistent Training Paradigm
Using the model's own encoder as a feature loss improves perceptual quality and iterative stability of a speech enhancement model compared with a WavLM-based loss.
-
InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing
Sparse-frame dubbing with adjacent-chunk keyframe sampling lets a streaming audio-video model produce full-body motion synchronized to new audio while preserving identity and camera motion.
-
Seeing Voices: Generating A-Roll Video from Audio with Mirage
Mirage generates photorealistic A-roll videos of people speaking directly from audio, using only joint self-attention over audio, text, and video tokens.
-
Seamless Dysfluent Speech Text Alignment for Disordered Speech Analysis
Neural LCS uses learned phoneme and word similarity instead of exact matches to align dysfluent speech to intended text, and it outperforms DTW and Hard LCS on simulated benchmarks.
-
Automatic classification of stop realisation with wav2vec2.0
wav2vec2.0 models classify stop burst presence with 88 to 94 percent accuracy, and a Japanese-pretrained model reaches near-peak accuracy on English with only 500 annotated tokens.
-
Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes
ADD-GP, a Gaussian Process classifier with XLS-R speech embeddings, adapts to unseen TTS models with as few as 5 samples and achieves state-of-the-art low error rates on the new LibriFake benchmark.
-
Contextualized Token Discrimination for Speech Search Query Correction
CTD uses BERT token representations plus a composition layer to correct Chinese spelling errors in ASR queries, but the reported gains lack matched baselines and released data.
-
EmoSLLM: Parameter-Efficient Adaptation of LLMs for Speech Emotion Recognition
EmoSLLM, a LoRA-fine-tuned 3B LLM with an audio mapper, beats several 7B speech-text LLMs on MSP-Podcast emotion recognition, but only when given the ground-truth transcript at inference.
-
Weak Supervision Techniques towards Enhanced ASR Models in Industry-level CRM Systems
Using LLM- and TTS-generated synthetic speech to fine-tune Whisper models cuts character error rates by up to 63% in a luxury retail CRM transcription task, though the evaluation has several methodological weaknesses.
-
LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models
Using consistency distillation, INT8 quantization, and pipeline parallelism, the LLIA system generates portrait video from audio at 78 FPS, with 140 ms initial latency.
-
Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques
Multimodal LLMs can be organized by fusion mechanism, fusion level, representation paradigm, and training paradigm, with 125 models classified accordingly.
-
SALF-MOS: Speaker Agnostic Latent Features Downsampled for MOS Prediction
SALF-MOS, a compact U-Net-style model using frozen wav2vec features, claims state-of-the-art MOS prediction on four benchmarks with only 1,574 parameters.
-
DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec
DS-Codec improves low-bitrate speech codec quality by first training a mirrored codec and then switching to a non-mirrored decoder, while using product quantization to form one large codebook.
-
Amplifying Emotional Signals: Data-Efficient Deep Learning for Robust Speech Emotion Recognition
A pretrained ResNet34 with augmentation reaches 66.7% accuracy on a combined RAVDESS/SAVEE emotion set, but only on a validation split, so the claimed new benchmark is unverified.
-
DeepEmoNet: Building Machine Learning Models for Automatic Emotion Recognition in Human Speeches
A ResNet34 pretrained on ImageNet and fine-tuned on log-mel spectrograms, with data augmentation, classifies eight speech emotions at 66.7% accuracy and F1 0.631 on the pooled RAVDESS/SAVEE validation set.
Discussion (0). Sign in to comment.