REVIEW 3 cited by
Opening the Black Box of wav2vec Feature Encoder
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Self-supervised models, namely, wav2vec and its variants, have shown promising results in various downstream tasks in the speech domain. However, their inner workings are poorly understood, calling for in-depth analyses on what the model learns. In this paper, we concentrate on the convolutional feature encoder where its latent space is often speculated to represent discrete acoustic units. To analyze the embedding space in a reductive manner, we feed the synthesized audio signals, which is the summation of simple sine waves. Through extensive experiments, we conclude that various information is embedded inside the feature encoder representations: (1) fundamental frequency, (2) formants, and (3) amplitude, packed with (4) sufficient temporal detail. Further, the information incorporated inside the latent representations is analogous to spectrograms but with a fundamental difference: latent representations construct a metric space so that closer representations imply acoustic similarity.
Forward citations
Cited by 3 Pith papers
-
Phone Segmentation and Recognition through Phonological Activation Mapping
SPAM projects S3M frames onto phonological vectors and uses gradient-free heads to jointly segment and recognize phones from under a minute of labels, generalizing to unseen phones and languages.
-
Automatic classification of stop realisation with wav2vec2.0
wav2vec2.0 models classify stop burst presence with 88 to 94 percent accuracy, and a Japanese-pretrained model reaches near-peak accuracy on English with only 500 annotated tokens.
-
Leveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation Assessment
Modeling each phoneme as a 32-component Gaussian mixture of self-supervised speech features improves atypical pronunciation scoring on four of five datasets, with S3Ms showing stronger allophonic structure than MFCCs ...
Discussion (0). Continue with ORCID to comment.