Pith. sign in

REVIEW 1 cited by

An Information-Theoretic Analysis of Self-supervised Discrete Representations of Speech

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.02405 v1 pith:WETM7BN7 submitted 2023-06-04 cs.CL

classification cs.CL
keywords discretephoneticspeechunitsself-supervisedcategoriesdistributionsframework
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Self-supervised representation learning for speech often involves a quantization step that transforms the acoustic input into discrete units. However, it remains unclear how to characterize the relationship between these discrete units and abstract phonetic categories such as phonemes. In this paper, we develop an information-theoretic framework whereby we represent each phonetic category as a distribution over discrete units. We then apply our framework to two different self-supervised models (namely wav2vec 2.0 and XLSR) and use American English speech as a case study. Our study demonstrates that the entropy of phonetic distributions reflects the variability of the underlying speech sounds, with phonetically similar sounds exhibiting similar distributions. While our study confirms the lack of direct, one-to-one correspondence, we find an intriguing, indirect relationship between phonetic categories and discrete units.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improved Intelligibility of Dysarthric Speech using Conditional Flow Matching

    cs.SD 2025-06 conditional novelty 6.0 of 10

    A conditional flow matching model using WavLM-derived discrete units converts dysarthric speech to a synthesized clean voice with 31.3% WER and 3.9 MOS, outperforming a mel-spectrogram model (84.1% WER).

Pith tools