Pith. sign in

REVIEW 1 cited by

Vector-quantized neural networks for acoustic unit discovery in the ZeroSpeech 2020 challenge

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2005.09409 v2 pith:S5ZDM7S7 submitted 2020-05-19 eess.AS cs.CL

classification eess.AScs.CL
keywords modelsquantizationvectoracousticchallengespeechdatadiscovery
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we explore vector quantization for acoustic unit discovery. Leveraging unlabelled data, we aim to learn discrete representations of speech that separate phonetic content from speaker-specific details. We propose two neural models to tackle this challenge - both use vector quantization to map continuous features to a finite set of codes. The first model is a type of vector-quantized variational autoencoder (VQ-VAE). The VQ-VAE encodes speech into a sequence of discrete units before reconstructing the audio waveform. Our second model combines vector quantization with contrastive predictive coding (VQ-CPC). The idea is to learn a representation of speech by predicting future acoustic units. We evaluate the models on English and Indonesian data for the ZeroSpeech 2020 challenge. In ABX phone discrimination tests, both models outperform all submissions to the 2019 and 2020 challenges, with a relative improvement of more than 30%. The models also perform competitively on a downstream voice conversion task. Of the two, VQ-CPC performs slightly better in general and is simpler and faster to train. Finally, probing experiments show that vector quantization is an effective bottleneck, forcing the models to discard speaker information.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On Parallelism in Music and Language: A Perspective from Symbol Emergence Systems based on Probabilistic Generative Models

    cs.HC 2025-01 conditional novelty 4.0 of 10

    This paper proposes that the meaning of music emerges from interoceptive predictive coding within a multi-agent symbol emergence system, parallel to language.

Pith tools