Pith. sign in

REVIEW 1 cited by

A Framework for Generative and Contrastive Learning of Audio Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.11459 v2 pith:U2OWL7YT submitted 2020-10-22 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords audiocontrastivelabelslearningrepresentationssupervisedaccessgenerative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we present a framework for contrastive learning for audio representations, in a self supervised frame work without access to any ground truth labels. The core idea in self supervised contrastive learning is to map an audio signal and its various augmented versions (representative of salient aspects of audio like pitch, timbre etc.) to a space where they are close together, and are separated from other different signals. In addition we also explore generative models based on state of the art transformer based architectures for learning latent spaces for audio signals, without access to any labels. Here, we map audio signals on a smaller scale to discrete dictionary elements and train transformers to predict the next dictionary element. We only use data as a method of supervision, bypassing the need of labels needed to act as a supervision for training the deep neural networks. We then use a linear classifier head in order to evaluate the performance of our models, for both self supervised contrastive and generative transformer based representations that are learned. Our system achieves considerable performance, compared to a fully supervised method, with access to ground truth labels to train the neural network model. These representations, with avail-ability of large scale audio data show promise in various tasks for audio understanding tasks

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music

    cs.SD 2024-12 conditional novelty 6.0 of 10

    A hybrid causal transformer that combines mel-spectrogram frames with EnCodec acoustic tokens matches or beats a 10-times larger token-only GPT on next-token likelihood for speech and music.

Pith tools