Pith. sign in

REVIEW 6 cited by

RAVE: A variational autoencoder for fast and high-quality neural audio synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.05011 v2 pith:HGVEQG3T submitted 2021-11-09 cs.LG cs.SDeess.AS

classification cs.LGcs.SDeess.AS
keywords audiomodelssynthesiscontrolvariationalwaveformautoencoderfast
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep generative models applied to audio have improved by a large margin the state-of-the-art in many speech and music related tasks. However, as raw waveform modelling remains an inherently difficult task, audio generative models are either computationally intensive, rely on low sampling rates, are complicated to control or restrict the nature of possible signals. Among those models, Variational AutoEncoders (VAE) give control over the generation by exposing latent variables, although they usually suffer from low synthesis quality. In this paper, we introduce a Realtime Audio Variational autoEncoder (RAVE) allowing both fast and high-quality audio waveform synthesis. We introduce a novel two-stage training procedure, namely representation learning and adversarial fine-tuning. We show that using a post-training analysis of the latent space allows a direct control between the reconstruction fidelity and the representation compactness. By leveraging a multi-band decomposition of the raw waveform, we show that our model is the first able to generate 48kHz audio signals, while simultaneously running 20 times faster than real-time on a standard laptop CPU. We evaluate synthesis quality using both quantitative and qualitative subjective experiments and show the superiority of our approach compared to existing models. Finally, we present applications of our model for timbre transfer and signal compression. All of our source code and audio examples are publicly available.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Structural Bottlenecks on Frequency Representation in End-to-End Audio Models

    cs.SD 2026-07 conditional novelty 7.0 of 10

    State-of-the-art strided audio encoders impose predictable alias-collapse and resolution bottlenecks on frequency primitives; Gabor Latent Refactorization recovers much of the lost separability post-hoc.

  2. MusGO: A Community-Driven Framework For Assessing Openness in Music-Generative AI

    cs.SD 2025-07 conditional novelty 6.0 of 10

    MusGO is a community-refined framework with 13 openness categories, applied to 16 music-generative models to produce a public openness leaderboard.

  3. ANIRA: An Architecture for Neural Network Inference in Real-Time Audio Applications

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Anira, a new library for real-time audio neural network inference, is benchmarked across three engines, finding ONNX Runtime fastest for stateless models and LibTorch fastest for stateful models.

  4. Learning to Upsample and Upmix Audio in the Latent Domain

    cs.SD 2025-05 conditional novelty 6.0 of 10

    Lightweight networks trained only on autoencoder latent codes can do bandwidth extension and mono-to-stereo upmixing at a fraction of the FLOPS of raw-audio models, but match those models only when the baselines are a...

  5. Workflow-Based Evaluation of Music Generation Systems

    eess.AS 2025-06 conditional novelty 5.0 of 10

    A single-producer workflow evaluation of eight music AI tools finds they work as idea and sound generators but not as complete composers, and proposes a reusable framework.

  6. Two Sonification Methods for the MindCube

    cs.HC 2025-06 conditional novelty 4.0 of 10

    The MindCube, a handheld fidget cube, is turned into a musical interface through a generative-AI mapping that conditions music loudness on sensor activity, and a direct VCV Rack synthesizer mapping.

Pith tools