Pith. sign in

REVIEW 10 cited by

High-Fidelity Audio Compression with Improved RVQGAN

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.06546 v2 pith:GYQ33UGC submitted 2023-06-11 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords audiocompressionhigh-fidelitymodelcompressgenerationimprovedmodeling
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Language models have been successfully used to model natural signals, such as images, speech, and music. A key component of these models is a high quality neural compression model that can compress high-dimensional natural signals into lower dimensional discrete tokens. To that end, we introduce a high-fidelity universal neural audio compression algorithm that achieves ~90x compression of 44.1 KHz audio into tokens at just 8kbps bandwidth. We achieve this by combining advances in high-fidelity audio generation with better vector quantization techniques from the image domain, along with improved adversarial and reconstruction losses. We compress all domains (speech, environment, music, etc.) with a single universal model, making it widely applicable to generative modeling of all audio. We compare with competing audio compression algorithms, and find our method outperforms them significantly. We provide thorough ablations for every design choice, as well as open-source code and trained model weights. We hope our work can lay the foundation for the next generation of high-fidelity audio modeling.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 24 citations worldwide. Full citation record

  1. On the Geometry of Music Bandwidth Extension in Latent Spaces of Audio Codecs

    cs.SD 2026-08 conditional novelty 6.0 of 10

    A single mean shift in the latent space of several neural codecs achieves competitive music bandwidth extension on some metrics, implying a largely linear structure.

  2. SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

    eess.AS 2026-08 conditional novelty 6.0 of 10

    SwanTale unifies instruction-driven and zero-shot speech and audio generation in one 48 kHz model, with a large captioning pipeline, and reports leading scores on several expressiveness and instruction-following benchmarks.

  3. OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Jointly training audio and video VAEs with segment contrastive loss and semantic distillation yields more learnable, cross-aligned latents that improve downstream joint generation quality and sync.

  4. Investigating Codec-Internal Latent Audio Watermarking for Neural Codec Robustness

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Embedding watermarks inside a codec-like autoencoder's continuous latent space improves EnCodec-24k bit accuracy to ~95–97%, but the gain is in-distribution and does not transfer to EnCodec-16k.

  5. ITGPT: A Transformer Based Architecture for the Generation of Dance Dance Revolution and In the Groove Charts

    cs.SD 2026-07 conditional novelty 6.0 of 10

    ITGPT, a transformer pipeline, generates DDR/ITG arrow charts from audio with better accuracy and roughly 7x lower generation time than the prior ConvLSTM approach.

  6. AI-Generated Song Detection via Lyrics Transcripts

    cs.SD 2025-06 conditional novelty 6.0 of 10

    Transcribing audio with Whisper and classifying the transcript with LLM2Vec detects AI-generated songs from audio alone, nearly matching clean-lyrics accuracy and beating audio-based detectors under perturbations and ...

  7. SEMamba++: A General Speech Restoration Framework Leveraging Global, Local, and Periodic Spectral Patterns

    eess.AS 2026-03 conditional novelty 5.5 of 10

    SEMamba++ combines Frequency GLP (FAN-based global-periodic + local conv) with multi-resolution parallel TFDP and learnable softplus mapping to outperform GSR baselines on VCTK, URGENT and AATC while remaining efficient.

  8. Balancing Information Preservation and Disentanglement in Self-Supervised Music Representation Learning

    cs.SD 2025-07 conditional novelty 5.0 of 10

    A multi-view SSL framework with combined reconstruction and separation-based contrastive losses obtains disentangled pitch and instrument subspaces without the accuracy loss seen with contrastive-only training on NSynth.

  9. Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine

    cs.SD 2025-07 conditional novelty 5.0 of 10

    Quantizing an intermediate layer of a pretrained audio model with residual vector quantization, and finetuning the model with task and codebook losses, preserves ASR and audio classification accuracy at bitrates near ...

  10. Analysis of Speaker Verification Performance Trade-offs with Neural Audio Codec Transmission

    cs.SD 2025-09 conditional novelty 4.0 of 10

    Neural audio codecs match or beat Opus for speaker verification on VoxCeleb1 below 12 kbps and stay within about 1.5 percentage points EER above it.

Pith tools