Pith. sign in

REVIEW 5 cited by

SNAC: Multi-Scale Neural Audio Codec

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.14411 v1 pith:U42ARJZ6 submitted 2024-10-18 cs.SD cs.LGeess.AS

classification cs.SDcs.LGeess.AS
keywords audioneuralcodeccompressionmulti-scalequantizerssnacacross
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Neural audio codecs have recently gained popularity because they can represent audio signals with high fidelity at very low bitrates, making it feasible to use language modeling approaches for audio generation and understanding. Residual Vector Quantization (RVQ) has become the standard technique for neural audio compression using a cascade of VQ codebooks. This paper proposes the Multi-Scale Neural Audio Codec, a simple extension of RVQ where the quantizers can operate at different temporal resolutions. By applying a hierarchy of quantizers at variable frame rates, the codec adapts to the audio structure across multiple timescales. This leads to more efficient compression, as demonstrated by extensive objective and subjective evaluations. The code and model weights are open-sourced at https://github.com/hubertsiuzdak/snac.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate

    eess.AS 2025-05 conditional novelty 7.0 of 10

    A neural speech codec that dynamically varies frame rate per segment using waveform entropy achieves competitive or better reconstruction quality at lower average frame rates than constant frame rate baselines.

  2. Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy

    cs.SD 2025-06 conditional novelty 6.0 of 10

    DCAR dynamically schedules chunk-wise token prediction in AR TTS, improving WER by up to 72.27% relative and speeding up inference by up to 2.89x over next-token baselines.

  3. StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion

    cs.MM 2025-06 conditional novelty 6.0 of 10

    StarVC is an autoregressive voice conversion model that generates text tokens before acoustic tokens, improving linguistic fidelity while retaining speaker similarity.

  4. Spectrogram Patch Codec: A 2D Block-Quantized VQ-VAE and HiFi-GAN for Neural Speech Coding

    cs.SD 2025-09 conditional novelty 5.0 of 10

    A 2D patch-quantized VQ-VAE with a single 4096-entry codebook plus a HiFi-GAN vocoder reaches ~7.5 kbits/s for 16 kHz speech, with intelligibility below EnCodec and DAC but a simpler, non-residual architecture.

  5. MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation

    cs.SD 2025-05 conditional novelty 5.0 of 10

    A single-layer streaming Transformer codec with masked Gaussian noise injection during training reports state-of-the-art reconstruction and better downstream generation and understanding in 16 kHz English speech.

Pith tools