REVIEW 5 cited by
SNAC: Multi-Scale Neural Audio Codec
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Neural audio codecs have recently gained popularity because they can represent audio signals with high fidelity at very low bitrates, making it feasible to use language modeling approaches for audio generation and understanding. Residual Vector Quantization (RVQ) has become the standard technique for neural audio compression using a cascade of VQ codebooks. This paper proposes the Multi-Scale Neural Audio Codec, a simple extension of RVQ where the quantizers can operate at different temporal resolutions. By applying a hierarchy of quantizers at variable frame rates, the codec adapts to the audio structure across multiple timescales. This leads to more efficient compression, as demonstrated by extensive objective and subjective evaluations. The code and model weights are open-sourced at https://github.com/hubertsiuzdak/snac.
Forward citations
Cited by 5 Pith papers
-
Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate
A neural speech codec that dynamically varies frame rate per segment using waveform entropy achieves competitive or better reconstruction quality at lower average frame rates than constant frame rate baselines.
-
Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy
DCAR dynamically schedules chunk-wise token prediction in AR TTS, improving WER by up to 72.27% relative and speeding up inference by up to 2.89x over next-token baselines.
-
StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion
StarVC is an autoregressive voice conversion model that generates text tokens before acoustic tokens, improving linguistic fidelity while retaining speaker similarity.
-
Spectrogram Patch Codec: A 2D Block-Quantized VQ-VAE and HiFi-GAN for Neural Speech Coding
A 2D patch-quantized VQ-VAE with a single 4096-entry codebook plus a HiFi-GAN vocoder reaches ~7.5 kbits/s for 16 kHz speech, with intelligibility below EnCodec and DAC but a simpler, non-residual architecture.
-
MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation
A single-layer streaming Transformer codec with masked Gaussian noise injection during training reports state-of-the-art reconstruction and better downstream generation and understanding in 16 kHz English speech.
Discussion (0). Sign in to comment.