Pith. sign in

REVIEW 20 cited by

HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.02765 v2 pith:HYPSWSNL submitted 2023-05-04 cs.SD eess.AS

classification cs.SDeess.AS
keywords audiocodecmodelsgenerationhifi-codecmodeltrainingacademicodec
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Audio codec models are widely used in audio communication as a crucial technique for compressing audio into discrete representations. Nowadays, audio codec models are increasingly utilized in generation fields as intermediate representations. For instance, AudioLM is an audio generation model that uses the discrete representation of SoundStream as a training target, while VALL-E employs the Encodec model as an intermediate feature to aid TTS tasks. Despite their usefulness, two challenges persist: (1) training these audio codec models can be difficult due to the lack of publicly available training processes and the need for large-scale data and GPUs; (2) achieving good reconstruction performance requires many codebooks, which increases the burden on generation models. In this study, we propose a group-residual vector quantization (GRVQ) technique and use it to develop a novel \textbf{Hi}gh \textbf{Fi}delity Audio Codec model, HiFi-Codec, which only requires 4 codebooks. We train all the models using publicly available TTS data such as LibriTTS, VCTK, AISHELL, and more, with a total duration of over 1000 hours, using 8 GPUs. Our experimental results show that HiFi-Codec outperforms Encodec in terms of reconstruction performance despite requiring only 4 codebooks. To facilitate research in audio codec and generation, we introduce AcademiCodec, the first open-source audio codec toolkit that offers training codes and pre-trained models for Encodec, SoundStream, and HiFi-Codec. Code and pre-trained model can be found on: \href{https://github.com/yangdongchao/AcademiCodec}{https://github.com/yangdongchao/AcademiCodec}

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Optimising Neural Speech Codecs for 300bps Communication using Reinforcement Learning

    cs.SD 2026-05 unverdicted novelty 7.0 of 10

    ClariCodec applies GRPO reinforcement learning to a 300 bps neural speech codec, using ASR word-error rate as reward to cut LibriSpeech test-clean WER from 4.64% to 3.55%.

  2. ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models

    cs.SD 2026-07 conditional novelty 6.5 of 10

    Hierarchical multi-prompt representation generation plus generalized flow matching yields high-quality single-stage waveform diffusion from 12.5 Hz latents and efficient LDM TTS.

  3. Investigating Codec-Internal Latent Audio Watermarking for Neural Codec Robustness

    cs.SD 2026-07 conditional novelty 6.0 of 10

    Embedding watermarks inside a codec-like autoencoder's continuous latent space improves EnCodec-24k bit accuracy to ~95–97%, but the gain is in-distribution and does not transfer to EnCodec-16k.

  4. Unified Audio Intelligence Without Regressing on Text Intelligence

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A unified 30B MoE audio-text LLM achieves state-of-the-art audio understanding, generation, and speech tasks while preserving text reasoning comparable to its text-only backbone.

  5. The Equalizer: Introducing Shape-Gain Decomposition in Neural Audio Codecs

    cs.SD 2026-02 conditional novelty 6.0 of 10

    Adding classical shape-gain decomposition to a neural audio codec makes it invariant to input gain and improves bitrate-distortion performance.

  6. DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation Learners

    cs.SD 2025-09 conditional novelty 6.0 of 10

    DeCodec learns a single neural codec that disentangles speech, background sound, semantic content, and paralinguistic style into orthogonal quantized streams, enabling reconstruction, enhancement, voice conversion, AS...

  7. AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation

    cs.SD 2025-09 conditional novelty 6.0 of 10

    A four-part benchmark plus a semantic/acoustic token taxonomy for comparing audio codecs, with correlation analysis across ten models.

  8. MDD: a Mask Diffusion Detector to Protect Speaker Verification Systems from Adversarial Perturbations

    eess.AS 2025-08 conditional novelty 6.0 of 10

    MDD, a text-conditioned masked diffusion detector, reports 98% detection of PGD adversarial audio and lowers ASV EER from 73.2% to 18.0%.

  9. Representing Speech Through Autoregressive Prediction of Cochlear Tokens

    cs.CL 2025-08 conditional novelty 6.0 of 10

    Autoregressive prediction over discrete cochlear tokens yields a speech representation that beats prior self-supervised models on lexical-semantic similarity and is competitive on SUPERB tasks.

  10. Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A dexterous VLA pretrained on a 2.5M-instance human hand motion dataset transfers skills to a real robot hand, outperforming baselines in manipulation tasks.

  11. Autoregressive Speech Enhancement via Acoustic Tokens

    cs.SD 2025-07 conditional novelty 6.0 of 10

    Acoustic tokens outperform semantic tokens on speaker identity in speech enhancement, an autoregressive transducer helps in some settings, but discrete representations still lag continuous ones.

  12. Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation

    eess.AS 2025-05 conditional novelty 6.0 of 10

    Compressed-to-fine language modeling improves speech token prediction by retaining prompt and local tokens while compressing long-range token spans into compact summaries.

  13. Probing the Robustness Properties of Neural Speech Codecs

    eess.AS 2025-05 conditional novelty 6.0 of 10

    DAC is the most noise-robust neural codec at high bitrates, but at 3 kbps EnCodec wins, and measured non-linearity correlates with robustness.

  14. Vision-Integrated High-Quality Neural Speech Coding

    eess.AS 2025-05 conditional novelty 6.0 of 10

    A vision-integrated neural speech codec that uses lip images, via explicit fusion or distillation, improves decoded speech quality and noise robustness at 6 kbps.

  15. PAST: Phonetic-Acoustic Speech Tokenizer

    cs.SD 2025-05 conditional novelty 6.0 of 10

    PAST jointly optimizes an EnCodec-style codec with CTC and phoneme classification losses to produce hybrid phonetic-acoustic speech tokens that outperform SpeechTokenizer and X-Codec.

  16. Analysing the Language of Neural Audio Codecs

    cs.CL 2025-09 conditional novelty 5.0 of 10

    NAC token sequences, especially 3-grams, fit Zipf's and Heaps' laws, and their fitted statistics correlate, though weakly and with confounds, with ASR error rates and UTMOS scores.

  17. Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges

    eess.AS 2025-06 conditional novelty 5.0 of 10

    A speech-deepfake dataset for ten public figures built with transcription-based segmentation reports high synthetic naturalness (NISQA 3.69) and a human misclassification rate of 61.9%.

  18. UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information

    cs.SD 2025-05 conditional novelty 5.0 of 10

    The authors propose DistilCodec, a 32,768-code single-codebook audio codec, and UniTTS, a Qwen2.5-7B TTS model trained with audio, text, and cross-modal autoregressive tasks on interleaved prompts.

  19. EASY: Emotion-aware Speaker Anonymization via Factorized Distillation

    eess.AS 2025-05 conditional novelty 5.0 of 10

    EASY separates speaker identity, linguistic content, and emotion through sequential factorized distillation, and reports better privacy and emotion preservation than prior VoicePrivacy 2024 systems.

  20. DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec

    cs.SD 2025-05 conditional novelty 4.0 of 10

    DS-Codec improves low-bitrate speech codec quality by first training a mirrored codec and then switching to a non-mirrored decoder, while using product quantization to form one large codebook.

Pith tools