Pith. sign in

REVIEW 14 cited by

BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.05377 v1 pith:345VMX6L submitted 2024-09-09 eess.AS cs.SD

classification eess.AScs.SD
keywords bigcodeccodecslow-bitrateneuralperformancespeechbetterbitrate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present BigCodec, a low-bitrate neural speech codec. While recent neural speech codecs have shown impressive progress, their performance significantly deteriorates at low bitrates (around 1 kbps). Although a low bitrate inherently restricts performance, other factors, such as model capacity, also hinder further improvements. To address this problem, we scale up the model size to 159M parameters that is more than 10 times larger than popular codecs with about 10M parameters. Besides, we integrate sequential models into traditional convolutional architectures to better capture temporal dependency and adopt low-dimensional vector quantization to ensure a high code utilization. Comprehensive objective and subjective evaluations show that BigCodec, with a bitrate of 1.04 kbps, significantly outperforms several existing low-bitrate codecs. Furthermore, BigCodec achieves objective performance comparable to popular codecs operating at 4-6 times higher bitrates, and even delivers better subjective perceptual quality than the ground truth.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LILAC: An Idempotent Neural Speech Codec

    cs.SD 2026-08 conditional novelty 8.0 of 10

    LILAC is a fully convolutional 24 kHz speech codec whose invertible analysis transform and finite scalar quantization make every decode-re-encode cycle return the same token stream by construction.

  2. Optimising Neural Speech Codecs for 300bps Communication using Reinforcement Learning

    cs.SD 2026-05 unverdicted novelty 7.0 of 10

    ClariCodec applies GRPO reinforcement learning to a 300 bps neural speech codec, using ASR word-error rate as reward to cut LibriSpeech test-clean WER from 4.64% to 3.55%.

  3. ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models

    cs.SD 2026-07 conditional novelty 6.5 of 10

    Hierarchical multi-prompt representation generation plus generalized flow matching yields high-quality single-stage waveform diffusion from 12.5 Hz latents and efficient LDM TTS.

  4. The Equalizer: Introducing Shape-Gain Decomposition in Neural Audio Codecs

    cs.SD 2026-02 conditional novelty 6.0 of 10

    Adding classical shape-gain decomposition to a neural audio codec makes it invariant to input gain and improves bitrate-distortion performance.

  5. Aliasing-Free Neural Audio Synthesis

    cs.SD 2025-12 conditional novelty 6.0 of 10

    Pupu-Vocoder and Pupu-Codec use a closed-form anti-aliased SnakeBeta activation and resampling-based upsampling to reduce aliasing and improve singing, music, and audio synthesis.

  6. TTS-1 Technical Report

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A new TTS system family combines pre-training, supervised fine-tuning, and GRPO reinforcement learning with a 48 kHz codec to produce multilingual speech with in-context voice cloning.

  7. XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs

    cs.SD 2025-06 conditional novelty 6.0 of 10

    XY-Tokenizer is a 1 kbps dual-channel speech codec that reports simultaneously strong text alignment and high speaker similarity, comparable to specialized codecs at similar bitrates.

  8. Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation

    eess.AS 2025-05 conditional novelty 6.0 of 10

    Compressed-to-fine language modeling improves speech token prediction by retaining prompt and local tokens while compressing long-range token spans into compact summaries.

  9. Absorbing Discrete Diffusion for Speech Enhancement

    cs.SD 2026-02 conditional novelty 5.0 of 10

    ADDSE performs speech enhancement by absorbing discrete diffusion over neural audio codec tokens, reaching competitive non-intrusive quality and low-SNR robustness in few sampling steps.

  10. Finite Scalar Quantization Enables Redundant and Transmission-Robust Neural Audio Compression at Low Bit-rates

    cs.SD 2025-09 conditional novelty 5.0 of 10

    FSQ-based audio codecs outperform RVQ-based codecs in simulated bit-error transmission, preserving intelligibility at bit-flip rates up to 10%.

  11. CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation

    eess.AS 2025-08 conditional novelty 5.0 of 10

    CodecBench ranks 14 audio codecs on acoustic fidelity and semantic preservation across 19 datasets and four audio domains, revealing a reconstruction-versus-semantics tradeoff.

  12. MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation

    cs.SD 2025-05 conditional novelty 5.0 of 10

    A single-layer streaming Transformer codec with masked Gaussian noise injection during training reports state-of-the-art reconstruction and better downstream generation and understanding in 16 kHz English speech.

  13. UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information

    cs.SD 2025-05 conditional novelty 5.0 of 10

    The authors propose DistilCodec, a 32,768-code single-codebook audio codec, and UniTTS, a Qwen2.5-7B TTS model trained with audio, text, and cross-modal autoregressive tasks on interleaved prompts.

  14. DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec

    cs.SD 2025-05 conditional novelty 4.0 of 10

    DS-Codec improves low-bitrate speech codec quality by first training a mirrored codec and then switching to a non-mirrored decoder, while using product quantization to form one large codebook.

Pith tools