Pith. sign in

REVIEW 8 cited by

Scaling Transformers for Low-Bitrate High-Quality Speech Coding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.19842 v1 pith:ZTLYMXND submitted 2024-11-29 eess.AS cs.AIcs.LGcs.SDeess.SP

classification eess.AScs.AIcs.LGcs.SDeess.SP
keywords speechmodelsscalingtokenizationaloneapplyingarchitecturearchitectures
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

The tokenization of speech with neural audio codec models is a vital part of modern AI pipelines for the generation or understanding of speech, alone or in a multimodal context. Traditionally such tokenization models have concentrated on low parameter-count architectures using only components with strong inductive biases. In this work we show that by scaling a transformer architecture with large parameter count to this problem, and applying a flexible Finite Scalar Quantization (FSQ) based bottleneck, it is possible to reach state-of-the-art speech quality at extremely low bit-rates of $400$ or $700$ bits-per-second. The trained models strongly out-perform existing baselines in both objective and subjective tests.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Next Tokens Denoising for Speech Synthesis

    cs.SD 2025-07 conditional novelty 6.0 of 10

    Dragon-FM generates speech autoregressively over two-second chunks while using flow matching inside each chunk, achieving fast synthesis at 12.5 discrete audio tokens per second.

  2. Probing the Robustness Properties of Neural Speech Codecs

    eess.AS 2025-05 conditional novelty 6.0 of 10

    DAC is the most noise-robust neural codec at high bitrates, but at 3 kbps EnCodec wins, and measured non-linearity correlates with robustness.

  3. CoDiCodec: Unifying Continuous and Discrete Compressed Representations of Audio

    cs.SD 2025-09 conditional novelty 5.0 of 10

    CoDiCodec unifies continuous and discrete audio compression in one consistency-trained autoencoder, using FSQ-dropout to serve both continuous ~11 Hz embeddings and 2.38 kbps discrete tokens.

  4. CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation

    eess.AS 2025-08 conditional novelty 5.0 of 10

    CodecBench ranks 14 audio codecs on acoustic fidelity and semantic preservation across 19 datasets and four audio domains, revealing a reconstruction-versus-semantics tradeoff.

  5. NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference

    eess.AS 2025-08 conditional novelty 5.0 of 10

    NanoCodec achieves competitive speech quality at 12.5 frames per second and 0.6-1.78 kbps, with a causal decoder for low-latency speech LLM inference.

  6. Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations

    cs.SD 2025-07 reject novelty 5.0 of 10

    QTTS models speech as sequences from a multi-codebook RVQ audio codec whose first codebook is trained with ASR supervision, aiming for higher-fidelity TTS than single-codebook systems.

  7. UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information

    cs.SD 2025-05 conditional novelty 5.0 of 10

    The authors propose DistilCodec, a 32,768-code single-codebook audio codec, and UniTTS, a Qwen2.5-7B TTS model trained with audio, text, and cross-modal autoregressive tasks on interleaved prompts.

  8. Robust Residual Finite Scalar Quantization for Neural Compression

    eess.IV 2025-08 reject novelty 4.0 of 10

    RFSQ applies learned scaling or invertible LayerNorm to residual FSQ to prevent magnitude decay, reporting DNSMOS and image loss gains, though the LayerNorm variant has a reconstruction inconsistency.

Pith tools