Pith. sign in

REVIEW 8 cited by

Codec-SUPERB: An In-Depth Analysis of Sound Codec Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.13071 v3 pith:VANXWYVM submitted 2024-02-20 eess.AS cs.SD

classification eess.AScs.SD
keywords codecsoundmodelscodec-superbanalysisbenchmarkdatain-depth
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The sound codec's dual roles in minimizing data transmission latency and serving as tokenizers underscore its critical importance. Recent years have witnessed significant developments in codec models. The ideal sound codec should preserve content, paralinguistics, speakers, and audio information. However, the question of which codec achieves optimal sound information preservation remains unanswered, as in different papers, models are evaluated on their selected experimental settings. This study introduces Codec-SUPERB, an acronym for Codec sound processing Universal PERformance Benchmark. It is an ecosystem designed to assess codec models across representative sound applications and signal-level metrics rooted in sound domain knowledge.Codec-SUPERB simplifies result sharing through an online leaderboard, promoting collaboration within a community-driven benchmark database, thereby stimulating new development cycles for codecs. Furthermore, we undertake an in-depth analysis to offer insights into codec models from both application and signal perspectives, diverging from previous codec papers mainly concentrating on signal-level comparisons. Finally, we will release codes, the leaderboard, and data to accelerate progress within the community.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ViSAGe: Video-to-Spatial Audio Generation

    cs.SD 2025-06 conditional novelty 7.0 of 10

    A new model generates first-order ambisonics spatial audio directly from silent video and camera direction, with a new 102K-clip dataset and spatial evaluation metrics.

  2. AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation

    cs.SD 2025-09 conditional novelty 6.0 of 10

    A four-part benchmark plus a semantic/acoustic token taxonomy for comparing audio codecs, with correlation analysis across ten models.

  3. Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy

    cs.SD 2025-06 conditional novelty 6.0 of 10

    DCAR dynamically schedules chunk-wise token prediction in AR TTS, improving WER by up to 72.27% relative and speeding up inference by up to 2.89x over next-token baselines.

  4. Probing the Robustness Properties of Neural Speech Codecs

    eess.AS 2025-05 conditional novelty 6.0 of 10

    DAC is the most noise-robust neural codec at high bitrates, but at 3 kbps EnCodec wins, and measured non-linearity correlates with robustness.

  5. Towards General Discrete Speech Codec for Complex Acoustic Environments: A Study of Reconstruction and Downstream Task Consistency

    cs.SD 2025-05 conditional novelty 6.0 of 10

    A new benchmark, ERSB, shows that neural speech codecs in noisy environments degrade both reconstruction quality and downstream speech enhancement and recognition consistency.

  6. CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation

    eess.AS 2025-08 conditional novelty 5.0 of 10

    CodecBench ranks 14 audio codecs on acoustic fidelity and semantic preservation across 19 datasets and four audio domains, revealing a reconstruction-versus-semantics tradeoff.

  7. Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model

    cs.SD 2025-06 conditional novelty 5.0 of 10

    Step-Audio-AQAA, a 130B end-to-end audio language model using dual-codebook tokens, text-audio interleaving, masked DPO and weight merging, is claimed to outperform Kimi-Audio and Qwen-Omni on the authors' StepEval-Au...

  8. Analysis of Speaker Verification Performance Trade-offs with Neural Audio Codec Transmission

    cs.SD 2025-09 conditional novelty 4.0 of 10

    Neural audio codecs match or beat Opus for speaker verification on VoxCeleb1 below 12 kbps and stay within about 1.5 percentage points EER above it.

Pith tools