Pith. sign in

REVIEW 10 cited by

Towards audio language modeling -- an overview

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.13236 v1 pith:P3VZ24PI submitted 2024-02-20 eess.AS cs.SD

classification eess.AScs.SD
keywords audiocodecsneuralcodec-basedcodeslanguagemodelsoverview
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Neural audio codecs are initially introduced to compress audio data into compact codes to reduce transmission latency. Researchers recently discovered the potential of codecs as suitable tokenizers for converting continuous audio into discrete codes, which can be employed to develop audio language models (LMs). Numerous high-performance neural audio codecs and codec-based LMs have been developed. The paper aims to provide a thorough and systematic overview of the neural audio codec models and codec-based LMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Testing chatbots on the creation of encoders for audio conditioned image generation

    cs.SD 2025-09 conditional novelty 6.0 of 10

    All chatbot-designed audio encoders failed to align with CLIP text embeddings and produced incoherent images, while showing a surprising architectural similarity across chatbots.

  2. XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs

    cs.SD 2025-06 conditional novelty 6.0 of 10

    XY-Tokenizer is a 1 kbps dual-channel speech codec that reports simultaneously strong text alignment and high speaker similarity, comparable to specialized codecs at similar bitrates.

  3. GenVC: Self-Supervised Zero-Shot Voice Conversion

    eess.AS 2025-02 conditional novelty 6.0 of 10

    GenVC performs self-supervised zero-shot voice conversion with an autoregressive language model and a Perceiver style embedding, achieving strong similarity and anonymization.

  4. Analysing the Language of Neural Audio Codecs

    cs.CL 2025-09 conditional novelty 5.0 of 10

    NAC token sequences, especially 3-grams, fit Zipf's and Heaps' laws, and their fitted statistics correlate, though weakly and with confounds, with ASR error rates and UTMOS scores.

  5. CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis

    eess.AS 2025-08 conditional novelty 5.0 of 10

    CLEAR is a zero-shot TTS model that autoregressively predicts compact continuous audio latents with a per-token rectified flow head, reaching 1.88% WER on LibriSpeech Subset-B with an RTF of 0.29 and a 96 ms streaming delay.

  6. Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine

    cs.SD 2025-07 conditional novelty 5.0 of 10

    Quantizing an intermediate layer of a pretrained audio model with residual vector quantization, and finetuning the model with task and codebook losses, preserves ASR and audio classification accuracy at bitrates near ...

  7. Towards Generalized Source Tracing for Codec-Based Deepfake Speech

    cs.SD 2025-06 conditional novelty 5.0 of 10

    SASTNet, which fuses Whisper semantic features with Wav2Vec2 and AudioMAE acoustic features, improves source tracing for codec-based deepfake speech on CodecFake+, while exposing that prior models overfit to silence.

  8. Effectively obtaining acoustic, visual and textual data from videos

    cs.MM 2025-09 conditional novelty 4.0 of 10

    A video-processing pipeline created a 2.24 million-sample audio-image-text dataset, with text captions generated by BLIP from video frames.

  9. Learning Sparsity for Effective and Efficient Music Performance Question Answering

    cs.SD 2025-06 conditional novelty 4.0 of 10

    Sparsify reports state-of-the-art accuracy on Music AVQA benchmarks by borrowing three existing sparsification techniques, cutting training time by 28% and retaining 70-80% of accuracy on a 25% data subset.

  10. Foundation Model Driven Robotics: A Comprehensive Review

    cs.RO 2025-07 conditional novelty 2.0 of 10

    A review of foundation-model-driven robotics that synthesizes recent work across perception, planning, control, HRI, simulation, and sim-to-real transfer, and highlights open challenges.

Pith tools