Pith. sign in

Paper Citation Record · LEDGER

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation

As of 8 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2508.20660.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.20660 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:00:14.179533Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9045120-c6bd-4e02-a8ef-0072e4086007 · outbound

This paper cites Sdr–half-baked or well done? In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Sdr–half-baked or well done? In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:15.022586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:00:14.014177Z digest=sha256:81ddd53841d9a4d5cccf900f01062c0d5254b6457e8f0805c406efcb1a120b05

Observation 0c97b505-afe9-45f2-9ef1-596b508117c7 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Librispeech: an asr corpus based on public domain audio books

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:14.924421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:00:14.032179Z digest=sha256:f592410d5ceac28252a0b6322c25d428443282123e078a7b0d38a3b57095195e

Observation 8f3a2d7c-3c5c-4a82-ad09-b6f4e4ee3e4e · outbound

This paper cites MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.045603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.045603Z digest=sha256:dd5d678ff4ac466e37042f18293610ac1bf96041159bb9c21660bd242c895ed0

Observation 49e0a2e3-dd50-461a-a6b3-fc6499b242ff · outbound

This paper cites A short-time objective intelligibility measure for time-frequency weighted noisy speech.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation A short-time objective intelligibility measure for time-frequency weighted noisy speech

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:14.812907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:00:14.064215Z digest=sha256:eea5677f12909edc71dd0f29d6061cb40fc7a243f8ade3236c8bb17cc53a2fad

Observation 92059eaf-e1d3-4fd6-a32d-58e5de8d0d8c · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.073906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.073906Z digest=sha256:980edbf2279dd8c99f9781414c0fc1d215efd1f6d7bf9c1fb3e4f1a9f3c63476

Observation cd815554-0fb7-4877-a0b9-f2406c475e7b · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.080956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.080956Z digest=sha256:478ffc76d1456774d8d6a3c4445cc3579074b544773846efa7076fae774a4518

Observation 379d0d46-c6e6-48b0-b389-e657e0d22e09 · outbound

This paper cites FlowDec: A flow-based full-band general audio codec with high perceptual quality.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation FlowDec: A flow-based full-band general audio codec with high perceptual quality

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.089389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.089389Z digest=sha256:c2384bfc966916388086d25a8e901e274348b046facdf6c1a8ce48ef8e68cbe6

Observation dd009f05-b7c0-49a8-bca3-cdaa2670d4a1 · outbound

This paper cites Codec-SUPERB: An In-Depth Analysis of Sound Codec Models.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Codec-SUPERB: An In-Depth Analysis of Sound Codec Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.103442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.103442Z digest=sha256:d1dd436f19abfa411c034e9f7af751f35a9977d5cfc278beea3b3f5148e2b16d

Observation ef9e0fda-ef88-4633-b2cc-8b8da6e8ec72 · outbound

This paper cites Laughter Synthesis using Pseudo Phonetic Tokens with a Large-scale In-the-wild Laughter Corpus.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Laughter Synthesis using Pseudo Phonetic Tokens with a Large-scale In-the-wild Laughter Corpus

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-05T15:00:14.381991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:00:14.116555Z digest=sha256:a354db638beb4ae8eab8f27c4372700cee20a1afc21bc470a82173e1088e8b69

Observation d06df0a6-9d79-4011-803e-22139be9f85e · outbound

This paper cites BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.133170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.133170Z digest=sha256:166305cf9b38c6b435af994fca8e9512eb9a7a0c930184e90bbf27406cd10d7e

Observation ccf870b0-a18d-4c93-abe1-94eff0dcee0c · outbound

This paper cites Qwen2.5 Technical Report.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Qwen2.5 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.143419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.143419Z digest=sha256:6eb31a6f69af3bfa6ceefd47d548c5cfcd25bb102e46d952db0b15f898fee7a6

Observation 31ade6f1-8007-4a2f-84af-848fe68826cc · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.150027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.150027Z digest=sha256:ffacac6c83f3493477b2eeedd5151fdbf188af9b67e5e0cbdc6f05ce44fce200

Observation c53f719a-b5e5-4271-b8d6-e98932066d8e · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.159489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.159489Z digest=sha256:292200b7f1d49934d8ab1858503cd952d3ad0e0ad8b9c3d81796441ca490e07a

Observation 93203482-b7d1-49f0-8ed7-7fd2812cbe20 · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.164561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.164561Z digest=sha256:ea05ba0aaee6e3be8a6e2c568bddb510e437caa91226f8e46eed88f500aa8eda

Observation 5e2992c9-eb93-4c69-8f52-ea94bc1765ae · outbound

This paper cites The clips within this dataset are manually selected from public field recordings compiled by the Freesound.org project.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation The clips within this dataset are manually selected from public field recordings compiled by the Freesound.org project

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:14.706512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:00:14.173358Z digest=sha256:875f4e45839e76932e93b09b9a1018da374d64adaf8268401e73257a0308f11d

Observation 5445decb-3782-43f6-ab4d-934660afc49d · outbound

This paper cites an unresolved cited work.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Unresolved cited work

Reference 28

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T15:00:14.675893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:00:14.179533Z digest=sha256:393ac4e6e7d85be4b8b5d3acf8efb96f5398aaf21bafc559d4401eefd6ca09d0

Observation 2fb3cf35-e48e-468f-83b3-8bc50d89a366 · outbound

This paper cites VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.058277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.058277Z digest=sha256:c1fb2a7ba16a1b21986a964e95eb2ba5ed72a6edfeaf999946deb71682ceb7b6

Observation 71afe80c-116d-42cc-b965-6cffa81d88dc · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Audio set: An ontology and human-labeled dataset for audio events

Reference 2005

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:15.077517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:00:13.991843Z digest=sha256:b64f005e22a21f90540704c9647726b858c44abfe324364efa53edb19964fcb2

Observation e4178831-f504-4363-8fbf-12f1d8439c64 · outbound

This paper cites Visqol v3: An open source production ready objective speech and audio metric.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Visqol v3: An open source production ready objective speech and audio metric

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:15.100047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:00:13.973695Z digest=sha256:88519837630093b7c3cdb07fd30d850a34b9fd7dac229bfb02e537f9fb6241f6

Observation a1fb8b1a-4fd4-46d0-a271-3b7e7aeec344 · outbound

This paper cites Scaling Transformers for Low-Bitrate High-Quality Speech Coding.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Scaling Transformers for Low-Bitrate High-Quality Speech Coding

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.039324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.039324Z digest=sha256:fe50940c76dd69800d325863604057163c3b34e7fbcf49df69c3a82e2c58b633

Observation 3190eddb-5666-4dde-8754-2cc8bf56d42b · outbound

This paper cites 1632–1636,.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation 1632–1636,

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:14.741535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:00:14.069461Z digest=sha256:e69725f3ab964e9fff627dfe8b6deee8b1db0eeb87e259a77cf842fe98d20e50

Observation ef66f28f-a91b-48ba-9091-de9ec5c49eb6 · outbound

This paper cites V ocalsound: A dataset for improving human vocal sounds recognition.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation V ocalsound: A dataset for improving human vocal sounds recognition

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:15.059689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:00:13.999297Z digest=sha256:5dfc0262dd35678ea02c5b828ac136578be7183e9463d81e837ac1f8c1cc2731

Observation 40b1610e-2ed8-40ec-b53c-be989790466a · outbound

This paper cites Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.020393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.020393Z digest=sha256:df2ff135ef15d2480653be614891678778e5a768abbde72b99341df69fc98c62

Observation 9db2a1b1-70dc-40ce-a056-72781458b97b · outbound

This paper cites LibriMix: An Open-Source Dataset for Generalizable Speech Separation.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:13.980353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:13.980353Z digest=sha256:c3b5103a32bf62b1088a6c42aed6c6c4216901be313051c3c34f679c78dc7710

Observation 69c6417a-e655-405f-ad0b-77ab70ec46da · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Robust Speech Recognition via Large-Scale Weak Supervision

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.052024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.052024Z digest=sha256:296d6dc20b0f137325b08e0060ead6968bf3688836f91bcb64ff9730b89acd60

Observation 05f7c0ff-f8eb-47aa-8c09-462361b66ef4 · outbound

This paper cites Benchmarking representations for speech, music, and acoustic events.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Benchmarking representations for speech, music, and acoustic events

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:15.040602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:00:14.008369Z digest=sha256:b5a2b1118fe7786205cef83dba7c69f04bedb70ae56bde5114bf32db01b4ca92

Observation d8fd45dc-73ae-4d33-b7fc-36a93b950a4c · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Moshi: a speech-text foundation model for real-time dialogue

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:13.986387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:13.986387Z digest=sha256:9ee0c8cd965d895b9eb5e44ccb49349f645aaa7495d889491a131d396c7a0d29

Observation c2a2d476-99a5-47f2-b5ce-84675ebd4f29 · outbound

This paper cites Clotho- aqa: A crowdsourced dataset for audio question answering.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Clotho- aqa: A crowdsourced dataset for audio question answering

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:15.001969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T15:00:14.026878Z digest=sha256:d1987a2ef9223bb637727108c1bf38ed4beb555ca21b08da9a823ae3c50b9b0d

Pith citing papers

No inbound Pith citation observations are available.