Pith. sign in

Paper Citation Record · LEDGER

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation

As of 20 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2508.20660.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.20660 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:00:14.179533Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9045120-c6bd-4e02-a8ef-0072e4086007 · outbound

This paper cites Sdr–half-baked or well done? In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Sdr–half-baked or well done? In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:15.022586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:00:14.014177Z digest=sha256:16b501a88c7fca81f15c40b300bba29312e244411e314fd84801bd467e00b7b3

Observation 0c97b505-afe9-45f2-9ef1-596b508117c7 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Librispeech: an asr corpus based on public domain audio books

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:14.924421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:00:14.032179Z digest=sha256:b6a2288d72b858fa88fd076daad2e89bfeaaabe76690242407e03e2032270b1e

Observation 8f3a2d7c-3c5c-4a82-ad09-b6f4e4ee3e4e · outbound

This paper cites MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.045603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.045603Z digest=sha256:6477a2ec992c175f76f102a04db97a155772dea2e9892f6744b2057f2d5b680b

Observation 49e0a2e3-dd50-461a-a6b3-fc6499b242ff · outbound

This paper cites A short-time objective intelligibility measure for time-frequency weighted noisy speech.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation A short-time objective intelligibility measure for time-frequency weighted noisy speech

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:14.812907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:00:14.064215Z digest=sha256:8235164a3394e42d89c2b3813225fc9db61b6329b7a6d6d5d4525d312b77f038

Observation 92059eaf-e1d3-4fd6-a32d-58e5de8d0d8c · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.073906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.073906Z digest=sha256:6944b1b00962063db948c93e5b92be4ff72fe856973eff96a99f863f28692d87

Observation cd815554-0fb7-4877-a0b9-f2406c475e7b · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.080956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.080956Z digest=sha256:0df1cf46e9129c144d9b97b5ad76c57996f6b8771d30f52caa1471b06deb6921

Observation 379d0d46-c6e6-48b0-b389-e657e0d22e09 · outbound

This paper cites FlowDec: A flow-based full-band general audio codec with high perceptual quality.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation FlowDec: A flow-based full-band general audio codec with high perceptual quality

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.089389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.089389Z digest=sha256:34c88e9e39a0275b5c5f084a041e168bc775f61c4ab6f95767833c4c0b4a8abf

Observation dd009f05-b7c0-49a8-bca3-cdaa2670d4a1 · outbound

This paper cites Codec-SUPERB: An In-Depth Analysis of Sound Codec Models.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Codec-SUPERB: An In-Depth Analysis of Sound Codec Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.103442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.103442Z digest=sha256:c5e32ea5aa283d60c9351dc2c064700c7c2705c58528b8111b831a03e3f89978

Observation ef9e0fda-ef88-4633-b2cc-8b8da6e8ec72 · outbound

This paper cites Laughter Synthesis using Pseudo Phonetic Tokens with a Large-scale In-the-wild Laughter Corpus.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Laughter Synthesis using Pseudo Phonetic Tokens with a Large-scale In-the-wild Laughter Corpus

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-05T15:00:14.381991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:00:14.116555Z digest=sha256:22cba1921ff6c0920bc9afa305873bacce167723500dd0700b1e7a4e306fd49d

Observation d06df0a6-9d79-4011-803e-22139be9f85e · outbound

This paper cites BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.133170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.133170Z digest=sha256:96a57c8bd0d3c2e9cd44effb439a028a4b234391fa131f65ca3649ae6af8ee88

Observation ccf870b0-a18d-4c93-abe1-94eff0dcee0c · outbound

This paper cites Qwen2.5 Technical Report.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Qwen2.5 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.143419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.143419Z digest=sha256:4acb6c09410732978638bdcb75bc11d0e3b937c7077fd4859a46ea25be9db95f

Observation 31ade6f1-8007-4a2f-84af-848fe68826cc · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.150027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.150027Z digest=sha256:491d894d1d1e2d872a8d747fd7f123952aa36aec62387fa766f185a097a4d038

Observation c53f719a-b5e5-4271-b8d6-e98932066d8e · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.159489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.159489Z digest=sha256:ff1155249f14522c2273c6beb0ad9cdb5d0dc65eb13311988bd06447d03ef15d

Observation 93203482-b7d1-49f0-8ed7-7fd2812cbe20 · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.164561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.164561Z digest=sha256:c01a374382d507777ad33345149e8e17af213b518ffcb3c7684bfcde73a3c37c

Observation 5e2992c9-eb93-4c69-8f52-ea94bc1765ae · outbound

This paper cites The clips within this dataset are manually selected from public field recordings compiled by the Freesound.org project.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation The clips within this dataset are manually selected from public field recordings compiled by the Freesound.org project

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:14.706512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:00:14.173358Z digest=sha256:3d8637e89d4eee9c082efad41eae4387c7466842d97c63bd9115e2ff9d9656f3

Observation 5445decb-3782-43f6-ab4d-934660afc49d · outbound

This paper cites an unresolved cited work.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Unresolved cited work

Reference 28

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T15:00:14.675893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:00:14.179533Z digest=sha256:552c0dfd102434317fd84424b20ad473d596a2f1e68262f7f75afaf0be860787

Observation 2fb3cf35-e48e-468f-83b3-8bc50d89a366 · outbound

This paper cites VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.058277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.058277Z digest=sha256:dbe4dd167c48fbac4dd90c2945d2136e17037ca0d425891928e46d3d51c5aa3e

Observation 71afe80c-116d-42cc-b965-6cffa81d88dc · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Audio set: An ontology and human-labeled dataset for audio events

Reference 2005

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:15.077517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:00:13.991843Z digest=sha256:18692361f0afed578ee3bf81bf2f0afc0f30406554eec06d8015b5578100573d

Observation e4178831-f504-4363-8fbf-12f1d8439c64 · outbound

This paper cites Visqol v3: An open source production ready objective speech and audio metric.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Visqol v3: An open source production ready objective speech and audio metric

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:15.100047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:00:13.973695Z digest=sha256:35ce814660b23f7fb1a9d398e8d148d1376806a881377b32317f669e5284fe3a

Observation a1fb8b1a-4fd4-46d0-a271-3b7e7aeec344 · outbound

This paper cites Scaling Transformers for Low-Bitrate High-Quality Speech Coding.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Scaling Transformers for Low-Bitrate High-Quality Speech Coding

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.039324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.039324Z digest=sha256:cd1f7241cf0ab49e9a6fe7bb68fffbe0b2629c0bb9a90e69a96beabaf59eec5c

Observation 3190eddb-5666-4dde-8754-2cc8bf56d42b · outbound

This paper cites 1632–1636,.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation 1632–1636,

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:14.741535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:00:14.069461Z digest=sha256:bb1c6306cbd849e8d15354ab2bf6ad96423be3d2bb253fdf0e08ba6d5d7f109f

Observation ef66f28f-a91b-48ba-9091-de9ec5c49eb6 · outbound

This paper cites V ocalsound: A dataset for improving human vocal sounds recognition.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation V ocalsound: A dataset for improving human vocal sounds recognition

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:15.059689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:00:13.999297Z digest=sha256:051989f671108906fb9e4926646d19afaccec143a09d6f7fd8aac05f8bd8edf1

Observation 40b1610e-2ed8-40ec-b53c-be989790466a · outbound

This paper cites Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.020393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.020393Z digest=sha256:ccc1ca0f03667634610ec88341166bfc33ab91e3cb71e371648e4fcd627ee7dc

Observation 9db2a1b1-70dc-40ce-a056-72781458b97b · outbound

This paper cites LibriMix: An Open-Source Dataset for Generalizable Speech Separation.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:13.980353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:13.980353Z digest=sha256:8263c68a841bda7b8af6f65f193a7ead086da2dfefbacc850fbc296b87f12296

Observation 69c6417a-e655-405f-ad0b-77ab70ec46da · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Robust Speech Recognition via Large-Scale Weak Supervision

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.052024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.052024Z digest=sha256:67b5c3fb81b6286bd619c551d0904ca1ea3277c77334c05491bc44aef2c1f50c

Observation 05f7c0ff-f8eb-47aa-8c09-462361b66ef4 · outbound

This paper cites Benchmarking representations for speech, music, and acoustic events.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Benchmarking representations for speech, music, and acoustic events

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:15.040602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:00:14.008369Z digest=sha256:7b9339e992f054dda94edc2c59cee0d3c36caf84cec432d37ea4fb9fb61dd9dc

Observation d8fd45dc-73ae-4d33-b7fc-36a93b950a4c · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Moshi: a speech-text foundation model for real-time dialogue

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:13.986387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:13.986387Z digest=sha256:33e24c9bc48b77a65d38f7322d9bd8d39d2341893275ef03be61d08406712034

Observation c2a2d476-99a5-47f2-b5ce-84675ebd4f29 · outbound

This paper cites Clotho- aqa: A crowdsourced dataset for audio question answering.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Clotho- aqa: A crowdsourced dataset for audio question answering

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:15.001969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T15:00:14.026878Z digest=sha256:267cf10118cc931264843576f1bb7fdb2b27d467d24508a8b539b192be51a58f

Pith citing papers

No inbound Pith citation observations are available.