Pith. sign in

Paper Citation Record · LEDGER

ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 4 inbound Pith citation observations for arXiv:2409.15897.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.15897 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:57:08.449289Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T05:47:14.348285Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ed60375e-15fe-446a-be80-5f813c7d875b · inbound

TS3-Codec: Transformer-Based Simple Streaming Single Codec cites this paper.

TS3-Codec: Transformer-Based Simple Streaming Single Codec ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T10:57:08.449289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:57:08.449289Z digest=sha256:8f3f19b129fee3e40286e147b30b811de512bb495d74427983ad076639e9a508

Observation 5966c653-7880-4e04-9fb2-7cd00e391d2c · inbound

VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music cites this paper.

VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T05:20:45.095799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:20:45.095799Z digest=sha256:3bb1e451e790c9c43021d892fa79573aeecc4e42f45c61de272617137a7dfadc

Observation 8a5f30cd-9596-4953-bc6d-a8f17d175927 · inbound

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems cites this paper.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:33.706834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:04:33.706834Z digest=sha256:520ef73342a0ae7a578bf8c14fb206f3349131184e3b02d82ba1380e3e6baa51

Observation 6595c748-2f80-4c2f-b947-401edb609c36 · inbound

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training cites this paper.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:47:14.407787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T05:47:12.794875Z digest=sha256:8233d835667aa01013e27cc962d1a0f286e8f6b443954200a15dfa404d018955