Pith. sign in

Paper Citation Record · LEDGER

Streaming Neural Speech Codecs through Time-Invariant Representations

As of 9 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2607.05250.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.05250 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-07T21:48:16.190698Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact4
  • verified fuzzy18
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6849589c-c30b-4e99-9e4b-53081c61f036 · outbound

This paper cites Sound- stream: An end-to-end neural audio codec,.

Streaming Neural Speech Codecs through Time-Invariant Representations Sound- stream: An end-to-end neural audio codec,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.074657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:8a48ed85be81f522e22632e3feb6db7ab522c179bce98dabd0a87646527bde8e

Observation 78c86203-f501-4e74-a872-0f4a01d41b19 · outbound

This paper cites High Fidelity Neural Audio Compression.

Streaming Neural Speech Codecs through Time-Invariant Representations High Fidelity Neural Audio Compression

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-07T21:54:08.069398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:9afcbf7cd2aded5cea6beb9df47f9d00b9b6da3a158f74ed2a5588a2897aaf87

Observation be4136e8-11d0-4dd0-8ccc-d962206e0e8b · outbound

This paper cites High-fidelity audio compression with improved rvqgan,.

Streaming Neural Speech Codecs through Time-Invariant Representations High-fidelity audio compression with improved rvqgan,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.132492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:c865f8eac6fbac7aeef329056d3f0670831835929764a31ed53321f4a19ab034

Observation c59295a2-bd93-4612-a9fd-66c72afd31c0 · outbound

This paper cites Audiolm: a language modeling approach to audio generation,.

Streaming Neural Speech Codecs through Time-Invariant Representations Audiolm: a language modeling approach to audio generation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.150077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:d826038405133cff15a287fea5061e1e8453ef26791afc486550f8918dc4ebed

Observation fab15ee6-9b62-46db-8766-9b4052239170 · outbound

This paper cites Text-free prosody-aware generative spoken language modeling,.

Streaming Neural Speech Codecs through Time-Invariant Representations Text-free prosody-aware generative spoken language modeling,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.082110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:f5565d7476d55df44c1da9c648ac3da8df28ccd7ac199d702444d98492d99e47

Observation dcc5ce31-16dc-4aff-a1ad-50e7ead15444 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Streaming Neural Speech Codecs through Time-Invariant Representations Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-07T21:54:08.050084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:476d45acf47332f68216793ee0b90c214f08088d08e72406b2bc9960d0900457

Observation e1b89ce9-937c-44b8-86a6-07168a63a81a · outbound

This paper cites Voicebox: Text-guided multilingual universal speech generation at scale,.

Streaming Neural Speech Codecs through Time-Invariant Representations Voicebox: Text-guided multilingual universal speech generation at scale,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.088285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:1a52c1b35ffd42f65bbed1302b6671d9014247b90cd7e298ac62f0dd954e7e75

Observation 0453ff9d-64a3-4eb4-aa09-49d67e6126ed · outbound

This paper cites Naturalspeech 3: zero-shot speech synthesis with factorized codec and diffusion models,.

Streaming Neural Speech Codecs through Time-Invariant Representations Naturalspeech 3: zero-shot speech synthesis with factorized codec and diffusion models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.153703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:d7d5a74e285f2105f4d30328ed6657cdbe4f1ae86899368aa36a3851581cb23c

Observation e43c97b1-6d6a-4f52-9a81-99ab19ff2e45 · outbound

This paper cites Fewer-token neural speech codec with time-invariant codes,.

Streaming Neural Speech Codecs through Time-Invariant Representations Fewer-token neural speech codec with time-invariant codes,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.079235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:f3859ccb4dffb3cf76106c326b824e1ff324a08d88e02a3756cd36795e1eb157

Observation 8982a116-943c-414b-a523-2911e3702c8d · outbound

This paper cites Voxceleb: Large-scale speaker verification in the wild,.

Streaming Neural Speech Codecs through Time-Invariant Representations Voxceleb: Large-scale speaker verification in the wild,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.156563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:3c4d4f53bfaf7b505e4f34311bf9bca19a5e47006731c951dbb5a3a8d97328d9

Observation eb208972-c090-4373-9e54-af31c0cbd41d · outbound

This paper cites A multi-device dataset for urban acoustic scene classification.

Streaming Neural Speech Codecs through Time-Invariant Representations A multi-device dataset for urban acoustic scene classification

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-07T21:54:08.055552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:bc6d2e8343f496cfaefc915d5b0aaa7d6b5ba7350c40249fc40cf376ef0bc0a1

Observation 46242ae5-9ddb-498b-938a-fe153b2aab14 · outbound

This paper cites Meld: A multimodal multi-party dataset for emotion recognition in conversations,.

Streaming Neural Speech Codecs through Time-Invariant Representations Meld: A multimodal multi-party dataset for emotion recognition in conversations,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.091351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:6426e2179b4b47694894ca67c8a61fa59683d9351ca8b4340f9775545acf3305

Observation 6657d6c3-88e3-4d5f-b942-4bed9a19ce91 · outbound

This paper cites Common voice: A massively-multilingual speech corpus.

Streaming Neural Speech Codecs through Time-Invariant Representations Common voice: A massively-multilingual speech corpus

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.121905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:c60fda51cf6175ab6a69d1f219d742a239ddd1e70ea70e41926ed98d0281903b

Observation 8f66738f-fff3-4ab0-b563-2e7ebaaaff1f · outbound

This paper cites Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition.

Streaming Neural Speech Codecs through Time-Invariant Representations Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-07T21:54:08.064118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:99abe858b2e7c3721a5a5a9fe0236ede17bb8b0cafd7a106f1b39e48daedf597

Observation b68d61df-a0e8-4828-9410-17114b733321 · outbound

This paper cites Libritts: A corpus derived from librispeech for text-to-speech,.

Streaming Neural Speech Codecs through Time-Invariant Representations Libritts: A corpus derived from librispeech for text-to-speech,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.085034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:939e943416c8f36145be33b3250f6560697aaf9c09813b12682903d529fa1db0

Observation 92ff019a-cd92-4b03-be42-662a0dc68772 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

Streaming Neural Speech Codecs through Time-Invariant Representations Librispeech: an asr corpus based on public domain audio books,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.139053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:f9aa565f7baecf08c2291cacee1965b984820739dfbf922a35d362d7074114f8

Observation ef1fb1e5-725b-4b93-8930-8f83393fb3d0 · outbound

This paper cites Cstr vctk corpus: English multi- speaker corpus for cstr voice cloning toolkit,.

Streaming Neural Speech Codecs through Time-Invariant Representations Cstr vctk corpus: English multi- speaker corpus for cstr voice cloning toolkit,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.095817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:4d40296b5aec00a2ed988e2300630b611ad479bfb1ca18cbeb4a7f4724fd8afd

Observation aebc5aff-f4c4-4e6f-b2ef-6258ce01bd88 · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large- scale speech generation,.

Streaming Neural Speech Codecs through Time-Invariant Representations Emilia: An extensive, multilingual, and diverse speech dataset for large- scale speech generation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.146661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:409ad46619355f3fef042e9ebfc874fbb345189f1efbf0d3179496211dce68f1

Observation 52622d88-a962-44ca-9052-5ef07a58dc80 · outbound

This paper cites Visqol v3: An open source production ready objective speech and audio metric,.

Streaming Neural Speech Codecs through Time-Invariant Representations Visqol v3: An open source production ready objective speech and audio metric,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.142840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:7ddc8e3df296cdb8cdbf44a442d781383873121b98d8d66b5ed886ef4088fc3b

Observation 0b78137b-71b7-46ca-aa55-3d8743543fe5 · outbound

This paper cites ViSQOL v3: An Open Source Production Ready Objective Speech and Audio Metric.

Streaming Neural Speech Codecs through Time-Invariant Representations ViSQOL v3: An Open Source Production Ready Objective Speech and Audio Metric

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T21:54:08.059960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:2b1a96e8fdfdf1278b0bd8a620e66af9dddd023c14e51350501c3e94e5b74870

Observation e0e86f0f-3330-474b-b7dc-85c2868b54b7 · outbound

This paper cites Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,.

Streaming Neural Speech Codecs through Time-Invariant Representations Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.135969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:93fe365ffa429d91c90dcca7fb167fc1c86d51808c23d8df0a904524335c934d

Observation 95d1be5e-f152-4144-b8ea-d745241a06a3 · outbound

This paper cites An algorithm for intel- ligibility prediction of time–frequency weighted noisy speech,.

Streaming Neural Speech Codecs through Time-Invariant Representations An algorithm for intel- ligibility prediction of time–frequency weighted noisy speech,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.099575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:c8a6f0b683f3e9f84ff5f6db59a7de1c64d4fdce671b408d585db382e7b8708f

Observation 1c1e85c2-67a1-4052-9fae-e14661be0e2d · outbound

This paper cites Mel-cepstral distance measure for objective speech quality assess- ment,.

Streaming Neural Speech Codecs through Time-Invariant Representations Mel-cepstral distance measure for objective speech quality assess- ment,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.111036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:23a7c15932a04684028b43616a106a6ca48f709ac53a50e7266c6ef138e70231

Pith citing papers

No inbound Pith citation observations are available.