Pith. sign in

Paper Citation Record · LEDGER

Streaming Neural Speech Codecs through Time-Invariant Representations

As of 12 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2607.05250.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.05250 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-07T21:48:16.190698Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact4
  • verified fuzzy18
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6849589c-c30b-4e99-9e4b-53081c61f036 · outbound

This paper cites Sound- stream: An end-to-end neural audio codec,.

Streaming Neural Speech Codecs through Time-Invariant Representations Sound- stream: An end-to-end neural audio codec,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.074657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:18e6c7762930e2def5b0e8601c903d15c2e4ece7fb2bd04ce845afccbb7c83bf

Observation 78c86203-f501-4e74-a872-0f4a01d41b19 · outbound

This paper cites High Fidelity Neural Audio Compression.

Streaming Neural Speech Codecs through Time-Invariant Representations High Fidelity Neural Audio Compression

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-07T21:54:08.069398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:481273fc4c0710cd62c7ac4b296a5da6290b058c66cfdce32892b47a1b01cf92

Observation be4136e8-11d0-4dd0-8ccc-d962206e0e8b · outbound

This paper cites High-fidelity audio compression with improved rvqgan,.

Streaming Neural Speech Codecs through Time-Invariant Representations High-fidelity audio compression with improved rvqgan,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.132492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:6e6f9f2cda2f55f4e46162383d57ba20b0a82c24edda2400dabea714e1f10c33

Observation c59295a2-bd93-4612-a9fd-66c72afd31c0 · outbound

This paper cites Audiolm: a language modeling approach to audio generation,.

Streaming Neural Speech Codecs through Time-Invariant Representations Audiolm: a language modeling approach to audio generation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.150077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:e17b072532b523e53a9eae290a5cc737ad98cc8d18bff1ee55667ae24ee54f91

Observation fab15ee6-9b62-46db-8766-9b4052239170 · outbound

This paper cites Text-free prosody-aware generative spoken language modeling,.

Streaming Neural Speech Codecs through Time-Invariant Representations Text-free prosody-aware generative spoken language modeling,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.082110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:f60bdf06b707c5ada9fb5355330acf758f1631cb8390fa2bcf3eb1f99550678a

Observation dcc5ce31-16dc-4aff-a1ad-50e7ead15444 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Streaming Neural Speech Codecs through Time-Invariant Representations Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-07T21:54:08.050084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:0f967f4d57bab0484a1a9866e63e5deec933a5c0be3dd84a5bd7b4d0b2271175

Observation e1b89ce9-937c-44b8-86a6-07168a63a81a · outbound

This paper cites Voicebox: Text-guided multilingual universal speech generation at scale,.

Streaming Neural Speech Codecs through Time-Invariant Representations Voicebox: Text-guided multilingual universal speech generation at scale,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.088285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:d58abb1507a0ed23876f0893705d2032ec70ff923aa64a73702d98896b18d46f

Observation 0453ff9d-64a3-4eb4-aa09-49d67e6126ed · outbound

This paper cites Naturalspeech 3: zero-shot speech synthesis with factorized codec and diffusion models,.

Streaming Neural Speech Codecs through Time-Invariant Representations Naturalspeech 3: zero-shot speech synthesis with factorized codec and diffusion models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.153703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:ffae356fc84a936fd81f5c2317093162bba04f64c07a593ce2acc4140348fa4c

Observation e43c97b1-6d6a-4f52-9a81-99ab19ff2e45 · outbound

This paper cites Fewer-token neural speech codec with time-invariant codes,.

Streaming Neural Speech Codecs through Time-Invariant Representations Fewer-token neural speech codec with time-invariant codes,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.079235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:e3b1a066f1733f72b6ccd6d18c2fa168dfd562390b99a6ca9a82cd820339fe32

Observation 8982a116-943c-414b-a523-2911e3702c8d · outbound

This paper cites Voxceleb: Large-scale speaker verification in the wild,.

Streaming Neural Speech Codecs through Time-Invariant Representations Voxceleb: Large-scale speaker verification in the wild,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.156563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:e4e0454e38677074bf85949f4e25d6535d2b255fe5a5c1605137b2d621f465eb

Observation eb208972-c090-4373-9e54-af31c0cbd41d · outbound

This paper cites A multi-device dataset for urban acoustic scene classification.

Streaming Neural Speech Codecs through Time-Invariant Representations A multi-device dataset for urban acoustic scene classification

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-07T21:54:08.055552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:639a9392923142b3ff38069b5e01c4f81be6c4db1883b8f1545557cdc5764519

Observation 46242ae5-9ddb-498b-938a-fe153b2aab14 · outbound

This paper cites Meld: A multimodal multi-party dataset for emotion recognition in conversations,.

Streaming Neural Speech Codecs through Time-Invariant Representations Meld: A multimodal multi-party dataset for emotion recognition in conversations,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.091351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:7b3011e3f407ecf820d61b8cbcf1b78e485a000b2189f61474e7ad5236af356c

Observation 6657d6c3-88e3-4d5f-b942-4bed9a19ce91 · outbound

This paper cites Common voice: A massively-multilingual speech corpus.

Streaming Neural Speech Codecs through Time-Invariant Representations Common voice: A massively-multilingual speech corpus

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.121905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:1f3a7140354e23236be9c35a186282befa4fffc20ee8ff648768f8a91a5de112

Observation 8f66738f-fff3-4ab0-b563-2e7ebaaaff1f · outbound

This paper cites Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition.

Streaming Neural Speech Codecs through Time-Invariant Representations Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-07T21:54:08.064118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:49ab1f50eae8ab43fa27e5ab72daa996b1b5181b18d418cdee3678c71f37e369

Observation b68d61df-a0e8-4828-9410-17114b733321 · outbound

This paper cites Libritts: A corpus derived from librispeech for text-to-speech,.

Streaming Neural Speech Codecs through Time-Invariant Representations Libritts: A corpus derived from librispeech for text-to-speech,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.085034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:2de16569b8ecbc31af4babeed3fb34d3f46bbf54cbc81ba7d23b6d7404d534d5

Observation 92ff019a-cd92-4b03-be42-662a0dc68772 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

Streaming Neural Speech Codecs through Time-Invariant Representations Librispeech: an asr corpus based on public domain audio books,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.139053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:f0e7f6a599a61248611fe2406450979a8e7aa11fd0198d55e74cec63292cb95b

Observation ef1fb1e5-725b-4b93-8930-8f83393fb3d0 · outbound

This paper cites Cstr vctk corpus: English multi- speaker corpus for cstr voice cloning toolkit,.

Streaming Neural Speech Codecs through Time-Invariant Representations Cstr vctk corpus: English multi- speaker corpus for cstr voice cloning toolkit,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.095817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:7cea51a42138a9af303fc3dc16bac7ab466794dcf75d3ebf59b2d539688c9b5e

Observation aebc5aff-f4c4-4e6f-b2ef-6258ce01bd88 · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large- scale speech generation,.

Streaming Neural Speech Codecs through Time-Invariant Representations Emilia: An extensive, multilingual, and diverse speech dataset for large- scale speech generation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.146661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:53f167b54ac84440ff90aefd524bcb22a5dced19ba37121f0d796ec2b27b1d6e

Observation 52622d88-a962-44ca-9052-5ef07a58dc80 · outbound

This paper cites Visqol v3: An open source production ready objective speech and audio metric,.

Streaming Neural Speech Codecs through Time-Invariant Representations Visqol v3: An open source production ready objective speech and audio metric,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.142840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:d2e1ce423b0185c77f62da1386a7a879209a09e6b87f1361635f06cf843bd0b3

Observation 0b78137b-71b7-46ca-aa55-3d8743543fe5 · outbound

This paper cites ViSQOL v3: An Open Source Production Ready Objective Speech and Audio Metric.

Streaming Neural Speech Codecs through Time-Invariant Representations ViSQOL v3: An Open Source Production Ready Objective Speech and Audio Metric

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T21:54:08.059960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:bc5c806f180774738da00d7dff51f4c17e60c5bbcd18965fc2c39c58dac36d67

Observation e0e86f0f-3330-474b-b7dc-85c2868b54b7 · outbound

This paper cites Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,.

Streaming Neural Speech Codecs through Time-Invariant Representations Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.135969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:47dc2b15fa0ac44e050a4b0bcc6d36f92446a46058fa1e95d0e9c084b6787f67

Observation 95d1be5e-f152-4144-b8ea-d745241a06a3 · outbound

This paper cites An algorithm for intel- ligibility prediction of time–frequency weighted noisy speech,.

Streaming Neural Speech Codecs through Time-Invariant Representations An algorithm for intel- ligibility prediction of time–frequency weighted noisy speech,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.099575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:a7edae05eb33553a7a660012443e98c5ad09a402421497c3f4bb9a5a83e158a3

Observation 1c1e85c2-67a1-4052-9fae-e14661be0e2d · outbound

This paper cites Mel-cepstral distance measure for objective speech quality assess- ment,.

Streaming Neural Speech Codecs through Time-Invariant Representations Mel-cepstral distance measure for objective speech quality assess- ment,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T21:54:08.111036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-07T21:48:16.190698Z digest=sha256:41964f2d5934b42bcc5a0c5422d725de9016e62232dacb2badf271b19e3af6a7

Pith citing papers

No inbound Pith citation observations are available.