Pith. sign in

Paper Citation Record · LEDGER

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding

As of 12 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.22362.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22362 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:12:00.395703Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:11:56.263021Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T22:12:00.605940Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy28
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5b9800ea-3637-49ad-8f7e-0333b5d11c53 · outbound

This paper cites DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:12:00.657415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:56.263021Z digest=sha256:754cb75afdd44fdee9ad19e671fd5ec7518848a38439247449fa2e6bdbe9d715

Observation 8df8c223-4cfd-4c6b-82d9-88e276c36f6b · outbound

This paper cites an unresolved cited work.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:12:05.884534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:56.378331Z digest=sha256:b493f471e199e02b363355e2325fc6fd3beb20eb3709cf96edd24c03e3befa4e

Observation 61ed7bf7-03e8-4cb2-92e4-ad7026f60aaf · outbound

This paper cites SS-SC and SS-CL are optimized with 1e−4 learn- ing rate for 1 million steps.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding SS-SC and SS-CL are optimized with 1e−4 learn- ing rate for 1 million steps

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:05.545419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:56.883055Z digest=sha256:babb50843ede92bd69cdad52afde057419887c9988f20c6f53f18312157fdb8c

Observation 58d2b895-c15f-4d5c-979d-86e1565358a0 · outbound

This paper cites There are two major limitations: it only supports non- 2The 10 audio clips are sampled uniformly from LibriTTS test-clean omitting those less than 5s.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding There are two major limitations: it only supports non- 2The 10 audio clips are sampled uniformly from LibriTTS test-clean omitting those less than 5s

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:05.432611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:57.068976Z digest=sha256:f27c00d060ef344360474bd022a29ff5737fe8c11c0fa11e9948c9ce3fa9e0f7

Observation 697cd690-a421-48a1-b57a-f10af5163824 · outbound

This paper cites High Fidelity Neural Audio Compression.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding High Fidelity Neural Audio Compression

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:57.819506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:57.819506Z digest=sha256:ef79b7f7edf71eda91c67cdf63b1713444a2d36159ce0bfe56b9db9aad50c298

Observation 470f8e5c-fd4b-44f3-a4ed-73491a7869a3 · outbound

This paper cites an unresolved cited work.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:12:05.785845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:56.556586Z digest=sha256:48be8cad26ad2e7f0d5f35bc0681d51af1db9f406b6747f6ca997a5b31a312dc

Observation cb89ca61-b5f7-4f40-bbcf-3a46db86612f · outbound

This paper cites HuBERT: Self-Supervised Speech Rep- resentation Learning by Masked Prediction of Hidden Units,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding HuBERT: Self-Supervised Speech Rep- resentation Learning by Masked Prediction of Hidden Units,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:05.326298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:57.215975Z digest=sha256:84e62821596718004ad43e56746fca3e6a2b04b1e35a82784f5b12abfa725bfe

Observation 7f71171f-daf1-4a7f-be44-40d65d974f96 · outbound

This paper cites wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Represen- tations,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Represen- tations,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:05.192372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:57.401222Z digest=sha256:74fc8205aa915d2c9968bd230f8709341b91e73ece7c72401f9dfe90d9da6e66

Observation d5a9493a-cb89-4734-9dbf-e5e0eeb6fddf · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Wavlm: Large-scale self-supervised pre-training for full stack speech processing,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:57.528724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:57.528724Z digest=sha256:353977a5c7c6deba241b728c7070ab07ee5faa6d4bd07be475a484aebc080a93

Observation fd77f7af-be89-4f03-bdd7-70e59f5c7aab · outbound

This paper cites w2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre- Training,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding w2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre- Training,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:05.084487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:57.697694Z digest=sha256:96b46bcf90032d75bde3496de40a6ec71bf1fbf12a806af6691c53ea5bd73eeb

Observation 28cb9e28-d715-4793-8603-bff255369fec · outbound

This paper cites PolyV oice: Language Models for Speech to Speech Translation,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding PolyV oice: Language Models for Speech to Speech Translation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:04.407814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:58.332796Z digest=sha256:43586a15ba0f8dce4e627c6fb7c848c76a3c70f281f5b41c649a65cce52803b7

Observation 2ae3cbb9-8dce-49cb-8963-43e3a5d44fe8 · outbound

This paper cites Soundstream: An end-to-end neural audio codec,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Soundstream: An end-to-end neural audio codec,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:04.950584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:57.887238Z digest=sha256:43b5c23f7a6c6286d8f07cfd1cb12b14e33263f1882307d8f16a4de96044a770

Observation 5616e741-3b16-4a35-b52a-46ffd11ba287 · outbound

This paper cites AudioLM: A Language Modeling Approach to Audio Generation,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding AudioLM: A Language Modeling Approach to Audio Generation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:04.753500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:57.968067Z digest=sha256:b36d58eb7ea95866b69f77626e9db5e2bf3fe111d9c844ae24700ee1dfad50ca

Observation f371b50b-589e-4282-903e-299acaccadd5 · outbound

This paper cites Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:58.039438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:58.039438Z digest=sha256:1eae4285914415362eda2d95369065863853bc714ab71afd0952110bb4a384ef

Observation f6314508-c797-4b3c-b0c6-6201c39e2af5 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:58.112443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:58.112443Z digest=sha256:faa476d0461e4285348efa48020c610bba72dc30e7d2b1d33425df96ea68fdbf

Observation 53225d5e-261b-47b4-b570-6f689c6a2011 · outbound

This paper cites TokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript- Conditioned Speech Separation and Recognition,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding TokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript- Conditioned Speech Separation and Recognition,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:04.590748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:58.204048Z digest=sha256:531f4b9f280d42e64f9ceb4a15847b425fa9cb44dcf39671cb07d79583918173

Observation 741d6f2e-ded0-471f-8981-7d2a8d44a1cd · outbound

This paper cites Denoising Diffusion Probabilistic Models,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Denoising Diffusion Probabilistic Models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:03.658852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:58.800630Z digest=sha256:6ca7540a6c0ba56b8e6dfeff23ddc72b7df74be0dfabc860b365748a1d61af04

Observation 541cf954-a92b-44d0-ac2b-02138a4b836a · outbound

This paper cites High-Fidelity Simultaneous Speech-To-Speech Translation.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding High-Fidelity Simultaneous Speech-To-Speech Translation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:58.428052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:58.428052Z digest=sha256:476c6ae30401c94a58a5e0ee41b5fe35d5f84dd76615d2676d470a11093ba6fb

Observation 346f092b-0acc-4437-bcc2-9fd0777a651b · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Moshi: a speech-text foundation model for real-time dialogue

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:58.489098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:58.489098Z digest=sha256:ffe6af4e1c99a8fb4398ee78034d32e3d3586b37717ac90521cc6172bcfdbfaf

Observation ba766575-ff7d-4c67-8a2f-1b753275e311 · outbound

This paper cites Autoregressive Image Generation Using Residual Quantization,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Autoregressive Image Generation Using Residual Quantization,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:04.193477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:58.579941Z digest=sha256:9ce8313cebcf552e6bf31879b712aaa89637046051518d22652513a9d55cde8c

Observation 0503de73-1854-4765-b019-aa31e7982554 · outbound

This paper cites HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:04.004761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:58.652586Z digest=sha256:696f6179d4a3266999cc24a7bec91f1671c30dc5e79a8e6bf62568cc3d149052

Observation 2eee440f-7147-4fdf-b57b-56e8a377ca85 · outbound

This paper cites Mel- GAN: Generative Adversarial Networks for Conditional Wave- form Synthesis,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Mel- GAN: Generative Adversarial Networks for Conditional Wave- form Synthesis,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:03.825848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:58.731676Z digest=sha256:281ecaa00f72ee26547fd25ba3bf012c0dd5f21e48b42db5df2a0b89cfad85d3

Observation b518cc89-7d79-492f-b39e-9406b9a98ced · outbound

This paper cites Auto-Encoding Variational Bayes,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Auto-Encoding Variational Bayes,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:02.760597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:59.303877Z digest=sha256:12d3fa6a37ff1fbf85aa8bc2815d9773a02b5bad01f8701ee18a1cb8d7a757b7

Observation 29dbd204-1178-4dc2-aa33-9055e61bcc79 · outbound

This paper cites Deep Unsupervised Learning using Nonequilibrium Thermody- namics,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Deep Unsupervised Learning using Nonequilibrium Thermody- namics,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:03.492841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:58.879689Z digest=sha256:8b426cec1a330875c72b83713aac17b7e8ff8fc07313b0b3348a20496b9555fa

Observation 415f794e-88e2-4cbb-9868-005b8a65a511 · outbound

This paper cites an unresolved cited work.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:12:05.654095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:56.737595Z digest=sha256:335e78f43c2091491e9bdd0ff4bb288ada6864d1c7f349e9c66d187f7655706c

Observation 75e26e44-6ab8-478a-9520-2502c3b003aa · outbound

This paper cites Multi- step Distillation of Diffusion Models via Moment Matching,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Multi- step Distillation of Diffusion Models via Moment Matching,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:03.345517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:58.962209Z digest=sha256:d16b1bc7f6c389133f74fb84f6fced68e294ff86e0b95587e8166a5855779566

Observation 09f42aa1-4ebc-4c2d-a76a-6ec7a3930a25 · outbound

This paper cites SpeechTok- enizer: Unified Speech Tokenizer for Speech Language Models,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding SpeechTok- enizer: Unified Speech Tokenizer for Speech Language Models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:03.154376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:59.044535Z digest=sha256:36bb77dc9351fa8b42e6d520bbb205c2c31f5799fabd7be7d623a326fb7aaaf1

Observation 1a30a0b0-92f1-4a23-8d32-43abdcd71bd9 · outbound

This paper cites Simple and Controllable Music Gen- eration,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Simple and Controllable Music Gen- eration,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:02.959308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:59.138687Z digest=sha256:5ee0e66995bf0d82c415bab3b3588a899a46e834bcfbdfd1c514400d780d88be

Observation d6b0c266-afa8-44f9-8036-d6e1b059ab5c · outbound

This paper cites SoundStorm: Efficient Parallel Audio Generation.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding SoundStorm: Efficient Parallel Audio Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:59.228454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:59.228454Z digest=sha256:b88096ada8e8f989169213da6753db7bc8ae93602d620b0c637cfe161eee8c03

Observation 57ce0ac9-6869-4c4a-ad8d-8a7f7314c980 · outbound

This paper cites FiLM: Visual Reasoning with a General Conditioning Layer,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding FiLM: Visual Reasoning with a General Conditioning Layer,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:02.583977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:59.401405Z digest=sha256:d14391fed7c1205b9a21c1f7bb96da706917fed32735053002b6bb50007fa240

Observation 5ff8a0bc-abcd-4e0b-bcd1-eb5534083d2f · outbound

This paper cites High-Resolution Image Synthesis With Latent Diffusion Mod- els,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding High-Resolution Image Synthesis With Latent Diffusion Mod- els,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:02.376140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:59.467113Z digest=sha256:7c462c0142ea410261bacdde2940975dbc75f7853bbde739612891455ed07a84

Observation 77eab515-24e4-4ea3-b1e9-cd5e767fa542 · outbound

This paper cites Improved Denoising Diffusion Probabilistic Models,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Improved Denoising Diffusion Probabilistic Models,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:02.220833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:59.562721Z digest=sha256:7d1355d20207fafd8419418471252a0a40c2b50040847532c9f2440b02530348

Observation 67e37efe-fc40-4e9b-992e-6b321fc40388 · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Progressive Distillation for Fast Sampling of Diffusion Models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:02.030022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:59.627762Z digest=sha256:9e3c606a8f4d7f5b0cd7c85b62210d1ac57c7782e6b66ccdaebd0554b5209bae

Observation 43f4d6bd-f5e9-4c83-a00c-da998cabf0cc · outbound

This paper cites Understanding Diffusion Objectives as the ELBO with Simple Data Augmentation,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Understanding Diffusion Objectives as the ELBO with Simple Data Augmentation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:01.835199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:59.723665Z digest=sha256:55e6123bc3aeaf3f3941e7c35a45fb8c6a5934e5c1636cc4506e3635e26b0d8c

Observation 98e74125-b74a-4b75-8179-05ebc6103f4f · outbound

This paper cites WaveNet: A Generative Model for Raw Audio.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding WaveNet: A Generative Model for Raw Audio

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:59.809510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:59.809510Z digest=sha256:8c0a5f3bbe7e40832104c59f645a8b88a4c164827e74276e02e1de60553cdfe0

Observation ee08011d-db9e-4b47-af8a-a1ced36515b6 · outbound

This paper cites Improved Distribution Matching Distillation for Fast Image Synthesis,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Improved Distribution Matching Distillation for Fast Image Synthesis,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:01.628591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:59.912833Z digest=sha256:7084b81f5aa5766827c802b5f6c6a795e4dfa8b73a286910dbe68340adf66114

Observation 98d1f1e6-947e-45f9-8bd9-2e14b9a64f87 · outbound

This paper cites Adam: A Method for Stochastic Opti- mization,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Adam: A Method for Stochastic Opti- mization,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:01.443509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:59.998448Z digest=sha256:783865e263b8499276d9f1a9c3955f05a627c1077075efeae3501971c9cfaa39

Observation 660927e8-3baa-4905-bf20-2df453f72676 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:01.166687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:12:00.112002Z digest=sha256:6586008103b3b434ddd94cc1a42c9b5270d60bc312432433f96c1cd948c7698b

Observation bc7abb8f-32a5-404e-a102-33693344d727 · outbound

This paper cites DNSMOS P.835: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding DNSMOS P.835: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:01.013883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:12:00.221210Z digest=sha256:6713357364d9e0400ab5a2f41fb914fc6beaa94f4b2061ea67d6d4fc7f39f95f

Observation 941ab650-d1d2-4991-8545-9bba2435104b · outbound

This paper cites Method for the subjective assessment of intermediate quality level of audio systems,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Method for the subjective assessment of intermediate quality level of audio systems,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:00.848285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:12:00.325633Z digest=sha256:83b29035df18425dd9da5659ed5fa03e528b0da5433a6d7bcf05fe8ce29bec17

Observation 5ef30e46-e58e-47f3-97e6-20c11c41a23b · outbound

This paper cites From Slow Bidirectional to Fast Autoregressive Video Diffusion Models,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding From Slow Bidirectional to Fast Autoregressive Video Diffusion Models,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:12:00.395703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:12:00.395703Z digest=sha256:3d5dd88c2f31ccaed0a60c102be3b962699c42802636946914026a72cb274f08

Pith citing papers

Observation 5b9800ea-3637-49ad-8f7e-0333b5d11c53 · inbound

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding cites this paper.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:12:00.657415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T22:11:56.263021Z digest=sha256:754cb75afdd44fdee9ad19e671fd5ec7518848a38439247449fa2e6bdbe9d715