Pith. sign in

Paper Citation Record · LEDGER

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding

As of 14 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 1 inbound Pith citation observation for arXiv:2506.22362.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22362 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:12:00.395703Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:11:56.263021Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T22:12:00.605940Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy28
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5b9800ea-3637-49ad-8f7e-0333b5d11c53 · outbound

This paper cites DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:12:00.657415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:56.263021Z digest=sha256:ee10de48e2086b322737d99c4b801c9d087fe7bfe81a2c0416568ef3d4eab046

Observation 8df8c223-4cfd-4c6b-82d9-88e276c36f6b · outbound

This paper cites an unresolved cited work.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:12:05.884534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:56.378331Z digest=sha256:877cc4e3ff0bfaf9bef1f7161d817dd1c0b2c1336cb0255819b3061c1513b346

Observation 61ed7bf7-03e8-4cb2-92e4-ad7026f60aaf · outbound

This paper cites SS-SC and SS-CL are optimized with 1e−4 learn- ing rate for 1 million steps.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding SS-SC and SS-CL are optimized with 1e−4 learn- ing rate for 1 million steps

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:05.545419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:56.883055Z digest=sha256:dbb983465e24b74953f28f8323e22e8b07d4669d32fa1a28fb4bda35d282100f

Observation 58d2b895-c15f-4d5c-979d-86e1565358a0 · outbound

This paper cites There are two major limitations: it only supports non- 2The 10 audio clips are sampled uniformly from LibriTTS test-clean omitting those less than 5s.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding There are two major limitations: it only supports non- 2The 10 audio clips are sampled uniformly from LibriTTS test-clean omitting those less than 5s

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:05.432611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:57.068976Z digest=sha256:8199a8d9d125bda55cbb450b58c609c09a5eebb632f31e441c21ab41310d8f54

Observation 697cd690-a421-48a1-b57a-f10af5163824 · outbound

This paper cites High Fidelity Neural Audio Compression.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding High Fidelity Neural Audio Compression

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:57.819506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:57.819506Z digest=sha256:ef79b7f7edf71eda91c67cdf63b1713444a2d36159ce0bfe56b9db9aad50c298

Observation 470f8e5c-fd4b-44f3-a4ed-73491a7869a3 · outbound

This paper cites an unresolved cited work.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:12:05.785845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:56.556586Z digest=sha256:d761c0a11ff9adf0288a36d5dcc3dee6f92932c6d186f01a292d3adc303556c9

Observation cb89ca61-b5f7-4f40-bbcf-3a46db86612f · outbound

This paper cites HuBERT: Self-Supervised Speech Rep- resentation Learning by Masked Prediction of Hidden Units,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding HuBERT: Self-Supervised Speech Rep- resentation Learning by Masked Prediction of Hidden Units,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:05.326298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:57.215975Z digest=sha256:05be476a2205b8d7506d5e97211b9a7e0d35d307d2fbc69625445776d7f2bb2c

Observation 7f71171f-daf1-4a7f-be44-40d65d974f96 · outbound

This paper cites wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Represen- tations,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Represen- tations,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:05.192372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:57.401222Z digest=sha256:e81f87592bb008153dc8da8b7fceec54bdadd38008b69668b137a8a684706d22

Observation d5a9493a-cb89-4734-9dbf-e5e0eeb6fddf · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Wavlm: Large-scale self-supervised pre-training for full stack speech processing,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:57.528724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:57.528724Z digest=sha256:353977a5c7c6deba241b728c7070ab07ee5faa6d4bd07be475a484aebc080a93

Observation fd77f7af-be89-4f03-bdd7-70e59f5c7aab · outbound

This paper cites w2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre- Training,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding w2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre- Training,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:05.084487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:57.697694Z digest=sha256:f5abb82611aeae5289892f3f7c91b70c4d57b039fd1c28822782b926fa964490

Observation 28cb9e28-d715-4793-8603-bff255369fec · outbound

This paper cites PolyV oice: Language Models for Speech to Speech Translation,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding PolyV oice: Language Models for Speech to Speech Translation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:04.407814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:58.332796Z digest=sha256:b8d88877226bfca6ddea1e2bb9e40499f4708ab5b04ed4efc830945425124677

Observation 2ae3cbb9-8dce-49cb-8963-43e3a5d44fe8 · outbound

This paper cites Soundstream: An end-to-end neural audio codec,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Soundstream: An end-to-end neural audio codec,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:04.950584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:57.887238Z digest=sha256:72ab2715600fa54e5cbf00c8f2a2aefeebfcbc4dbdc64c3cba3d9fc22d759e7c

Observation 5616e741-3b16-4a35-b52a-46ffd11ba287 · outbound

This paper cites AudioLM: A Language Modeling Approach to Audio Generation,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding AudioLM: A Language Modeling Approach to Audio Generation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:04.753500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:57.968067Z digest=sha256:d4c488a0b7ebb836abbc056d8171a9144140f1dabfaae41661857765e7a51685

Observation f371b50b-589e-4282-903e-299acaccadd5 · outbound

This paper cites Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:58.039438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:58.039438Z digest=sha256:52db9defd075e46051a9303a2cd02024c1fb0d3061a6f93a5733d1065e2ac5aa

Observation f6314508-c797-4b3c-b0c6-6201c39e2af5 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:58.112443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:58.112443Z digest=sha256:faa476d0461e4285348efa48020c610bba72dc30e7d2b1d33425df96ea68fdbf

Observation 53225d5e-261b-47b4-b570-6f689c6a2011 · outbound

This paper cites TokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript- Conditioned Speech Separation and Recognition,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding TokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript- Conditioned Speech Separation and Recognition,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:04.590748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:58.204048Z digest=sha256:97c1075414187c749eafdac642398c0c82c1e13bcf61200c311507a34d35dbe5

Observation 741d6f2e-ded0-471f-8981-7d2a8d44a1cd · outbound

This paper cites Denoising Diffusion Probabilistic Models,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Denoising Diffusion Probabilistic Models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:03.658852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:58.800630Z digest=sha256:0c5706d4f32055777437a87ddd2622cc850b7aa2904d1aa2943e234dfff15f2e

Observation 541cf954-a92b-44d0-ac2b-02138a4b836a · outbound

This paper cites High-Fidelity Simultaneous Speech-To-Speech Translation.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding High-Fidelity Simultaneous Speech-To-Speech Translation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:58.428052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:58.428052Z digest=sha256:a9208ded646f61c3dc9c7d86806d70f9b86f5f9a2c866f2c654cde1d0a8787c1

Observation 346f092b-0acc-4437-bcc2-9fd0777a651b · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Moshi: a speech-text foundation model for real-time dialogue

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:58.489098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:58.489098Z digest=sha256:ffe6af4e1c99a8fb4398ee78034d32e3d3586b37717ac90521cc6172bcfdbfaf

Observation ba766575-ff7d-4c67-8a2f-1b753275e311 · outbound

This paper cites Autoregressive Image Generation Using Residual Quantization,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Autoregressive Image Generation Using Residual Quantization,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:04.193477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:58.579941Z digest=sha256:e3e4d3b19278f48f36fd7031fe71020c66d5ed571166b62edb5b87358330c0eb

Observation 0503de73-1854-4765-b019-aa31e7982554 · outbound

This paper cites HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:04.004761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:58.652586Z digest=sha256:ee6a5d17c2d26c94995fd910eba9e2d118c4a4626a0d6948fc936b6957aaa530

Observation 2eee440f-7147-4fdf-b57b-56e8a377ca85 · outbound

This paper cites Mel- GAN: Generative Adversarial Networks for Conditional Wave- form Synthesis,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Mel- GAN: Generative Adversarial Networks for Conditional Wave- form Synthesis,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:03.825848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:58.731676Z digest=sha256:4f7ee2d4da773b7c34cdef001d884e34dade80c409f1ba43329c25c9f60d4300

Observation b518cc89-7d79-492f-b39e-9406b9a98ced · outbound

This paper cites Auto-Encoding Variational Bayes,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Auto-Encoding Variational Bayes,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:02.760597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:59.303877Z digest=sha256:6c475dc6100ccad3502c6eb3e752a1bf8d97ed3b0f4b3d1ab4733ca6f87e330a

Observation 29dbd204-1178-4dc2-aa33-9055e61bcc79 · outbound

This paper cites Deep Unsupervised Learning using Nonequilibrium Thermody- namics,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Deep Unsupervised Learning using Nonequilibrium Thermody- namics,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:03.492841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:58.879689Z digest=sha256:8b48e1cf9eacbb09b8ad68d095f58539954fd95303a278025fbc8418dadec439

Observation 415f794e-88e2-4cbb-9868-005b8a65a511 · outbound

This paper cites an unresolved cited work.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-06T22:12:05.654095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:56.737595Z digest=sha256:2753186f4d3f19e7fffce28732175a27186bfd62d87fde3cf6977e13ba0e5e1a

Observation 75e26e44-6ab8-478a-9520-2502c3b003aa · outbound

This paper cites Multi- step Distillation of Diffusion Models via Moment Matching,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Multi- step Distillation of Diffusion Models via Moment Matching,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:03.345517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:58.962209Z digest=sha256:c1e0d487a7f7b20adb9dee6750f16edef034e9c2679ccc8668cf16d2a35ba01c

Observation 09f42aa1-4ebc-4c2d-a76a-6ec7a3930a25 · outbound

This paper cites SpeechTok- enizer: Unified Speech Tokenizer for Speech Language Models,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding SpeechTok- enizer: Unified Speech Tokenizer for Speech Language Models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:03.154376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:59.044535Z digest=sha256:54a6e43b8cff5a28070f6ceefd3d2a90446b3d05a403500dc44c451542aa220e

Observation 1a30a0b0-92f1-4a23-8d32-43abdcd71bd9 · outbound

This paper cites Simple and Controllable Music Gen- eration,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Simple and Controllable Music Gen- eration,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:02.959308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:59.138687Z digest=sha256:6c71f11895b5732195866786c3f1016e5d432759e214c7cf4d47936f36f7dc59

Observation d6b0c266-afa8-44f9-8036-d6e1b059ab5c · outbound

This paper cites SoundStorm: Efficient Parallel Audio Generation.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding SoundStorm: Efficient Parallel Audio Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:59.228454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:59.228454Z digest=sha256:fd45a0fcf242ecae3eeb0451cd2a0b46c7b6d517db0dac6e42e4f1e06114ed35

Observation 57ce0ac9-6869-4c4a-ad8d-8a7f7314c980 · outbound

This paper cites FiLM: Visual Reasoning with a General Conditioning Layer,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding FiLM: Visual Reasoning with a General Conditioning Layer,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:02.583977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:59.401405Z digest=sha256:628913933beb3c837ccc38db0a6229bc2790d4ef2fc9ab9bf739ae1471d5d11e

Observation 5ff8a0bc-abcd-4e0b-bcd1-eb5534083d2f · outbound

This paper cites High-Resolution Image Synthesis With Latent Diffusion Mod- els,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding High-Resolution Image Synthesis With Latent Diffusion Mod- els,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:02.376140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:59.467113Z digest=sha256:8e5cfd95f166f44fbd6f967bc3c3b22bb57679f5decce9e1f4156f718f709d75

Observation 77eab515-24e4-4ea3-b1e9-cd5e767fa542 · outbound

This paper cites Improved Denoising Diffusion Probabilistic Models,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Improved Denoising Diffusion Probabilistic Models,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:02.220833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:59.562721Z digest=sha256:61befc54d4912dbde650121b340b05bdbaee0212c26766752996fc1fb9b08eed

Observation 67e37efe-fc40-4e9b-992e-6b321fc40388 · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Progressive Distillation for Fast Sampling of Diffusion Models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:02.030022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:59.627762Z digest=sha256:2c8bef07e12eb26d841bc744bea4e97a3ccb977e20512097435bfd9ce6487188

Observation 43f4d6bd-f5e9-4c83-a00c-da998cabf0cc · outbound

This paper cites Understanding Diffusion Objectives as the ELBO with Simple Data Augmentation,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Understanding Diffusion Objectives as the ELBO with Simple Data Augmentation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:01.835199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:59.723665Z digest=sha256:e06ff70da3661415c539063293690ef26b405b247fe36f101388842bee273b13

Observation 98e74125-b74a-4b75-8179-05ebc6103f4f · outbound

This paper cites WaveNet: A Generative Model for Raw Audio.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding WaveNet: A Generative Model for Raw Audio

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:59.809510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:59.809510Z digest=sha256:8158c887f64fe87bb56186359c2d7c487fa6df56c052f074cc4d629ee3c25db6

Observation ee08011d-db9e-4b47-af8a-a1ced36515b6 · outbound

This paper cites Improved Distribution Matching Distillation for Fast Image Synthesis,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Improved Distribution Matching Distillation for Fast Image Synthesis,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:01.628591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:59.912833Z digest=sha256:1eccce804dff3ecb3e808c37eb599143af47d1f7ad8318d2bf4ef94d3fa4d830

Observation 98d1f1e6-947e-45f9-8bd9-2e14b9a64f87 · outbound

This paper cites Adam: A Method for Stochastic Opti- mization,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Adam: A Method for Stochastic Opti- mization,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:01.443509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:59.998448Z digest=sha256:c9554327ce8e171d4833c047b5db2f93424db729bc76645677fb551c2b783ffe

Observation 660927e8-3baa-4905-bf20-2df453f72676 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:01.166687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:12:00.112002Z digest=sha256:fc1187a48047de5e84f9d42ad110e48519481bb65bb7fba5e846cd5e558180ad

Observation bc7abb8f-32a5-404e-a102-33693344d727 · outbound

This paper cites DNSMOS P.835: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding DNSMOS P.835: A Non-Intrusive Perceptual Objective Speech Quality Metric to Evaluate Noise Suppressors,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:01.013883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:12:00.221210Z digest=sha256:8c6349cbe70253063758c75dbeb876b85d38450d39436e1520c8d49060fc24a4

Observation 941ab650-d1d2-4991-8545-9bba2435104b · outbound

This paper cites Method for the subjective assessment of intermediate quality level of audio systems,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Method for the subjective assessment of intermediate quality level of audio systems,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:12:00.848285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:12:00.325633Z digest=sha256:08d7f42ccaf685bd05a0430c1c564a793efbdab9d2193e158279b8777449c3c4

Observation 5ef30e46-e58e-47f3-97e6-20c11c41a23b · outbound

This paper cites From Slow Bidirectional to Fast Autoregressive Video Diffusion Models,.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding From Slow Bidirectional to Fast Autoregressive Video Diffusion Models,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:12:00.395703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:12:00.395703Z digest=sha256:3d5dd88c2f31ccaed0a60c102be3b962699c42802636946914026a72cb274f08

Pith citing papers

Observation 5b9800ea-3637-49ad-8f7e-0333b5d11c53 · inbound

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding cites this paper.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:12:00.657415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T22:11:56.263021Z digest=sha256:ee10de48e2086b322737d99c4b801c9d087fe7bfe81a2c0416568ef3d4eab046