Pith. sign in

Paper Citation Record · LEDGER

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing

As of 10 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2608.06424.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06424 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:29:50.822607Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact4
  • verified fuzzy7
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 493e87f9-4f2a-400b-9ce8-e63e37ce9c30 · outbound

This paper cites Mapache: Masked parallel transformer for advanced speech editing and synthesis.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Mapache: Masked parallel transformer for advanced speech editing and synthesis

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.183649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:29:50.730382Z digest=sha256:b3ca1abbd092f0aec5ff52a05d5de17242f516f3f11b06b0153c9b5f1762b4fa

Observation a92a2c7b-0e8b-4805-942e-2373fc3c6952 · outbound

This paper cites High Fidelity Neural Audio Compression.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing High Fidelity Neural Audio Compression

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.739671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.739671Z digest=sha256:069e738e0c1d03a68a36a07c775b1db8e00ebecb8a14aafc935e140b62bcfe79

Observation e1e26ccf-f13c-4ce4-8a66-ff0e77653a3f · outbound

This paper cites For substitution and insertion operations, the system must determine the number of codec frames allocated to the edited span before generation.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing For substitution and insertion operations, the system must determine the number of codec frames allocated to the edited span before generation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.112349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:29:50.822607Z digest=sha256:dfae1be77fffb8736990187564eaf9292d859a4a2db7298435a69c494387acb0

Observation 870693b7-d270-4e07-be0a-43b7cf31c432 · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.753985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.753985Z digest=sha256:28e8940b50a58fe43d536221e4d6667c00e4f7919af76fb00c8dff1a5f0403a2

Observation 13167d88-46d8-4538-a913-7a85fd875129 · outbound

This paper cites Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.768197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.768197Z digest=sha256:42b088d3ee13fc64d2ca64b872437e5f16c9af2f2c55c7c340660d0e908c81e3

Observation 306adf4f-176d-4628-89a5-b0a44f456ae9 · outbound

This paper cites Speech inpainting: Context-based speech synthesis guided by video.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Speech inpainting: Context-based speech synthesis guided by video

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-10T04:29:50.991444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:29:50.772653Z digest=sha256:a87131a1ee6241dd10413e99a0c3ee70f44a59021258acc82788551cf35b4c10

Observation f294eefb-200d-46d8-8e4c-7adfd03ee704 · outbound

This paper cites Transient Noise Removal via Diffusion-based Speech Inpainting.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Transient Noise Removal via Diffusion-based Speech Inpainting

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-10T04:29:50.970902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:29:50.777180Z digest=sha256:50b234a38a42ce986831d18eca8c22b799cd6f0aa9479302d1956b683e2dae4c

Observation 68e1f938-7d18-4350-9453-50e95ab20872 · outbound

This paper cites Large Language Diffusion Models.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Large Language Diffusion Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.786020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.786020Z digest=sha256:eca351d41f2f1b69635632504f971615637c9679a28aa8dd00b9100667b48e33

Observation cff0339d-8e44-4e09-9fed-bc787adb64e1 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Robust Speech Recognition via Large-Scale Weak Supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.790349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.790349Z digest=sha256:89af19701c1c3e12bbeb509979a203d72e092dd53e6fea8cbf035724fbf5fc5a

Observation 4460e2a5-e242-44b2-ade8-5b385c264709 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Score-Based Generative Modeling through Stochastic Differential Equations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.799726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.799726Z digest=sha256:5e5b81ab01e4b428ab50e080efeef0fa1a0e609f0a983e86e9da4244151f79aa

Observation 695dd7ba-2dbb-40ff-bdaa-3e17869b4416 · outbound

This paper cites Score-based Continuous-time Discrete Diffusion Models.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Score-based Continuous-time Discrete Diffusion Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.804323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.804323Z digest=sha256:ffd2b265ea0a6f2410f81b0bde06ab5e31921553e46d8cc67312265ef6c9e2ca

Observation ec8c193e-abee-4711-86df-568a36703bc2 · outbound

This paper cites usee: Unified speech enhancement and editing with conditional diffusion models.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing usee: Unified speech enhancement and editing with conditional diffusion models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.126709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:29:50.813283Z digest=sha256:ef635b95bd62843a280e98486c4f5603d6a2d1107e326dd6aa65d3c7e282aaac

Observation 496d843a-deeb-4b8e-bc93-2d289e9be7c4 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.817951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.817951Z digest=sha256:d88c0af9443a969d479f02c51c64f10ee13d8aa29808dd32bbc638a1d021cf55

Observation ae67ffad-d65e-4189-a488-2f004e558dde · outbound

This paper cites Discrete diffusion for generative modeling of text-aligned speech tokens.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Discrete diffusion for generative modeling of text-aligned speech tokens

Reference 1981

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.155056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:29:50.758646Z digest=sha256:662b41efd5b921b4fd6fa8a5346153b4741af733c851d72af0a9de890b890c28

Observation aad7ddf9-75d8-4283-804b-d82883370553 · outbound

This paper cites Block diffusion: Interpolating between autoregressive and diffusion language models.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Block diffusion: Interpolating between autoregressive and diffusion language models

Reference 2011

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.197422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:29:50.720825Z digest=sha256:78671f720a4463e62b395c8e55d3240387552d700225af67044d1379f8a5148b

Observation 8231d0da-6716-4abe-887c-ea922dd5d965 · outbound

This paper cites FluentEditor: Text-based Speech Editing by Considering Acoustic and Prosody Consistency.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing FluentEditor: Text-based Speech Editing by Considering Acoustic and Prosody Consistency

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.763054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.763054Z digest=sha256:d30eb48139b1de1651e0b22a56940e128f18cec02585a17ffc197c11ac7d160c

Observation 5f7e274b-4338-4875-be56-f1de170b042d · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.795291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.795291Z digest=sha256:02f94f9eeb8a12d5fc78fffd0e38b3e8ccc29163814a90653d82d4a289199cc5

Observation 377a4f7b-9223-4b5a-b392-376cf5a50bb7 · outbound

This paper cites Token-based audio inpainting via discrete diffusion.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Token-based audio inpainting via discrete diffusion

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.169124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:29:50.744507Z digest=sha256:ed7e2e5fceeb314b1ae84344133e0da24390d5e3966c01f84dd6aefc46e92b2b

Observation 7b322f63-27e0-4658-895b-cc44b133efae · outbound

This paper cites SpeechPainter: Text-conditioned Speech Inpainting.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing SpeechPainter: Text-conditioned Speech Inpainting

Reference 2022

Resolution
verified exact
local_arxiv, observed 2026-08-10T04:29:51.098061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:29:50.725712Z digest=sha256:2bad576b0b6a1bf18e08e80ba25f14c5a915bde579bed074687573290518528a

Observation 433897b6-e47c-4e2e-a46f-19001d27301c · outbound

This paper cites Classifier-Free Diffusion Guidance.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Classifier-Free Diffusion Guidance

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.749133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.749133Z digest=sha256:a53f17366dadbb6117e902ee79463f37ba0bc234a12c8a38799dd20b26cb32eb

Observation a755dd70-14c7-4da9-a1b6-0ed318841856 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.734798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.734798Z digest=sha256:ee63d6bc7c77ba60efc420b79208c271502a0b9c2d31972b14344c436165284d

Observation e7f8c054-2f42-463c-98ce-d9cd246d7ac6 · outbound

This paper cites XPhoneBERT: A Pre-trained Multilingual Model for Phoneme Representations for Text-to-Speech.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing XPhoneBERT: A Pre-trained Multilingual Model for Phoneme Representations for Text-to-Speech

Reference 2025

Resolution
verified exact
local_arxiv, observed 2026-08-10T04:29:50.950424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:29:50.781601Z digest=sha256:4906c83d6abb48c21fb37dacc695a16c95c65f98cedb195e12fdd94f280a4d48

Observation 0d2d4c22-7347-4b15-8700-235ef0533c0e · outbound

This paper cites Ssr-speech: Towards stable, safe and robust zero-shot text-based speech editing and synthesis.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Ssr-speech: Towards stable, safe and robust zero-shot text-based speech editing and synthesis

Reference 2026

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.141461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:29:50.808904Z digest=sha256:e12456472c58e2183b82f160afb3afea268901ed98642178f22b9530cc59436a

Pith citing papers

No inbound Pith citation observations are available.