Pith. sign in

Paper Citation Record · LEDGER

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing

As of 15 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2608.06424.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06424 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:29:50.822607Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact4
  • verified fuzzy7
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 493e87f9-4f2a-400b-9ce8-e63e37ce9c30 · outbound

This paper cites Mapache: Masked parallel transformer for advanced speech editing and synthesis.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Mapache: Masked parallel transformer for advanced speech editing and synthesis

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.183649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T04:29:50.730382Z digest=sha256:db6a4270681a79b73e124187e5017eb457be9b25193fb831708d7dbabf558b65

Observation a92a2c7b-0e8b-4805-942e-2373fc3c6952 · outbound

This paper cites High Fidelity Neural Audio Compression.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing High Fidelity Neural Audio Compression

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.739671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.739671Z digest=sha256:fed12818da88f72f1058aed7235b7513910311371ec7d134545a1f880a0c034f

Observation e1e26ccf-f13c-4ce4-8a66-ff0e77653a3f · outbound

This paper cites For substitution and insertion operations, the system must determine the number of codec frames allocated to the edited span before generation.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing For substitution and insertion operations, the system must determine the number of codec frames allocated to the edited span before generation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.112349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T04:29:50.822607Z digest=sha256:239cbb7c4cbbe5b2c2b3c8a133ebaa453dba1c73e2aade4e84223db15010c3c2

Observation 870693b7-d270-4e07-be0a-43b7cf31c432 · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.753985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.753985Z digest=sha256:896b20e45d2d11777695c6db1203298a34c7d5c9d163ee93a9951cb83d952f4f

Observation 13167d88-46d8-4538-a913-7a85fd875129 · outbound

This paper cites Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.768197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.768197Z digest=sha256:ead2f59278e82eda70b79a35afa8826d83e3295424d621bc12ff41057906a702

Observation 306adf4f-176d-4628-89a5-b0a44f456ae9 · outbound

This paper cites Speech inpainting: Context-based speech synthesis guided by video.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Speech inpainting: Context-based speech synthesis guided by video

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-10T04:29:50.991444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T04:29:50.772653Z digest=sha256:7530dc6c6135ee235199fa71c2bc4bad1e59e6c90c82fa3b3763571ac47e66eb

Observation f294eefb-200d-46d8-8e4c-7adfd03ee704 · outbound

This paper cites Transient Noise Removal via Diffusion-based Speech Inpainting.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Transient Noise Removal via Diffusion-based Speech Inpainting

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-10T04:29:50.970902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T04:29:50.777180Z digest=sha256:fc0b3abaac545c6a3bd3e7c9ed202c43080ed8d7512b9f692f79b0ed44925387

Observation 68e1f938-7d18-4350-9453-50e95ab20872 · outbound

This paper cites Large Language Diffusion Models.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Large Language Diffusion Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.786020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.786020Z digest=sha256:ab34a5866edb976c935be17bf0cefe1123ad1f336b7f1617578f964cf3efd69e

Observation cff0339d-8e44-4e09-9fed-bc787adb64e1 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Robust Speech Recognition via Large-Scale Weak Supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.790349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.790349Z digest=sha256:1b18ea6a4a01d99838d265c361f8a2a8bd97b270f2fb6c5d0aa5fd04c5a2cbcc

Observation 4460e2a5-e242-44b2-ade8-5b385c264709 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Score-Based Generative Modeling through Stochastic Differential Equations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.799726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.799726Z digest=sha256:e7ae86b4e56cb4e95ac31ab8e12f295158de20df274057f47399c9ca31f93de8

Observation 695dd7ba-2dbb-40ff-bdaa-3e17869b4416 · outbound

This paper cites Score-based Continuous-time Discrete Diffusion Models.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Score-based Continuous-time Discrete Diffusion Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.804323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.804323Z digest=sha256:f4fbdd2ce6cfbe676c8720912668d2f3a881103cf19f34a1dad127ac1a2803bf

Observation ec8c193e-abee-4711-86df-568a36703bc2 · outbound

This paper cites usee: Unified speech enhancement and editing with conditional diffusion models.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing usee: Unified speech enhancement and editing with conditional diffusion models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.126709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T04:29:50.813283Z digest=sha256:9f62d58dcdc38f76068d02eab0032c153b88133487f5b21e365eaf5bcedddaab

Observation 496d843a-deeb-4b8e-bc93-2d289e9be7c4 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.817951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.817951Z digest=sha256:1eec182948d948d00429cec895c7587a5502a2c9fdc8c6b16b595b20df4a8782

Observation ae67ffad-d65e-4189-a488-2f004e558dde · outbound

This paper cites Discrete diffusion for generative modeling of text-aligned speech tokens.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Discrete diffusion for generative modeling of text-aligned speech tokens

Reference 1981

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.155056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T04:29:50.758646Z digest=sha256:eacf9f60a81899c995f670fc21ece64d415a7d18128125f294e5f5c7d4546e06

Observation aad7ddf9-75d8-4283-804b-d82883370553 · outbound

This paper cites Block diffusion: Interpolating between autoregressive and diffusion language models.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Block diffusion: Interpolating between autoregressive and diffusion language models

Reference 2011

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.197422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T04:29:50.720825Z digest=sha256:ab5532279fd6df4c121bd2f007190fc1d3e4132fa265dba2037286594fc616a5

Observation 8231d0da-6716-4abe-887c-ea922dd5d965 · outbound

This paper cites FluentEditor: Text-based Speech Editing by Considering Acoustic and Prosody Consistency.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing FluentEditor: Text-based Speech Editing by Considering Acoustic and Prosody Consistency

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.763054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.763054Z digest=sha256:87d9d5662f63c4581d107c2c1ceed9af8cb47d68f366c354362f54a25cc80e17

Observation 5f7e274b-4338-4875-be56-f1de170b042d · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.795291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.795291Z digest=sha256:cd702132b8e223e66d6b34a82f5ac9eec814187df0d50a4b4b4293f91a94aa97

Observation 377a4f7b-9223-4b5a-b392-376cf5a50bb7 · outbound

This paper cites Token-based audio inpainting via discrete diffusion.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Token-based audio inpainting via discrete diffusion

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.169124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T04:29:50.744507Z digest=sha256:c1f2c3c5c27852f0f6a9f4b369612a3b30d238e4362f28e0604eb26bf3caf79a

Observation 7b322f63-27e0-4658-895b-cc44b133efae · outbound

This paper cites SpeechPainter: Text-conditioned Speech Inpainting.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing SpeechPainter: Text-conditioned Speech Inpainting

Reference 2022

Resolution
verified exact
local_arxiv, observed 2026-08-10T04:29:51.098061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T04:29:50.725712Z digest=sha256:3cda88cf9bc92bca48c36a15d0f468c7a60149fc7636e376016e1cb1412b5c10

Observation 433897b6-e47c-4e2e-a46f-19001d27301c · outbound

This paper cites Classifier-Free Diffusion Guidance.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Classifier-Free Diffusion Guidance

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.749133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.749133Z digest=sha256:ffe2095c8abd20f574c19c60b6f513df92a935b4a82a7eceb462cee416c27cc6

Observation a755dd70-14c7-4da9-a1b6-0ed318841856 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.734798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.734798Z digest=sha256:5e61419458c092d2469b7428e3ec1967b2b036a8a3b0aeab1fefcdd51b837077

Observation e7f8c054-2f42-463c-98ce-d9cd246d7ac6 · outbound

This paper cites XPhoneBERT: A Pre-trained Multilingual Model for Phoneme Representations for Text-to-Speech.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing XPhoneBERT: A Pre-trained Multilingual Model for Phoneme Representations for Text-to-Speech

Reference 2025

Resolution
verified exact
local_arxiv, observed 2026-08-10T04:29:50.950424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T04:29:50.781601Z digest=sha256:26cc2e5baecdf7029308b34ed5840a026daab08b473eeb2d535827f315d68a8c

Observation 0d2d4c22-7347-4b15-8700-235ef0533c0e · outbound

This paper cites Ssr-speech: Towards stable, safe and robust zero-shot text-based speech editing and synthesis.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Ssr-speech: Towards stable, safe and robust zero-shot text-based speech editing and synthesis

Reference 2026

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.141461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T04:29:50.808904Z digest=sha256:829837ce18a6e3bc53af73a1fde9f196f92366e0549dfdf4938048696a2b3a99

Pith citing papers

No inbound Pith citation observations are available.