Pith. sign in

Paper Citation Record · LEDGER

Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2109.13821.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2109.13821 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:12:09.662888Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T12:26:37.498741Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c7ee5465-f4e3-43ff-a7f0-10f80a53ac2a · inbound

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models cites this paper.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.501347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:36060869dbb532bfb181fe6354921a80364b09caa455b7377b93c608d22ff8b3

Observation abe0a4ea-5cc0-4129-a4a5-d4c0a310b386 · inbound

Zero-shot Voice Conversion with Diffusion Transformers cites this paper.

Zero-shot Voice Conversion with Diffusion Transformers Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T20:12:09.662888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:12:09.662888Z digest=sha256:e9a5e9d2b41223bb017d7ca2e170a1a63c7fa4b0f4ccda359d3a2fc43c80593b

Observation dbdcfbb5-1d16-4517-8d61-f925e6c12d26 · inbound

EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion cites this paper.

EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T23:28:04.210795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:28:04.210795Z digest=sha256:893dc07a803e7c5617ba541ce23e44bdcd7d549e615e9511acb36876f48868d1

Observation 57383cb9-df52-4d0a-abda-828db85abc72 · inbound

DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification cites this paper.

DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T21:21:12.112357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:21:12.112357Z digest=sha256:eed7b7c1adab62da1a1345a5addf8e016b88ad38d696854922a5ea9619bce604

Observation a7708550-21c0-46a7-999f-be4b9a4bb678 · inbound

Stepback: Enhanced Disentanglement for Voice Conversion via Multi-Task Learning cites this paper.

Stepback: Enhanced Disentanglement for Voice Conversion via Multi-Task Learning Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T14:09:29.209538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:09:29.209538Z digest=sha256:290d4c230ee297ed2c2b7ab0c584ca1e7f742ac260c71e242cbb4075c66b1d22

Observation 47aa4c1a-aabe-44ef-9e16-7d104fc9573b · inbound

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion cites this paper.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:57.976936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:57.976936Z digest=sha256:8b7d0676b2c84f2e0b8867142173ff692b46851065de7ee049c0a37d5fe196fb

Observation 0e065f7e-7662-4794-9666-584dcf6a3aff · inbound

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching cites this paper.

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:56.791796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:56.791796Z digest=sha256:c2693110673ca5bc0bd089a8148c9bae71bbdf13ea479341f1a7ad4931d7dc50

Observation e638bdbd-8655-4859-a950-a8396c317d62 · inbound

ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization cites this paper.

ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:25.936184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:25.936184Z digest=sha256:dc34e3a5af03f6e3d7d7df6adba4746c4e976bb1c0576502099966fa213dfd0a

Observation bf6a37d3-8d93-4b24-8718-48335111518a · inbound

Approximate Borderline Sampling using Granular-Ball for Classification Tasks cites this paper.

Approximate Borderline Sampling using Granular-Ball for Classification Tasks Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:30:34.794851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:30:34.794851Z digest=sha256:34d5f2161dca3cb4690554114540780c4547e37bb305557551b1a4ed8754ca36

Observation b8bb26c0-d798-4ad7-830a-f0d1443dacdc · inbound

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion cites this paper.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:39.352136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:39.352136Z digest=sha256:d973cde9268d968ea0d36cd8593d7a424b1a96bbe503c878afcdaa1dc743d1c7

Observation 53fff485-e278-4622-b773-7695c7021ff4 · inbound

MoLEx: Mixture of LoRA Experts in Speech Self-Supervised Models for Audio Deepfake Detection cites this paper.

MoLEx: Mixture of LoRA Experts in Speech Self-Supervised Models for Audio Deepfake Detection Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:28.960970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:28.960970Z digest=sha256:fcdfc1675e73a517c25488aee1e38f958ea4932b3daebdbe19a2c3c57905b616

Observation 34b47704-6210-4aa4-804a-daf81940ccf7 · inbound

ProSDD: Learning Prosodic Representations for Speech Deepfake Detection against Expressive and Emotional Attacks cites this paper.

ProSDD: Learning Prosodic Representations for Speech Deepfake Detection against Expressive and Emotional Attacks Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:35:26.586414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T13:31:12.802897Z digest=sha256:3940798dc9f4bb701a00bac7828c5952ae2c9b13af85644bf590bb49f172b4c1

Observation 96c9e229-312f-48f2-9802-64bb7c550a23 · inbound

Remix the Timbre: Diffusion-Based Style Transfer Across Polyphonic Stems cites this paper.

Remix the Timbre: Diffusion-Based Style Transfer Across Polyphonic Stems Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:46:27.986754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T04:58:26.634355Z digest=sha256:1d33ec9436d1b391984582fc8091d537082391d40cbd4c9e5d0acbb6e168ca35

Observation e7b1ade3-0c46-4c20-95d6-f9752b32c64f · inbound

What You Train Is What You Get: Gender Bias, Training Composition, and Post-Hoc Mitigation in Audio Deepfake Detection cites this paper.

What You Train Is What You Get: Gender Bias, Training Composition, and Post-Hoc Mitigation in Audio Deepfake Detection Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-14T14:47:58.592181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:47:58.592181Z digest=sha256:4d470dc662b7856b2e8e181dc5e7d4666cbbc465312bf7bc5b5dca7a014643a6

Observation 07650bfa-c611-42d6-88ef-957cae3bd642 · inbound

AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks cites this paper.

AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-08T11:50:21.166263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:50:21.166263Z digest=sha256:33cecc588b7fe0f8f941db0390442cb8f6e0b387b3cd7d54582930e8505c08b3