Pith. sign in

Paper Citation Record · LEDGER

Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2109.13821.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2109.13821 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:12:09.662888Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T12:26:37.498741Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c7ee5465-f4e3-43ff-a7f0-10f80a53ac2a · inbound

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models cites this paper.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.501347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:17ecb3635537d36fa87ebcfca7e32dc7792e0e5b4dbbbf2cc7fa1b4c4dcaeedc

Observation abe0a4ea-5cc0-4129-a4a5-d4c0a310b386 · inbound

Zero-shot Voice Conversion with Diffusion Transformers cites this paper.

Zero-shot Voice Conversion with Diffusion Transformers Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T20:12:09.662888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:12:09.662888Z digest=sha256:e9a5e9d2b41223bb017d7ca2e170a1a63c7fa4b0f4ccda359d3a2fc43c80593b

Observation dbdcfbb5-1d16-4517-8d61-f925e6c12d26 · inbound

EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion cites this paper.

EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T23:28:04.210795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:28:04.210795Z digest=sha256:ab75508db250c57057daca8e14ebc6357533950a83798e9c0c040b63ed43f977

Observation 57383cb9-df52-4d0a-abda-828db85abc72 · inbound

DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification cites this paper.

DiffAttack: Diffusion-based Timbre-reserved Adversarial Attack in Speaker Identification Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T21:21:12.112357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:21:12.112357Z digest=sha256:eed7b7c1adab62da1a1345a5addf8e016b88ad38d696854922a5ea9619bce604

Observation a7708550-21c0-46a7-999f-be4b9a4bb678 · inbound

Stepback: Enhanced Disentanglement for Voice Conversion via Multi-Task Learning cites this paper.

Stepback: Enhanced Disentanglement for Voice Conversion via Multi-Task Learning Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T14:09:29.209538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:09:29.209538Z digest=sha256:d36e9fb0b6fa449fea73e3dffe3c9779cfec7ed082a07e9599caeadfc9b2a150

Observation 47aa4c1a-aabe-44ef-9e16-7d104fc9573b · inbound

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion cites this paper.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:57.976936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:57.976936Z digest=sha256:8b7d0676b2c84f2e0b8867142173ff692b46851065de7ee049c0a37d5fe196fb

Observation 0e065f7e-7662-4794-9666-584dcf6a3aff · inbound

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching cites this paper.

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:56.791796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:56.791796Z digest=sha256:a34829ebf345c01a76bc4e9013b19240ba0d1a714ae9a8dc7540a65e085a46cd

Observation e638bdbd-8655-4859-a950-a8396c317d62 · inbound

ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization cites this paper.

ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:25.936184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:25.936184Z digest=sha256:dc34e3a5af03f6e3d7d7df6adba4746c4e976bb1c0576502099966fa213dfd0a

Observation bf6a37d3-8d93-4b24-8718-48335111518a · inbound

Approximate Borderline Sampling using Granular-Ball for Classification Tasks cites this paper.

Approximate Borderline Sampling using Granular-Ball for Classification Tasks Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:30:34.794851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:30:34.794851Z digest=sha256:34d5f2161dca3cb4690554114540780c4547e37bb305557551b1a4ed8754ca36

Observation b8bb26c0-d798-4ad7-830a-f0d1443dacdc · inbound

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion cites this paper.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:39.352136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:39.352136Z digest=sha256:d973cde9268d968ea0d36cd8593d7a424b1a96bbe503c878afcdaa1dc743d1c7

Observation 53fff485-e278-4622-b773-7695c7021ff4 · inbound

MoLEx: Mixture of LoRA Experts in Speech Self-Supervised Models for Audio Deepfake Detection cites this paper.

MoLEx: Mixture of LoRA Experts in Speech Self-Supervised Models for Audio Deepfake Detection Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:28.960970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:28.960970Z digest=sha256:e8de61f39d253ddf4d1fd753ea17de87588bb77c6449e2fb51f6e17f39edb170

Observation 34b47704-6210-4aa4-804a-daf81940ccf7 · inbound

ProSDD: Learning Prosodic Representations for Speech Deepfake Detection against Expressive and Emotional Attacks cites this paper.

ProSDD: Learning Prosodic Representations for Speech Deepfake Detection against Expressive and Emotional Attacks Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:35:26.586414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T13:31:12.802897Z digest=sha256:6a3a60a7fd83b331687e44e8a2eda85c3a09950790cb56f3677721fa0698ae0f

Observation 96c9e229-312f-48f2-9802-64bb7c550a23 · inbound

Remix the Timbre: Diffusion-Based Style Transfer Across Polyphonic Stems cites this paper.

Remix the Timbre: Diffusion-Based Style Transfer Across Polyphonic Stems Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:46:27.986754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T04:58:26.634355Z digest=sha256:65a9423399c172b2ce7e653a6675368c523dfd8c3b15ed7b34de823af3ea046d

Observation e7b1ade3-0c46-4c20-95d6-f9752b32c64f · inbound

What You Train Is What You Get: Gender Bias, Training Composition, and Post-Hoc Mitigation in Audio Deepfake Detection cites this paper.

What You Train Is What You Get: Gender Bias, Training Composition, and Post-Hoc Mitigation in Audio Deepfake Detection Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-14T14:47:58.592181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:47:58.592181Z digest=sha256:fb774f75ef9ccab4666d078905d9b17037c5f463e45480d53d81c362de725fb3

Observation 07650bfa-c611-42d6-88ef-957cae3bd642 · inbound

AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks cites this paper.

AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-08T11:50:21.166263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:50:21.166263Z digest=sha256:9dcfdc937621d354391140b26288ccbf0af965bab9f72ea8b705eb970d21b740