Pith. sign in

Paper Citation Record · LEDGER

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model

As of 14 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 3 inbound Pith citation observations for arXiv:2502.05471.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05471 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:15:52.190529Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:58:58.610208Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T21:00:08.120813Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact1
  • verified fuzzy15
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5d84f3fd-9b4c-4204-b572-9064d89a5838 · outbound

This paper cites Autovc: Zero-shot voice style transfer with only autoencoder loss,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Autovc: Zero-shot voice style transfer with only autoencoder loss,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.887514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:15:52.034234Z digest=sha256:c9eaf932c5c6770a392b0087c38e2b3b0677bfc3a9e70ffd3839a31b2f54af78

Observation 93a89f50-f21d-4408-b5ae-cf42c10a674c · outbound

This paper cites V oicemixer: Adver- sarial voice style mixup,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model V oicemixer: Adver- sarial voice style mixup,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.873972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:15:52.039255Z digest=sha256:c9c5ff82ffe6aa126a62481a37a834f8c229290d45fd9bfcbb9ed307d5afe2d7

Observation d8af58c6-fcea-4d97-8b41-58ab72fbedfd · outbound

This paper cites One-shot Voice Conversion by Separating Speaker and Content Representations with Instance Normalization.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model One-shot Voice Conversion by Separating Speaker and Content Representations with Instance Normalization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.044146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.044146Z digest=sha256:2cc9b7f1e17dcf82471c09313eb5e007a2a8284920f93c5f40dd7fffb3ea4dba

Observation 43547107-ed32-4a3a-b534-064397af18dc · outbound

This paper cites Again-vc: A one- shot voice conversion using activation guidance and adaptive instance normalization,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Again-vc: A one- shot voice conversion using activation guidance and adaptive instance normalization,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.048846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.048846Z digest=sha256:af423a3d07a0defcc27542d322fe5cc27afc03947dd1a5209e2e4f324dbd5bde

Observation 4cb787d0-464c-4fcd-8407-52687a4eabd1 · outbound

This paper cites Unsupervised speech decomposition via triple information bottleneck,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Unsupervised speech decomposition via triple information bottleneck,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.852583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:15:52.053187Z digest=sha256:04178330838e450ba24fd7854fb74061169f0ad1ed2eea83acaf506563238ed6

Observation b245516c-ce93-4249-bb02-3e955d060569 · outbound

This paper cites Neural analysis and synthesis: Reconstructing speech from self-supervised representa- tions,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Neural analysis and synthesis: Reconstructing speech from self-supervised representa- tions,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.839279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:15:52.057935Z digest=sha256:a4a56b40296452d8b7616d874dd71be151255210c67b6aaa851676957071c692

Observation a2649297-bf2f-4482-9e64-dad6d626e485 · outbound

This paper cites Expressive-vc: Highly expressive voice conversion with attention fusion of bottleneck and perturbation features,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Expressive-vc: Highly expressive voice conversion with attention fusion of bottleneck and perturbation features,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.062671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.062671Z digest=sha256:b828511d2773e69c190e1233faebf0b2ea5ff0d97111f752d1bb3abbc078167d

Observation 6441b8f8-fa2c-4e16-b2ed-17aafe308749 · outbound

This paper cites Contentvec: An improved self-supervised speech representation by disentangling speakers,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Contentvec: An improved self-supervised speech representation by disentangling speakers,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.817906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:15:52.066933Z digest=sha256:879094a57d8dbbba52916eee9e7dfc5ac25d53f8be7d79bb43605d59cf3b913d

Observation 3a85f7a5-9e41-456a-ba72-9a176f078019 · outbound

This paper cites UnitSpeech: Speaker-adaptive Speech Synthesis with Untranscribed Data.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model UnitSpeech: Speaker-adaptive Speech Synthesis with Untranscribed Data

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.070918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.070918Z digest=sha256:ef0c328945ec137625010c40a8add3e7cd30fd87887fec344dbae0e2d9a59b44

Observation 38b95a36-cb90-4933-9d2a-9d3ab2c38828 · outbound

This paper cites A comparison of discrete and soft speech units for improved voice conversion,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model A comparison of discrete and soft speech units for improved voice conversion,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.804918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:15:52.075200Z digest=sha256:0a611c2b7536bc5efade42957ef97e7e5cf0eb9f1d1ac75e1052b429515bd1b7

Observation 1523f89e-a3ba-433c-910f-ad5e671b769b · outbound

This paper cites OpenSR: Open-Modality Speech Recognition via Maintaining Multi-Modality Alignment.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model OpenSR: Open-Modality Speech Recognition via Maintaining Multi-Modality Alignment

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-08T19:15:52.599589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:15:52.079358Z digest=sha256:1fe487964be2f1d7816d4bceeae5ed545f761dc1ce15c1fa015eef39656c3e51

Observation 87f5e7dc-3889-406c-a99a-02e538b44f82 · outbound

This paper cites Diff-hiervc: Diffusion-based hier- archical voice conversion with robust pitch generation and masked prior for zero-shot speaker adaptation,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Diff-hiervc: Diffusion-based hier- archical voice conversion with robust pitch generation and masked prior for zero-shot speaker adaptation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.790811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:15:52.083533Z digest=sha256:b03a69d1389118b9e28085b588d5a0952f265cdab3ee962e8422b6f3acab386d

Observation 82359265-8bcb-441c-92b6-2f423f7bf549 · outbound

This paper cites TransFace: Unit-Based Audio-Visual Speech Synthesizer for Talking Head Translation.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model TransFace: Unit-Based Audio-Visual Speech Synthesizer for Talking Head Translation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.087678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.087678Z digest=sha256:ac75d0c0662ad4a8674a3f5e323cef9570b709a60a35d004f084601d8678f468

Observation 6aa15f87-6d3e-42c3-835e-378f81d778c7 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.091850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.091850Z digest=sha256:b0024ef73a64765df0e8e6d818428c53e7dca8b69fc8c02aa0c8efb2d81ae7e1

Observation 7aaf16de-cdf7-47dd-9d8b-0e43c321bd87 · outbound

This paper cites XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.095800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.095800Z digest=sha256:9f0218f14911a84a20d14d6bc4a3637cbed2cc42d25e615f7f435ed7bbff2a2d

Observation ad0c4ef2-79d1-43d7-8272-972afc4e003f · outbound

This paper cites Dgc-vector: A new speaker embedding for zero-shot voice conversion,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Dgc-vector: A new speaker embedding for zero-shot voice conversion,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.769014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:15:52.100086Z digest=sha256:2bab431cd9f41879244a31bece8e86202dc2038be99939b39174b0c8a3c44e50

Observation d57c9866-9809-4f2d-9c37-af4da096c778 · outbound

This paper cites Zero-shot multi-speaker text-to-speech with state-of-the- art neural speaker embeddings,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Zero-shot multi-speaker text-to-speech with state-of-the- art neural speaker embeddings,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.756026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:15:52.104052Z digest=sha256:aa72679e65876d74ad7a9c516e655efb66c52de7266d91e294b03d07cc0f0a6f

Observation 9f88f6c7-6892-4eeb-8dbd-d307c5e12eac · outbound

This paper cites Sef-vc: Speaker embedding free zero-shot voice conversion with cross attention,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Sef-vc: Speaker embedding free zero-shot voice conversion with cross attention,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.107919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.107919Z digest=sha256:dad0bd7f6081f12c01b7cac632ff87397dcc753a98f1d3e48e3f9d18268cd833

Observation 2fda4c95-d68c-4d21-82d7-5bafa7c93a60 · outbound

This paper cites Refxvc: Cross-lingual voice conversion with enhanced reference leveraging,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Refxvc: Cross-lingual voice conversion with enhanced reference leveraging,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.733770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:15:52.111782Z digest=sha256:91f28b091dd1a5b47fbb0fd77ae9a4100b5cbf4a8581ee82a34878bdba43bdb5

Observation d393707e-dfb3-4b75-a6b8-12d9dc40ff9a · outbound

This paper cites Zero-shot voice conversion via self-supervised prosody representation learning,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Zero-shot voice conversion via self-supervised prosody representation learning,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.720572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:15:52.116371Z digest=sha256:d72037fd7f0acb7c22bb05e0cbddd435a00f5d6c6acffcf79f901aec3892feb9

Observation e4cb2835-65ba-4e4a-839b-281ff92ae0f6 · outbound

This paper cites Promptvc: Flexible stylistic voice conversion in latent space driven by natural language prompts,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Promptvc: Flexible stylistic voice conversion in latent space driven by natural language prompts,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.707301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:15:52.120562Z digest=sha256:ac872877daf715b3a429db7f339cc700978e8c9a752987b3f7b8691ae2e3d9f0

Observation e194f579-f819-468f-bf1b-ed4792528fbc · outbound

This paper cites FluentSpeech: Stutter-Oriented Automatic Speech Editing with Context-Aware Diffusion Models.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model FluentSpeech: Stutter-Oriented Automatic Speech Editing with Context-Aware Diffusion Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.124370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.124370Z digest=sha256:9d9c021be4ee467f29dc1813985532b9aa998c4fe83bbf199c2183102554cd5c

Observation 93fd6717-fcae-4e11-8b67-9220276b82f5 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.128771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.128771Z digest=sha256:d2b07f92d57906cba46940c1c1ef1a0b61319947f72d0e83cba2131b022b5a52

Observation 6340e8b2-c522-4926-9e68-620a185f2abc · outbound

This paper cites Speech Resynthesis from Discrete Disentangled Self-Supervised Representations.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Speech Resynthesis from Discrete Disentangled Self-Supervised Representations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.133057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.133057Z digest=sha256:554a3fa177317a093a2ef281d3bcb564aa843d3537442d54d523f4f3d52a98fa

Observation 38995372-dcbe-4898-9df2-0961d064be26 · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.137166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.137166Z digest=sha256:2d40714b6d541d51984ab71751a8ca826eb18635a9c035b3a33976101741cc2c

Observation bc9bc02f-b76f-46c1-8358-ec84892ab68e · outbound

This paper cites Ace: A generative cross-modal re- trieval framework with coarse-to-fine semantic modeling,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Ace: A generative cross-modal re- trieval framework with coarse-to-fine semantic modeling,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.141263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.141263Z digest=sha256:e34706d51d923f20661394697630d656df739080271ea1a4cfd297047105b5a0

Observation eaa02832-2131-4664-832e-6c0440586611 · outbound

This paper cites Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.145073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.145073Z digest=sha256:d060f9a2ab672b65802817005fd4892defdfea3c0c26631defe3b6055df6d64e

Observation 287fa4db-112d-461a-bf0b-212ed2e208de · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.149153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.149153Z digest=sha256:a7f6045414b9c74551df48e2ad39f21b38268af429a1f3132f1b1fc66ffc65bb

Observation 75f19251-7957-4dfc-8a97-d9e9c0bf4ed3 · outbound

This paper cites Neural ordinary differential equations,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Neural ordinary differential equations,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.685090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:15:52.153410Z digest=sha256:edd600ae08ba2a8fb640b0fcb183af6aeac8aaf485f7c26f40df4435e83fbb75

Observation 3d6a598e-5eb3-4808-a4fa-9a2d80d7b681 · outbound

This paper cites Flow Matching for Generative Modeling.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Flow Matching for Generative Modeling

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.157201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.157201Z digest=sha256:238ca66a7bbe12286107047bf878ac2f019f3e1e16f74bbdd66cee9ded0110f8

Observation 39624ee8-c5e4-474c-bd57-5436f2d96644 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.161357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.161357Z digest=sha256:e11e2973664782afe599ead9f12e089f4ed69db74954926342f6876eba25aa45

Observation 5a9459da-28bb-4a7e-9ac9-3aa110f5e83f · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.165447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.165447Z digest=sha256:81b7ed09056e9e46f3968eca820c51941eae3a7b83291386f2efa6f2b12428a7

Observation d0efcc18-4f6e-4ecb-ae43-e1f71c48e62b · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Librispeech: an asr corpus based on public domain audio books,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.169472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.169472Z digest=sha256:3964a3b0968800a80573f5d1b49bde39df93c958b93f346114c1073b11416b5c

Observation ce3c3db8-3abc-4b9d-b97d-7b59327a41d8 · outbound

This paper cites Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.174166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.174166Z digest=sha256:62fdac2289be59b0548b2366c800b6ad367a22338c7c9ab2accd584b6e3a60b0

Observation 7a5f6dbd-87cb-464d-b939-b4990f0428ed · outbound

This paper cites Matcha-tts: A fast tts architecture with conditional flow matching,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Matcha-tts: A fast tts architecture with conditional flow matching,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.654476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:15:52.178133Z digest=sha256:3193a86eddf4d3bf45ab3c9ee788727cc5ee1f3a940dca43468f9390a910ff02

Observation 143e4c8e-f340-4307-b93c-822ef83fa9be · outbound

This paper cites emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.182473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.182473Z digest=sha256:6bc182f3b9ae281b22af5a1d0adfadc6a44763936e3f25b2b4e931caeb95b977

Observation 7c2ebe9d-6be7-41f6-a18d-7d3b966b2fb1 · outbound

This paper cites F0- consistent many-to-many non-parallel voice conversion via conditional autoencoder,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model F0- consistent many-to-many non-parallel voice conversion via conditional autoencoder,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.640937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T19:15:52.186723Z digest=sha256:13a7427b077f4709ecae243ba2ef51415972be81bf1b6303eebad53ddea848a7

Observation e7e8479f-ed8a-4ef3-ab44-15efc7b8ad70 · outbound

This paper cites Text-Free Prosody-Aware Generative Spoken Language Modeling.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Text-Free Prosody-Aware Generative Spoken Language Modeling

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.190529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.190529Z digest=sha256:aa3113cc1db36bfa74ce7cf7b596fd33aba3c9a3f516a3d812e2e7d14ff359e7

Pith citing papers

Observation 74f6896e-0381-4b69-8a07-941b1548c3bc · inbound

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion cites this paper.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:58.610208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:58.610208Z digest=sha256:81de5998fa979ca329ed82ed566d8e04017d43d75439a1ac19ce7fc179a56dc8

Observation 69fa6331-1fd0-4995-8696-417777240f33 · inbound

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching cites this paper.

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:58.568919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:58.568919Z digest=sha256:5e458e197c3aaaaf4f677c53db212e00b500452a90ea47c968b9ddb87a030508

Observation 8432b24e-8c02-4f75-883e-c008952b50d2 · inbound

SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech cites this paper.

SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:00:08.183010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T21:00:08.017244Z digest=sha256:82625b4c051bb63315ac4bff883c71b341d43bfe744175c46d7169a6a954c642