Pith. sign in

Paper Citation Record · LEDGER

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model

As of 9 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 3 inbound Pith citation observations for arXiv:2502.05471.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05471 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:15:52.190529Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:58:58.610208Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T21:00:08.120813Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact1
  • verified fuzzy15
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5d84f3fd-9b4c-4204-b572-9064d89a5838 · outbound

This paper cites Autovc: Zero-shot voice style transfer with only autoencoder loss,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Autovc: Zero-shot voice style transfer with only autoencoder loss,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.887514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:15:52.034234Z digest=sha256:1e7584e1ae3eb2971220d8900b594e11ce20159c4798c2d1070abce5d9bf0528

Observation 93a89f50-f21d-4408-b5ae-cf42c10a674c · outbound

This paper cites V oicemixer: Adver- sarial voice style mixup,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model V oicemixer: Adver- sarial voice style mixup,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.873972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:15:52.039255Z digest=sha256:67e12aa801670bb671f97a75011fce3adc8686447f92629e454381d676315ae6

Observation d8af58c6-fcea-4d97-8b41-58ab72fbedfd · outbound

This paper cites One-shot Voice Conversion by Separating Speaker and Content Representations with Instance Normalization.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model One-shot Voice Conversion by Separating Speaker and Content Representations with Instance Normalization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.044146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.044146Z digest=sha256:1f49265ccca4dc42d5102bc544ca759c5904980c279c67b70d54aa117bcb1cfc

Observation 43547107-ed32-4a3a-b534-064397af18dc · outbound

This paper cites Again-vc: A one- shot voice conversion using activation guidance and adaptive instance normalization,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Again-vc: A one- shot voice conversion using activation guidance and adaptive instance normalization,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.048846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.048846Z digest=sha256:c91cd802943732fc33773cff86d8963dbf7e48ca898a9f09c8aa617e34344058

Observation 4cb787d0-464c-4fcd-8407-52687a4eabd1 · outbound

This paper cites Unsupervised speech decomposition via triple information bottleneck,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Unsupervised speech decomposition via triple information bottleneck,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.852583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:15:52.053187Z digest=sha256:b156093390a1d150f6d08c50ad7bf32689a172f3ce53bfdc90aca9f6030454fc

Observation b245516c-ce93-4249-bb02-3e955d060569 · outbound

This paper cites Neural analysis and synthesis: Reconstructing speech from self-supervised representa- tions,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Neural analysis and synthesis: Reconstructing speech from self-supervised representa- tions,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.839279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:15:52.057935Z digest=sha256:7c7b4a6582fe91aaec7e57c9f6a8fe007f0063517d513a402bece6d65236051b

Observation a2649297-bf2f-4482-9e64-dad6d626e485 · outbound

This paper cites Expressive-vc: Highly expressive voice conversion with attention fusion of bottleneck and perturbation features,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Expressive-vc: Highly expressive voice conversion with attention fusion of bottleneck and perturbation features,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.062671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.062671Z digest=sha256:7bef424568557d04213d4d58b2979a65b2b0c3245713c78f3429470090145977

Observation 6441b8f8-fa2c-4e16-b2ed-17aafe308749 · outbound

This paper cites Contentvec: An improved self-supervised speech representation by disentangling speakers,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Contentvec: An improved self-supervised speech representation by disentangling speakers,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.817906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:15:52.066933Z digest=sha256:7bbe302bb5ea03b57be76b54b3bc95ac95ccaba31771144eb3d4403255df3fab

Observation 3a85f7a5-9e41-456a-ba72-9a176f078019 · outbound

This paper cites UnitSpeech: Speaker-adaptive Speech Synthesis with Untranscribed Data.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model UnitSpeech: Speaker-adaptive Speech Synthesis with Untranscribed Data

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.070918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.070918Z digest=sha256:e8c2e03f159fedd8c6762ff150d023eda07fc06314c03baafb626dff37ad6d88

Observation 38b95a36-cb90-4933-9d2a-9d3ab2c38828 · outbound

This paper cites A comparison of discrete and soft speech units for improved voice conversion,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model A comparison of discrete and soft speech units for improved voice conversion,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.804918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:15:52.075200Z digest=sha256:b11d0762ed2e5c737ece4c1ce66bd37aa71b9ceeab9bca2e9144d591b4579340

Observation 1523f89e-a3ba-433c-910f-ad5e671b769b · outbound

This paper cites OpenSR: Open-Modality Speech Recognition via Maintaining Multi-Modality Alignment.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model OpenSR: Open-Modality Speech Recognition via Maintaining Multi-Modality Alignment

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-08T19:15:52.599589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:15:52.079358Z digest=sha256:ed3ef8d56276dc5f77c812a5baad4164c7c95593b03beef16b98dda6fdb5cdd1

Observation 87f5e7dc-3889-406c-a99a-02e538b44f82 · outbound

This paper cites Diff-hiervc: Diffusion-based hier- archical voice conversion with robust pitch generation and masked prior for zero-shot speaker adaptation,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Diff-hiervc: Diffusion-based hier- archical voice conversion with robust pitch generation and masked prior for zero-shot speaker adaptation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.790811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:15:52.083533Z digest=sha256:4c3d6d649016ea9a34ebc43a729c9484bde1b4876d0c6715bb3b439a64b896fc

Observation 82359265-8bcb-441c-92b6-2f423f7bf549 · outbound

This paper cites TransFace: Unit-Based Audio-Visual Speech Synthesizer for Talking Head Translation.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model TransFace: Unit-Based Audio-Visual Speech Synthesizer for Talking Head Translation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.087678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.087678Z digest=sha256:8e1b26b85024e136a6b2cc9913ce5fafd4c7b9f0d9d191c930dec7f9d684b7a8

Observation 6aa15f87-6d3e-42c3-835e-378f81d778c7 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.091850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.091850Z digest=sha256:caf52d15c6d55e0e1b232a70b6ef26894a15fda19cdcb7839ed91c6d04743644

Observation 7aaf16de-cdf7-47dd-9d8b-0e43c321bd87 · outbound

This paper cites XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.095800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.095800Z digest=sha256:fe9e26088b61e9df82d29ffada4ca766f8ba68feae9050f81094f91fc7e97a0c

Observation ad0c4ef2-79d1-43d7-8272-972afc4e003f · outbound

This paper cites Dgc-vector: A new speaker embedding for zero-shot voice conversion,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Dgc-vector: A new speaker embedding for zero-shot voice conversion,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.769014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:15:52.100086Z digest=sha256:e48bf4f0b28202971a933a2b2fdd63ba2060c26256364e93feee94272a0d488c

Observation d57c9866-9809-4f2d-9c37-af4da096c778 · outbound

This paper cites Zero-shot multi-speaker text-to-speech with state-of-the- art neural speaker embeddings,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Zero-shot multi-speaker text-to-speech with state-of-the- art neural speaker embeddings,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.756026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:15:52.104052Z digest=sha256:b59d53e17c7e7f5f287e4b744229886ce4f1296fcd0753d986ef165aa6b34666

Observation 9f88f6c7-6892-4eeb-8dbd-d307c5e12eac · outbound

This paper cites Sef-vc: Speaker embedding free zero-shot voice conversion with cross attention,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Sef-vc: Speaker embedding free zero-shot voice conversion with cross attention,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.107919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.107919Z digest=sha256:82b0261069c8a4962f22faf5507e8885ce781958db5edb39c353fbb8a15d43e3

Observation 2fda4c95-d68c-4d21-82d7-5bafa7c93a60 · outbound

This paper cites Refxvc: Cross-lingual voice conversion with enhanced reference leveraging,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Refxvc: Cross-lingual voice conversion with enhanced reference leveraging,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.733770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:15:52.111782Z digest=sha256:7dd5bb0811091aa17e944319960a56c6b49797fd77e4aa5a3ba559be7b98e8e7

Observation d393707e-dfb3-4b75-a6b8-12d9dc40ff9a · outbound

This paper cites Zero-shot voice conversion via self-supervised prosody representation learning,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Zero-shot voice conversion via self-supervised prosody representation learning,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.720572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:15:52.116371Z digest=sha256:c82742b149a8741e7d1f4a3109db4a620835cdf9fffd0e2d71df3143eba4b91b

Observation e4cb2835-65ba-4e4a-839b-281ff92ae0f6 · outbound

This paper cites Promptvc: Flexible stylistic voice conversion in latent space driven by natural language prompts,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Promptvc: Flexible stylistic voice conversion in latent space driven by natural language prompts,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.707301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:15:52.120562Z digest=sha256:414cc7f189deadd98fbb53b5b05b6a03c8deaf636952a8c10db56b0c8400be43

Observation e194f579-f819-468f-bf1b-ed4792528fbc · outbound

This paper cites FluentSpeech: Stutter-Oriented Automatic Speech Editing with Context-Aware Diffusion Models.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model FluentSpeech: Stutter-Oriented Automatic Speech Editing with Context-Aware Diffusion Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.124370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.124370Z digest=sha256:db63b6426986b0a830832350ed0c09c99a79682da7e886d82df3697c09e475ad

Observation 93fd6717-fcae-4e11-8b67-9220276b82f5 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.128771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.128771Z digest=sha256:8805991db136b86a69bdefe51af5094134a6a451d75db0f8ccbe5fa3d4b94c61

Observation 6340e8b2-c522-4926-9e68-620a185f2abc · outbound

This paper cites Speech Resynthesis from Discrete Disentangled Self-Supervised Representations.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Speech Resynthesis from Discrete Disentangled Self-Supervised Representations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.133057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.133057Z digest=sha256:424c2e54671a9461b538d193bae08a1c35314bf523f209273d2e3f3bf1be8748

Observation 38995372-dcbe-4898-9df2-0961d064be26 · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.137166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.137166Z digest=sha256:1ad4507486f55713a0e0970ee0ab703c814601ac350e2d94336de996e6749fbf

Observation bc9bc02f-b76f-46c1-8358-ec84892ab68e · outbound

This paper cites Ace: A generative cross-modal re- trieval framework with coarse-to-fine semantic modeling,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Ace: A generative cross-modal re- trieval framework with coarse-to-fine semantic modeling,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.141263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.141263Z digest=sha256:75695cb76faae550f31e7e9e41f15a12db4a0f84ef9350e2bfbad0118871b326

Observation eaa02832-2131-4664-832e-6c0440586611 · outbound

This paper cites Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.145073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.145073Z digest=sha256:8146131a7c90180fcce546b727ed9e5a53f548e26fa220b30301f7e8c9247ad1

Observation 287fa4db-112d-461a-bf0b-212ed2e208de · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.149153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.149153Z digest=sha256:b93ab1a22e8a63b32460bcc6956cc501c4519305c10306aac7d06cfc1a6d9edb

Observation 75f19251-7957-4dfc-8a97-d9e9c0bf4ed3 · outbound

This paper cites Neural ordinary differential equations,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Neural ordinary differential equations,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.685090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:15:52.153410Z digest=sha256:9b0d01716d81563f2647ee8db440ae05294abdafd28f7c4429c0adc39e1416a8

Observation 3d6a598e-5eb3-4808-a4fa-9a2d80d7b681 · outbound

This paper cites Flow Matching for Generative Modeling.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Flow Matching for Generative Modeling

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.157201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.157201Z digest=sha256:79d9cdda9cd2e894dbeb4d000ea247c19f4fc12325addcbb6382235b4f6f1e1e

Observation 39624ee8-c5e4-474c-bd57-5436f2d96644 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.161357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.161357Z digest=sha256:1cfd105197086e7316992f837b06f8ea1f4cd2a3a0fee52388b6242589e765a8

Observation 5a9459da-28bb-4a7e-9ac9-3aa110f5e83f · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.165447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.165447Z digest=sha256:962ec20f2ed9672d5b09bfb41d69312d92da96364f718ce14c0a85d462ec0d6d

Observation d0efcc18-4f6e-4ecb-ae43-e1f71c48e62b · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Librispeech: an asr corpus based on public domain audio books,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.169472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.169472Z digest=sha256:f380cf5380dea917553b0606d32bc59f7fffdddb38b0c873ca37f5028560e170

Observation ce3c3db8-3abc-4b9d-b97d-7b59327a41d8 · outbound

This paper cites Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.174166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.174166Z digest=sha256:8288dc53b7d5c91ca5457862aeee1d06a02211dbd2740a47ce5a976bfad59f93

Observation 7a5f6dbd-87cb-464d-b939-b4990f0428ed · outbound

This paper cites Matcha-tts: A fast tts architecture with conditional flow matching,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Matcha-tts: A fast tts architecture with conditional flow matching,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.654476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:15:52.178133Z digest=sha256:ba43f5fa1c82c6012bfcc8fef31c254b1ef99cf172e3f4af265cd628b63ce91d

Observation 143e4c8e-f340-4307-b93c-822ef83fa9be · outbound

This paper cites emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.182473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.182473Z digest=sha256:f3f9159e60318de7d6ee4b25c480b87e7e82822a7b4ed3927c13732ebcc80133

Observation 7c2ebe9d-6be7-41f6-a18d-7d3b966b2fb1 · outbound

This paper cites F0- consistent many-to-many non-parallel voice conversion via conditional autoencoder,.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model F0- consistent many-to-many non-parallel voice conversion via conditional autoencoder,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:15:52.640937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T19:15:52.186723Z digest=sha256:632707744f09ce5595abd604d9273fdaf72e4c6e544f3c44aa5bf22910859664

Observation e7e8479f-ed8a-4ef3-ab44-15efc7b8ad70 · outbound

This paper cites Text-Free Prosody-Aware Generative Spoken Language Modeling.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model Text-Free Prosody-Aware Generative Spoken Language Modeling

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.190529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.190529Z digest=sha256:b1506055b0d31d7a7a45434889a0c6ba368bc42003b5cca842d9bf240edf6552

Pith citing papers

Observation 74f6896e-0381-4b69-8a07-941b1548c3bc · inbound

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion cites this paper.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:58.610208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:58.610208Z digest=sha256:8b00a00e157e73e93092133c54b4ca5d61478654ab6862a0b5f0cbb908cabf3b

Observation 69fa6331-1fd0-4995-8696-417777240f33 · inbound

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching cites this paper.

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:58.568919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:58.568919Z digest=sha256:37139d8c423cc69ec91c6fd17265c498a72352edc416d951c70b182e6eb91169

Observation 8432b24e-8c02-4f75-883e-c008952b50d2 · inbound

SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech cites this paper.

SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:00:08.183010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:00:08.017244Z digest=sha256:50d90ea9034d91d6cdf64279031fc62f39be75b4a0aec3a5866ce94307b27518