Pith. sign in

Paper Citation Record · LEDGER

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech

As of 16 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2505.13805.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13805 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:13:46.876521Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:13:45.498553Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T20:13:47.342474Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact4
  • verified fuzzy11
  • unresolved24
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e225211b-9720-4903-aab3-b578015a7cea · outbound

This paper cites an unresolved cited work.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:13:48.512929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:13:45.442385Z digest=sha256:d820055d32a1ba309e2cbf282f64b462c685789547a731626686f77745a01504

Observation a5221266-f0da-4dd7-afef-5fb8a4d37a81 · outbound

This paper cites System Overview As illustrated in Fig.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech System Overview As illustrated in Fig

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:48.462992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:13:45.493913Z digest=sha256:a451da9749df8d74b9337e82552242f424b8660a812cc47ab98e62cfbf4b074c

Observation 9b5b0957-6312-44d7-835a-56e2f64bd8fc · outbound

This paper cites Experimental Setups 3.1.1.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Experimental Setups 3.1.1

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T20:13:48.345570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:13:45.503785Z digest=sha256:1f73010bc01422cf5d5fd767db77483fe24602972572fe2608b2fbc0f58828e2

Observation 36356220-d537-4e83-90ae-d44eec721d11 · outbound

This paper cites Specifically, the proposed ClapFM-EVC initially employs EVC-CLAP to extract and align emotional elements across audio-text modalities.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Specifically, the proposed ClapFM-EVC initially employs EVC-CLAP to extract and align emotional elements across audio-text modalities

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:48.210044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:13:45.586564Z digest=sha256:d40ac8745d8b52d59e826c16f58a98d4ab35afa7aaa09c9b9596059fa9ab4453

Observation dcc29f3a-ab75-4dee-9b30-7d87e0617701 · outbound

This paper cites MSAC: Multiple Speech Attribute Control Method for Reliable Speech Emotion Recognition.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech MSAC: Multiple Speech Attribute Control Method for Reliable Speech Emotion Recognition

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:13:47.287643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:13:45.799903Z digest=sha256:ece57a04a797401222a51bfa25fd2b94bdd2264f3d62cd7cb4d824690e9199e9

Observation 4f8ee533-e6dc-4821-9b95-67253c353882 · outbound

This paper cites Mixed emotion mod- elling for emotional voice conversion,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Mixed emotion mod- elling for emotional voice conversion,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:48.065329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:13:45.591330Z digest=sha256:a1ac8917753b3874a61874e06636fb5f8a56190052ec9cd4511124e5f9fe9a9e

Observation 7f36b60e-55e2-4bf2-99c5-ad443aaa2219 · outbound

This paper cites A pre-training based personalized dialogue generation model with persona- sparse data,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech A pre-training based personalized dialogue generation model with persona- sparse data,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:47.995333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:13:45.596315Z digest=sha256:cec482add8bad019f523283446de5fb5261c9bfb6aaf2a34b0ff41e97ca5f690

Observation 572cb274-12a8-4bd3-a3cf-a6a90c9fb3b1 · outbound

This paper cites Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:45.693164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:45.693164Z digest=sha256:ab0cf7dba8cbaa5b84867b5a8a8f2f501791edadc310cdfb800100c30fc7d213

Observation db7f7ba3-305b-40c8-9e01-dd72d1ae1960 · outbound

This paper cites From speaker to dubber: movie dubbing with prosody and duration consistency learning,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech From speaker to dubber: movie dubbing with prosody and duration consistency learning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:45.699418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:45.699418Z digest=sha256:f8bfda69666fb9a98d999e637a44e76470165d7172563f0db3a01d8629fc05ef

Observation d27ae634-a04c-4d92-9f2e-e31d9bb82bb5 · outbound

This paper cites Promptvc: Flexible stylistic voice conversion in latent space driven by natural language prompts,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Promptvc: Flexible stylistic voice conversion in latent space driven by natural language prompts,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.039171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.039171Z digest=sha256:a05a3d72128de2b816283b7e71bbc38dd8fbfc8f6e09b9806a993abd8b337c2b

Observation 343f80d5-8607-47fd-bba7-895e25f00d8b · outbound

This paper cites GMP-TL: Gender-augmented Multi-scale Pseudo-label Enhanced Transfer Learning for Speech Emotion Recognition.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech GMP-TL: Gender-augmented Multi-scale Pseudo-label Enhanced Transfer Learning for Speech Emotion Recognition

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:13:47.217133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:13:45.805882Z digest=sha256:924a0e7b82b387d92be983d03fc2ba59171ce0af3867d955af9c1f825abe7962

Observation 9cb69022-47e4-49ed-ad31-9af27db019a2 · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:45.889714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:45.889714Z digest=sha256:51ea2ef7ede0200c91e03515790f0301c16af69fcac1ed59dcd8473fcf85ea39

Observation 67282a1e-8bc9-4ff5-be50-1f12d0dd59d3 · outbound

This paper cites Squeezeformer: An effi- cient transformer for automatic speech recognition,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Squeezeformer: An effi- cient transformer for automatic speech recognition,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:47.969688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:13:45.895867Z digest=sha256:57aae3958497e5a1eab46a9d2e5f2a520b9c59679e713d7b5d330ccab659a936

Observation 77fcd420-0f8a-4bde-86a8-049a87d4ebe4 · outbound

This paper cites CAM++: A Fast and Efficient Network for Speaker Verification Using Context-Aware Masking.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech CAM++: A Fast and Efficient Network for Speaker Verification Using Context-Aware Masking

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:45.900121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:45.900121Z digest=sha256:b7e1b4a83902f2b56b31c47a284d0b311b0aa391c211883cc91d7bb8be0df3a2

Observation de513d13-5a8a-4c4e-9746-049ca636d280 · outbound

This paper cites One-shot emotional voice conversion based on feature separa- tion,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech One-shot emotional voice conversion based on feature separa- tion,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:47.702923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:13:46.160856Z digest=sha256:e23c0877f3b13a2c21573b566c98b0384d2f180ec56468576b4118f7b10a334c

Observation 71042821-bd59-4101-8f0e-e8a89c26795d · outbound

This paper cites Stargan for emo- tional speech conversion: Validated by data augmentation of end- to-end emotion recognition,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Stargan for emo- tional speech conversion: Validated by data augmentation of end- to-end emotion recognition,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:47.901699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:13:46.044411Z digest=sha256:3eeaab6ca77d91ca9a8d7b58ce75520bc5a909d5a952ea1aaf0aff0949a8cf2d

Observation a43d34f8-1843-4420-a877-dd568f13b7dc · outbound

This paper cites Transforming Spectrum and Prosody for Emotional Voice Conversion with Non-Parallel Training Data.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Transforming Spectrum and Prosody for Emotional Voice Conversion with Non-Parallel Training Data

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.049395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.049395Z digest=sha256:b50b11ad258241516622993e822411aa69f7bc8bd2f00c6467b04be708c9a1f8

Observation 7f5725d5-c7bf-41ee-bc3f-4253e987efec · outbound

This paper cites Non-parallel Emotion Conversion using a Deep-Generative Hybrid Network and an Adversarial Pair Discriminator.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Non-parallel Emotion Conversion using a Deep-Generative Hybrid Network and an Adversarial Pair Discriminator

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:13:47.135146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:13:46.107143Z digest=sha256:71fd4394efe6aeccf665685fca5eb400b83f16b208e2873b35de8894c7320211

Observation c66c6c6c-e88d-401a-be4f-c386f049d6e6 · outbound

This paper cites Emotional voice conversion using multitask learning with text-to-speech,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Emotional voice conversion using multitask learning with text-to-speech,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:47.886352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:13:46.156042Z digest=sha256:20b8fe4e3bca60abfcf2aa202ed77ed8d98d1538758a39fed1c12a67b0f63c9c

Observation 18ec0fb6-68be-4911-9794-c4920a1846e6 · outbound

This paper cites ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:13:47.399807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:13:45.498553Z digest=sha256:561dfa9083d126782c356a5bcfd113a585a8cf63ac016cc66a18f9e030d9e7a6

Observation ee0fd1e1-347e-4c0a-96a4-1581b67766f6 · outbound

This paper cites Attention-based Interactive Disentangling Network for Instance-level Emotional Voice Conversion.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Attention-based Interactive Disentangling Network for Instance-level Emotional Voice Conversion

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.218658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.218658Z digest=sha256:5b2af307d1ca6b9f63ce0209db02ccf74035ce46b7c347f2ffd8c94530b81ea2

Observation 9d556e4e-b33c-4914-a42e-bdf1427f02a9 · outbound

This paper cites Towards Realistic Emotional Voice Conversion using Controllable Emotional Intensity.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Towards Realistic Emotional Voice Conversion using Controllable Emotional Intensity

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.229017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.229017Z digest=sha256:e3b979bff01f5e2baae41b2958feec03524f08fb8fee3c4df332d1163c4656b3

Observation eda8fe2d-6d4e-47e0-a564-abf7164d6c1a · outbound

This paper cites Emotion intensity and its control for emotional voice conversion,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Emotion intensity and its control for emotional voice conversion,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.316657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.316657Z digest=sha256:ba343abe0c371dbdc4798ff970aed7c9bb5e5b4b589c105fae45d2704ca7309c

Observation de4f306b-fbef-4b9d-b260-a03ce487a8d3 · outbound

This paper cites To- ward any-to-any emotion voice conversion using disentangled dif- fusion framework,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech To- ward any-to-any emotion voice conversion using disentangled dif- fusion framework,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.337942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.337942Z digest=sha256:a2e1d9a7f542d2f26d412a335f9cdb82dd5b05e9f021c72467da345997099be2

Observation 87b908d3-bc05-4ccb-92a5-a7a781e46890 · outbound

This paper cites Hybridformer: Improving squeezeformer with hybrid attention and nsr mecha- nism,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Hybridformer: Improving squeezeformer with hybrid attention and nsr mecha- nism,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:47.672151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:13:46.342987Z digest=sha256:926941905cb615d62258b600084ea2ae5b42bae6d9c42201f9b55c758bc48b41

Observation 3f78e272-81a0-4e33-a8ec-0120324f6373 · outbound

This paper cites StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.346843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.346843Z digest=sha256:8d8033567bbb60669fcb0915aeb0a5bcd42bf8f478e82be361f0db3fc8425dae

Observation fcac195d-39af-48a9-8470-e423da9c24ec · outbound

This paper cites Zero-Shot Voice Conversion via Content-Aware Timbre Ensemble and Conditional Flow Matching.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Zero-Shot Voice Conversion via Content-Aware Timbre Ensemble and Conditional Flow Matching

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.457188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.457188Z digest=sha256:21bf0775499698650c51c2a54214a751ef7f1b4f44bcbfc45b38a750031040d7

Observation d8a77d93-ae6d-4436-84ee-fd87dfd2d419 · outbound

This paper cites BigVGAN: A Universal Neural Vocoder with Large-Scale Training.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.462263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.462263Z digest=sha256:a44a9e9a8a3871b878159d811d28b3096f41bc127186055fdfb2d42a51a67f50

Observation 60b06010-5f73-406b-9a45-8a5e76942f93 · outbound

This paper cites Clap learning audio concepts from natural language supervision,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Clap learning audio concepts from natural language supervision,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.466574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.466574Z digest=sha256:de8f2527fa1701703c238c157730bbcdb64e578abf56c4bba9ada3f7c3935a74

Observation f47d558e-0cda-4286-a880-dbe20b5016da · outbound

This paper cites Gemo-clap: Gender-attribute-enhanced contrastive language- audio pretraining for accurate speech emotion recognition,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Gemo-clap: Gender-attribute-enhanced contrastive language- audio pretraining for accurate speech emotion recognition,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.471290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.471290Z digest=sha256:719ee7e79f559bbcba8b6c116b7349023c92276ab564e010afaeacf1a0b5f975

Observation f8a13fed-b1f9-425f-958a-f2e2de5e6900 · outbound

This paper cites Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.569730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.569730Z digest=sha256:b9c269a5a42053cc5b2de1c874659cdfa19c24574b9d849a8bf586687621bcdc

Observation 25d12f18-ecc6-4787-bf02-9ed03b3b20d3 · outbound

This paper cites Unsupervised Cross-lingual Representation Learning at Scale.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Unsupervised Cross-lingual Representation Learning at Scale

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.608629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.608629Z digest=sha256:8ea6504364ff143530b9f2f08437f13d2e3e9b3afe9e086c1c1d690b3274a198

Observation f7039d9a-33ef-4008-b2bc-6cc3986e72e5 · outbound

This paper cites Meta-stylespeech: Multi-speaker adaptive text-to-speech generation,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Meta-stylespeech: Multi-speaker adaptive text-to-speech generation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:47.601158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:13:46.613085Z digest=sha256:6e3a1c24cdcaf55ff91b4910fbe613b2f0f9e445cf3721177ac718d4e2affc5d

Observation 86b80600-f4a1-4adf-9fba-e1b2b416a31e · outbound

This paper cites Deep residual learning for image recognition,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Deep residual learning for image recognition,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.618221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.618221Z digest=sha256:acf1babac778e6fc0a5add11f801d465712155bf8beedfd2de3ae00f39b46102

Observation 1a256fe6-cb64-4b85-96aa-17964e925c52 · outbound

This paper cites Attention is all you need,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Attention is all you need,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.750931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.750931Z digest=sha256:17c2a800c776b9a75d08019f4c9f277945ca3d6d15e5af4543b65f1b213b6e26

Observation bd721654-28f8-4062-9e97-d1991f8419f9 · outbound

This paper cites Film: Visual reasoning with a general conditioning layer,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Film: Visual reasoning with a general conditioning layer,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.815744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.815744Z digest=sha256:886acbe640a8c46359ce186882d91672f74b44194c145ab719c7c175451db9b5

Observation 895f996a-c5b1-4561-95c8-0133e93f0769 · outbound

This paper cites Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.821102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.821102Z digest=sha256:627ca4af3363b7eed2239ed558d8c2e40d477593185bc5f78b8d8ed18dea1931

Observation 4211e738-a583-401c-b5d9-771ac1f3df12 · outbound

This paper cites emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.850108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.850108Z digest=sha256:501b3d5bd832cbf3798e8b7a18f2f8aa4dbaa2a85b67aeb2928c9ec7b65132d4

Observation 24ecd4be-fbc1-45e2-8776-e81bf6a49384 · outbound

This paper cites Limited Data Emotional Voice Conversion Leveraging Text-to-Speech: Two-stage Sequence-to-Sequence Training.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Limited Data Emotional Voice Conversion Leveraging Text-to-Speech: Two-stage Sequence-to-Sequence Training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.872351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.872351Z digest=sha256:908faa9b7560936ce987ab65999ef301b8a6ca46e6b6c2cb1f0b73e6b4083d56

Observation ad3a691b-9d2c-41b3-9bd8-35d457b16ff0 · outbound

This paper cites Speech synthesis with mixed emotions,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Speech synthesis with mixed emotions,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:47.459852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:13:46.876521Z digest=sha256:4fec513313d82abd0cf932624b18002758aea6fcb9081d7a269fc9ccb5c3e5d3

Pith citing papers

Observation 18ec0fb6-68be-4911-9794-c4920a1846e6 · inbound

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech cites this paper.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:13:47.399807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:13:45.498553Z digest=sha256:561dfa9083d126782c356a5bcfd113a585a8cf63ac016cc66a18f9e030d9e7a6