Pith. sign in

Paper Citation Record · LEDGER

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech

As of 18 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2505.13805.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13805 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:13:46.876521Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:13:45.498553Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T20:13:47.342474Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact4
  • verified fuzzy11
  • unresolved24
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e225211b-9720-4903-aab3-b578015a7cea · outbound

This paper cites an unresolved cited work.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:13:48.512929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:13:45.442385Z digest=sha256:00091706a87ccfd37052c7a4d8fe232e0c6a3a58aee962b8c1dc8c6151b844bd

Observation a5221266-f0da-4dd7-afef-5fb8a4d37a81 · outbound

This paper cites System Overview As illustrated in Fig.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech System Overview As illustrated in Fig

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:48.462992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:13:45.493913Z digest=sha256:47df16dfc246f991e02f82547da2885e28f84a158f95ee780287a746c9a8979b

Observation 9b5b0957-6312-44d7-835a-56e2f64bd8fc · outbound

This paper cites Experimental Setups 3.1.1.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Experimental Setups 3.1.1

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T20:13:48.345570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:13:45.503785Z digest=sha256:9e11c2e95a6a2831f49714b3bc55103b504c580b4a141f9b40d8e28b8ae974db

Observation 36356220-d537-4e83-90ae-d44eec721d11 · outbound

This paper cites Specifically, the proposed ClapFM-EVC initially employs EVC-CLAP to extract and align emotional elements across audio-text modalities.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Specifically, the proposed ClapFM-EVC initially employs EVC-CLAP to extract and align emotional elements across audio-text modalities

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:48.210044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:13:45.586564Z digest=sha256:14cbc1ee675c3f367ed2d5c5f5c3a97a1c224fdc4ff955e5f31c865675c8776b

Observation dcc29f3a-ab75-4dee-9b30-7d87e0617701 · outbound

This paper cites MSAC: Multiple Speech Attribute Control Method for Reliable Speech Emotion Recognition.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech MSAC: Multiple Speech Attribute Control Method for Reliable Speech Emotion Recognition

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:13:47.287643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:13:45.799903Z digest=sha256:231431f7231bc9207fbaa07090f8531a287be13c534d282ceaee919949338746

Observation 4f8ee533-e6dc-4821-9b95-67253c353882 · outbound

This paper cites Mixed emotion mod- elling for emotional voice conversion,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Mixed emotion mod- elling for emotional voice conversion,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:48.065329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:13:45.591330Z digest=sha256:7720f33f20eb50803a724c918e819bed1bb06603e8995d9ba184be2ab38e6802

Observation 7f36b60e-55e2-4bf2-99c5-ad443aaa2219 · outbound

This paper cites A pre-training based personalized dialogue generation model with persona- sparse data,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech A pre-training based personalized dialogue generation model with persona- sparse data,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:47.995333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:13:45.596315Z digest=sha256:add9e4a2252a576c8df30a37d9dc47c678e0f125414ef17cd52c25fa1c012284

Observation 572cb274-12a8-4bd3-a3cf-a6a90c9fb3b1 · outbound

This paper cites Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:45.693164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:45.693164Z digest=sha256:f9fb2098fa6e6944d4466d15ab948e80fff4dec68dba643817bc4466ab9ee2af

Observation db7f7ba3-305b-40c8-9e01-dd72d1ae1960 · outbound

This paper cites From speaker to dubber: movie dubbing with prosody and duration consistency learning,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech From speaker to dubber: movie dubbing with prosody and duration consistency learning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:45.699418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:45.699418Z digest=sha256:4c0a73d3cb4cf1caea02881f9becc94376baaf48396abbe24ba5d307ec7881fb

Observation d27ae634-a04c-4d92-9f2e-e31d9bb82bb5 · outbound

This paper cites Promptvc: Flexible stylistic voice conversion in latent space driven by natural language prompts,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Promptvc: Flexible stylistic voice conversion in latent space driven by natural language prompts,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.039171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.039171Z digest=sha256:f19087349ebf7c6e08957d71fb83b52364bbc09998af40d7391ac5d8d0d9e536

Observation 343f80d5-8607-47fd-bba7-895e25f00d8b · outbound

This paper cites GMP-TL: Gender-augmented Multi-scale Pseudo-label Enhanced Transfer Learning for Speech Emotion Recognition.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech GMP-TL: Gender-augmented Multi-scale Pseudo-label Enhanced Transfer Learning for Speech Emotion Recognition

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:13:47.217133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:13:45.805882Z digest=sha256:3f739e8b76983afae1d4d5e9bf2ae800c09893e5ca055f886ab33ec05dbbff45

Observation 9cb69022-47e4-49ed-ad31-9af27db019a2 · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:45.889714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:45.889714Z digest=sha256:f96f42499009339a5991e8b1a33502d9e81667d2b513d9407a2e0708edc9533d

Observation 67282a1e-8bc9-4ff5-be50-1f12d0dd59d3 · outbound

This paper cites Squeezeformer: An effi- cient transformer for automatic speech recognition,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Squeezeformer: An effi- cient transformer for automatic speech recognition,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:47.969688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:13:45.895867Z digest=sha256:6b3d9caf676c7d5ed48b78a76b13cf1024d64e1c4de60ee5ff45769a03044f90

Observation 77fcd420-0f8a-4bde-86a8-049a87d4ebe4 · outbound

This paper cites CAM++: A Fast and Efficient Network for Speaker Verification Using Context-Aware Masking.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech CAM++: A Fast and Efficient Network for Speaker Verification Using Context-Aware Masking

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:45.900121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:45.900121Z digest=sha256:7016dc0778961a30863ad370046f55772b74c61fe51d6ebcdae61b9b3b5f912e

Observation de513d13-5a8a-4c4e-9746-049ca636d280 · outbound

This paper cites One-shot emotional voice conversion based on feature separa- tion,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech One-shot emotional voice conversion based on feature separa- tion,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:47.702923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:13:46.160856Z digest=sha256:6c40269e826f2025d5bdd5abfd5b3b146c2151519d8a4a894c056a5cf4bd5e21

Observation 71042821-bd59-4101-8f0e-e8a89c26795d · outbound

This paper cites Stargan for emo- tional speech conversion: Validated by data augmentation of end- to-end emotion recognition,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Stargan for emo- tional speech conversion: Validated by data augmentation of end- to-end emotion recognition,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:47.901699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:13:46.044411Z digest=sha256:40572f9b957339e218fe7915146e0bdfd494f1eb41769c57b1b10f3b0a63a253

Observation a43d34f8-1843-4420-a877-dd568f13b7dc · outbound

This paper cites Transforming Spectrum and Prosody for Emotional Voice Conversion with Non-Parallel Training Data.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Transforming Spectrum and Prosody for Emotional Voice Conversion with Non-Parallel Training Data

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.049395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.049395Z digest=sha256:56975aed3bce9b190dae54c7153dd9c1c6a507faecd7e47b51eaa177e2cbf80c

Observation 7f5725d5-c7bf-41ee-bc3f-4253e987efec · outbound

This paper cites Non-parallel Emotion Conversion using a Deep-Generative Hybrid Network and an Adversarial Pair Discriminator.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Non-parallel Emotion Conversion using a Deep-Generative Hybrid Network and an Adversarial Pair Discriminator

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:13:47.135146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:13:46.107143Z digest=sha256:1f8d7d57323a61a1a0d1656cf106bb3d558b64c7b9ba2ffd6cf8d904a63e0035

Observation c66c6c6c-e88d-401a-be4f-c386f049d6e6 · outbound

This paper cites Emotional voice conversion using multitask learning with text-to-speech,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Emotional voice conversion using multitask learning with text-to-speech,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:47.886352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:13:46.156042Z digest=sha256:dcdc4f0f64067e2b9d2591448d08c1d355148555e8edc5a9680a558fb93cb447

Observation 18ec0fb6-68be-4911-9794-c4920a1846e6 · outbound

This paper cites ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:13:47.399807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:13:45.498553Z digest=sha256:b6f8fadc1eff3924028abb7a763d1e4fe9dcf0a5a42129b12bd1bffa1d2053b2

Observation ee0fd1e1-347e-4c0a-96a4-1581b67766f6 · outbound

This paper cites Attention-based Interactive Disentangling Network for Instance-level Emotional Voice Conversion.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Attention-based Interactive Disentangling Network for Instance-level Emotional Voice Conversion

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.218658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.218658Z digest=sha256:74eacbefa82ef61a26967e22f99d2dcdae3fb60529c2654b2fcd73e573961ae2

Observation 9d556e4e-b33c-4914-a42e-bdf1427f02a9 · outbound

This paper cites Towards Realistic Emotional Voice Conversion using Controllable Emotional Intensity.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Towards Realistic Emotional Voice Conversion using Controllable Emotional Intensity

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.229017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.229017Z digest=sha256:dc2e2012e68ac2398e1c8389d9d77d631a15421aa32770b5c6b948d63ec97b9b

Observation eda8fe2d-6d4e-47e0-a564-abf7164d6c1a · outbound

This paper cites Emotion intensity and its control for emotional voice conversion,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Emotion intensity and its control for emotional voice conversion,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.316657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.316657Z digest=sha256:244e49afb01a13559ee870f0c5da1d3fa8d7e5251c1ed2c20b72c9de20f643e8

Observation de4f306b-fbef-4b9d-b260-a03ce487a8d3 · outbound

This paper cites To- ward any-to-any emotion voice conversion using disentangled dif- fusion framework,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech To- ward any-to-any emotion voice conversion using disentangled dif- fusion framework,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.337942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.337942Z digest=sha256:87a2db4cdeaa3d44f7abae36ede0a9154e4df26025dcb979f35f73a68b41d5e4

Observation 87b908d3-bc05-4ccb-92a5-a7a781e46890 · outbound

This paper cites Hybridformer: Improving squeezeformer with hybrid attention and nsr mecha- nism,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Hybridformer: Improving squeezeformer with hybrid attention and nsr mecha- nism,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:47.672151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:13:46.342987Z digest=sha256:3323935d8acfc2d1c5d472c6f8eeb503b1a5d2053abb08595fbcf01dbdf572ef

Observation 3f78e272-81a0-4e33-a8ec-0120324f6373 · outbound

This paper cites StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.346843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.346843Z digest=sha256:99b0bc26daf0888354438cabfd8aaaf57fbf86e7ad3ee744a66669a9ac85dd75

Observation fcac195d-39af-48a9-8470-e423da9c24ec · outbound

This paper cites Zero-Shot Voice Conversion via Content-Aware Timbre Ensemble and Conditional Flow Matching.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Zero-Shot Voice Conversion via Content-Aware Timbre Ensemble and Conditional Flow Matching

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.457188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.457188Z digest=sha256:f21a9ec7f0d781fa3a0b4c2a6578dfaa33278fed08303008dcfea88384946934

Observation d8a77d93-ae6d-4436-84ee-fd87dfd2d419 · outbound

This paper cites BigVGAN: A Universal Neural Vocoder with Large-Scale Training.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.462263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.462263Z digest=sha256:bae05c5869d465cd72cd17e25011ba11e2118424e0084db500e99d19a33aa504

Observation 60b06010-5f73-406b-9a45-8a5e76942f93 · outbound

This paper cites Clap learning audio concepts from natural language supervision,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Clap learning audio concepts from natural language supervision,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.466574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.466574Z digest=sha256:323a371d3d089004c3a8566fd731ca2f11411e0cefe4824c4bbfd0845155d1fc

Observation f47d558e-0cda-4286-a880-dbe20b5016da · outbound

This paper cites Gemo-clap: Gender-attribute-enhanced contrastive language- audio pretraining for accurate speech emotion recognition,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Gemo-clap: Gender-attribute-enhanced contrastive language- audio pretraining for accurate speech emotion recognition,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.471290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.471290Z digest=sha256:52dc555704a6d10c6341b039300fb66cf03935d0bb19d66af5dd3f8de81bbb6c

Observation f8a13fed-b1f9-425f-958a-f2e2de5e6900 · outbound

This paper cites Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.569730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.569730Z digest=sha256:849236c2ec6ee0301a8f0c537821c1be8f12a3106b8ecef04b8c99f1e7180cd0

Observation 25d12f18-ecc6-4787-bf02-9ed03b3b20d3 · outbound

This paper cites Unsupervised Cross-lingual Representation Learning at Scale.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Unsupervised Cross-lingual Representation Learning at Scale

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.608629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.608629Z digest=sha256:fa7cbe5b44ca10caa853e999e13612189fdfbf092904eabfbad40c3e7bfc3eda

Observation f7039d9a-33ef-4008-b2bc-6cc3986e72e5 · outbound

This paper cites Meta-stylespeech: Multi-speaker adaptive text-to-speech generation,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Meta-stylespeech: Multi-speaker adaptive text-to-speech generation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:47.601158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:13:46.613085Z digest=sha256:b2f0a9b66eadeef20fca620ed613aac5e034064c27d9561aecabdbbaa3b7b0ad

Observation 86b80600-f4a1-4adf-9fba-e1b2b416a31e · outbound

This paper cites Deep residual learning for image recognition,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Deep residual learning for image recognition,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.618221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.618221Z digest=sha256:6608ab580005233396b9d3f5cf124a7d6a1eec4a028286dc7240c1d0881ba5a3

Observation 1a256fe6-cb64-4b85-96aa-17964e925c52 · outbound

This paper cites Attention is all you need,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Attention is all you need,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.750931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.750931Z digest=sha256:d41061c4c4106abcd1e6f4b4865ef19aee9a6d2ebddd54d1a019fa3995933331

Observation bd721654-28f8-4062-9e97-d1991f8419f9 · outbound

This paper cites Film: Visual reasoning with a general conditioning layer,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Film: Visual reasoning with a general conditioning layer,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.815744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.815744Z digest=sha256:76ca3c74cc9bdb819aa78aa52da2670a7f595923b011bd1216df1d8ad45d7ace

Observation 895f996a-c5b1-4561-95c8-0133e93f0769 · outbound

This paper cites Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.821102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.821102Z digest=sha256:9a175b2d175ca8467b077966d293dfd7b4833551f78b6fb0cf32b3eb55283fdd

Observation 4211e738-a583-401c-b5d9-771ac1f3df12 · outbound

This paper cites emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.850108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.850108Z digest=sha256:382c4e35ba68768ca2509d33abeb14cc6e26de5ae72a67101debfb5b41a0fdb2

Observation 24ecd4be-fbc1-45e2-8776-e81bf6a49384 · outbound

This paper cites Limited Data Emotional Voice Conversion Leveraging Text-to-Speech: Two-stage Sequence-to-Sequence Training.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Limited Data Emotional Voice Conversion Leveraging Text-to-Speech: Two-stage Sequence-to-Sequence Training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:46.872351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:13:46.872351Z digest=sha256:a22a8bedf0915d5272db278ca5e7b39b679473469494dc167d524b98668b8516

Observation ad3a691b-9d2c-41b3-9bd8-35d457b16ff0 · outbound

This paper cites Speech synthesis with mixed emotions,.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech Speech synthesis with mixed emotions,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:13:47.459852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:13:46.876521Z digest=sha256:3e8fdf2e2be342ffb513c815de6eec170fffc6f49666afdc530bc6b250f57b7a

Pith citing papers

Observation 18ec0fb6-68be-4911-9794-c4920a1846e6 · inbound

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech cites this paper.

ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:13:47.399807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:13:45.498553Z digest=sha256:b6f8fadc1eff3924028abb7a763d1e4fe9dcf0a5a42129b12bd1bffa1d2053b2