Pith. sign in

Paper Citation Record · LEDGER

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion

As of 12 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2511.01056.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.01056 v3

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T00:27:32.622659Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T00:27:30.285078Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2033c47d-d8b4-407a-bf84-b43e0798d82e · outbound

This paper cites WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:30.285078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:30.285078Z digest=sha256:8195372624ba85ed1e9036654d505ada124268e771584f334722e3a2e8db7a7f

Observation 6327563c-8ea8-4817-9ce9-72857de7159c · outbound

This paper cites Overview The proposed whisper-to-speech (W2S) framework comprises three stages, as illustrated in Fig.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion Overview The proposed whisper-to-speech (W2S) framework comprises three stages, as illustrated in Fig

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:30.440254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:30.440254Z digest=sha256:ce189c7e84c488a212f7a58ac773664e8f5ab0c7a890520aba558ca10afe3c32

Observation 4bb93aa6-22ec-457e-958d-9c59077066d8 · outbound

This paper cites an unresolved cited work.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:30.539087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:30.539087Z digest=sha256:95e2998d259c89872ffccada710c6cda1afc0de46b817bc2e892e25b6c355769

Observation 83ea74e0-dfff-4670-89f7-f5e78a4d8f60 · outbound

This paper cites Objective evaluations show consistent gains over whispered inputs and performance approaching that of ground-truth recordings in terms of naturalness (DNSMOS 3.11, UTMOS 2.52vs.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion Objective evaluations show consistent gains over whispered inputs and performance approaching that of ground-truth recordings in terms of naturalness (DNSMOS 3.11, UTMOS 2.52vs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:30.633764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:30.633764Z digest=sha256:128773376ea1f870ff40004c69789f02cbe1544861916e87a02724a735368a5a

Observation 93ae1124-b765-4d9a-b87f-b3b786023bad · outbound

This paper cites Attention-Guided Generative Adversarial Network for Whisper to Normal Speech Conversion.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion Attention-Guided Generative Adversarial Network for Whisper to Normal Speech Conversion

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:30.742369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:30.742369Z digest=sha256:9417beca648fce39d48d3d84e12ad992a77b71eaa8e4b075d59f5913027e1d81

Observation 030bb5cc-66fb-4cd9-bbca-eee124b96177 · outbound

This paper cites A novel attention-guided generative ad- versarial network for whisper-to-normal speech conversion,.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion A novel attention-guided generative ad- versarial network for whisper-to-normal speech conversion,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:30.877727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:30.877727Z digest=sha256:ec76453b96ab7e50e4b347d94f464bb587a37076c09e98f77e34d40e64aa75e2

Observation 1a6215b3-b46f-4579-b6ce-c1914bed3cea · outbound

This paper cites End-to-End Whisper to Natural Speech Conversion using Modified Transformer Network.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion End-to-End Whisper to Natural Speech Conversion using Modified Transformer Network

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:31.077349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:31.077349Z digest=sha256:71ce8c0dc670815677e1ba8c1c96f564be57391137ad63c986f372e492d3910b

Observation 2660c0f7-146d-46da-8c7c-55de5fa4a870 · outbound

This paper cites Gener- ative adversarial networks for whispered to voiced speech con- version: a comparative study,.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion Gener- ative adversarial networks for whispered to voiced speech con- version: a comparative study,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:31.209083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:31.209083Z digest=sha256:d0a2c64a865c6d07ac5a90e8b0d6e0d470ce6bdec1264aad3fd4c98056c5650b

Observation ee10939d-fc61-4a92-bc9e-324173ed39a2 · outbound

This paper cites Maskcyclegan-based whisper to normal speech conversion,.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion Maskcyclegan-based whisper to normal speech conversion,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:31.302379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:31.302379Z digest=sha256:2f1a13c19cecd687c827a8affde026f1e957490d3e520d9ae01b5a3d6a862aea

Observation e08590eb-faab-4165-879b-1d715e8a37cb · outbound

This paper cites V ocoder-free non-parallel conversion of whispered speech with masked cycle-consistent generative adversarial net- works,.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion V ocoder-free non-parallel conversion of whispered speech with masked cycle-consistent generative adversarial net- works,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:31.471258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:31.471258Z digest=sha256:c933b8254cf88ff544b6cd617068edadc8c9a67324c82e542b89ab858b5e1a5e

Observation bd4b45c9-b9b1-4ea9-a473-2fb6394a360c · outbound

This paper cites Wesper: Zero-shot and realtime whisper to normal voice conversion for whisper-based speech interac- tions,.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion Wesper: Zero-shot and realtime whisper to normal voice conversion for whisper-based speech interac- tions,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:31.592935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:31.592935Z digest=sha256:fe0473eb6c9891c8fa081cf6da7d38f0a93feee06e567d1e48ce839ba4497fc6

Observation 01832f1e-6f80-4f34-b96a-89feb01029e7 · outbound

This paper cites Distillw2n: A lightweight one-shot whisper to normal voice conversion model using distillation of self- supervised features,.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion Distillw2n: A lightweight one-shot whisper to normal voice conversion model using distillation of self- supervised features,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:31.640744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:31.640744Z digest=sha256:f314bf2848a62c228e0c63ce6510cefbd8aa99edc38f57385a994ffc5e1e2f93

Observation 1a2a0a42-ddca-4696-8201-fbc699ef57aa · outbound

This paper cites Improvement Speaker Similarity for Zero-Shot Any-to-Any Voice Conversion of Whispered and Regular Speech.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion Improvement Speaker Similarity for Zero-Shot Any-to-Any Voice Conversion of Whispered and Regular Speech

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:31.662280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:31.662280Z digest=sha256:e7b9779a49ffee08b23e18bcc9e25066e50c0e46fb07a51c9ee2a0c3a438974b

Observation 894eef0e-9f7d-48dd-887c-61da2d95b4e9 · outbound

This paper cites Whis- pered speech conversion based on the inversion of mel fre- quency cepstral coefficient features,.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion Whis- pered speech conversion based on the inversion of mel fre- quency cepstral coefficient features,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:31.799075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:31.799075Z digest=sha256:5dad7e88bfbc5766d6f08dc4c60004ca5a0b4fbee2bc58d987d8f94ceb13b4c1

Observation 31dae4b7-00d6-4f15-b7a9-e23d9e3303e6 · outbound

This paper cites Glottal flow synthesis for whisper-to-speech conversion,.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion Glottal flow synthesis for whisper-to-speech conversion,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:31.932344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:31.932344Z digest=sha256:a1cd6f6330576d52dd51956cda9af6a258e1b72984e6d007ed904d1cc18d5602

Observation 27428f30-a187-4aef-8e21-e3c211174b1b · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion Robust speech recognition via large-scale weak supervision,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:32.088017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:32.088017Z digest=sha256:08fcea43a6a1bdbea66d750eb99dee1018cd8b229e44cbf0d7239984c3a2bf6a

Observation 467a1e4a-1c9b-4d7a-86b9-ec7f070fe0ce · outbound

This paper cites Aishell6-whisper: A chinese mandarin audio-visual whisper speech dataset with speech recognition baselines,.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion Aishell6-whisper: A chinese mandarin audio-visual whisper speech dataset with speech recognition baselines,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:32.208797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:32.208797Z digest=sha256:60c2f8857ec02535d41990fda3f038cfcff0b0d731c4bc61b95a2b875db441ff

Observation 1aa43ace-e3cb-442d-840a-997ac9e62c34 · outbound

This paper cites Soft-dtw: a differentiable loss function for time-series,.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion Soft-dtw: a differentiable loss function for time-series,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:32.325364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:32.325364Z digest=sha256:d3164354ab97061522f197a441f782bf31f47079bf1c8baa3420e643c4b44c09

Observation 7dac97be-17ba-447e-ad96-257d4e781f18 · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:32.435469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:32.435469Z digest=sha256:1c4d026ba4f065629e3ce0d97e5b360e4ed462909c2e2120a283085f7aca5cea

Observation 6d8a2234-0ca5-4a03-b47e-b1c4c88de3c2 · outbound

This paper cites Wespeaker: A research and production oriented speaker embedding learning toolkit,.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion Wespeaker: A research and production oriented speaker embedding learning toolkit,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:32.459345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:32.459345Z digest=sha256:5ab46ee58ff772e87a3d5830e0f8ee9e0ac8427c7d057026ddda1bc2095e13d2

Observation 4ea14fc2-563a-4b30-9675-108021d84dff · outbound

This paper cites VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:32.487566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:32.487566Z digest=sha256:f2613af4e3de49fdbaea8a89e008be5faf7f4274c7a5975bf217defd9adb6d5d

Observation 3945a79a-ecfe-4eac-8e1c-4e45dd8250e5 · outbound

This paper cites VoxCeleb2: Deep Speaker Recognition.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion VoxCeleb2: Deep Speaker Recognition

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:32.599880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:32.599880Z digest=sha256:d46bc98acc835e1e2dd342719a7a6fbf7d70da15c9fd39b1c7c986bcde0fe8e3

Observation de3b0585-e9fd-4f91-b6fe-3b758095457c · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:32.617233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:32.617233Z digest=sha256:9c4fd7db3a4bd0f4bcb5ed623369b6143a61102d01c062e03fc645163920bf85

Observation 5cbee1b4-df98-485e-891e-deffc5ba766a · outbound

This paper cites Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:32.620011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:32.620011Z digest=sha256:13c12a9b8126da96b44e65c7d398827e4bf5895f7a86a050bcae66b4b5053852

Observation 1a38ff9c-55ca-4e97-9ee8-0db8bbd969e2 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:32.622659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:32.622659Z digest=sha256:2a072ebd8e8ccb0f0c24672949d3002b937e8fd28669b00c54d12405fe400c87

Pith citing papers

Observation 2033c47d-d8b4-407a-bf84-b43e0798d82e · inbound

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion cites this paper.

WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T00:27:30.285078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:27:30.285078Z digest=sha256:8195372624ba85ed1e9036654d505ada124268e771584f334722e3a2e8db7a7f