Pith. sign in

Paper Citation Record · LEDGER

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion

As of 12 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 3 inbound Pith citation observations for arXiv:2506.04013.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04013 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:55:20.365775Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:55:20.110093Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T19:07:18.246563Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy33
  • unresolved14
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 73dae815-3c55-4526-84c8-912be14e0f2d · outbound

This paper cites Conventional VC models perform well in replicating speaker identity but struggle when the target speech is highly expressive.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Conventional VC models perform well in replicating speaker identity but struggle when the target speech is highly expressive

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.855135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.100974Z digest=sha256:e82f4bcf634042b61489a35b6bacf8e9e13689447392f8f8030cd2577d9ee176

Observation 60d15bb1-2e4b-439f-9dfc-c7c6aa33ce69 · outbound

This paper cites These models typically required text supervi- sion.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion These models typically required text supervi- sion

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.845772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.105713Z digest=sha256:c02c0a930dcbfdfe266b3a6ba7f282411954325045fd7d477505e35fd5a9ccb2

Observation 2b05af34-6da4-4dab-be57-1f5d46dc096c · outbound

This paper cites an unresolved cited work.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:55:20.837000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.113786Z digest=sha256:20b87cd4efe17c55332e74184ad363d4dcda53140d3ef85f8088ef908ef674cd

Observation d66535c6-1087-46f6-aa49-3738da5f3ba0 · outbound

This paper cites All datasets are English and total duration is around 228 hours with more than 920 speakers.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion All datasets are English and total duration is around 228 hours with more than 920 speakers

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T10:55:20.825514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.117679Z digest=sha256:3da42403ffa4ff8c69804073f06f2454957db9a82a67bb89042f686e29cad832

Observation 89c442f2-6f5c-4f5a-8dd8-d074ee68083e · outbound

This paper cites an unresolved cited work.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:55:20.816313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.121966Z digest=sha256:04054d45fe12f6cfe835771d02a4b7e3926e69ce1e4ebf4c43fa3cfc830bf610

Observation e262b5f6-fc04-414b-a4fc-43b864e1852d · outbound

This paper cites an unresolved cited work.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:55:20.805968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.125796Z digest=sha256:ce6ceccfdf3eb24a460ba75e6d07891175afeaadf78758cd0c06c20d3c1d9ae2

Observation 92367e0f-9cda-4004-997d-3c19c61642e7 · outbound

This paper cites Styles2st: Zero-shot style transfer for direct speech-to- speech translation,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Styles2st: Zero-shot style transfer for direct speech-to- speech translation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.760631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.151933Z digest=sha256:d21216805190243417da72dedf100f1af941fb545d76cbab7a9466cfabb89433

Observation c660e27c-596c-48f2-9e38-a14f95eef46e · outbound

This paper cites Chil: Computers in the human interaction loop,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Chil: Computers in the human interaction loop,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.796781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.129443Z digest=sha256:bd5222cd087c8e76f1a388d65b2011a879a62f3a1169be40bea0076a2e6b8369

Observation ad0ef1ef-5d7f-465f-b9f1-e912a536f3a0 · outbound

This paper cites Towards an open-domain social dialog system,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Towards an open-domain social dialog system,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.788033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.132609Z digest=sha256:8cc7be86134db1c34462946a0967257f59956861e3dff2f3402f7ecbeac6f0e0

Observation 25550670-15b0-48b7-9b2b-280828d1941b · outbound

This paper cites Face-dubbing++: Lip-synchronous, voice preserv- ing translation of videos,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Face-dubbing++: Lip-synchronous, voice preserv- ing translation of videos,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.779015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.136544Z digest=sha256:738e767a330be2294ba1964311b9fdff55fed6c7f8c71481fd78b980b76e35aa

Observation c9a77ed3-3ba0-4d54-8e9e-3620d532acd9 · outbound

This paper cites Findings of the IWSLT 2024 Evaluation Campaign.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Findings of the IWSLT 2024 Evaluation Campaign

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.139760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.139760Z digest=sha256:03d530868b653508fc31528dc8638a40062f8eb7523b445dbce8e4fa247f54c9

Observation ba314579-786e-4c6b-a2b0-691e19b159d0 · outbound

This paper cites Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.110093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.110093Z digest=sha256:ba707fdbe7d3dd92c925a11d30ebdc94f0cd03ce002f0d59cf02309f2273d913

Observation 0ab7ae64-8fa2-4efe-ba84-aec402de6f5e · outbound

This paper cites Simultaneous translation of open do- main lectures and speeches,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Simultaneous translation of open do- main lectures and speeches,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.770020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.144697Z digest=sha256:61c27395de07d74d7bb57451ff06628b836243069caf760274320064a1fec69e

Observation c9f3b311-7eea-48d9-9bb0-b849850a81b0 · outbound

This paper cites Seamless: Multilingual Expressive and Streaming Speech Translation.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.148186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.148186Z digest=sha256:9865ef6cab6aedd4c1cd92bcb5cf091f531b4aa4f2598aea703f8950f28f238f

Observation 5a98814b-9b96-409e-8a88-75df122dd784 · outbound

This paper cites Limited data emotional voice conversion leveraging text-to-speech: Two-stage sequence-to- sequence training,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Limited data emotional voice conversion leveraging text-to-speech: Two-stage sequence-to- sequence training,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.750351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.155032Z digest=sha256:acb6b85ccae32cf3f1d5679f7fa8e5b708a4bad0b2795d4ed281fe86e157d08e

Observation 28febc5b-ccf0-467a-95d5-5131f4e549e8 · outbound

This paper cites Nonparallel emotional speech conversion,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Nonparallel emotional speech conversion,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.740558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.158385Z digest=sha256:8e49350626e050c7591afffee7782a4985d2ca5636087b50abaf96c604339c09

Observation c9294002-d084-43b3-b41f-78731dc9f955 · outbound

This paper cites Nonpar- allel emotional speech conversion using vae-gan.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Nonpar- allel emotional speech conversion using vae-gan

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.730923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.161449Z digest=sha256:f4ac42fc95cca85a99232ca54a65ec005b6a0cc2395ba3b473bbc728551072b0

Observation 90cacd4e-433b-4991-b40f-d1fb14fc9821 · outbound

This paper cites Textless speech emotion conversion using discrete and decom- posed representations,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Textless speech emotion conversion using discrete and decom- posed representations,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.721511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.165248Z digest=sha256:0bb2831ce6e83fdf43d38b715ec5fa56eb494ead3f243faddf675313bc5fd005

Observation 3e6b91b7-04d0-456e-8398-c26d7a5e7f41 · outbound

This paper cites Hiervst: Hierar- chical adaptive zero-shot voice style transfer,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Hiervst: Hierar- chical adaptive zero-shot voice style transfer,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.711516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.168157Z digest=sha256:1feae4fc499de39aaff6e4201dc6adf419edc297042316475b561df0264d07a4

Observation c5850056-bcc9-495f-af8a-9b5d27e67111 · outbound

This paper cites Expressive-vc: Highly expressive voice conversion with attention fusion of bottleneck and perturbation features,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Expressive-vc: Highly expressive voice conversion with attention fusion of bottleneck and perturbation features,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.701184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.172054Z digest=sha256:1d70602ee71e0cd71ff2a2a8f4344381608a6821b606ca9c08990eb78dbfb871

Observation aec52e0c-9e09-44f4-8d98-c38be0db8dcc · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.690149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.176032Z digest=sha256:8cb7b035bbf9c90d6673b65b9c59c096ce9c513c55ae488d1cd265be69442326

Observation da85ec13-5aee-4693-b038-fcfb0199455f · outbound

This paper cites Freevc: Towards high-quality text-free one-shot voice conversion,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Freevc: Towards high-quality text-free one-shot voice conversion,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.680603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.179658Z digest=sha256:84c650520a9236fa352dff9e8adcfdfb33c793ffc99d5e7c0c40d839753bcaec

Observation c3e35a3f-36c7-449b-b665-11c3fdf4c2da · outbound

This paper cites mhubert-147: A compact multilingual hubert model,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion mhubert-147: A compact multilingual hubert model,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.670869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.183562Z digest=sha256:6431a46e034735f798e590c012f63246071ec8d90b080f38e2fea15f907e56bb

Observation 10c8d12d-7494-4972-a593-9ecf7d8b2fb6 · outbound

This paper cites V oice privacy- investigating voice conversion architecture with different bottle- neck features,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion V oice privacy- investigating voice conversion architecture with different bottle- neck features,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.661268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.186998Z digest=sha256:b0c89041d9629f37d0156b29c7179931c203c54c3ddfd81d74f12e29c63d198b

Observation 92ec07ed-625a-468a-9f3f-fbd34e7cb781 · outbound

This paper cites Generspeech: Towards style transfer for generalizable out-of-domain text-to- speech,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Generspeech: Towards style transfer for generalizable out-of-domain text-to- speech,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.651408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.190765Z digest=sha256:a84c0bfec8366065b2b4b98850d7234da8351c11a59fb1eb2063dd6fbcb5c1d8

Observation 4d272469-5516-4d6f-8533-831cf5137352 · outbound

This paper cites Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.194507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.194507Z digest=sha256:cad8fa6496ad6d5d79011cbf97034691810af5b39318d04ed104b5fc575cafeb

Observation 5910ef84-2008-41a0-a9c6-2d41bbdbac57 · outbound

This paper cites Emotion intensity and its control for emotional voice conversion,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Emotion intensity and its control for emotional voice conversion,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.636704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.197913Z digest=sha256:112d9a53ae82a69aec69ecd42058ba68cb6a23e1b02a17366aa7a678229b7083

Observation 20b2f702-bbf1-430c-b7e3-baaee31197d3 · outbound

This paper cites Accent conversion using pre-trained model and synthesized data from voice conver- sion.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Accent conversion using pre-trained model and synthesized data from voice conver- sion

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.627065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.294744Z digest=sha256:cc6f31f5a16ca6db858d82d19ae48b7a5d1ab4c7c9ae6bde845535a0d03a5a06

Observation 2e0dc039-9d98-47b0-96f2-9f98dbf965b3 · outbound

This paper cites Improving pronunciation and accent conversion through knowledge distilla- tion and synthetic ground-truth from native tts,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Improving pronunciation and accent conversion through knowledge distilla- tion and synthetic ground-truth from native tts,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.616973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.298990Z digest=sha256:b41d1fe2d94a2d6d8e433e064decc49d4ad4ce665a4d5b7b3d7e0644677db652

Observation e8cf2831-0a5b-4443-aa8a-5b00fb0bde59 · outbound

This paper cites Stargan for emo- tional speech conversion: Validated by data augmentation of end- to-end emotion recognition,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Stargan for emo- tional speech conversion: Validated by data augmentation of end- to-end emotion recognition,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.607359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.302169Z digest=sha256:060065e79487f612b17eec3344190fb70453e79f9d354507bdb407402eed5502

Observation 9d68f57f-7d99-40b1-9cb3-0fcb5d3b92f9 · outbound

This paper cites EmoCat: Language-agnostic Emotional Voice Conversion.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion EmoCat: Language-agnostic Emotional Voice Conversion

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.306160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.306160Z digest=sha256:49cb1961b2f31359a5cc9956444fbf399a70d0d566469b0143340f69fcabea2a

Observation 95c0fc03-f874-48e2-8517-fe0d1d91aed8 · outbound

This paper cites V oice conversion with just nearest neighbors,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion V oice conversion with just nearest neighbors,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.597489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.309572Z digest=sha256:584f6f781f6c11772118a426933fc7382bce02abca0542088cc4e418c3be55d2

Observation 03777a17-2dc4-4fa3-8d3d-4eb9d74b8b6a · outbound

This paper cites Disentangling prosody representations with unsupervised speech reconstruction,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Disentangling prosody representations with unsupervised speech reconstruction,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.587981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.312653Z digest=sha256:1030c15d18a11e68dd02ebe7b0653ddd889b01e1b1b349c6a61b9b634f45e9d9

Observation b804b6f8-c9b6-4dea-a3cf-412b9b014202 · outbound

This paper cites Using joint train- ing speaker encoder with consistency loss to achieve cross-lingual voice conversion and expressive voice conversion,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Using joint train- ing speaker encoder with consistency loss to achieve cross-lingual voice conversion and expressive voice conversion,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.577501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.315921Z digest=sha256:4af91d308d81a12af34f335e045c07d29e8578b8d88e31e70c8a0436cc36456e

Observation 70f5d6af-002d-4d63-b689-720df816fb73 · outbound

This paper cites X-e-speech: Joint training framework of non- autoregressive cross-lingual emotional text-to-speech and voice conversion,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion X-e-speech: Joint training framework of non- autoregressive cross-lingual emotional text-to-speech and voice conversion,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.566994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.319244Z digest=sha256:967cc425b370c69e67e0c8859ba11991291b13eb472e411c5af57779ec7528d5

Observation 2613bdb2-d8e8-4fd1-acf3-8c897598eabc · outbound

This paper cites Zse-vits: A zero-shot expressive voice cloning method based on vits,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Zse-vits: A zero-shot expressive voice cloning method based on vits,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.556372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.322515Z digest=sha256:dac135fd6f84d727364f82151041ef6d25fcd0e53115bfabdc28e2ca7a1ae48b

Observation 465a0c8c-bb1e-4e75-8021-eb5edcc90a2c · outbound

This paper cites HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.325815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.325815Z digest=sha256:b926c777e7f84d6f6b0eb104fb019833c19d63cc909b2e9833f06369e619b086

Observation da1d2bd3-e7b9-4952-b276-cfb28b0a60f6 · outbound

This paper cites VITS2: Improving Quality and Efficiency of Single-Stage Text-to-Speech with Adversarial Learning and Architecture Design.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion VITS2: Improving Quality and Efficiency of Single-Stage Text-to-Speech with Adversarial Learning and Architecture Design

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:55:20.427901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.329181Z digest=sha256:cedf5367d462f0d651704d21c7b5cf097b85f0df1e3ff93b2faf809d3dd4ae62

Observation cdd4abe5-1b3c-4335-b5e3-418d8ea4ea1a · outbound

This paper cites Hifi-gan: Generative adversarial net- works for efficient and high fidelity speech synthesis,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Hifi-gan: Generative adversarial net- works for efficient and high fidelity speech synthesis,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.545193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.333182Z digest=sha256:7c9169a28ce6dceca1ac7d01f175bba0c6d740a761090c3e55bea9414d8e7ad3

Observation 2f69c47b-1579-421b-8de8-6ab66a58c7f5 · outbound

This paper cites Libritts: A corpus derived from librispeech for text- to-speech,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Libritts: A corpus derived from librispeech for text- to-speech,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.534957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.336222Z digest=sha256:c59855587d07a48206da45b88c15297e130f718a162d6915457ada8d4b2e2002

Observation fe8a54c9-162f-4886-a27f-b26498b66e42 · outbound

This paper cites Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.524960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.339190Z digest=sha256:02dadeb77dc2728649ed413e493c704dccc8e9e7e1faedd1bc5f6460bdbf95eb

Observation b1d1b093-dbc2-468e-83ea-b92226d3d12a · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.342184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.342184Z digest=sha256:92a1f6ffa5ffe572acf6a915b93f21da6e54720be78b6427534813ddb4e2365a

Observation a5e2b151-c0e1-4e2c-beb8-922275a56e54 · outbound

This paper cites EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.345371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.345371Z digest=sha256:e6a4a96e1a8d094a2100289a7456f30f75051acfb96d94e2181c1e6e1235fa37

Observation 76475748-f2ec-45c9-b57d-617c56837d54 · outbound

This paper cites Robust speech recognition via large-scale weak su- pervision,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Robust speech recognition via large-scale weak su- pervision,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.349892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.349892Z digest=sha256:336baac5f74d7695e12a3dab1b3db1705779c7b008ebf914736b50f8811e4171

Observation 2c0e4eed-aa8b-4d57-84ec-3030fd4b73d1 · outbound

This paper cites emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.353059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.353059Z digest=sha256:be72ecfef44554144f4b5cbd237a6b2fd82b092d67aeaff3c3853e3497c565f5

Observation 4933fb26-7b1b-41ba-a702-96cae9263b1a · outbound

This paper cites The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.508696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.356312Z digest=sha256:53ffb7f87144cfc2a5d73fd6648f117ff056c5caaa3ae05ba4533f8397eb3737

Observation 90739c25-a623-4060-a345-6fcf22a17a06 · outbound

This paper cites A comparison of discrete and soft speech units for improved voice conversion,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion A comparison of discrete and soft speech units for improved voice conversion,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.498394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.359499Z digest=sha256:4603c1b003ddbc2ad094ceab78826ca94e5266b6cdb809be139b7d88d3766ab5

Observation 48e54d4e-5efa-4b12-8455-e53b9a662115 · outbound

This paper cites Scaling speech technology to 1,000+ languages,.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Scaling speech technology to 1,000+ languages,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.362733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.362733Z digest=sha256:83fd5d229fd2e46b53ee5419da662e6402ee9f215e01cb72c19744f5fdd406ec

Observation f0698fca-027e-419d-a48e-4d851c7edf2b · outbound

This paper cites A database of german emotional speech.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion A database of german emotional speech

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:20.482825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T10:55:20.365775Z digest=sha256:53d27df8c963e001cc2abaa867bdbfcfe40bd25c78340029d3d6edd71d8d1057

Pith citing papers

Observation ba314579-786e-4c6b-a2b0-691e19b159d0 · inbound

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion cites this paper.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.110093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.110093Z digest=sha256:ba707fdbe7d3dd92c925a11d30ebdc94f0cd03ce002f0d59cf02309f2273d913

Observation ca35d848-c3ba-4ab9-bd33-71cd9388e510 · inbound

Mixed-Precision Information Bottlenecks for On-Device Trait-State Disentanglement in Bipolar Agitation Detection cites this paper.

Mixed-Precision Information Bottlenecks for On-Device Trait-State Disentanglement in Bipolar Agitation Detection Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:00:36.608042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-08T19:11:30.638672Z digest=sha256:a6d94be377a366d66535003992f40f22c0e239a85dbeb30d70164f89c44bdf44

Observation ad6d5a5a-38a2-4068-a167-03f84c166bfe · inbound

KIT's Submission to Cross-Lingual Voice Cloning in IWSLT 2026 cites this paper.

KIT's Submission to Cross-Lingual Voice Cloning in IWSLT 2026 Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T19:07:18.248088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T21:39:11.265338Z digest=sha256:0c595e6c93cceb25f4656e134dd7ff6af7754e7e4710ebfc4b75cce2674b34f9