Pith. sign in

Paper Citation Record · LEDGER

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody

As of 7 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 2 inbound Pith citation observations for arXiv:2508.06890.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.06890 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T22:34:43.671787Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T15:21:28.848595Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T09:11:27.116919Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact2
  • verified fuzzy19
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 20926e1c-a02b-4a11-8558-dec56b9999a8 · outbound

This paper cites Emotional voice conversion: Theory, databases and esd,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Emotional voice conversion: Theory, databases and esd,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:47.562774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:34:41.120573Z digest=sha256:78c6efe87571e2374cd489573cd41f966575d638c3fa8fd28af19afae4e473fb

Observation 9875e92b-cf54-4384-84e3-fa3d1fe0ddf1 · outbound

This paper cites Training socially engaging robots: Modeling backchannel behaviors with batch reinforce- ment learning,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Training socially engaging robots: Modeling backchannel behaviors with batch reinforce- ment learning,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:47.302187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:34:41.210040Z digest=sha256:7ebe5b257f7c06bb209a976fbf0db45ee5476bd3bff1b540284ea5e124e84954

Observation 569bbde1-c154-4d48-a233-e9dd585397c9 · outbound

This paper cites Real-time speech emotion analysis for smart home assistants,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Real-time speech emotion analysis for smart home assistants,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:47.140382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:34:41.279725Z digest=sha256:88e18c4f97be7f2547a023695eaba49a5bbbf07c434f77524c753041f5c01219

Observation 799fd7fa-82b8-47cf-bfe5-052bae3554bf · outbound

This paper cites Pittermann, A.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Pittermann, A

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:46.934170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:34:41.356713Z digest=sha256:f4c17b46ee4a306ec1a0e9e70de83a444221c441ab2b058facad4883d2d37eed

Observation 8056ebe5-a7fe-434a-88d9-2bf869e08d2b · outbound

This paper cites Toward artificial emotional intelligence for cooperative social human–machine interaction,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Toward artificial emotional intelligence for cooperative social human–machine interaction,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:46.714801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:34:41.416266Z digest=sha256:bb3e0ad5f4239bbcf065bdec9db2b41efbed1063e93027608597bc44a1d3d0d4

Observation d084a43a-46cc-4540-9571-8507a4873817 · outbound

This paper cites Pavits: Exploring prosody-aware vits for end-to-end emotional voice conversion,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Pavits: Exploring prosody-aware vits for end-to-end emotional voice conversion,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:46.519963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:34:41.491550Z digest=sha256:7bc1aea89cec303038275ecf78a36803b0d9858edb05239e67587cdc0368af43

Observation 96917dc9-93b3-41d0-8a24-bdb7d82010c2 · outbound

This paper cites Converting Anyone's Emotion: Towards Speaker-Independent Emotional Voice Conversion.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Converting Anyone's Emotion: Towards Speaker-Independent Emotional Voice Conversion

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:41.548368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:41.548368Z digest=sha256:eed719172643a2d9f39eb63ac8ceb0a2c6924ebd92ea3ae24a3a756101cea280

Observation ad683cd8-97c3-4caa-a811-83c1199a57a7 · outbound

This paper cites Limited Data Emotional Voice Conversion Leveraging Text-to-Speech: Two-stage Sequence-to-Sequence Training.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Limited Data Emotional Voice Conversion Leveraging Text-to-Speech: Two-stage Sequence-to-Sequence Training

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-05T22:34:44.023541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:34:41.607221Z digest=sha256:6b7b1fa63d55a1186db20d3d80435e1fb6e53e95e7a07c803a540910c6950c1c

Observation aad03414-d5f0-44cd-9fcf-45afb8ed88bf · outbound

This paper cites Emotion inten- sity and its control for emotional voice conversion,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Emotion inten- sity and its control for emotional voice conversion,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:46.358372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:34:41.694838Z digest=sha256:97fc560e8ccf65e90f9e5aca8b23fa2be5e50748526cfc7fde99e40e25589635

Observation 6b2b43b0-48cb-49d9-9acc-10fe9346ae3e · outbound

This paper cites Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:46.193432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:34:41.760599Z digest=sha256:252dc1f5dd7da45c50684649344edb6419e3d70ccd0b4eda823eef646b4692b3

Observation 211ee3b7-2bc4-4751-9f7a-1e7d739a6f16 · outbound

This paper cites Emotional voice conversion with semi-supervised generative modeling,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Emotional voice conversion with semi-supervised generative modeling,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:46.024775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:34:41.783931Z digest=sha256:7232dd6634f2237720f0b324213d962066f68b46fb35c77f1506a008fa0d0973

Observation 7bd4e329-4d2b-4c5a-909d-00d088e033f0 · outbound

This paper cites Speaker-independent emotional voice conversion via disentangled rep- resentations,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Speaker-independent emotional voice conversion via disentangled rep- resentations,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:45.844240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:34:41.859084Z digest=sha256:1b24ce00c5dd94c5526fe242af981e4b66dbdb024dfc0063180e32647f8c282d

Observation e9bc8915-48d1-43f4-8959-3a690a7847f4 · outbound

This paper cites Nonparallel emotional voice conversion for unseen speaker-emotion pairs using dual domain adversarial network & virtual domain pairing,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Nonparallel emotional voice conversion for unseen speaker-emotion pairs using dual domain adversarial network & virtual domain pairing,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:45.664702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:34:41.941417Z digest=sha256:f87adb792992660c0ff65225620c8a2c3ef82481d2875a865c5e705605b77490

Observation ba414906-ff95-4c66-8f02-5c0bf5a27a81 · outbound

This paper cites Enhancing zero- shot emotional voice conversion via speaker adaptation and duration prediction,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Enhancing zero- shot emotional voice conversion via speaker adaptation and duration prediction,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:45.493477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:34:42.000189Z digest=sha256:d447b804c2a12d19b1b7e22772b77519e8ac8b987e2ee4c04cd5c56be157d46f

Observation 20bdcbc9-46bd-41a1-b92b-881aaaa60d92 · outbound

This paper cites Zero shot audio to audio emotion trans- fer with speaker disentanglement,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Zero shot audio to audio emotion trans- fer with speaker disentanglement,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:45.240270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:34:42.115007Z digest=sha256:4b2aa4ef2a6761e9c9714a86ca5b75bec385b6bce016fba1514a8e9184ad2ddb

Observation 100315df-b02e-4db2-aff9-654018ffebb6 · outbound

This paper cites Multi-speaker emotional speech synthesis with fine-grained prosody modeling,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Multi-speaker emotional speech synthesis with fine-grained prosody modeling,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:45.027825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:34:42.182480Z digest=sha256:6dd770315728481e6522f47d56bf8e49cf3d47af844fc24d1961e4e77abc2bee

Observation 9f1cda4b-3ec9-4fe2-86f8-a70cf40257b4 · outbound

This paper cites Attention is all you need,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Attention is all you need,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:42.258853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:42.258853Z digest=sha256:29091e3be92622316e844c282e0c859cf38fefd4e48f0c0a6f36f1a3d9d18b82

Observation a11c8ce5-6338-410c-bcf9-cee006ee5a8a · outbound

This paper cites Unsupervised domain adaptation by back- propagation,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Unsupervised domain adaptation by back- propagation,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:42.325027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:42.325027Z digest=sha256:f1c7bb3fecca7f2d77cd9c34e224859ca111ea7f536114ec2a5601a17bd27a47

Observation bf6675bf-e40a-4b2a-8e66-d8c90f8f04e5 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:42.391019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:42.391019Z digest=sha256:83d6bcb8a69f07c829189905f23f8001cab172e7cfc4a1cfab418142319ac3cf

Observation 801da7cb-fa6c-41af-8f11-f25e6968e1ea · outbound

This paper cites Sef-vc: Speaker embedding free zero-shot voice conversion with cross attention,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Sef-vc: Speaker embedding free zero-shot voice conversion with cross attention,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:44.833308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:34:42.453515Z digest=sha256:2b946fb3e161e3f590f828528c529e3358ef4e38158fb54106a6a5fe22bc8273

Observation be03c332-48be-41df-aaaf-f72c62f3b4fe · outbound

This paper cites Textless Speech Emotion Conversion using Discrete and Decomposed Representations.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Textless Speech Emotion Conversion using Discrete and Decomposed Representations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:42.520181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:42.520181Z digest=sha256:bf445e49becaf3151bd2fd7d72d4de8825c99598cfc15506ee6e093fa1bedfe5

Observation fb2b3f2e-cacf-44d5-b267-4036f96cc59d · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:42.593208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:42.593208Z digest=sha256:e5cff942b518461ff65330ce630c05cc083e251d30412075f2390aaaebd76573

Observation 00cac1e2-1a7e-484f-9e59-5dc1273d5a1d · outbound

This paper cites Speech emotion diarization: Which emotion appears when?.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Speech emotion diarization: Which emotion appears when?

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:44.612588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:34:42.701389Z digest=sha256:6d69eb1194c4e8a58408cf60a4be7cdc00067105431c606889d3b3275f7995c3

Observation 4643191f-1620-45d5-be46-ab5cccd2a862 · outbound

This paper cites Smoothing and differentiation of data by simplified least squares procedures.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Smoothing and differentiation of data by simplified least squares procedures

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:44.501286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:34:42.785584Z digest=sha256:597f14e8d5a6c12204480aa7e080582eb522e52b3484a0cdd03cadad69df9143

Observation 9b6dc289-b4f0-4549-9aa6-dc52a1638dda · outbound

This paper cites Disentanglement of Emotional Style and Speaker Identity for Expressive Voice Conversion.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Disentanglement of Emotional Style and Speaker Identity for Expressive Voice Conversion

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-05T22:34:43.870586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:34:42.870526Z digest=sha256:0048ed05068b5b22c08edc25ae563f067141775a3a953a4e52b3977f91b90863

Observation a65bb56f-2d89-4a1d-acc0-5cc8eac11491 · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:42.929455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:42.929455Z digest=sha256:7b47cbf9b15d243194cee80ba9af282e08eec08f987f970f4104d3ee91b4f634

Observation 0a3eed7f-de8c-4bad-a880-dc9c14addb56 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Librispeech: an asr corpus based on public domain audio books,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:42.988717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:42.988717Z digest=sha256:2d93a7143f505aec3ffc315119a28281a5fb29d422916a62e4de9a6e7d2d0bb5

Observation 8194931f-49f6-4562-b8c8-9f7815849fbf · outbound

This paper cites VoxCeleb: a large-scale speaker identification dataset.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody VoxCeleb: a large-scale speaker identification dataset

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:43.077499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:43.077499Z digest=sha256:8bb3b2db49d2f980ebbebd98d9e994215d7f7d35b966ca0dbf7bd9823752fd97

Observation e316a941-445a-4437-aca8-b5ff72e5b0ce · outbound

This paper cites World: a vocoder-based high-quality speech synthesis system for real-time applications,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody World: a vocoder-based high-quality speech synthesis system for real-time applications,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:43.132364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:43.132364Z digest=sha256:29a0ad8caa317245027df55e72aa21dfded439c68aaa6701e9a7aa3f5748e6ec

Observation 8545e167-d0ec-48e0-8495-0f55808c1266 · outbound

This paper cites CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:44.378234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:34:43.226294Z digest=sha256:300fb5efe53d3a84a94f809e7c6d0c3c00c9365f66d5e5a1a487f89089c43f62

Observation 6e0497cd-f0cf-4150-acd0-997ad90397e5 · outbound

This paper cites Crema-d: Crowd-sourced emotional multimodal actors dataset,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Crema-d: Crowd-sourced emotional multimodal actors dataset,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:43.313933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:43.313933Z digest=sha256:899f0bb6febbda76b3416e87cbdc9fbbcab6c369d41ba936d10293fc1feaa73b

Observation f8e74ea5-68c6-46db-b60d-59dcc73f5f55 · outbound

This paper cites Iemocap: Interactive emotional dyadic motion capture database,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Iemocap: Interactive emotional dyadic motion capture database,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:43.377951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:43.377951Z digest=sha256:5bf5b8cb5fe05be639a2d29e963a7f01380ad36099aa6f060de1cb7256543250

Observation 76398b06-ef06-4cd6-8449-401c4d65e55c · outbound

This paper cites Robust speech recognition via large-scale weak supervi- sion,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Robust speech recognition via large-scale weak supervi- sion,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:43.467839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:43.467839Z digest=sha256:6983c0e26b4782a88973d39282421552991eb386e67008fc07f5ac9c1c70493c

Observation f5dcc674-87c2-4a7e-a7b8-417df35ec02d · outbound

This paper cites emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:43.524538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:43.524538Z digest=sha256:d74d339fe3b08bc6c1e2208a1f6c454eedd757af34e72605cbf38d30ecd31405

Observation 806a96a9-e1d7-4f86-9e2e-4f1b10d5b70f · outbound

This paper cites Pearson correlation coefficient,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Pearson correlation coefficient,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:43.593809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:43.593809Z digest=sha256:4aed11685267aa7ac2d2bb982b0ebdcb9027768efd7a0cb54dc28692b89c04bf

Observation b15cd2ad-0f72-4282-99ab-f5e65f92e585 · outbound

This paper cites Dynamic time warping,.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody Dynamic time warping,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:34:44.198017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T22:34:43.671787Z digest=sha256:cc8ec8dbd81f516f140be480ceb2fcb214afe31b4f2e64e8f80272077512683e

Pith citing papers

Observation eddc379d-1f05-4fca-a08a-7a160911202c · inbound

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation cites this paper.

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:11:27.119312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T12:34:40.888089Z digest=sha256:47e269d3b769e675e87efd9467f1b8bc705540fdf695db7e87158bea4607b0dd

Observation 0bae5d3c-20b9-4be2-9f76-b6bed6dfe19f · inbound

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation cites this paper.

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T15:21:28.848595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:21:28.848595Z digest=sha256:7a65d90a88ec324768d0f543504abb75be25bab60f48800dc5a0e39ccb993a9f