Pith. sign in

Paper Citation Record · LEDGER

Hierarchical Control of Emotion Rendering in Speech Synthesis

As of 11 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 1 inbound Pith citation observation for arXiv:2412.12498.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12498 v3

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:06:08.843002Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:21:09.698405Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T11:21:12.229573Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact2
  • verified fuzzy62
  • unresolved19
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 08d3e8ac-5307-4189-8835-5b7965c3d74e · outbound

This paper cites An overview of affective speech synthesis and conversion in the deep learning era,.

Hierarchical Control of Emotion Rendering in Speech Synthesis An overview of affective speech synthesis and conversion in the deep learning era,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.529362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.529362Z digest=sha256:51477ed09f81f1330ac3d8994960db5bd774ef39ecf60e32245b9633b80623e6

Observation 7ea42af2-61d1-49a6-a472-6ccc45c4f23b · outbound

This paper cites The age of artificial emotional intelligence,.

Hierarchical Control of Emotion Rendering in Speech Synthesis The age of artificial emotional intelligence,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.534138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.534138Z digest=sha256:517bd31ddcf908c64258027f3e2b840ee64beb264ea283aa650f6985acd9ccb8

Observation f5a65bdb-402c-4b67-8377-f3a677b3e779 · outbound

This paper cites Pittermann, A.

Hierarchical Control of Emotion Rendering in Speech Synthesis Pittermann, A

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.537992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.537992Z digest=sha256:3cf4f14f4a7fbe3103f0e85658a43765d0a157d67aa345f893b8d380c84a4b3a

Observation de745932-fa9b-48d9-8ed3-98cb491f9fe0 · outbound

This paper cites Emotion modelling for speech generation,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Emotion modelling for speech generation,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.542142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.542142Z digest=sha256:8a156c769b21261a2def50fe958aa94268f529a932f1e70c882b61595688f45a

Observation d8e9decb-2849-4676-b019-2d95944d4b8d · outbound

This paper cites An overview of affective speech synthesis and conversion in the deep learning era,.

Hierarchical Control of Emotion Rendering in Speech Synthesis An overview of affective speech synthesis and conversion in the deep learning era,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.545989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.545989Z digest=sha256:1765c58a49bb6e156e9ac4466b8d047a2f674eb5981e095d5a779654794cbb6d

Observation bbb3fce4-6774-4428-86de-191eef892933 · outbound

This paper cites Fastspeech 2: Fast and high-quality end-to-end text to speech,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Fastspeech 2: Fast and high-quality end-to-end text to speech,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.554067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.554067Z digest=sha256:ad7e6d272650520864c3f3a456f442280598cebc10de8d4d3b9b50a89ace7407

Observation 8566ab21-8c13-4f2a-8c06-cf4144af8221 · outbound

This paper cites Phonetic enhanced language modeling for text-to-speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Phonetic enhanced language modeling for text-to-speech synthesis,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.740594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.557786Z digest=sha256:e8766595152cf019c4799b590b89edb3f329346200f693c0f462a69ac6c7486c

Observation a7099819-e79b-4cad-a1d3-31fbe54e0631 · outbound

This paper cites Text-to-speech for low-resource agglutinative language with morphology-aware language model pre-training,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Text-to-speech for low-resource agglutinative language with morphology-aware language model pre-training,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.729060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.561903Z digest=sha256:dc9bc464dea9aaf0a4fad53fbd53cc546fe7f4b69463a0f70bf051a54f020e15

Observation 2bb81916-c203-42e1-9bbb-fdd053892d84 · outbound

This paper cites Emotional speech synthesis: A review,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Emotional speech synthesis: A review,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.718028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.565933Z digest=sha256:69b334fd4f21a36b6a4b728a34b7950bcc1b780d977c90fe35b385b6e1c24803

Observation 1251d4f1-acd2-460b-87af-466245531d99 · outbound

This paper cites A Survey on Neural Speech Synthesis.

Hierarchical Control of Emotion Rendering in Speech Synthesis A Survey on Neural Speech Synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.569719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.569719Z digest=sha256:f4d887a15e130d6c82fe8550f9e25b4ed21e2abb36e4ff6004b42853aa863c35

Observation d8e301cf-7f52-4092-9ec3-5cf01d6801e8 · outbound

This paper cites Pragmatics and intonation,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Pragmatics and intonation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.705759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.574210Z digest=sha256:4d834b04dcabdd6032093ee5f23d145983280f2896d1a4cc78257fa317c96120

Observation 72648521-38b0-47c5-afc8-d4598095268d · outbound

This paper cites Perception of affective and linguistic prosody: an ale meta-analysis of neuroimaging studies,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Perception of affective and linguistic prosody: an ale meta-analysis of neuroimaging studies,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.693377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.577953Z digest=sha256:99e7f83b53ed4a1c63c9e6fa763b6789497c0894c208084d74ef11d9a9afd7f7

Observation 0940779d-18ad-40c5-9d9c-a35dededc153 · outbound

This paper cites Laukka,Vocal Communication of Emotion.

Hierarchical Control of Emotion Rendering in Speech Synthesis Laukka,Vocal Communication of Emotion

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.581669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.581669Z digest=sha256:76b6dbab4d502f87c5d8f5353f5361fb4275d7ac6ca81e54fba898918c9356e1

Observation 553e88ce-f7e0-4a08-8928-aaea88dd6c48 · outbound

This paper cites Vaw-gan for disentanglement and recomposition of emotional elements in speech,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Vaw-gan for disentanglement and recomposition of emotional elements in speech,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.681139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.585131Z digest=sha256:278263d115d5ad669490da65695f37eedf52390435ffb2cd9b2db27b45fcf39a

Observation 86230675-45ef-4fb5-aef9-c97c60873202 · outbound

This paper cites Converting anyone’s emotion: Towards speaker-independent emotional voice conversion,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Converting anyone’s emotion: Towards speaker-independent emotional voice conversion,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.669949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.588437Z digest=sha256:0cdc12e91b5e20e8f5c0308ddba305ff4deaf59fda0b0cfe2760462546bfd67a

Observation 5cf6d6f2-b8ca-4c5c-b2c8-ae8e739f3d18 · outbound

This paper cites opensmile – the munich versatile and fast open-source audio feature extractor,.

Hierarchical Control of Emotion Rendering in Speech Synthesis opensmile – the munich versatile and fast open-source audio feature extractor,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.658229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.591916Z digest=sha256:251058b8b114ea55049a9b005b7abd905124530a5d016c17cbac889064f29f0d

Observation af631296-ce47-423f-8982-1fb9d22ec81a · outbound

This paper cites emotion2vec: Self-supervised pre-training for speech emotion repre- sentation,.

Hierarchical Control of Emotion Rendering in Speech Synthesis emotion2vec: Self-supervised pre-training for speech emotion repre- sentation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.647402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.595316Z digest=sha256:a9b8eb77072ca787a7dcaf7eca842a515e68bb564785a8d1c7d3837cf44e9725

Observation f375cf66-d0ae-4c25-bd10-e11b0e70e90c · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Wavlm: Large-scale self-supervised pre-training for full stack speech processing,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.636109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.598849Z digest=sha256:6c20bcb9a6906750c9e1077f9547aedbc494d028a5ae2dd3ea7dfc829e74c10b

Observation 6b3de789-6f1f-46f8-9513-141b08bd1b1d · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.624187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.602516Z digest=sha256:0629857c3d69c3eeb1ca5ba82cc012b1ac0fd371a179e11056c1075ed2cb40f9

Observation 51bc479e-69e0-4c27-9e2d-5146282bcce9 · outbound

This paper cites Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.611419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.606036Z digest=sha256:a4805f52a962caa35b8454c1c49fc8548d74d75446173dde1ca8c5cee49f9b86

Observation b4b56386-85af-455d-b678-41e89379a2d3 · outbound

This paper cites End-to-end emotional speech synthesis using style tokens and semi-supervised training,.

Hierarchical Control of Emotion Rendering in Speech Synthesis End-to-end emotional speech synthesis using style tokens and semi-supervised training,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.599773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.609465Z digest=sha256:03b778ffcae2dc52bc78b047ee78b61d9c5970a80a5dd0dc2f402c1b8337f15e

Observation 9e75c06d-5f43-4436-9fad-9df94ae907c3 · outbound

This paper cites Controllable emotion transfer for end-to-end speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Controllable emotion transfer for end-to-end speech synthesis,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.587506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.613031Z digest=sha256:4bd5be99ef29d9be145750fb323b5501fb5ec17a9ff5c9a223f77875c7209918

Observation fbf80c48-966e-4927-9fe7-771385373f4e · outbound

This paper cites Semi-supervised learning for contin- uous emotional intensity controllable speech synthesis with disentangled representations,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Semi-supervised learning for contin- uous emotional intensity controllable speech synthesis with disentangled representations,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.576093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.616553Z digest=sha256:c7b8522635610687fe7d91612d54d9c8416ada9cc25e7e9b59698713c2ac53b7

Observation 58483e6a-fdcb-424b-8ea5-20ef5f380240 · outbound

This paper cites iemotts: Toward robust cross-speaker emotion transfer and control for speech synthesis based on disentanglement between prosody and timbre,.

Hierarchical Control of Emotion Rendering in Speech Synthesis iemotts: Toward robust cross-speaker emotion transfer and control for speech synthesis based on disentanglement between prosody and timbre,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.563496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.620007Z digest=sha256:0183532f2de75dbd759d25f2f469dd4578432612bdbdf35257fea786ad6dee6d

Observation df59d7bb-c414-461c-930c-afd84cc6d1fb · outbound

This paper cites Cross-speaker emotion disentangling and transfer for end-to-end speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Cross-speaker emotion disentangling and transfer for end-to-end speech synthesis,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.551617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.623521Z digest=sha256:ddfa3802dbcc5c6c0bce14a75375484d26ce54f45d86fd8a2817b888792c155e

Observation 1bd9a0a4-7126-4814-98ee-cc77c2016b49 · outbound

This paper cites Relative attributes,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Relative attributes,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.540602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.626913Z digest=sha256:503f3bc8d481a4c6cbc58887cfd9786c431bba5c1a83fa47b3de4fb874900318

Observation e53eaca7-7203-4ded-ae55-c3c69f3a4bfe · outbound

This paper cites Emotion inten- sity and its control for emotional voice conversion,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Emotion inten- sity and its control for emotional voice conversion,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.529277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.630492Z digest=sha256:04570fc9ccddfd34054bbeee65170fa64fc85ed4ebe2436a57356235d9dc1286

Observation 007f184e-c906-4cc2-8aea-4c71f99e0347 · outbound

This paper cites Controlling emotion strength with relative attribute for end-to-end speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Controlling emotion strength with relative attribute for end-to-end speech synthesis,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.517716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.633927Z digest=sha256:ab38645281af333f9417a3576d73c4bcecf87f75ef56f29136f2102205c11d8e

Observation 62d61b50-b565-476d-9b3f-ee3540482272 · outbound

This paper cites Speech synthe- sis with mixed emotions,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Speech synthe- sis with mixed emotions,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.505413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.637550Z digest=sha256:fca2b7a0135020dda86e444f9b197ffb52ded2a23a322708242452ca85f6ec34

Observation a7917f73-2356-4f6a-951d-d65bd4af3bd7 · outbound

This paper cites Msemotts: Multi-scale emotion transfer, prediction, and control for emotional speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Msemotts: Multi-scale emotion transfer, prediction, and control for emotional speech synthesis,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.494097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.641090Z digest=sha256:f3c054e3ad0f3172435ee3201b9ba2a18245bb91b425615a3669d6fdb603c334

Observation bd28a430-d85f-4c44-a79f-783e1804141a · outbound

This paper cites Hierarchical emotion prediction and control in text-to-speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Hierarchical emotion prediction and control in text-to-speech synthesis,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.482373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.644651Z digest=sha256:70bf6be0d65175aa0435bbcfecccd1b80a95c365a205f653cef9e513c2863d09

Observation 29abbc49-652d-41bf-8193-ac8a04c500ec · outbound

This paper cites Fine-grained quantitative emotion editing for speech generation,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Fine-grained quantitative emotion editing for speech generation,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.470470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.648316Z digest=sha256:5f32214cdb9588e3c2f5cc8a4164f13dabf7e34ed4caaed448577b8d1c419c9e

Observation c6ec9f19-999d-4bc1-a840-16411c1d93cf · outbound

This paper cites an unresolved cited work.

Hierarchical Control of Emotion Rendering in Speech Synthesis Unresolved cited work

Reference 33

Resolution
verified exact
doi, observed 2026-08-11T14:06:08.897451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.656176Z digest=sha256:e9747bbf9a72bab1837a128b9a00eb6f9c36ea2b3c07ec2b258f09060577163a

Observation 49a28819-f679-415d-a143-4435751171f9 · outbound

This paper cites Transforming spectrum and prosody for emotional voice conversion with non-parallel training data,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Transforming spectrum and prosody for emotional voice conversion with non-parallel training data,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.458206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.660221Z digest=sha256:2da030ce0438a7ae06f80183d4d514e4a67105d7d919f6490afb3e8b256205e8

Observation 4e76db84-c4fc-481c-bb83-acf416247de6 · outbound

This paper cites Intonation and emotion: Influence of pitch levels and contour type on creating emotions,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Intonation and emotion: Influence of pitch levels and contour type on creating emotions,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.446239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.664011Z digest=sha256:eb6ced72d80b3d48e0789a705e12ec4a826b57845c763ccec0b397b250f28255

Observation d171ff62-628a-42d7-a24d-56491c56f374 · outbound

This paper cites Norms of valence, arousal, and dominance for 13,915 english lemmas,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Norms of valence, arousal, and dominance for 13,915 english lemmas,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.434523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.667486Z digest=sha256:236b106e860e27029969c4358e15a8c4900451dd075a183cb2e6feee12c926ec

Observation 7a113808-64e3-4cfd-94b5-99c826fbe3c7 · outbound

This paper cites The influence of pitch range, duration, amplitude and spectral features on the interpretation of the rise-fall-rise intonation contour in english,.

Hierarchical Control of Emotion Rendering in Speech Synthesis The influence of pitch range, duration, amplitude and spectral features on the interpretation of the rise-fall-rise intonation contour in english,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.422051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.671081Z digest=sha256:241833649c1604da26ab5d04304119a4ce957ace94276c2676d175c913d5138d

Observation 74cbc6ab-b04a-4be1-8957-d538deadb96b · outbound

This paper cites Using prosody to avoid ambiguity: Effects of speaker awareness and referential context,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Using prosody to avoid ambiguity: Effects of speaker awareness and referential context,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.408439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.674691Z digest=sha256:93d04ec60d979f3b05c6ed9ef56307edef4666d0c8badb3d2fda1e441913644c

Observation 846906b9-1916-4720-af85-33c619be4bbb · outbound

This paper cites Analysis of emotionally salient aspects of fundamental frequency for emotion detection,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Analysis of emotionally salient aspects of fundamental frequency for emotion detection,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.396219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.678149Z digest=sha256:59e5cb080be70db4be140bce3ae99f2b81f36734815c5a727da5e5d0fc3a2c69

Observation 6f37876d-3586-4f80-821e-57263b1036d0 · outbound

This paper cites Survey on speech emotion recognition: Features, classification schemes, and databases,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Survey on speech emotion recognition: Features, classification schemes, and databases,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.384194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.681733Z digest=sha256:b8a7265a68a99d99ab74bb066814cab2418b7e723065a4b156a7eefa1e3a73ac

Observation 507fc731-cad1-408c-bd43-699b8498c694 · outbound

This paper cites Expression of emotional–motivational connotations with a one-word utterance,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Expression of emotional–motivational connotations with a one-word utterance,

Reference 41

Resolution
verified exact
doi, observed 2026-08-11T14:06:08.885442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.685558Z digest=sha256:e0c1c6716a125d9059a939c2c6473c87659549cb24c5d16fa7a18efa6ecebbfa

Observation 04ad7066-1700-451b-93f0-d389f3b42c8b · outbound

This paper cites Controllable accented text- to-speech synthesis with fine and coarse-grained intensity rendering,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Controllable accented text- to-speech synthesis with fine and coarse-grained intensity rendering,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.372004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.689285Z digest=sha256:316965619698e6d2eb946fd64a5015f6057141fa0af9f9a247e262970b46b561

Observation 3a71807e-85b6-4a64-acb3-087cbf41ba31 · outbound

This paper cites Connecting cross-modal rep- resentations for compact and robust multimodal sentiment analysis with sentiment word substitution error,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Connecting cross-modal rep- resentations for compact and robust multimodal sentiment analysis with sentiment word substitution error,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.359568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.693128Z digest=sha256:d62d6361453bbe7707b6d8d59c2fbfe972b465b70439d5afcb153136afd7847d

Observation 7adf3645-7708-4a2d-9e70-c517a3ce8768 · outbound

This paper cites Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.697089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.697089Z digest=sha256:a886a0789b1e2f70b682a6b8745b751048d1e92f4e50119788a7ffd1db53c6b7

Observation f055f3e1-1a1f-4e76-8ad5-c13e09ec8813 · outbound

This paper cites Improve emotional speech synthesis quality by learning explicit and im- plicit representations with semi-supervised training.

Hierarchical Control of Emotion Rendering in Speech Synthesis Improve emotional speech synthesis quality by learning explicit and im- plicit representations with semi-supervised training

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.339091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.700600Z digest=sha256:3d1f43037b75b5d4a9d71eb944e02a53d959b612e8e6d3390d65cee2c6582853

Observation 18323115-3318-40cb-a572-c9f6c843ff6b · outbound

This paper cites Multi-speaker emotional speech synthesis with fine-grained prosody modeling,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Multi-speaker emotional speech synthesis with fine-grained prosody modeling,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.326666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.704111Z digest=sha256:daa0815e6b5e7c6877310d4c22988410acc1aa5153dd3aa371b4fc943dabe997

Observation 3451055b-47ef-40eb-a121-e6c3f51d4ca1 · outbound

This paper cites Language model-based emotion prediction methods for emotional speech synthesis systems,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Language model-based emotion prediction methods for emotional speech synthesis systems,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.313000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.707638Z digest=sha256:c06464bf37356fe024c7aa623807d7ed1f4f1d9b5d1969eba1bb4a72c63fd27c

Observation 12a6575f-239c-484a-8120-57e33d627da4 · outbound

This paper cites Fine-grained style modeling, transfer and prediction in text-to-speech synthesis via phone-level content-style disentanglement,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Fine-grained style modeling, transfer and prediction in text-to-speech synthesis via phone-level content-style disentanglement,

Reference 48

Resolution
malformed identifier
no resolver link, observed 2026-08-11T14:06:08.711157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.711157Z digest=sha256:fd72ec5a94cd6172030fc45911d4cf6a373ce37313d534892b7fb8946fcaad9e

Observation ff31eb97-8b64-437d-9424-ac8e06022f16 · outbound

This paper cites Hierarchical multi-grained generative model for ex- pressive speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Hierarchical multi-grained generative model for ex- pressive speech synthesis,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.300941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.714934Z digest=sha256:ae1e4dabde4f165088b80bbc18ddf9d36644177e9225912e9c282db80a0ac98b

Observation f030e8e2-866b-4215-9605-ade282db58e8 · outbound

This paper cites Towards multi-scale style control for expressive speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Towards multi-scale style control for expressive speech synthesis,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.289100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.718757Z digest=sha256:191013e0df4431cfaad9d5c25cfcaa89c0acb0cc1a7e6f11bbd6fdd7d18339e3

Observation b9119c76-aa80-461b-b5eb-8da500f6d1df · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.276828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.722320Z digest=sha256:6421960aa5e4ac37a8cddc2d1dc2e306a916057db65b0fe71091accbb030d2d2

Observation 415ea64f-bcf3-4a3d-8b6d-6ca1ffc4c713 · outbound

This paper cites An emotion speech synthesis method based on vits,.

Hierarchical Control of Emotion Rendering in Speech Synthesis An emotion speech synthesis method based on vits,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.265049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.725748Z digest=sha256:af88b87a346b9420b9da91aa8e16a826d3d4182b51a8295d09c9d4f55ddc60e3

Observation 0da5d858-9530-45f4-9673-0adb61c6a828 · outbound

This paper cites Generative adversarial networks,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Generative adversarial networks,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.729316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.729316Z digest=sha256:7ba1a8167de225ca6fbb2fc59e37dc360031107be173395c40dffdb29c67d02b

Observation 890f59db-3461-4d35-81dc-43fdaa98605c · outbound

This paper cites FastDiff 2: Revisiting and incorporating GANs and diffusion models in high-fidelity speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis FastDiff 2: Revisiting and incorporating GANs and diffusion models in high-fidelity speech synthesis,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.246376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.732675Z digest=sha256:f873e0ec88e4c6daf04924a0db9673011f34af61514189abbaef63af659ceb60

Observation 04b22edc-df59-45d9-9737-0b869f81c49e · outbound

This paper cites Bddm: Bilateral denoising diffusion models for fast and high-quality speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Bddm: Bilateral denoising diffusion models for fast and high-quality speech synthesis,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.234257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.736092Z digest=sha256:910a14ff0d896aea8f05162e3186100ae9e3635f52aa468730522bace27d0f62

Observation 96ff01ce-bd93-44f2-8a9a-31ac9b5725e2 · outbound

This paper cites Matcha-tts: A fast tts architecture with conditional flow matching,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Matcha-tts: A fast tts architecture with conditional flow matching,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.221999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.739604Z digest=sha256:53a487b644063b45cb64e5283dac847a2562af93652e34d594dc6dded1b407a2

Observation 91f90457-53fd-40e5-b56e-c505c5bab4e0 · outbound

This paper cites Flow matching for generative modeling,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Flow matching for generative modeling,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.743172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.743172Z digest=sha256:6ad76c3ea3b9c4c38a7cd1934ba33c4756b32bead9847cd0e639fb8649f4610e

Observation e82fc72d-6b5d-42c9-b002-1ce0533b03ba · outbound

This paper cites Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.190533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.752366Z digest=sha256:87f8b3d19a5ff8f8f4f25972517ba985e7ef05acddede404bbf0ba5640a12ec9

Observation e6c0991c-45fc-4f31-b163-f22bef0a4add · outbound

This paper cites Emotional voice conversion: Theory, databases and esd,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Emotional voice conversion: Theory, databases and esd,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.179120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.755963Z digest=sha256:7431a87949ddce7b3a487144215224888bfd547802089a8be6e92b4812c9a8f8

Observation d33e773c-9774-419f-8da0-05ba17a9d5b9 · outbound

This paper cites Attention is all you need,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Attention is all you need,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.759549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.759549Z digest=sha256:8e1f41c943688923ed02d6417edb02cca6baae76267b2489907ed64f4745454b

Observation 813f499e-a0c9-4543-b437-e51b04cbfddc · outbound

This paper cites Glow-tts: A generative flow for text-to-speech via monotonic alignment search,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Glow-tts: A generative flow for text-to-speech via monotonic alignment search,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.160130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.763267Z digest=sha256:91bdc46c50d5b227e77a9cd87c2428af1a60652e69e57bfc7ec07cfafc829e7b

Observation dba3efa7-c20e-480c-ba76-366bc30dde28 · outbound

This paper cites Generalized end-to-end loss for speaker verification,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Generalized end-to-end loss for speaker verification,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.202501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.766815Z digest=sha256:993fd46ac77afa85afafc429fb1e4e15c6618f2324644dda62753556e30fd25e

Observation 132e3ca8-96e5-4561-a8a9-7fec55ef931e · outbound

This paper cites V ocos: Closing the gap between time-domain and fourier- based neural vocoders for high-quality audio synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis V ocos: Closing the gap between time-domain and fourier- based neural vocoders for high-quality audio synthesis,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.148508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.770234Z digest=sha256:0eed0870833ead6c327aa722949ade2c0d17bf40eebba4c2fcbb1c2db28d4e26

Observation f9939591-14e3-46db-9ec5-f21518a35f09 · outbound

This paper cites Adam: A method for stochastic optimization,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Adam: A method for stochastic optimization,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.773936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.773936Z digest=sha256:44fccebab82137e826082d7799ef8ecc1700893140584decd0450a4d8eb26c87

Observation 753e61bd-483a-4493-8e97-ca4741da447b · outbound

This paper cites Unsupervised domain adaptation by backpropagation,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Unsupervised domain adaptation by backpropagation,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.777748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.777748Z digest=sha256:8aec119110a6c739fecd48b0ece29b114a0ecb8d643362a63b1f7f5be0b46ffb

Observation a3b2ab0a-2312-4a12-9903-23778bf20223 · outbound

This paper cites Adversarial domain generalized transformer for cross-corpus speech emotion recognition,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Adversarial domain generalized transformer for cross-corpus speech emotion recognition,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.123636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.781281Z digest=sha256:e3f99e32c19d764a03492ef4a68d2d2941ecdfec78bd01a687a1482425c6db59

Observation f68cf9c0-3cba-4c10-b9e7-08500f99dc13 · outbound

This paper cites Mel-cepstral distance measure for objective speech qual- ity assessment,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Mel-cepstral distance measure for objective speech qual- ity assessment,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.111777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.785019Z digest=sha256:a344cae2ee9f4d06c6cce27b8fa64edd53c63cdebb31b222669b81e536bea168

Observation ea67377c-8dd6-4316-b004-1b36882f95a1 · outbound

This paper cites 3-d convolutional recurrent neural networks with attention model for speech emotion recognition,.

Hierarchical Control of Emotion Rendering in Speech Synthesis 3-d convolutional recurrent neural networks with attention model for speech emotion recognition,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.099517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.788524Z digest=sha256:cbf0f15db78487fd0c1ba551e1866b0bc15ef61be3a2942544bdc1824fff07a9

Observation c64804da-90f8-4de1-acfa-c3e2a175ff24 · outbound

This paper cites Multi-conditioning and data augmentation using generative noise model for speech emotion recognition in noisy conditions,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Multi-conditioning and data augmentation using generative noise model for speech emotion recognition in noisy conditions,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.792294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.792294Z digest=sha256:e2a84ea08bc2b90ab7760d0bbd02a0432b32f84a5dc8f15789d67fb38d06bf65

Observation 27af8343-f844-40ba-8bd9-e22be958a71a · outbound

This paper cites Speech emotion recognition in noisy and reverberant environments,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Speech emotion recognition in noisy and reverberant environments,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.080298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.795971Z digest=sha256:90ff2da84472e879c8c25b2ca1013e0fc3eb51f215ee71368887970739c34f88

Observation 736ee137-ad4e-43ef-9ad9-991c1a8b1e36 · outbound

This paper cites Deep learning techniques for speech emotion recognition,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Deep learning techniques for speech emotion recognition,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.068317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.799554Z digest=sha256:c93595aefa3847a2d47517604eec5449d508f7e1349ae7f0aa6181250a7a74c8

Observation a122c003-a410-4541-8099-139f18cbf217 · outbound

This paper cites Improved emotion recognition using gaussian mixture model and extreme learning machine in speech and glottal signals,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Improved emotion recognition using gaussian mixture model and extreme learning machine in speech and glottal signals,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.055938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.803083Z digest=sha256:0e6ef57a3b1d2d0725553fc70edc40296bcceee7b2eda7cf5e0a5f1234cf6e5b

Observation 82bef523-36d6-4aa2-824a-553e97ed47a7 · outbound

This paper cites Specaugment: A simple data augmentation method for automatic speech recognition,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Specaugment: A simple data augmentation method for automatic speech recognition,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.806926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.806926Z digest=sha256:d35090400557e2fa160d7ecec25a029781f2b9acd81ed8e4c18889c155cdca77

Observation 75740fd4-f651-4a51-b9ef-0a835da7b6e3 · outbound

This paper cites Wespeaker: A research and production oriented speaker embedding learning toolkit,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Wespeaker: A research and production oriented speaker embedding learning toolkit,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.036690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.810492Z digest=sha256:7c1de7e7907b4a084789f6e15428b5ffa7168f5fbe4c4a13be8dabe7e3f9eeed

Observation a6217b18-5345-455d-99fe-21903b1531c3 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Hierarchical Control of Emotion Rendering in Speech Synthesis Robust Speech Recognition via Large-Scale Weak Supervision

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.814123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.814123Z digest=sha256:3a30b59d6e3e258f16849d578ebe9400c614f151bcdabe3255f5c7ba15e673ce

Observation 8fa8ec19-2b99-47ff-a705-448931d98a54 · outbound

This paper cites Best-worst scaling more reliable than rating scales: A case study on sentiment intensity annotation,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Best-worst scaling more reliable than rating scales: A case study on sentiment intensity annotation,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.024616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.818127Z digest=sha256:a8dc303665dd495ceb190175cc7bc711c31104bb6cfdaaad632d407f5eb385f6

Observation bdd7e523-0143-4a73-8f5d-990a8eaa6e35 · outbound

This paper cites Measuring disentanglement: A review of metrics,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Measuring disentanglement: A review of metrics,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.013442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.821706Z digest=sha256:c1c9e878039cebd1039896c05d00fe38b0bd493fdeefd193d6a25f9584a8a367

Observation e5f835b3-95da-444b-81f4-e42c7df6c66e · outbound

This paper cites Isolating sources of disentanglement in vaes,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Isolating sources of disentanglement in vaes,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.001787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.825361Z digest=sha256:56a7b996e3e3e81483ba8834fe9269ebb021d78a376389d470f2c0e84e3a0877

Observation a51c0188-1dce-41ab-9916-5432d91f5b93 · outbound

This paper cites A framework for the quantitative evaluation of disentangled representations,.

Hierarchical Control of Emotion Rendering in Speech Synthesis A framework for the quantitative evaluation of disentangled representations,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:08.989499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.829068Z digest=sha256:c282b4694da9406a085effe065ca44f49d54aff83d8da187bb1a8dd967afc633

Observation aae9dfeb-e6b1-4f37-8c91-5e959dccd5d8 · outbound

This paper cites Learning deep disentangled embeddings with the f-statistic loss,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Learning deep disentangled embeddings with the f-statistic loss,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:08.977717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.832368Z digest=sha256:a14ff0dc2a479f40bb1fcae2827515fdaf7c741846329c99cd9a6baa7ddbea78

Observation 2007afa7-cd4a-4a23-8ff2-adc7d7eea17f · outbound

This paper cites Speech emotion recognition: Two decades in a nutshell, benchmarks, and ongoing trends,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Speech emotion recognition: Two decades in a nutshell, benchmarks, and ongoing trends,

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.835917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.835917Z digest=sha256:08406569a041535ce03d37c37cfda91be0be8fa83b0c3bf5cedc861636ea6432

Observation 7058b55d-b386-4746-87bc-7f3790192bbd · outbound

This paper cites Dawn of the transformer era in speech emotion recognition: Closing the valence gap,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Dawn of the transformer era in speech emotion recognition: Closing the valence gap,

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.839539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.839539Z digest=sha256:9479b7093e338eb9eeb84a9ef6e11f390dfe827b14b894be21efa80e3625ea19

Observation 77a008ac-5af3-470c-9ab4-4fc7747fd84a · outbound

This paper cites Contrastive learn- ing based modality-invariant feature acquisition for robust multimodal emotion recognition with missing modalities,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Contrastive learn- ing based modality-invariant feature acquisition for robust multimodal emotion recognition with missing modalities,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:08.951996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.843002Z digest=sha256:48ead69fa2e95cab4079eac725a0284fe13bd7b66b621abc1a955f9d43223727

Observation ebcb22a9-6a66-4103-b963-4c4c2beeed13 · outbound

This paper cites Available: https://api.semanticscholar.org/CorpusID: 252762216.

Hierarchical Control of Emotion Rendering in Speech Synthesis Available: https://api.semanticscholar.org/CorpusID: 252762216

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.759306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.550085Z digest=sha256:62dcfc711c8db5c359c52481a7845d35af25f0bba0505f69aa66465cd85597f5

Observation de2b0f40-376a-4d8e-8abf-e3ef02beb990 · outbound

This paper cites Fine-Grained Quantitative Emotion Editing for Speech Generation.

Hierarchical Control of Emotion Rendering in Speech Synthesis Fine-Grained Quantitative Emotion Editing for Speech Generation

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T14:06:08.929237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T14:06:08.652070Z digest=sha256:b6f1873964ed680a1dcc81ad80ffc5d3487cbe0fa863ffd762c09f3accf80a76

Pith citing papers

Observation 4473eeaa-47f2-4a02-bada-4098ec1aec00 · inbound

Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions cites this paper.

Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions Hierarchical Control of Emotion Rendering in Speech Synthesis

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:21:12.319458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:21:09.698405Z digest=sha256:311652a0b473bb3a0152f60112f4fe0c0fe3235281176dbcfea68f3fec95c1fe