Pith. sign in

Paper Citation Record · LEDGER

Hierarchical Control of Emotion Rendering in Speech Synthesis

As of 12 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 2 inbound Pith citation observations for arXiv:2412.12498.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12498 v3

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:06:08.843002Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:32:24.055265Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T11:21:12.229573Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact2
  • verified fuzzy62
  • unresolved19
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 08d3e8ac-5307-4189-8835-5b7965c3d74e · outbound

This paper cites An overview of affective speech synthesis and conversion in the deep learning era,.

Hierarchical Control of Emotion Rendering in Speech Synthesis An overview of affective speech synthesis and conversion in the deep learning era,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.529362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.529362Z digest=sha256:51477ed09f81f1330ac3d8994960db5bd774ef39ecf60e32245b9633b80623e6

Observation 7ea42af2-61d1-49a6-a472-6ccc45c4f23b · outbound

This paper cites The age of artificial emotional intelligence,.

Hierarchical Control of Emotion Rendering in Speech Synthesis The age of artificial emotional intelligence,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.534138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.534138Z digest=sha256:517bd31ddcf908c64258027f3e2b840ee64beb264ea283aa650f6985acd9ccb8

Observation f5a65bdb-402c-4b67-8377-f3a677b3e779 · outbound

This paper cites Pittermann, A.

Hierarchical Control of Emotion Rendering in Speech Synthesis Pittermann, A

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.537992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.537992Z digest=sha256:3cf4f14f4a7fbe3103f0e85658a43765d0a157d67aa345f893b8d380c84a4b3a

Observation de745932-fa9b-48d9-8ed3-98cb491f9fe0 · outbound

This paper cites Emotion modelling for speech generation,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Emotion modelling for speech generation,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.542142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.542142Z digest=sha256:8a156c769b21261a2def50fe958aa94268f529a932f1e70c882b61595688f45a

Observation d8e9decb-2849-4676-b019-2d95944d4b8d · outbound

This paper cites An overview of affective speech synthesis and conversion in the deep learning era,.

Hierarchical Control of Emotion Rendering in Speech Synthesis An overview of affective speech synthesis and conversion in the deep learning era,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.545989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.545989Z digest=sha256:1765c58a49bb6e156e9ac4466b8d047a2f674eb5981e095d5a779654794cbb6d

Observation bbb3fce4-6774-4428-86de-191eef892933 · outbound

This paper cites Fastspeech 2: Fast and high-quality end-to-end text to speech,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Fastspeech 2: Fast and high-quality end-to-end text to speech,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.554067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.554067Z digest=sha256:ad7e6d272650520864c3f3a456f442280598cebc10de8d4d3b9b50a89ace7407

Observation 8566ab21-8c13-4f2a-8c06-cf4144af8221 · outbound

This paper cites Phonetic enhanced language modeling for text-to-speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Phonetic enhanced language modeling for text-to-speech synthesis,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.740594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.557786Z digest=sha256:843ac51ff6ad599c54e1115a27b0c1e8fafe45d35de5187306872fd162e15bad

Observation a7099819-e79b-4cad-a1d3-31fbe54e0631 · outbound

This paper cites Text-to-speech for low-resource agglutinative language with morphology-aware language model pre-training,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Text-to-speech for low-resource agglutinative language with morphology-aware language model pre-training,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.729060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.561903Z digest=sha256:b485ca60443e8fbe5db7d6cea56177d14456dc0a3866779180d28cf6d2e94a21

Observation 2bb81916-c203-42e1-9bbb-fdd053892d84 · outbound

This paper cites Emotional speech synthesis: A review,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Emotional speech synthesis: A review,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.718028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.565933Z digest=sha256:55d07847c0a369ca6a7d26a4a1c1cbdafcf2747cd7f1d250e4fd596a8465e76e

Observation 1251d4f1-acd2-460b-87af-466245531d99 · outbound

This paper cites A Survey on Neural Speech Synthesis.

Hierarchical Control of Emotion Rendering in Speech Synthesis A Survey on Neural Speech Synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.569719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.569719Z digest=sha256:f4d887a15e130d6c82fe8550f9e25b4ed21e2abb36e4ff6004b42853aa863c35

Observation d8e301cf-7f52-4092-9ec3-5cf01d6801e8 · outbound

This paper cites Pragmatics and intonation,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Pragmatics and intonation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.705759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.574210Z digest=sha256:53c9543de5a380910a2a55e855014b0a4c8be906cbb8c828016ca5c130d3167b

Observation 72648521-38b0-47c5-afc8-d4598095268d · outbound

This paper cites Perception of affective and linguistic prosody: an ale meta-analysis of neuroimaging studies,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Perception of affective and linguistic prosody: an ale meta-analysis of neuroimaging studies,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.693377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.577953Z digest=sha256:71d47a92953de47849a380fd1da57b82ec70f620bffa3208d912c9633f023a5a

Observation 0940779d-18ad-40c5-9d9c-a35dededc153 · outbound

This paper cites Laukka,Vocal Communication of Emotion.

Hierarchical Control of Emotion Rendering in Speech Synthesis Laukka,Vocal Communication of Emotion

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.581669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.581669Z digest=sha256:76b6dbab4d502f87c5d8f5353f5361fb4275d7ac6ca81e54fba898918c9356e1

Observation 553e88ce-f7e0-4a08-8928-aaea88dd6c48 · outbound

This paper cites Vaw-gan for disentanglement and recomposition of emotional elements in speech,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Vaw-gan for disentanglement and recomposition of emotional elements in speech,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.681139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.585131Z digest=sha256:3f1af7e76b2ebb1122f40234368e0e6095f23ce67a76f4a35ab9195e73027978

Observation 86230675-45ef-4fb5-aef9-c97c60873202 · outbound

This paper cites Converting anyone’s emotion: Towards speaker-independent emotional voice conversion,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Converting anyone’s emotion: Towards speaker-independent emotional voice conversion,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.669949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.588437Z digest=sha256:7c6109297c1f648356abb6403cc2eb0e1f039fe7e8590496909f32100e47c1a5

Observation 5cf6d6f2-b8ca-4c5c-b2c8-ae8e739f3d18 · outbound

This paper cites opensmile – the munich versatile and fast open-source audio feature extractor,.

Hierarchical Control of Emotion Rendering in Speech Synthesis opensmile – the munich versatile and fast open-source audio feature extractor,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.658229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.591916Z digest=sha256:654f3fd5cc2af3ce3b2d67c7d822ef8651ac63be61da0b3297a98afbb747aa36

Observation af631296-ce47-423f-8982-1fb9d22ec81a · outbound

This paper cites emotion2vec: Self-supervised pre-training for speech emotion repre- sentation,.

Hierarchical Control of Emotion Rendering in Speech Synthesis emotion2vec: Self-supervised pre-training for speech emotion repre- sentation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.647402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.595316Z digest=sha256:4c57933e06af2fdd6a41dccd8b2e30b9d24e5c4f1212a41a59156aaef0cb5003

Observation f375cf66-d0ae-4c25-bd10-e11b0e70e90c · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Wavlm: Large-scale self-supervised pre-training for full stack speech processing,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.636109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.598849Z digest=sha256:207426368dc987b5ce3b566dcfdb35c688885c6a565c72824639d384da303ab0

Observation 6b3de789-6f1f-46f8-9513-141b08bd1b1d · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Hubert: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.624187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.602516Z digest=sha256:11c43f973098f2bccba0b5b3fa980d41e1843583bcd7ed9c1de7344e6d9ae857

Observation 51bc479e-69e0-4c27-9e2d-5146282bcce9 · outbound

This paper cites Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.611419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.606036Z digest=sha256:db78f8f4bb01f8ce4cfc5a8945f39163e6eba39c3ca060cb1f32b99ae274b698

Observation b4b56386-85af-455d-b678-41e89379a2d3 · outbound

This paper cites End-to-end emotional speech synthesis using style tokens and semi-supervised training,.

Hierarchical Control of Emotion Rendering in Speech Synthesis End-to-end emotional speech synthesis using style tokens and semi-supervised training,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.599773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.609465Z digest=sha256:cb7e94124bc979e23dae6bcaa8ec7532c6343fd0d660ebbd0ca088d3f45da351

Observation 9e75c06d-5f43-4436-9fad-9df94ae907c3 · outbound

This paper cites Controllable emotion transfer for end-to-end speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Controllable emotion transfer for end-to-end speech synthesis,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.587506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.613031Z digest=sha256:d0548f0a51d959944f73347ad966f10ac1c436002ea6caff204e4e681b3db418

Observation fbf80c48-966e-4927-9fe7-771385373f4e · outbound

This paper cites Semi-supervised learning for contin- uous emotional intensity controllable speech synthesis with disentangled representations,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Semi-supervised learning for contin- uous emotional intensity controllable speech synthesis with disentangled representations,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.576093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.616553Z digest=sha256:892def37fc991ef95d3ce593e204ac58f8ba9a0b0205971b4a1cc68ee05a09cb

Observation 58483e6a-fdcb-424b-8ea5-20ef5f380240 · outbound

This paper cites iemotts: Toward robust cross-speaker emotion transfer and control for speech synthesis based on disentanglement between prosody and timbre,.

Hierarchical Control of Emotion Rendering in Speech Synthesis iemotts: Toward robust cross-speaker emotion transfer and control for speech synthesis based on disentanglement between prosody and timbre,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.563496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.620007Z digest=sha256:41c03cfc32a82b5924877ef0e26a6558d4b1f5b455be389d85537c57077abc9f

Observation df59d7bb-c414-461c-930c-afd84cc6d1fb · outbound

This paper cites Cross-speaker emotion disentangling and transfer for end-to-end speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Cross-speaker emotion disentangling and transfer for end-to-end speech synthesis,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.551617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.623521Z digest=sha256:d6ab8b8e1140dffa0297a803692006128c5b1401782cef870cb341b9fb37a796

Observation 1bd9a0a4-7126-4814-98ee-cc77c2016b49 · outbound

This paper cites Relative attributes,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Relative attributes,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.540602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.626913Z digest=sha256:30af0f68a0695f967e5f9c877a0b5699e713f6a6342b97f2ebbd112f0ce96b50

Observation e53eaca7-7203-4ded-ae55-c3c69f3a4bfe · outbound

This paper cites Emotion inten- sity and its control for emotional voice conversion,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Emotion inten- sity and its control for emotional voice conversion,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.529277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.630492Z digest=sha256:846de6c37abfc1add54e474f50c5f757a911d6aef276cf70f32718f9edaa96d3

Observation 007f184e-c906-4cc2-8aea-4c71f99e0347 · outbound

This paper cites Controlling emotion strength with relative attribute for end-to-end speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Controlling emotion strength with relative attribute for end-to-end speech synthesis,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.517716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.633927Z digest=sha256:8d7b33686fc1c0e22e6fdc4b30ce9e792addac79d9938af02c485f5891fbab4e

Observation 62d61b50-b565-476d-9b3f-ee3540482272 · outbound

This paper cites Speech synthe- sis with mixed emotions,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Speech synthe- sis with mixed emotions,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.505413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.637550Z digest=sha256:d04c81200b296106edb1ff31972ec9c257e54653a52214fe71d94c49ca5e8fd3

Observation a7917f73-2356-4f6a-951d-d65bd4af3bd7 · outbound

This paper cites Msemotts: Multi-scale emotion transfer, prediction, and control for emotional speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Msemotts: Multi-scale emotion transfer, prediction, and control for emotional speech synthesis,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.494097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.641090Z digest=sha256:cf9269f12d75d69cc7a0e042e46314f44937256ea68ad98ebf49782ad7289eb6

Observation bd28a430-d85f-4c44-a79f-783e1804141a · outbound

This paper cites Hierarchical emotion prediction and control in text-to-speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Hierarchical emotion prediction and control in text-to-speech synthesis,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.482373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.644651Z digest=sha256:c5f5baec7c6f7d3da4a7d3d9a980fcf80bd6ffdfeb45af2e3386fe5b503c3b16

Observation 29abbc49-652d-41bf-8193-ac8a04c500ec · outbound

This paper cites Fine-grained quantitative emotion editing for speech generation,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Fine-grained quantitative emotion editing for speech generation,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.470470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.648316Z digest=sha256:62ada28f772813d92cd687a573dc0249e5a34e6a7e5dc211e001cd4d4cc2ec23

Observation c6ec9f19-999d-4bc1-a840-16411c1d93cf · outbound

This paper cites an unresolved cited work.

Hierarchical Control of Emotion Rendering in Speech Synthesis Unresolved cited work

Reference 33

Resolution
verified exact
doi, observed 2026-08-11T14:06:08.897451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.656176Z digest=sha256:f3345db35d69dd59f060fd0061ea0579c404af7a340f4d5a19675bd77c6c7efd

Observation 49a28819-f679-415d-a143-4435751171f9 · outbound

This paper cites Transforming spectrum and prosody for emotional voice conversion with non-parallel training data,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Transforming spectrum and prosody for emotional voice conversion with non-parallel training data,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.458206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.660221Z digest=sha256:ac77cc327551874cf37efec525e3e6b75604d3d7036852c726d751dd10d9cdfa

Observation 4e76db84-c4fc-481c-bb83-acf416247de6 · outbound

This paper cites Intonation and emotion: Influence of pitch levels and contour type on creating emotions,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Intonation and emotion: Influence of pitch levels and contour type on creating emotions,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.446239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.664011Z digest=sha256:6ace4e593a457dc7e99fe6b263c2cf0bf293286a2038edcae3388a974c898eba

Observation d171ff62-628a-42d7-a24d-56491c56f374 · outbound

This paper cites Norms of valence, arousal, and dominance for 13,915 english lemmas,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Norms of valence, arousal, and dominance for 13,915 english lemmas,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.434523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.667486Z digest=sha256:19a23f69360a2025410608a9daf44087daef691c86c25e55cd6cc8d3dbef8c4f

Observation 7a113808-64e3-4cfd-94b5-99c826fbe3c7 · outbound

This paper cites The influence of pitch range, duration, amplitude and spectral features on the interpretation of the rise-fall-rise intonation contour in english,.

Hierarchical Control of Emotion Rendering in Speech Synthesis The influence of pitch range, duration, amplitude and spectral features on the interpretation of the rise-fall-rise intonation contour in english,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.422051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.671081Z digest=sha256:16855435a1268a205837d08ad8a39d28b011c8ecbd5c2b93af86820202bd99f9

Observation 74cbc6ab-b04a-4be1-8957-d538deadb96b · outbound

This paper cites Using prosody to avoid ambiguity: Effects of speaker awareness and referential context,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Using prosody to avoid ambiguity: Effects of speaker awareness and referential context,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.408439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.674691Z digest=sha256:ae5ea592bf1a88cd808d0535b6d80a91d47687695965b7423506deee55324b54

Observation 846906b9-1916-4720-af85-33c619be4bbb · outbound

This paper cites Analysis of emotionally salient aspects of fundamental frequency for emotion detection,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Analysis of emotionally salient aspects of fundamental frequency for emotion detection,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.396219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.678149Z digest=sha256:30c552da48288e56f4a9a180cf3b1aa0dfec24b2cebfa105d57ff56b861dfd95

Observation 6f37876d-3586-4f80-821e-57263b1036d0 · outbound

This paper cites Survey on speech emotion recognition: Features, classification schemes, and databases,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Survey on speech emotion recognition: Features, classification schemes, and databases,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.384194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.681733Z digest=sha256:9f567341b0650994b1be642252e3fdebc3a33e383f9a8ea1ef6e954079ad542d

Observation 507fc731-cad1-408c-bd43-699b8498c694 · outbound

This paper cites Expression of emotional–motivational connotations with a one-word utterance,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Expression of emotional–motivational connotations with a one-word utterance,

Reference 41

Resolution
verified exact
doi, observed 2026-08-11T14:06:08.885442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.685558Z digest=sha256:0ca613eb529bb237a2e93d68c6aef474c011d153fdb2c5ab842bdf198923fb02

Observation 04ad7066-1700-451b-93f0-d389f3b42c8b · outbound

This paper cites Controllable accented text- to-speech synthesis with fine and coarse-grained intensity rendering,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Controllable accented text- to-speech synthesis with fine and coarse-grained intensity rendering,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.372004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.689285Z digest=sha256:5eb20a6a3f0394fadbf459b016ab02920a61e0e3a88c4c89d352e6cc924114c2

Observation 3a71807e-85b6-4a64-acb3-087cbf41ba31 · outbound

This paper cites Connecting cross-modal rep- resentations for compact and robust multimodal sentiment analysis with sentiment word substitution error,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Connecting cross-modal rep- resentations for compact and robust multimodal sentiment analysis with sentiment word substitution error,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.359568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.693128Z digest=sha256:91a85764a42056a719405569040b46013dbcdc3a96e0ba9363caf4c92926483c

Observation 7adf3645-7708-4a2d-9e70-c517a3ce8768 · outbound

This paper cites Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Style tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.697089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.697089Z digest=sha256:a886a0789b1e2f70b682a6b8745b751048d1e92f4e50119788a7ffd1db53c6b7

Observation f055f3e1-1a1f-4e76-8ad5-c13e09ec8813 · outbound

This paper cites Improve emotional speech synthesis quality by learning explicit and im- plicit representations with semi-supervised training.

Hierarchical Control of Emotion Rendering in Speech Synthesis Improve emotional speech synthesis quality by learning explicit and im- plicit representations with semi-supervised training

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.339091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.700600Z digest=sha256:3d395afb530578310263dd65a926bb951fb54220e42021ae52f24cc9bc18317e

Observation 18323115-3318-40cb-a572-c9f6c843ff6b · outbound

This paper cites Multi-speaker emotional speech synthesis with fine-grained prosody modeling,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Multi-speaker emotional speech synthesis with fine-grained prosody modeling,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.326666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.704111Z digest=sha256:6249290844946067791e08974d0c491c8c774353fb813ba29e88c2fa1d33ed1e

Observation 3451055b-47ef-40eb-a121-e6c3f51d4ca1 · outbound

This paper cites Language model-based emotion prediction methods for emotional speech synthesis systems,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Language model-based emotion prediction methods for emotional speech synthesis systems,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.313000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.707638Z digest=sha256:d4e38c1a28c08daa683a5db2a94f9d76a0770f4c8e2354d99a61220057d89bca

Observation 12a6575f-239c-484a-8120-57e33d627da4 · outbound

This paper cites Fine-grained style modeling, transfer and prediction in text-to-speech synthesis via phone-level content-style disentanglement,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Fine-grained style modeling, transfer and prediction in text-to-speech synthesis via phone-level content-style disentanglement,

Reference 48

Resolution
malformed identifier
no resolver link, observed 2026-08-11T14:06:08.711157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.711157Z digest=sha256:fd72ec5a94cd6172030fc45911d4cf6a373ce37313d534892b7fb8946fcaad9e

Observation ff31eb97-8b64-437d-9424-ac8e06022f16 · outbound

This paper cites Hierarchical multi-grained generative model for ex- pressive speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Hierarchical multi-grained generative model for ex- pressive speech synthesis,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.300941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.714934Z digest=sha256:b100ad02d57fac1ae912f014d26c7a044d213e3967900243211e283114c8080a

Observation f030e8e2-866b-4215-9605-ade282db58e8 · outbound

This paper cites Towards multi-scale style control for expressive speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Towards multi-scale style control for expressive speech synthesis,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.289100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.718757Z digest=sha256:0bf766109f886ba2a0e11f6a6e516d02f14ec21a097152899135d285920e2abc

Observation b9119c76-aa80-461b-b5eb-8da500f6d1df · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.276828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.722320Z digest=sha256:8b194c1bf0affc4686b7630786e5c4fce76fd137ec7f19986af6f4c010ad6c57

Observation 415ea64f-bcf3-4a3d-8b6d-6ca1ffc4c713 · outbound

This paper cites An emotion speech synthesis method based on vits,.

Hierarchical Control of Emotion Rendering in Speech Synthesis An emotion speech synthesis method based on vits,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.265049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.725748Z digest=sha256:8d543c1e6349ca86b662ce42336a758d8144e1153cd0591ce668c5d18f170991

Observation 0da5d858-9530-45f4-9673-0adb61c6a828 · outbound

This paper cites Generative adversarial networks,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Generative adversarial networks,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.729316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.729316Z digest=sha256:7ba1a8167de225ca6fbb2fc59e37dc360031107be173395c40dffdb29c67d02b

Observation 890f59db-3461-4d35-81dc-43fdaa98605c · outbound

This paper cites FastDiff 2: Revisiting and incorporating GANs and diffusion models in high-fidelity speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis FastDiff 2: Revisiting and incorporating GANs and diffusion models in high-fidelity speech synthesis,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.246376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.732675Z digest=sha256:d66267ad9cb03db425730129da918f4534d879ba0b5e79609382a112a3e0e998

Observation 04b22edc-df59-45d9-9737-0b869f81c49e · outbound

This paper cites Bddm: Bilateral denoising diffusion models for fast and high-quality speech synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Bddm: Bilateral denoising diffusion models for fast and high-quality speech synthesis,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.234257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.736092Z digest=sha256:0f69533eff10ae4fe63c933ed23fd96782901fdbe2d02b7ccdb1ac82305f9cd0

Observation 96ff01ce-bd93-44f2-8a9a-31ac9b5725e2 · outbound

This paper cites Matcha-tts: A fast tts architecture with conditional flow matching,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Matcha-tts: A fast tts architecture with conditional flow matching,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.221999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.739604Z digest=sha256:45dd2976bd1292a2a996e4c0ef7dff99c5a8c8933326f03b03ee170868b3ff6a

Observation 91f90457-53fd-40e5-b56e-c505c5bab4e0 · outbound

This paper cites Flow matching for generative modeling,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Flow matching for generative modeling,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.743172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.743172Z digest=sha256:6ad76c3ea3b9c4c38a7cd1934ba33c4756b32bead9847cd0e639fb8649f4610e

Observation e82fc72d-6b5d-42c9-b002-1ce0533b03ba · outbound

This paper cites Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.190533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.752366Z digest=sha256:11ece97360a4fd87ed55324d29b95b44a08c1e6596148f9839caf807ea5099d9

Observation e6c0991c-45fc-4f31-b163-f22bef0a4add · outbound

This paper cites Emotional voice conversion: Theory, databases and esd,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Emotional voice conversion: Theory, databases and esd,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.179120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.755963Z digest=sha256:e62b5acb129f7b2b011bac501a2b5f67c4508dc8810ed965ea177126d6f222d5

Observation d33e773c-9774-419f-8da0-05ba17a9d5b9 · outbound

This paper cites Attention is all you need,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Attention is all you need,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.759549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.759549Z digest=sha256:8e1f41c943688923ed02d6417edb02cca6baae76267b2489907ed64f4745454b

Observation 813f499e-a0c9-4543-b437-e51b04cbfddc · outbound

This paper cites Glow-tts: A generative flow for text-to-speech via monotonic alignment search,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Glow-tts: A generative flow for text-to-speech via monotonic alignment search,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.160130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.763267Z digest=sha256:cf7ebdccd5c785d0c4ce95af4ce9817df96dda53e9dff9963e123b5822e992d3

Observation dba3efa7-c20e-480c-ba76-366bc30dde28 · outbound

This paper cites Generalized end-to-end loss for speaker verification,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Generalized end-to-end loss for speaker verification,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.202501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.766815Z digest=sha256:91e82d61de1aaa3d96bc692c4b7c38c62ecfbd249552aa68a00f85a24ea83e45

Observation 132e3ca8-96e5-4561-a8a9-7fec55ef931e · outbound

This paper cites V ocos: Closing the gap between time-domain and fourier- based neural vocoders for high-quality audio synthesis,.

Hierarchical Control of Emotion Rendering in Speech Synthesis V ocos: Closing the gap between time-domain and fourier- based neural vocoders for high-quality audio synthesis,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.148508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.770234Z digest=sha256:8178ff029e3beaf4910d3c85d207d890d3b0fef3afd2f14047a5a1c5de3fd246

Observation f9939591-14e3-46db-9ec5-f21518a35f09 · outbound

This paper cites Adam: A method for stochastic optimization,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Adam: A method for stochastic optimization,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.773936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.773936Z digest=sha256:44fccebab82137e826082d7799ef8ecc1700893140584decd0450a4d8eb26c87

Observation 753e61bd-483a-4493-8e97-ca4741da447b · outbound

This paper cites Unsupervised domain adaptation by backpropagation,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Unsupervised domain adaptation by backpropagation,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.777748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.777748Z digest=sha256:8aec119110a6c739fecd48b0ece29b114a0ecb8d643362a63b1f7f5be0b46ffb

Observation a3b2ab0a-2312-4a12-9903-23778bf20223 · outbound

This paper cites Adversarial domain generalized transformer for cross-corpus speech emotion recognition,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Adversarial domain generalized transformer for cross-corpus speech emotion recognition,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.123636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.781281Z digest=sha256:9b1423163ddea8c3709260517a8d91ba929759cfb1589f1038dc4d1eac12b4d7

Observation f68cf9c0-3cba-4c10-b9e7-08500f99dc13 · outbound

This paper cites Mel-cepstral distance measure for objective speech qual- ity assessment,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Mel-cepstral distance measure for objective speech qual- ity assessment,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.111777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.785019Z digest=sha256:35b9ee50d44779d32c10fbde906f8fec17ba524f279196ee617ad80cb3b8dfb6

Observation ea67377c-8dd6-4316-b004-1b36882f95a1 · outbound

This paper cites 3-d convolutional recurrent neural networks with attention model for speech emotion recognition,.

Hierarchical Control of Emotion Rendering in Speech Synthesis 3-d convolutional recurrent neural networks with attention model for speech emotion recognition,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.099517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.788524Z digest=sha256:0de0a3cb29ccc55b5f2bb65d3911ec1615baa5fdc4112cf382ffabb79cfe8b93

Observation c64804da-90f8-4de1-acfa-c3e2a175ff24 · outbound

This paper cites Multi-conditioning and data augmentation using generative noise model for speech emotion recognition in noisy conditions,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Multi-conditioning and data augmentation using generative noise model for speech emotion recognition in noisy conditions,

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.792294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.792294Z digest=sha256:e2a84ea08bc2b90ab7760d0bbd02a0432b32f84a5dc8f15789d67fb38d06bf65

Observation 27af8343-f844-40ba-8bd9-e22be958a71a · outbound

This paper cites Speech emotion recognition in noisy and reverberant environments,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Speech emotion recognition in noisy and reverberant environments,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.080298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.795971Z digest=sha256:3fd5a7e5d4c486339abee56426cec541b4a559511735bcb9df6811811dd55382

Observation 736ee137-ad4e-43ef-9ad9-991c1a8b1e36 · outbound

This paper cites Deep learning techniques for speech emotion recognition,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Deep learning techniques for speech emotion recognition,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.068317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.799554Z digest=sha256:f09a82237b01e645ad7ffc167f5d1d7ab17187007c2e7984b7462cf818024a79

Observation a122c003-a410-4541-8099-139f18cbf217 · outbound

This paper cites Improved emotion recognition using gaussian mixture model and extreme learning machine in speech and glottal signals,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Improved emotion recognition using gaussian mixture model and extreme learning machine in speech and glottal signals,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.055938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.803083Z digest=sha256:f62faa9a8d6191b6524626c576978ef48e8d1fa7991208e2eadb4108b450cc11

Observation 82bef523-36d6-4aa2-824a-553e97ed47a7 · outbound

This paper cites Specaugment: A simple data augmentation method for automatic speech recognition,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Specaugment: A simple data augmentation method for automatic speech recognition,

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.806926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.806926Z digest=sha256:d35090400557e2fa160d7ecec25a029781f2b9acd81ed8e4c18889c155cdca77

Observation 75740fd4-f651-4a51-b9ef-0a835da7b6e3 · outbound

This paper cites Wespeaker: A research and production oriented speaker embedding learning toolkit,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Wespeaker: A research and production oriented speaker embedding learning toolkit,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.036690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.810492Z digest=sha256:2db7226ba4cc3c2917b397ebcc69850a64b8be62b8a0dc5425c9f0ee5973509a

Observation a6217b18-5345-455d-99fe-21903b1531c3 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Hierarchical Control of Emotion Rendering in Speech Synthesis Robust Speech Recognition via Large-Scale Weak Supervision

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.814123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.814123Z digest=sha256:3a30b59d6e3e258f16849d578ebe9400c614f151bcdabe3255f5c7ba15e673ce

Observation 8fa8ec19-2b99-47ff-a705-448931d98a54 · outbound

This paper cites Best-worst scaling more reliable than rating scales: A case study on sentiment intensity annotation,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Best-worst scaling more reliable than rating scales: A case study on sentiment intensity annotation,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.024616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.818127Z digest=sha256:f958603d25ecfdd4869ffa400105d45b0bdc2d735ae1098fc9935a221c8034eb

Observation bdd7e523-0143-4a73-8f5d-990a8eaa6e35 · outbound

This paper cites Measuring disentanglement: A review of metrics,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Measuring disentanglement: A review of metrics,

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.013442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.821706Z digest=sha256:3448e5890ffa1321c6887d7692abe4f71f901b8184b87a53a270b6e7036b35ee

Observation e5f835b3-95da-444b-81f4-e42c7df6c66e · outbound

This paper cites Isolating sources of disentanglement in vaes,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Isolating sources of disentanglement in vaes,

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.001787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.825361Z digest=sha256:9f85b850c7126a6328d388db858fff10597d9358bb3ce479efe90327e003408b

Observation a51c0188-1dce-41ab-9916-5432d91f5b93 · outbound

This paper cites A framework for the quantitative evaluation of disentangled representations,.

Hierarchical Control of Emotion Rendering in Speech Synthesis A framework for the quantitative evaluation of disentangled representations,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:08.989499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.829068Z digest=sha256:84c273ef3c723e21535a8902570a5bb9f8ac19dbee74e41161ca29edbfba64ca

Observation aae9dfeb-e6b1-4f37-8c91-5e959dccd5d8 · outbound

This paper cites Learning deep disentangled embeddings with the f-statistic loss,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Learning deep disentangled embeddings with the f-statistic loss,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:08.977717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.832368Z digest=sha256:ff7210532166a6168c7329b6f86008f8b8ebbfcdd10a91a72e6be24eb184587a

Observation 2007afa7-cd4a-4a23-8ff2-adc7d7eea17f · outbound

This paper cites Speech emotion recognition: Two decades in a nutshell, benchmarks, and ongoing trends,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Speech emotion recognition: Two decades in a nutshell, benchmarks, and ongoing trends,

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.835917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.835917Z digest=sha256:08406569a041535ce03d37c37cfda91be0be8fa83b0c3bf5cedc861636ea6432

Observation 7058b55d-b386-4746-87bc-7f3790192bbd · outbound

This paper cites Dawn of the transformer era in speech emotion recognition: Closing the valence gap,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Dawn of the transformer era in speech emotion recognition: Closing the valence gap,

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T14:06:08.839539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:06:08.839539Z digest=sha256:9479b7093e338eb9eeb84a9ef6e11f390dfe827b14b894be21efa80e3625ea19

Observation 77a008ac-5af3-470c-9ab4-4fc7747fd84a · outbound

This paper cites Contrastive learn- ing based modality-invariant feature acquisition for robust multimodal emotion recognition with missing modalities,.

Hierarchical Control of Emotion Rendering in Speech Synthesis Contrastive learn- ing based modality-invariant feature acquisition for robust multimodal emotion recognition with missing modalities,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:08.951996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.843002Z digest=sha256:d4b6ece746c64c51a1640e11c27ff334b8ed86c7283e3e119dbd623f29c28805

Observation ebcb22a9-6a66-4103-b963-4c4c2beeed13 · outbound

This paper cites Available: https://api.semanticscholar.org/CorpusID: 252762216.

Hierarchical Control of Emotion Rendering in Speech Synthesis Available: https://api.semanticscholar.org/CorpusID: 252762216

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:06:09.759306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.550085Z digest=sha256:e5943bd9d66964ff5761d331da39ea8efe3bfbcb6dd98cbcac852862dd59c20c

Observation de2b0f40-376a-4d8e-8abf-e3ef02beb990 · outbound

This paper cites Fine-Grained Quantitative Emotion Editing for Speech Generation.

Hierarchical Control of Emotion Rendering in Speech Synthesis Fine-Grained Quantitative Emotion Editing for Speech Generation

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T14:06:08.929237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T14:06:08.652070Z digest=sha256:6047b54e859cbced22c6ade9adba5730997705c3cd84085ebb7e76b30ba2f591

Pith citing papers

Observation 99ae3806-5571-4a12-a529-2c9c05b018ec · inbound

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey cites this paper.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey Hierarchical Control of Emotion Rendering in Speech Synthesis

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.055265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.055265Z digest=sha256:4bb32d0c930c7d19534f6e3e96bf27cb6e086355acfb6a38468d1afb5eba5559

Observation 4473eeaa-47f2-4a02-bada-4098ec1aec00 · inbound

Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions cites this paper.

Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions Hierarchical Control of Emotion Rendering in Speech Synthesis

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:21:12.319458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:21:09.698405Z digest=sha256:5c24905805d8f340b64357e66daa0312653ddf099731eaddcd2820cb47ee003c