Pith. sign in

Paper Citation Record · LEDGER

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet

As of 16 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 2 inbound Pith citation observations for arXiv:2507.04349.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04349 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:55:03.611900Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T15:21:28.820442Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T09:11:27.121428Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact4
  • verified fuzzy17
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f5764825-324b-4fac-a39b-750df4887754 · outbound

This paper cites Adding Conditional Control to Text-to-Image Diffusion Models.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Adding Conditional Control to Text-to-Image Diffusion Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:02.241628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:02.241628Z digest=sha256:e266e2fd65684ee9732419d9d3043c6fd4ed807d7666cfc1ce57e75de8d395cd

Observation e546cec3-0c0f-4291-8e58-d3284f1384f9 · outbound

This paper cites An emotion speech synthesis method based on VITS.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet An emotion speech synthesis method based on VITS

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.445789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:55:02.318699Z digest=sha256:82e4e749cac102da76548fa104ad0c233b8f6b6795fa0e625d58e4e31f3aa300

Observation 208baeef-84d1-4e7c-9ac4-f08ed652bcc1 · outbound

This paper cites V oiceBox: Text-guided multilingual universal speech generation at scale.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet V oiceBox: Text-guided multilingual universal speech generation at scale

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.309106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:55:02.438700Z digest=sha256:a9009a4ab763724df46e15f9343461a6d2f32bdefe66d2be5b8d84a653e38123

Observation c7f1d86e-2076-4a00-a0eb-69b5d39e2ef9 · outbound

This paper cites Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:02.574535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:02.574535Z digest=sha256:f62ada5241016b54b6060c57181133533e459028a7a9f949bf68b3395cb3f975

Observation 44608682-5498-460d-a2ca-1803573f58dd · outbound

This paper cites E2 tts: Embarrass- ingly easy fully non-autoregressive zero-shot tts.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet E2 tts: Embarrass- ingly easy fully non-autoregressive zero-shot tts

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.144160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:55:02.684793Z digest=sha256:de706824eca57f17e56d75092e9ab61f3b7ea06196c98f5472f3a73d70ffb4c1

Observation 2c96ced3-fec9-4a79-bac3-b9a38325340c · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:02.966453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:02.966453Z digest=sha256:04802b401e6be8ba2a3c955e8943a16632064ce088939efedb9ac3169e1e55bf

Observation 29734753-f174-4dcb-b602-7a45227fba33 · outbound

This paper cites Tacotron: Towards End-to-End Speech Synthesis.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Tacotron: Towards End-to-End Speech Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.102660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.102660Z digest=sha256:78a86eb0b11dca203ae6eeb2c0ab41208e521fe98def8706f4d11680d67f0550

Observation 9f6c56d8-d55d-4ea0-acde-d94a900578f6 · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Fastspeech: Fast, robust and controllable text to speech

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.137493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.137493Z digest=sha256:c78be4b3f3dc06d2cd6a3b03cc323a49c603c6cfd9b89d63d33f0b91383bde81

Observation 300d42a3-b2bf-4d08-a03f-d3579d2f27cf · outbound

This paper cites ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.233086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.233086Z digest=sha256:72c7502f6786e8eb6d7fba2d60eda088390037f29ea0d16c03fb982975a8bdb0

Observation b1df3dc2-cdee-42ec-a9e5-05ede80647d4 · outbound

This paper cites EmoDiff: Intensity controllable emotional text-to-speech with soft-label guidance.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet EmoDiff: Intensity controllable emotional text-to-speech with soft-label guidance

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.086938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:55:03.346441Z digest=sha256:f04e5f87da03f7fabb3779f53e1426d0ff43d8c8f39bee5b21d70ae66684e387

Observation aa2deb88-4ce8-4e88-be3e-a9407cfa5235 · outbound

This paper cites EmoMix: Emotion Mixing via Diffusion Models for Emotional Speech Synthesis.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet EmoMix: Emotion Mixing via Diffusion Models for Emotional Speech Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.419208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.419208Z digest=sha256:3789110ce1808e708f3f577f4869480da64a37818a40f96ac5312d53e31cf810

Observation 366e1ecd-8c4f-4c60-a42c-bed69eb75c84 · outbound

This paper cites Speech synthesis with mixed emotions.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Speech synthesis with mixed emotions

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.076900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:55:03.483681Z digest=sha256:e1631421e924493b34612ba033325004a66a1d410ced6c75a4b0e71fc2bb4e4a

Observation 26020b9b-e846-477d-b112-66bafcc54e32 · outbound

This paper cites QI-TTS: Questioning intonation control for emotional speech synthesis.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet QI-TTS: Questioning intonation control for emotional speech synthesis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.064363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:55:03.518055Z digest=sha256:216526e8599b94a1e089a2aaf3299fc28080794d2b2fe6ea65cb72d06602147c

Observation abe66048-9699-43e9-95d5-3f7a7d3ee0f5 · outbound

This paper cites MsEmoTTS: Multi-scale emotion transfer, prediction, and control for emotional speech synthesis.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet MsEmoTTS: Multi-scale emotion transfer, prediction, and control for emotional speech synthesis

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.053177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:55:03.523695Z digest=sha256:a288f308b4822e79b6e0ff81af2017f188a3aaefd60a089197865b12c1cd3169

Observation ca31ec89-b39c-4d57-84fa-e2ae6f8c2ab1 · outbound

This paper cites Text-driven Emotional Style Control and Cross-speaker Style Transfer in Neural TTS.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Text-driven Emotional Style Control and Cross-speaker Style Transfer in Neural TTS

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:55:03.856194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:55:03.526791Z digest=sha256:3357990759bf4e96082a786f32145386f78929590996e0b34df9c6c1fcbec746

Observation 18355188-c91f-46cf-89a8-32016829c550 · outbound

This paper cites Emotional End-to-End Neural Speech Synthesizer.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Emotional End-to-End Neural Speech Synthesizer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.529499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.529499Z digest=sha256:d0c0823055693fb8d50402d2cdac9e1d3142c5c8fdedbdd2d4d5b981e447f095

Observation 6a73c16a-4189-4908-8f33-4f0f697b58cc · outbound

This paper cites Controllable emotion transfer for end-to-end speech synthesis.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Controllable emotion transfer for end-to-end speech synthesis

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.043446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:55:03.532021Z digest=sha256:05dfbc08cc7dcb79b971a5653b6939569eb59b7c7e7708c08a0978b15f2ac933

Observation 9b569e5c-3eaa-4f8d-be37-b62c236fe297 · outbound

This paper cites Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.032509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:55:03.534804Z digest=sha256:4e8e0719f2032c157fc22de8fd66b9bae5a6bfaf075b9268ee47697e94249807

Observation cd5dac71-3209-466d-bf22-a250b1ae6ad1 · outbound

This paper cites Emosphere- tts: Emotional style and intensity modeling via spherical emotion vector for controllable emotional text-to-speech.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Emosphere- tts: Emotional style and intensity modeling via spherical emotion vector for controllable emotional text-to-speech

Reference 20

Resolution
verified exact
doi, observed 2026-08-06T19:55:03.658196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:55:03.537283Z digest=sha256:56ac3a68a59a73c1a295eb2d029d0ffc73b8c9c0307e0fa08a4a9abe77057c68

Observation 24f69278-2478-400a-b62a-f7038cf8050b · outbound

This paper cites Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-to-Speech.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-to-Speech

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:55:03.834991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:55:03.540269Z digest=sha256:a76f6d0757731a3e1531a11f9b881806c6b108413f6c80429061dfe6f71307b7

Observation c28b31f1-ae99-4660-a07b-f6aba23591c0 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.543011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.543011Z digest=sha256:14e34908401928ac4b4fba5cb8582d568a3be5a7f847ba7cf887e9efa5b8c503

Observation ae120038-da12-4821-b592-095af89548a3 · outbound

This paper cites High-resolution image synthesis with latent diffusion models, 2021.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet High-resolution image synthesis with latent diffusion models, 2021

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.545404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.545404Z digest=sha256:77c80edbc43d4926e3901abb0e1ab2fc55fb170870618d86b9ecf79f32ee5d85

Observation db240b9d-a81b-4cce-b5cb-3a8760823976 · outbound

This paper cites wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.548339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.548339Z digest=sha256:931431d824f6f88686004df1601a379c8e12f8631e0e1fc0c1a3498f1e5d0dfc

Observation 74ab91bd-8e53-4437-8b1b-443cd01ffbdb · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet High-resolution image synthesis with latent diffusion models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.553864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.553864Z digest=sha256:1e3ee4c81c88d8d165a93c5f30a96e3f0f1a68e7934f81e2343cfa47f1274c42

Observation 3aba256c-3a5c-46b0-bc54-58c020af47e5 · outbound

This paper cites Scalable Diffusion Models with Transformers.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Scalable Diffusion Models with Transformers

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.556597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.556597Z digest=sha256:47b0a9e3767b7471b342041c64d97d2fe5163e05e9604869922144be013bc1ab

Observation dd543283-742f-407e-8f66-e0181482b3e6 · outbound

This paper cites an unresolved cited work.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.558680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.558680Z digest=sha256:9a2f08688cacdb343a24cf9fdcd6c1f2b3738c523bb2f17563393a9fdfea9305

Observation 4722960f-6b47-4504-8095-59584de4ba11 · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.561139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.561139Z digest=sha256:96877dbf3076c6eff2d261f0491eda815c7a9706b3f50ee7a715cd01219f00b0

Observation b47965bb-61e1-41ba-9435-633e28f772fd · outbound

This paper cites Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.563909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.563909Z digest=sha256:aad9bcd41ca7b56decac654f55226f8d067dd8054ce20c4329351ca62a07e6c2

Observation aca95d79-9e1d-4b04-a981-57869bda2470 · outbound

This paper cites Stable Flow: Vital Layers for Training-Free Image Editing.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Stable Flow: Vital Layers for Training-Free Image Editing

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.566512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.566512Z digest=sha256:82f82a831e9f12507d7e82b5c57fda6c2ff76db9ad9dccc5304a6045fde2fae0

Observation 1cfc32c3-f63a-4151-ae45-23d83dec5180 · outbound

This paper cites an unresolved cited work.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:55:04.003235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:55:03.569359Z digest=sha256:39d48cf7c9d14a72485d18f1b238aa3c09c1d0470aa1cf4637e917e3cedfb323

Observation 7c0a4d10-0711-41e7-9bcb-e6d557b03047 · outbound

This paper cites Jointly predicting arousal, valence and dominance with multi-task learning.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Jointly predicting arousal, valence and dominance with multi-task learning

Reference 33

Resolution
verified exact
doi, observed 2026-08-06T19:55:03.647994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:55:03.572713Z digest=sha256:b6e123b3e029f3ed742c34410ba7b7f5c0d1ae6fff87efd69599d59de9b50609

Observation c0bef902-9eaf-4a1a-93a1-9fae01b79ccf · outbound

This paper cites Automatic speech emotion recognition using recurrent neural networks with local attention.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Automatic speech emotion recognition using recurrent neural networks with local attention

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.575345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.575345Z digest=sha256:9ef9058d336fb6253c51ef88ca612375bde970273a88673b1637e108b4b77eb4

Observation 57d46f65-0a4c-4e40-8a49-541b498ce322 · outbound

This paper cites DialogueGCN: A Graph Convolutional Neural Network for Emotion Recognition in Conversation.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet DialogueGCN: A Graph Convolutional Neural Network for Emotion Recognition in Conversation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.577522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.577522Z digest=sha256:9cb0995152305c44f1e99c4dd688c55bf9e19d12f0dd581bd43fae68d4311e1d

Observation 1c380437-f813-455d-95bc-6a23c1569fa7 · outbound

This paper cites Dawn of the transformer era in speech emotion recognition: Closing the valence gap.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Dawn of the transformer era in speech emotion recognition: Closing the valence gap

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:03.991773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:55:03.580775Z digest=sha256:25bdf66a8df94c64a9fbcedd937ecd81e273637d2efb4126e3c55408b537d649

Observation f0c04121-3ed6-4d39-b2ef-27e9337b95c5 · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.583254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.583254Z digest=sha256:a1604dd4d6bdf7df83246a0bcb1a4c80a40ac8e72370a8c00b9ade51f6e5911f

Observation 2a13b751-60f8-42e4-b698-360ce9de2ead · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.585947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.585947Z digest=sha256:3caaf99a7a2012b71f6fd5768e72f81ea975da321946e3d8cdd71080b09f0bda

Observation 9b0e4c01-87bd-4f7a-b487-cfea8791fdd4 · outbound

This paper cites Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:03.975384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:55:03.589132Z digest=sha256:4aa6d8055826d33912a29597309152695184ab1d1fe789e7419d27405f2e7472

Observation 76e9d925-af9f-4cdc-a122-cb4e78110d5a · outbound

This paper cites EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:03.964072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:55:03.591411Z digest=sha256:2c01cde80281a0d6fb4311f06248ed3fd39a3a490d9aca3046e23310d62fe7e9

Observation 04826e61-70a3-40b4-9541-418907ce9bcd · outbound

This paper cites Iemocap: Interactive emotional dyadic motion capture database.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Iemocap: Interactive emotional dyadic motion capture database

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:03.952788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:55:03.593862Z digest=sha256:db0164047fd2b39c4d5e3b633dde89ab5760db8c479c95ee280bc19ec7a0d9f8

Observation 786b6ee7-f2df-484f-adf6-d87af02a949e · outbound

This paper cites Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:03.942258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:55:03.596052Z digest=sha256:54d1d78560dbd9befef6146c2d0f8cf602db649df819f5d129a562f977374262

Observation 50fd03a3-26bf-4b1e-a145-3c47d6137f59 · outbound

This paper cites Expresso: A benchmark and analysis of discrete expressive speech resynthesis.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Expresso: A benchmark and analysis of discrete expressive speech resynthesis

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:03.930958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:55:03.598786Z digest=sha256:3a59e620e6a0356c55847093eac3fc1052f09bdd046fe96f2c5571c8db3fd94f

Observation d39eaf97-f758-4234-9303-b0fcedb68c25 · outbound

This paper cites The ryerson audio-visual database of emotional speech and song (RA VDESS): A dynamic, multimodal set of facial and vocal expressions in north american english.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet The ryerson audio-visual database of emotional speech and song (RA VDESS): A dynamic, multimodal set of facial and vocal expressions in north american english

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:03.918607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:55:03.603813Z digest=sha256:508553ec1a892f5bdaf46fd9045c84c5c50269ccb01fc0fb56c1288661b1db4d

Observation 57707053-24ee-4ccf-b3fa-4bc411d1448c · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Robust speech recognition via large-scale weak supervision

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:03.908017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T19:55:03.606423Z digest=sha256:2fb847eded58e04a6122106147ade0bd59ff13825d5355a392444d188f8a7988

Observation 39cf408f-2267-471a-813e-88f5f227bf17 · outbound

This paper cites Seamless: Multilingual Expressive and Streaming Speech Translation.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.608406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.608406Z digest=sha256:5fccb435b6b5b6c0ad3d972f2f101f356cf0e1f7d79a940c5ae78c35ae4ec41b

Observation ce67d789-761f-4502-8d41-2ab1bd58d570 · outbound

This paper cites supple_demo/index.html.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet supple_demo/index.html

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.611900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.611900Z digest=sha256:524479bcf32fa153bf94c3944196740dbe5273f871f212361c2d09b8bf015d64

Observation 94723fb0-b9a0-4101-abd0-d7076b1748b1 · outbound

This paper cites an unresolved cited work.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Unresolved cited work

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.600978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.600978Z digest=sha256:cf7ede022aceb1d70750c7701ec3c46c8496f6c51a0496e186e4a739e8b035a0

Pith citing papers

Observation 440da532-b107-4cfc-b485-61e310bb88d9 · inbound

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation cites this paper.

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:11:27.124142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-07T12:34:40.888089Z digest=sha256:2d391a15368dc5e29d4d5756ddd5c2ba47147656a180f360da99f404b109b8e3

Observation a0b54a62-c850-489d-8430-cbcb74ef4ad6 · inbound

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation cites this paper.

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T15:21:28.820442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:21:28.820442Z digest=sha256:65fa78faf435ea0169dd64d746bcca5385f763c8ccd8e432fc7f3c097d14941c