Pith. sign in

Paper Citation Record · LEDGER

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet

As of 9 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 2 inbound Pith citation observations for arXiv:2507.04349.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04349 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:55:03.611900Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T15:21:28.820442Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T09:11:27.121428Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact4
  • verified fuzzy17
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f5764825-324b-4fac-a39b-750df4887754 · outbound

This paper cites Adding Conditional Control to Text-to-Image Diffusion Models.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Adding Conditional Control to Text-to-Image Diffusion Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:02.241628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:02.241628Z digest=sha256:81859c5a2c16d3c988b4d32b2576a0b369ad60df1d5887f25110d5aa1329a45b

Observation e546cec3-0c0f-4291-8e58-d3284f1384f9 · outbound

This paper cites An emotion speech synthesis method based on VITS.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet An emotion speech synthesis method based on VITS

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.445789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:02.318699Z digest=sha256:b89ee4e2adab0eb760348e54c97c9820293b76aba50ff92986222724f1dfec3a

Observation 208baeef-84d1-4e7c-9ac4-f08ed652bcc1 · outbound

This paper cites V oiceBox: Text-guided multilingual universal speech generation at scale.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet V oiceBox: Text-guided multilingual universal speech generation at scale

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.309106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:02.438700Z digest=sha256:a905b9090a92da32d6f1cae385bb3e7d5f2cf0e3358ea74368fbdaf1b3607448

Observation c7f1d86e-2076-4a00-a0eb-69b5d39e2ef9 · outbound

This paper cites Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:02.574535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:02.574535Z digest=sha256:247d84fe35d3e0bd5d00627a6bec554f3671def59a01788bef028fa560374a7a

Observation 44608682-5498-460d-a2ca-1803573f58dd · outbound

This paper cites E2 tts: Embarrass- ingly easy fully non-autoregressive zero-shot tts.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet E2 tts: Embarrass- ingly easy fully non-autoregressive zero-shot tts

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.144160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:02.684793Z digest=sha256:b5903d38b5aa75e555118611900727617ff5777e25a6966fa8763173fd57e548

Observation 2c96ced3-fec9-4a79-bac3-b9a38325340c · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:02.966453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:02.966453Z digest=sha256:735621eb1bd05783bc37835a3c3872f23be7e39c94193eb93b007d45be1ec4fc

Observation 29734753-f174-4dcb-b602-7a45227fba33 · outbound

This paper cites Tacotron: Towards End-to-End Speech Synthesis.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Tacotron: Towards End-to-End Speech Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.102660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.102660Z digest=sha256:a8af9ced433958c8e2b6fbcddde57bafe5f1d053db386a7065eaa89c011d0889

Observation 9f6c56d8-d55d-4ea0-acde-d94a900578f6 · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Fastspeech: Fast, robust and controllable text to speech

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.137493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.137493Z digest=sha256:c78be4b3f3dc06d2cd6a3b03cc323a49c603c6cfd9b89d63d33f0b91383bde81

Observation 300d42a3-b2bf-4d08-a03f-d3579d2f27cf · outbound

This paper cites ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.233086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.233086Z digest=sha256:cee0e1886eca1339e28daa14561e988b9c0dfe1b38621b61f39b753cb57c5fe8

Observation b1df3dc2-cdee-42ec-a9e5-05ede80647d4 · outbound

This paper cites EmoDiff: Intensity controllable emotional text-to-speech with soft-label guidance.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet EmoDiff: Intensity controllable emotional text-to-speech with soft-label guidance

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.086938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:03.346441Z digest=sha256:4b74dfeab66ab19928af0f0dc507e6dfd1652ca2dbbec2a34e05fdfd240cd35e

Observation aa2deb88-4ce8-4e88-be3e-a9407cfa5235 · outbound

This paper cites EmoMix: Emotion Mixing via Diffusion Models for Emotional Speech Synthesis.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet EmoMix: Emotion Mixing via Diffusion Models for Emotional Speech Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.419208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.419208Z digest=sha256:801ea5c3bc8eabbc253b5e38fe5ed9bb7c5926ff38eb499d92230b2a56db7da6

Observation 366e1ecd-8c4f-4c60-a42c-bed69eb75c84 · outbound

This paper cites Speech synthesis with mixed emotions.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Speech synthesis with mixed emotions

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.076900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:03.483681Z digest=sha256:38144c158c937192c82f5a17b82821dae0b0147a8ab68f6f8cdaa5bb490e9c78

Observation 26020b9b-e846-477d-b112-66bafcc54e32 · outbound

This paper cites QI-TTS: Questioning intonation control for emotional speech synthesis.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet QI-TTS: Questioning intonation control for emotional speech synthesis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.064363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:03.518055Z digest=sha256:7fdc85bf90ffe173fe8e881daef09ca09ba325a2cf293a516e267b1493864754

Observation abe66048-9699-43e9-95d5-3f7a7d3ee0f5 · outbound

This paper cites MsEmoTTS: Multi-scale emotion transfer, prediction, and control for emotional speech synthesis.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet MsEmoTTS: Multi-scale emotion transfer, prediction, and control for emotional speech synthesis

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.053177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:03.523695Z digest=sha256:3f50ef710b55ebbd1daf3729b60b33878130c42091081cb8a8286c2d8bd877c3

Observation ca31ec89-b39c-4d57-84fa-e2ae6f8c2ab1 · outbound

This paper cites Text-driven Emotional Style Control and Cross-speaker Style Transfer in Neural TTS.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Text-driven Emotional Style Control and Cross-speaker Style Transfer in Neural TTS

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:55:03.856194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:03.526791Z digest=sha256:ef868be3a6e4374a1882f424a181f9db9019c0d2be1417d659e7f5980c9c8037

Observation 18355188-c91f-46cf-89a8-32016829c550 · outbound

This paper cites Emotional End-to-End Neural Speech Synthesizer.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Emotional End-to-End Neural Speech Synthesizer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.529499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.529499Z digest=sha256:07d463154390f74db524e7cda4ae0804b8c962f2bb283e1cf90fe48c3fefb8bd

Observation 6a73c16a-4189-4908-8f33-4f0f697b58cc · outbound

This paper cites Controllable emotion transfer for end-to-end speech synthesis.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Controllable emotion transfer for end-to-end speech synthesis

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.043446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:03.532021Z digest=sha256:af0c8f1c75cdd5c14af3872f9f18b8c18099fa31b11a5871a1055c2756b425ce

Observation 9b569e5c-3eaa-4f8d-be37-b62c236fe297 · outbound

This paper cites Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Emotion controllable speech synthesis using emotion-unlabeled dataset with the assistance of cross-domain speech emotion recognition

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:04.032509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:03.534804Z digest=sha256:55201b09d7a5ec15fdcef37edb291890400349b0732bcf3fb693d70e1ce2a5d9

Observation cd5dac71-3209-466d-bf22-a250b1ae6ad1 · outbound

This paper cites Emosphere- tts: Emotional style and intensity modeling via spherical emotion vector for controllable emotional text-to-speech.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Emosphere- tts: Emotional style and intensity modeling via spherical emotion vector for controllable emotional text-to-speech

Reference 20

Resolution
verified exact
doi, observed 2026-08-06T19:55:03.658196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:03.537283Z digest=sha256:1cd4c563a494c63cab09f5a60eba74aff7e0c9282047ba9033f600e3be284b45

Observation 24f69278-2478-400a-b62a-f7038cf8050b · outbound

This paper cites Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-to-Speech.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-to-Speech

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:55:03.834991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:03.540269Z digest=sha256:c53039b37258c9ad0b05d784b1d24dd9aba5f57a40532d0937a60e462408543d

Observation c28b31f1-ae99-4660-a07b-f6aba23591c0 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.543011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.543011Z digest=sha256:14e34908401928ac4b4fba5cb8582d568a3be5a7f847ba7cf887e9efa5b8c503

Observation ae120038-da12-4821-b592-095af89548a3 · outbound

This paper cites High-resolution image synthesis with latent diffusion models, 2021.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet High-resolution image synthesis with latent diffusion models, 2021

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.545404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.545404Z digest=sha256:77c80edbc43d4926e3901abb0e1ab2fc55fb170870618d86b9ecf79f32ee5d85

Observation db240b9d-a81b-4cce-b5cb-3a8760823976 · outbound

This paper cites wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.548339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.548339Z digest=sha256:62b700655e90d6643c8009822462b85d183cfeb0b10782e6aad2e63c918e5633

Observation 74ab91bd-8e53-4437-8b1b-443cd01ffbdb · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet High-resolution image synthesis with latent diffusion models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.553864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.553864Z digest=sha256:1e3ee4c81c88d8d165a93c5f30a96e3f0f1a68e7934f81e2343cfa47f1274c42

Observation 3aba256c-3a5c-46b0-bc54-58c020af47e5 · outbound

This paper cites Scalable Diffusion Models with Transformers.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Scalable Diffusion Models with Transformers

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.556597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.556597Z digest=sha256:47b0a9e3767b7471b342041c64d97d2fe5163e05e9604869922144be013bc1ab

Observation dd543283-742f-407e-8f66-e0181482b3e6 · outbound

This paper cites an unresolved cited work.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.558680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.558680Z digest=sha256:9a2f08688cacdb343a24cf9fdcd6c1f2b3738c523bb2f17563393a9fdfea9305

Observation 4722960f-6b47-4504-8095-59584de4ba11 · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.561139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.561139Z digest=sha256:3207b1dea58c9ebafbd3705eb6964385b4056bccf5b299fd8c471f220934fa05

Observation b47965bb-61e1-41ba-9435-633e28f772fd · outbound

This paper cites Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Ip-adapter: Text compatible image prompt adapter for text-to-image diffusion models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.563909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.563909Z digest=sha256:aad9bcd41ca7b56decac654f55226f8d067dd8054ce20c4329351ca62a07e6c2

Observation aca95d79-9e1d-4b04-a981-57869bda2470 · outbound

This paper cites Stable Flow: Vital Layers for Training-Free Image Editing.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Stable Flow: Vital Layers for Training-Free Image Editing

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.566512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.566512Z digest=sha256:ff54ac631e2a492861346f3dd9a3bdd7d9d447237388d25e0da4d5af9cb18a57

Observation 1cfc32c3-f63a-4151-ae45-23d83dec5180 · outbound

This paper cites an unresolved cited work.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:55:04.003235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:03.569359Z digest=sha256:00c7706082a7b8b4362681b0c88bfe46e6f43fb41aa484c6b97b83dc267e9d15

Observation 7c0a4d10-0711-41e7-9bcb-e6d557b03047 · outbound

This paper cites Jointly predicting arousal, valence and dominance with multi-task learning.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Jointly predicting arousal, valence and dominance with multi-task learning

Reference 33

Resolution
verified exact
doi, observed 2026-08-06T19:55:03.647994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:03.572713Z digest=sha256:096d7aa4c5269cff367a67dad314d8cd77a6ccae0dc4d40d9258e6a8975ed872

Observation c0bef902-9eaf-4a1a-93a1-9fae01b79ccf · outbound

This paper cites Automatic speech emotion recognition using recurrent neural networks with local attention.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Automatic speech emotion recognition using recurrent neural networks with local attention

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.575345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.575345Z digest=sha256:9ef9058d336fb6253c51ef88ca612375bde970273a88673b1637e108b4b77eb4

Observation 57d46f65-0a4c-4e40-8a49-541b498ce322 · outbound

This paper cites DialogueGCN: A Graph Convolutional Neural Network for Emotion Recognition in Conversation.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet DialogueGCN: A Graph Convolutional Neural Network for Emotion Recognition in Conversation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.577522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.577522Z digest=sha256:7ee99ca5bb77a74328978dc334180d5669f26ce9427446b581db221d0832b1a1

Observation 1c380437-f813-455d-95bc-6a23c1569fa7 · outbound

This paper cites Dawn of the transformer era in speech emotion recognition: Closing the valence gap.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Dawn of the transformer era in speech emotion recognition: Closing the valence gap

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:03.991773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:03.580775Z digest=sha256:1cb2e59b9523eddcb6f2d64b857b57d5559fc0d2259e66587e232854eb340c8d

Observation f0c04121-3ed6-4d39-b2ef-27e9337b95c5 · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.583254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.583254Z digest=sha256:a1604dd4d6bdf7df83246a0bcb1a4c80a40ac8e72370a8c00b9ade51f6e5911f

Observation 2a13b751-60f8-42e4-b698-360ce9de2ead · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.585947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.585947Z digest=sha256:9c46f478359c4be8c0e8808b9135d9ac4127cb5ff71612e6880ad754f77691d6

Observation 9b0e4c01-87bd-4f7a-b487-cfea8791fdd4 · outbound

This paper cites Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:03.975384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:03.589132Z digest=sha256:df04254e523d885cf6119b06d87102d17efee85cf1732be59a31408bc1e3331b

Observation 76e9d925-af9f-4cdc-a122-cb4e78110d5a · outbound

This paper cites EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet EARS: An anechoic fullband speech dataset benchmarked for speech enhancement and dereverberation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:03.964072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:03.591411Z digest=sha256:d074dba93cc6897d0d78602d83166dd7fcbd52db9bc483e7b63c644e2ddc48cd

Observation 04826e61-70a3-40b4-9541-418907ce9bcd · outbound

This paper cites Iemocap: Interactive emotional dyadic motion capture database.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Iemocap: Interactive emotional dyadic motion capture database

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:03.952788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:03.593862Z digest=sha256:bd3c5a157378d1244e15634b10931620fa66162f72d0a70020c7b4d6545cea45

Observation 786b6ee7-f2df-484f-adf6-d87af02a949e · outbound

This paper cites Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Seen and unseen emotional style transfer for voice conversion with a new emotional speech dataset

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:03.942258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:03.596052Z digest=sha256:5494733d8e4206c128c3a1377b0e31c5a8fb7d045a42ac29a34ebc09d4a24e2b

Observation 50fd03a3-26bf-4b1e-a145-3c47d6137f59 · outbound

This paper cites Expresso: A benchmark and analysis of discrete expressive speech resynthesis.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Expresso: A benchmark and analysis of discrete expressive speech resynthesis

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:03.930958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:03.598786Z digest=sha256:6da57a6be8f20009d704ee1a868f5d8f01e4e26d5d0e27b04d69dccb33a50ce3

Observation d39eaf97-f758-4234-9303-b0fcedb68c25 · outbound

This paper cites The ryerson audio-visual database of emotional speech and song (RA VDESS): A dynamic, multimodal set of facial and vocal expressions in north american english.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet The ryerson audio-visual database of emotional speech and song (RA VDESS): A dynamic, multimodal set of facial and vocal expressions in north american english

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:03.918607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:03.603813Z digest=sha256:6a9c15f84c32469e1b3482bc80ae0170c29cba5ff950af693304b3886ae6b3f7

Observation 57707053-24ee-4ccf-b3fa-4bc411d1448c · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Robust speech recognition via large-scale weak supervision

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:55:03.908017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:55:03.606423Z digest=sha256:d9da397828f49bbc3cb53bccf0df204e591a846405a40dec19b77971d794551e

Observation 39cf408f-2267-471a-813e-88f5f227bf17 · outbound

This paper cites Seamless: Multilingual Expressive and Streaming Speech Translation.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.608406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.608406Z digest=sha256:8f6b46b30002e01eb9593d409ae233a8602fcaea9a3d2b27602b7f3716454cc2

Observation ce67d789-761f-4502-8d41-2ab1bd58d570 · outbound

This paper cites supple_demo/index.html.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet supple_demo/index.html

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.611900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.611900Z digest=sha256:524479bcf32fa153bf94c3944196740dbe5273f871f212361c2d09b8bf015d64

Observation 94723fb0-b9a0-4101-abd0-d7076b1748b1 · outbound

This paper cites an unresolved cited work.

TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet Unresolved cited work

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T19:55:03.600978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:55:03.600978Z digest=sha256:cf7ede022aceb1d70750c7701ec3c46c8496f6c51a0496e186e4a739e8b035a0

Pith citing papers

Observation 440da532-b107-4cfc-b485-61e310bb88d9 · inbound

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation cites this paper.

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:11:27.124142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T12:34:40.888089Z digest=sha256:4b9262a01e83f0c9e366088cff6c2514b4192e0134e80a73190b51e87d81279d

Observation a0b54a62-c850-489d-8430-cbcb74ef4ad6 · inbound

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation cites this paper.

The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T15:21:28.820442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:21:28.820442Z digest=sha256:7e4c81d80f26b8961bf471a99773db8428ac4fc213362e8e5f6ad7af6276707e