Pith. sign in

Paper Citation Record · LEDGER

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control

As of 15 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 3 inbound Pith citation observations for arXiv:2501.06276.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.06276 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:09:35.082934Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:09:34.920841Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:08:37.819812Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ae09bece-a7cd-4d03-a640-8fc3ab0ff7cb · outbound

This paper cites an unresolved cited work.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:09:35.564020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:34.915567Z digest=sha256:636e349454c538d466c9dab754035fdeb728420a5bad2bfc58a235c7045924a8

Observation a4491153-9b71-43eb-8361-b68bab420ec7 · outbound

This paper cites PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:09:34.920841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:09:34.920841Z digest=sha256:388da9b878464c7877ebcf90d46dd16b40e4fed58377cb58a0147dfb2c9644f8

Observation a4e8761a-f39a-4a1a-bd6c-1cd9b5819efe · outbound

This paper cites Initially, we pre-train a multispeaker English TTS backbone model uti- lizing a large publicly accessible dataset.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Initially, we pre-train a multispeaker English TTS backbone model uti- lizing a large publicly accessible dataset

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.551954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:34.925804Z digest=sha256:9b7b6f39ae56c3a5467b28f683cfd2e9128665e2c4beb9e13dced9c0a53a32e4

Observation 28294d3a-8f25-45e5-8e5b-169636cb96ff · outbound

This paper cites Baselines and Datasets We utilize the LibriTTS [28] for pretraining, which includes 33, 236 training samples (equating to 53.78 hours) collected from 247 speakers.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Baselines and Datasets We utilize the LibriTTS [28] for pretraining, which includes 33, 236 training samples (equating to 53.78 hours) collected from 247 speakers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.527205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:34.934269Z digest=sha256:c780ba6180bc86326ad7dc6b52a5b2fae508137e2452689e058c4b4287ddc470

Observation f9a1b84b-2917-4371-b561-17c8f575e5aa · outbound

This paper cites Objective Evaluation To assess synthesized speech expressiveness, we analyze emo- tion recognition across different models.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Objective Evaluation To assess synthesized speech expressiveness, we analyze emo- tion recognition across different models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.515164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:34.938291Z digest=sha256:f526ec0c934e641fde26b36dcdf8b56277e7ed39f3be08ae99ba3f3461204e07

Observation 9ea96620-6330-4c69-9261-a154450227ab · outbound

This paper cites an unresolved cited work.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:09:35.503840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:34.942387Z digest=sha256:71e266dd471f1831c38e3ae561b4dd7bc7446e454f2576d6bc9e24dec2a608ea

Observation ec112fbb-13c0-4a28-b8b7-2c43f7111a50 · outbound

This paper cites Multi- speaker expressive speech synthesis via multiple factors decou- pling,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Multi- speaker expressive speech synthesis via multiple factors decou- pling,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.445273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:34.969860Z digest=sha256:1657b7d444c693b0c7527aed81cbb8a6ccdd1e67e618573542c42ee90527f2b3

Observation 49bb835a-2999-479a-8c7c-c9865de43a80 · outbound

This paper cites A review of deep learning techniques for speech processing,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control A review of deep learning techniques for speech processing,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.492681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:34.946189Z digest=sha256:7debc8a96e5b911ea36da74f5d5ec4d1613c89494db8583763ebc92ea68b54ba

Observation 304b3469-c214-49b8-b21e-bf8811bbf1c3 · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:09:34.949717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:09:34.949717Z digest=sha256:62ee36d9b24338e0895251f7bb22fedd96a8a4a46268f0f0ad63ab30357fb76b

Observation 1918cb47-b0eb-41d2-90e0-1ba2df8db0be · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T21:09:34.953521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:09:34.953521Z digest=sha256:b9c394c52df5bfce35c891b9a4b00c274ceec971e569682501a848d79ad58c93

Observation 6b4dce59-35db-4191-ad7e-d00eed9c84ca · outbound

This paper cites Expressive speech synthesis: a review,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Expressive speech synthesis: a review,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.475270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:34.957779Z digest=sha256:e0ef07a2a49b8f20f1f51f0528cb8adb0aa7c93990b5d8ab4b3f207dfceddd49

Observation 1773c3ea-f87f-4015-96e4-324cc75fc722 · outbound

This paper cites V oicebox: Text-guided multilingual universal speech generation at scale,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control V oicebox: Text-guided multilingual universal speech generation at scale,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:09:34.961669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:09:34.961669Z digest=sha256:8c4e99787c3ce21f0e14312671b1bcbd6bfc7f9cc49010d0da23b7e7ed43b990

Observation 22dcde4a-6bc9-45a8-8df5-d8a63b7a8ec0 · outbound

This paper cites Msstyletts: Multi-scale style modeling with hierarchical context information for expressive speech synthesis,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Msstyletts: Multi-scale style modeling with hierarchical context information for expressive speech synthesis,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.457079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:34.965764Z digest=sha256:7d950b8131f1f74a11b6af0a0ea51795770bec7922647b5b21cd16592550580c

Observation 171976f6-926b-42ec-ac39-073b22a02913 · outbound

This paper cites Style tokens: Un- supervised style modeling, control and transfer in end-to-end speech synthesis,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Style tokens: Un- supervised style modeling, control and transfer in end-to-end speech synthesis,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:09:34.995520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:09:34.995520Z digest=sha256:d3b20097a865e0c2ea2bc0fe993e84c4f8dff1e4497fe6f9ac8ec17cb872e0fe

Observation fb79fcc9-ca34-49e8-bc8b-7c603abeecbd · outbound

This paper cites Ensemble prosody prediction for expressive speech synthesis,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Ensemble prosody prediction for expressive speech synthesis,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.432987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:34.973993Z digest=sha256:ee987d04aff0620b5d7822024267e6e791a7839038336ca48082aefb26fd80d2

Observation e22642a2-96e3-4976-ad94-0c22a55faee0 · outbound

This paper cites An overview of affective speech synthesis and conversion in the deep learning era,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control An overview of affective speech synthesis and conversion in the deep learning era,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.420498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:34.977259Z digest=sha256:812a1295980ba82fef4a61bd21a8a5a5b811419a37da469c0e96f566d78209f2

Observation 1a956523-08bd-488b-8c75-57f198143293 · outbound

This paper cites Prosody-tts: An end-to-end speech synthesis system with prosody control,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Prosody-tts: An end-to-end speech synthesis system with prosody control,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T21:09:34.980661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:09:34.980661Z digest=sha256:4874d07500c006532cab9fa830a1df98373b93b50f1a6b22457e283705622946

Observation 22757c72-eb71-40d1-ad12-f7f566e29cd9 · outbound

This paper cites Speech Synthesis with Self-Supervisedly Learnt Prosodic Representations,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Speech Synthesis with Self-Supervisedly Learnt Prosodic Representations,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.402116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:34.984093Z digest=sha256:3be68ce802ecefcba68f0fb6fefdd3c9477c08ff9176e223b7713e3f3b1c2101

Observation fbf42999-012e-4ed7-9534-969779d28c25 · outbound

This paper cites Exploiting emotion information in speaker em- beddings for expressive text-to-speech,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Exploiting emotion information in speaker em- beddings for expressive text-to-speech,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.391318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:34.987587Z digest=sha256:751b7bf757ec6f6ccba232fe14be27c7de936f625db6dcbede91cfd469abe884

Observation 435a2ef0-aca9-4e44-8be1-67bc8d8d5223 · outbound

This paper cites EmoMix: Emotion Mixing via Diffusion Models for Emotional Speech Syn- thesis,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control EmoMix: Emotion Mixing via Diffusion Models for Emotional Speech Syn- thesis,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.381236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:34.991073Z digest=sha256:3f09abfb3e64d34c4e6f98f500514f2483b86b9f29df6bb40ba8e62cb8bc76e8

Observation 44ef7b31-f0db-49b0-b167-98dcaefe20d1 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.308417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:35.023124Z digest=sha256:28151278928b1b0013b35dc0e9d8652ed9de6454db6dccfd230b95c75204b72f

Observation 960dc8a4-2ea3-47fa-90a3-4a4a4d0ac8a8 · outbound

This paper cites Prompttts: Control- lable text-to-speech with text descriptions,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Prompttts: Control- lable text-to-speech with text descriptions,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:09:34.999491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:09:34.999491Z digest=sha256:a6d3e11c1459b2bb39d75dd61b4c6cf7f4ec7b487579e5b557b5c2f4a599edbe

Observation 7cb87d6c-6670-4bc8-aa3b-1bac18fe59d6 · outbound

This paper cites We intend to improve upon the algorithm proposed in [23] for predicting emotion intensity levels using a multispeaker emotive speech dataset.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control We intend to improve upon the algorithm proposed in [23] for predicting emotion intensity levels using a multispeaker emotive speech dataset

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.539109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:34.929998Z digest=sha256:beab7392dfbc94ab3f2cb5e0696bf86b7d5c5b6855f5281a755d1832ea48d19c

Observation ee9f89f8-5449-4ff9-8744-ab3313cd80af · outbound

This paper cites Controllable speaking styles using a large language model,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Controllable speaking styles using a large language model,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.357008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:35.003413Z digest=sha256:bc60b66202d363d299ad339881e69c85b8aab1f04fe8a622323fe788a5ce96e1

Observation fbc63428-e708-4095-a5cc-30197bd47213 · outbound

This paper cites Emodiff: Intensity con- trollable emotional text-to-speech with soft-label guidance,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Emodiff: Intensity con- trollable emotional text-to-speech with soft-label guidance,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.344939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:35.007091Z digest=sha256:c6292f6fa4867f2567bd9247ab45c9b26ea67646db9e74e62e9f4f1fb11fe73e

Observation 845a4112-e8c3-4d00-a9f5-27441f69db91 · outbound

This paper cites Fine-grained quantitative emotion editing for speech generation,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Fine-grained quantitative emotion editing for speech generation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.333132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:35.010668Z digest=sha256:beceb9d17551e0b964ff7833a6b550cd49e056901ff99441406d41ca142dc255

Observation c8c9ac93-f5ed-4e8f-8254-f0a9bf0fdcbd · outbound

This paper cites Train- ing language models to follow instructions with human feedback,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Train- ing language models to follow instructions with human feedback,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.320058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:35.014195Z digest=sha256:4913d2188437bf9f07f71d6e734bf1cc2c0c64cc099945953de10c0ff9d701d6

Observation 247c7768-c071-4474-85a1-1ce7ff51ca9c · outbound

This paper cites InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T21:09:35.018333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:09:35.018333Z digest=sha256:ce5c353fd9a862a69226cbdcc1f9a03906bd30c65c2e19762570b8c3fe3b592b

Observation b88e0b78-3d9d-436f-b3f1-680243dad3f6 · outbound

This paper cites Hubert: Self-supervised speech rep- resentation learning by masked prediction of hidden units,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Hubert: Self-supervised speech rep- resentation learning by masked prediction of hidden units,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T21:09:35.028120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:09:35.028120Z digest=sha256:3814800f4991d65a3833fd1345a08d8b8bec102d30085e6321d16fbdb73be6f4

Observation 860ced15-bb13-4538-b808-ec87e987fb56 · outbound

This paper cites Emotion intensity and its control for emotional voice conversion,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Emotion intensity and its control for emotional voice conversion,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.288765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:35.031787Z digest=sha256:0b7451e2f8e285b937ab60ea3e17ce1f515307a1a441dbe062d96a8672fe4964

Observation d85ad086-acba-49c4-af5e-db3e5c0e9f64 · outbound

This paper cites Learning visual attributes,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Learning visual attributes,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.276915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:35.035454Z digest=sha256:7598ff80093fe519995904fbc661f09aab4e475a405609fa3a67c47387d8af57

Observation 373f9213-95bb-42a0-b293-98f84ffc3ef0 · outbound

This paper cites opensmile: The mu- nich versatile and fast open-source audio feature extractor,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control opensmile: The mu- nich versatile and fast open-source audio feature extractor,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.264947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:35.039142Z digest=sha256:ea7044b68018104425e1a32b15f9238e508329260438e62a14ced2c9da5db92d

Observation 8ff08090-d94e-4596-ab52-f61aab1eedaa · outbound

This paper cites Relative attributes,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Relative attributes,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.251963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:35.042771Z digest=sha256:1c151d0b02e660cb011de70498b6d368a23e23723e506cdec6238811ba2a7a8a

Observation 742f6c2b-7447-41d3-9f2a-d10af22316f0 · outbound

This paper cites Gpt-4 system card,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Gpt-4 system card,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.239664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:35.046239Z digest=sha256:db4f34b9a4bad586ccf92201bf348abbbc22435f8f3895c45a02f0b10cb7f318

Observation 1e54675a-d36e-441c-be70-3ced3c72c5c1 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T21:09:35.050554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:09:35.050554Z digest=sha256:73d3f6344eb121091f41d95a43a6620e93da7f26705a73587805621c57b5de02

Observation f911a4dd-66ad-4704-a1b0-8f257ba21694 · outbound

This paper cites Daft-Exprt: Cross-Speaker Prosody Transfer on Any Text for Expressive Speech Synthesis.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Daft-Exprt: Cross-Speaker Prosody Transfer on Any Text for Expressive Speech Synthesis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T21:09:35.054553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:09:35.054553Z digest=sha256:f159d115f59b374513cc1bd90f9d42f97f2cd24ae6742367d22735a95c24f73d

Observation 345a9f60-4e17-407f-b52c-a48cfe91c329 · outbound

This paper cites ubisoft-laforge-daft-exprt,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control ubisoft-laforge-daft-exprt,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.227120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:35.059519Z digest=sha256:b6c35f4b9b4e0be984f1828726c433a8457fa754100bd7cec55d363cd214fc54

Observation 9b38db95-e8b5-4448-a571-ca6750f7bec0 · outbound

This paper cites Generalized end-to-end loss for speaker verification,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Generalized end-to-end loss for speaker verification,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T21:09:35.063108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:09:35.063108Z digest=sha256:e1269d674c7c69641af2b7b17609fdda7568874a3e1e084047d0b0b863f930b4

Observation 38bb2fbf-e097-4d0b-ac8b-a839e2f6fb03 · outbound

This paper cites Hubert-base,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Hubert-base,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.207952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:35.067053Z digest=sha256:95502d7fcfa4ce204615e106e734a7c39f9c17c006a35f109c830c25d4b04eff

Observation 881ad810-739b-4073-92b1-e5e29258e8b6 · outbound

This paper cites Synthesizer voice qual- ity of new languages calibrated with mean mel cepstral distor- tion.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Synthesizer voice qual- ity of new languages calibrated with mean mel cepstral distor- tion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.196402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:35.070693Z digest=sha256:e5f611fd4dc4f7976c151d1db0c94fedbf40d142e0c39973bb32477ba7adb580

Observation 90936019-a344-4f86-90ec-6b831b620153 · outbound

This paper cites Robust speech recognition via large-scale weak su- pervision,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Robust speech recognition via large-scale weak su- pervision,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T21:09:35.075180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:09:35.075180Z digest=sha256:d19ced380cef73245a34c3ac38f957b3105d9590bd6f0de21f4820ccd52a413d

Observation 3615b9af-3e54-4f7b-a0ad-3b6a80c95031 · outbound

This paper cites Mean opinion score (mos) revisited: Methods and applications, limitations and alter- natives,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Mean opinion score (mos) revisited: Methods and applications, limitations and alter- natives,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.178593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:35.079283Z digest=sha256:b2baf8a84ba2a6af78042bc95ba52a8771f0fad009a6fbcebf78adcd52aeb863

Observation acd42077-81cd-427b-a8e4-c0bb04f4ef43 · outbound

This paper cites Fine-grained emotional control of text-to-speech: Learning to rank inter-and intra- class emotion intensities,.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control Fine-grained emotional control of text-to-speech: Learning to rank inter-and intra- class emotion intensities,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:09:35.166810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T21:09:35.082934Z digest=sha256:9c23c759477c2c789dc11dd7b4968909fe98ae36ee95e2f81f5b742297802e44

Pith citing papers

Observation a4491153-9b71-43eb-8361-b68bab420ec7 · inbound

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control cites this paper.

PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:09:34.920841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:09:34.920841Z digest=sha256:388da9b878464c7877ebcf90d46dd16b40e4fed58377cb58a0147dfb2c9644f8

Observation 596feb8c-5bfb-4a1f-86f2-051aa62c4a9a · inbound

Intelligent Agents with Emotional Intelligence: Current Trends, Challenges, and Future Prospects cites this paper.

Intelligent Agents with Emotional Intelligence: Current Trends, Challenges, and Future Prospects PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control

Reference 150

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:12:30.260198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T08:11:31.181704Z digest=sha256:f649ef4c3ef9f893b89387baca5fef4bf0fe7017b16216ae55eedd2395d777d1

Observation 404d4c5d-a7f2-4f4c-8bca-76eff04cbb5a · inbound

Emo-LiPO: Listwise Preference Optimization for Fine-Grained Emotion Intensity Control in LLM-based Text-to-Speech cites this paper.

Emo-LiPO: Listwise Preference Optimization for Fine-Grained Emotion Intensity Control in LLM-based Text-to-Speech PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:08:37.821400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T06:01:06.123832Z digest=sha256:acbca6444d93a65ae29f408c7ebffebe8be2f2c696916f18c9db8c6fd673d763