Pith. sign in

Paper Citation Record · LEDGER

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis

As of 13 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2501.04904.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.04904 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:29:43.405193Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy36
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 455ed74b-dcef-479c-9e4e-35ffe4ac91b9 · outbound

This paper cites Natural TTS Synthesis by Conditioning Wavenet on Mel Spectrogram Predictions,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Natural TTS Synthesis by Conditioning Wavenet on Mel Spectrogram Predictions,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.878455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.250452Z digest=sha256:2e6823f62bf4f79a04839c3f4f03ea18e319e9a97f1e44c4e65c073b404b2b5d

Observation 5775c4de-2f6f-4eb9-989d-47f62652191c · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis FastSpeech 2: Fast and High-Quality End-to-End Text to Speech,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.866233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.254928Z digest=sha256:643ec61c2591cf256378999dd2b7b4e4cd6caba1ea8b302de5e38a7b5279ea61

Observation e7e85684-bfa0-4440-a5d5-4eb2e7cadf91 · outbound

This paper cites HierSpeech: Bridging the Gap between Text and Speech by Hierarchical Variational Inference using Self-supervised Representations for Speech Synthesis,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis HierSpeech: Bridging the Gap between Text and Speech by Hierarchical Variational Inference using Self-supervised Representations for Speech Synthesis,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.853674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.258540Z digest=sha256:522ebbea1c0d882c5c9d61af8786e3e0f67dee9501dcbdd08d9792d3021f307b

Observation 90b748e9-fa3a-4628-835d-8b599e50947f · outbound

This paper cites Matcha-TTS: A fast TTS Architecture with Conditional Flow Matching,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Matcha-TTS: A fast TTS Architecture with Conditional Flow Matching,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.841512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.262268Z digest=sha256:e11919e3e741a7863c2fd683e231136b5af021ad2683b130c1b48fd4639e9ad5

Observation e20bbb16-47d6-40d3-ba20-4e328af25522 · outbound

This paper cites V oice- Flow: Efficient Text-To-Speech with Rectified Flow Matching,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis V oice- Flow: Efficient Text-To-Speech with Rectified Flow Matching,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.829346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.266357Z digest=sha256:bd88ba9b61a2074e44034d6447ec61b813f6f333c8784a47b470e16e766bd3c9

Observation 32510b16-9547-40a3-8b5f-d1a683c714f9 · outbound

This paper cites Conversational End-to-End TTS for V oice Agents,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Conversational End-to-End TTS for V oice Agents,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.817053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.269897Z digest=sha256:915b21ef58503e8ecf8941d7a80a7cf8043d18c59fae29c9f78970ab191baed3

Observation ba724b39-767d-444f-9a78-f4f13c82626b · outbound

This paper cites Enhancing Speaking Styles in Conversational Text-to- Speech Synthesis with Graph-Based Multi-Modal Context Modeling,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Enhancing Speaking Styles in Conversational Text-to- Speech Synthesis with Graph-Based Multi-Modal Context Modeling,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.804447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.273627Z digest=sha256:abfc695c8461fcb08adb623faa07014f47def9efcabb8efbe467683ac53dbbf5

Observation ee3e98f9-7070-44e0-931e-9431589ab592 · outbound

This paper cites M2-CTTS: End-to-End Multi-Scale Multi-Modal Conversational Text-to-Speech Synthesis,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis M2-CTTS: End-to-End Multi-Scale Multi-Modal Conversational Text-to-Speech Synthesis,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.791811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.276720Z digest=sha256:031a86781467bdc554e090212d5253adc074e3f993b811f1092c20eba86de078

Observation 742097bf-bc69-4f58-8f98-1d2a851855ec · outbound

This paper cites Concss: Contrastive-based Context Comprehension for Dialogue-Appropriate Prosody in Conversational Speech Synthesis,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Concss: Contrastive-based Context Comprehension for Dialogue-Appropriate Prosody in Conversational Speech Synthesis,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.780401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.280053Z digest=sha256:498f8a1b6d1631f87726414c07176cbba690d3b76d96c3171d166602de8d5516

Observation 5ce4b5e3-cffa-4fc9-9454-d64d312ae447 · outbound

This paper cites Considering Temporal Connection between Turns for Conversational Speech Synthesis,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Considering Temporal Connection between Turns for Conversational Speech Synthesis,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.771439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.283619Z digest=sha256:43db8955464c8787884f25a91bcc7874bbc21aef710bb8f3d17aaf5827ad33e9

Observation fb2c1b8d-15ba-40ce-bab2-c8de5594d452 · outbound

This paper cites A new recurrent neural-network architecture for visual pattern recognition,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis A new recurrent neural-network architecture for visual pattern recognition,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.762313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.287141Z digest=sha256:fa0660bcd5000b60ce080fa1bfcb3a4966dd0c03ee6e3c9afb19196b54b2bb39

Observation 3a40749b-a6f4-4121-b242-d43c29bd8288 · outbound

This paper cites Classification of drowsiness levels based on a deep spatio-temporal convolutional bidirectional LSTM network using electroencephalogra- phy signals,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Classification of drowsiness levels based on a deep spatio-temporal convolutional bidirectional LSTM network using electroencephalogra- phy signals,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.751693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.290406Z digest=sha256:be46d3b8a0bd62f0835566241d5e5c3298e993f68c6c357c8360432f2e5100e0

Observation beffdf87-2a96-4cb6-8721-f96253610690 · outbound

This paper cites Towards an EEG-based intuitive BCI communication system using imagined speech and visual imagery,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Towards an EEG-based intuitive BCI communication system using imagined speech and visual imagery,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.741084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.293873Z digest=sha256:04ec10fe33be3da6c2cb38411a7699fb3b5f8fc34e2e08e292b42ac3bc50f4a1

Observation 164e2906-c334-4023-94c1-e87a15e31658 · outbound

This paper cites A multi-view cnn with novel variance layer for motor imagery brain computer interface,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis A multi-view cnn with novel variance layer for motor imagery brain computer interface,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.729368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.297415Z digest=sha256:e4e6edf1230a2337f1bad3ced9fa464c4302b937bfb07279e9a5d3a22f04d5d6

Observation fcafa22c-1ec9-4a12-a8e9-a0e61124b666 · outbound

This paper cites An adaptive deep reinforcement learning framework enables curling robots with human-like performance in real-world conditions,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis An adaptive deep reinforcement learning framework enables curling robots with human-like performance in real-world conditions,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.717871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.300931Z digest=sha256:833aa935a0ada0938edc5e02b1cd31e0ad5d221bd7de3ee47396119c204f00c0

Observation 9c98865d-edeb-494d-960a-ebc03960a0c9 · outbound

This paper cites PeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform Generation.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis PeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T21:29:43.304427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:29:43.304427Z digest=sha256:29724e1b8b9ab72bd329a4f6518932689866f72a047933699a18a3f964a863d9

Observation 6ab269ed-6f19-44d1-88c0-899b507cf3f8 · outbound

This paper cites Emoq-tts: Emotion intensity quantization for fine-grained controllable emotional text-to-speech,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Emoq-tts: Emotion intensity quantization for fine-grained controllable emotional text-to-speech,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.707472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.308697Z digest=sha256:41781aaf723ef586ebf269b24a4a33590afb1fbe2b62df549b51e6a9f81d563c

Observation 6996cda2-e285-40fb-830d-fb887bd96225 · outbound

This paper cites Diffprosody: Diffusion-based latent prosody generation for expressive speech synthe- sis with prosody conditional adversarial training,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Diffprosody: Diffusion-based latent prosody generation for expressive speech synthe- sis with prosody conditional adversarial training,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.696860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.312310Z digest=sha256:183978adf176364d9318217b50a36a407b5558235d434574afb59abaa12994eb

Observation 984f69b8-44a3-4d20-8b4a-8b5c772f9de0 · outbound

This paper cites EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.686283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.316220Z digest=sha256:be7db3319ec2814bb0e19072373aa0be419f7f78d9e7125c62499953382797bd

Observation 341a8fea-0dac-4930-8223-911bc5be0baf · outbound

This paper cites DurFlex-EVC: Duration-Flexible Emotional Voice Conversion Leveraging Discrete Representations without Text Alignment.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis DurFlex-EVC: Duration-Flexible Emotional Voice Conversion Leveraging Discrete Representations without Text Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T21:29:43.320033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:29:43.320033Z digest=sha256:7447d419c24dd71b85f66f3581ef01526075bc953a988a162923645e515eac4e

Observation 94e8086e-f874-41c0-a17f-f6512c75925b · outbound

This paper cites Emotion rendering for conversational speech synthesis with heterogeneous graph- based context modeling,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Emotion rendering for conversational speech synthesis with heterogeneous graph- based context modeling,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.675754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.324595Z digest=sha256:51abd638b4f2f310c67f9b446ec52ef5ba32f5c1292676649331852fcbc36b56

Observation 7e155f12-d29e-4db2-ac40-b730b9212873 · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:29:43.328351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:29:43.328351Z digest=sha256:df55b07c5f0bea3c857f11bb8e960a14e774120835fcc5cb672c77c732386b80

Observation d4b4056f-d47d-4147-a27a-19bb8f70547a · outbound

This paper cites BLSP-KD: Bootstrapping Language-Speech Pre-training via Knowledge Distillation.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis BLSP-KD: Bootstrapping Language-Speech Pre-training via Knowledge Distillation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T21:29:43.332430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:29:43.332430Z digest=sha256:3947d512ee6de157111c02d80906c1258b6528973720bca29f1638290c0e9ef6

Observation 6e31169c-973e-422b-8cda-896d7e88e3e7 · outbound

This paper cites Whisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Whisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong General Audio Event Taggers,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.665088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.336305Z digest=sha256:10c266ccd3c10abbed10b0351301886ee3164e49c104cda06b2d0c37bc0ec2d1

Observation 72786add-657c-47f6-a067-d0c711000053 · outbound

This paper cites BLIP-2: Boot- strapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis BLIP-2: Boot- strapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.654723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.340004Z digest=sha256:ef9227b2e57c652c2d5b88e4b6783254402a1aec892fbc269af6e2a08a06aaae

Observation d6b51203-c677-478f-bc65-b3d7ff262897 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Robust Speech Recognition via Large-Scale Weak Supervision,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.646248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.343722Z digest=sha256:8ed538de8695af09c038e440a1a81466be9000d601b8ee49a95d8638a3694bac

Observation 2b3da1c3-3ed5-40f3-b68a-f60507af2742 · outbound

This paper cites Joint Audio and Speech Understanding,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Joint Audio and Speech Understanding,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.637554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.347055Z digest=sha256:44ec2cc71eb9e33ee9f808bda8ba78316ce887876fae92cbe4fdafc560a0fdea

Observation 4677b17d-40d3-4301-8b47-3c7cf683dc18 · outbound

This paper cites HiFi-GAN: Gen- erative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis HiFi-GAN: Gen- erative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.627949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.350771Z digest=sha256:72178d3994332d1aee73e701852046787bed870d45ef6125361cb6b6b158b0e4

Observation 147b7bcf-18f5-4439-b418-170a3605770e · outbound

This paper cites DailyTalk: Spoken Dialogue Dataset for Conversational Text-to-Speech,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis DailyTalk: Spoken Dialogue Dataset for Conversational Text-to-Speech,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.617661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.354260Z digest=sha256:5217f6af074c52b62a94f16b141117a5c1607a93beb94ebbf5399d9370576cf5

Observation f0ae0780-a0db-4eb4-a9f3-39cd9a8dff17 · outbound

This paper cites DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.607229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.357664Z digest=sha256:3aec51625801cf4374623e085020a10b9d7105bd5de1ad8fcbf95c6d87501824

Observation 52c67a9b-2bec-4a87-9803-ea1867d58737 · outbound

This paper cites CREMA-D: Crowd-Sourced Emotional Multimodal Actors Dataset,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis CREMA-D: Crowd-Sourced Emotional Multimodal Actors Dataset,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.596339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.361632Z digest=sha256:d165fce563ae309676d58eb7bdb0b2cc8f45bcf578f2b07f1a52e36bb380b9b9

Observation 826908ef-0d70-4fe4-9dcb-0992498ea9d1 · outbound

This paper cites The Emotional Voices Database: Towards Controlling the Emotion Dimension in Voice Generation Systems.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis The Emotional Voices Database: Towards Controlling the Emotion Dimension in Voice Generation Systems

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T21:29:43.365529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:29:43.365529Z digest=sha256:5ad407d814eb01358b06bd7282575a79b061155a0be8e814f9b599e3a8bad676

Observation 641c9827-6c87-4460-9c1b-7c46aef6d3cc · outbound

This paper cites IEMOCAP: Interactive emotional dyadic motion capture database,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis IEMOCAP: Interactive emotional dyadic motion capture database,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.585460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.369873Z digest=sha256:828a9ba759f45853fea4255b437337f4a6ad99a3b066fe53364268d165315955

Observation 7171c184-92f1-43ca-8ef0-bf1fc0f118e5 · outbound

This paper cites MEAD: A Large-scale Audio-visual Dataset for Emotional Talking-face Generation,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis MEAD: A Large-scale Audio-visual Dataset for Emotional Talking-face Generation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.574845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.373630Z digest=sha256:9958a130c458f8a85667a574cc8e4d15348c00f85279cebcd2ddb177063dcb32

Observation 6c25328d-a94d-4d36-a1fe-da76096e92e1 · outbound

This paper cites Toronto emotional speech set (tess)-younger talker happy,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Toronto emotional speech set (tess)-younger talker happy,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.564236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.377953Z digest=sha256:ecf9d5ad5522487a048bb740483d9521ec9474a1586ca0b73ae4805e3114df6c

Observation 2c0d3525-bab5-4ed6-b275-f2e7825b8c83 · outbound

This paper cites Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.552689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.381638Z digest=sha256:611993e8a76091f0742f05014e3bf3140e084c3e9bc35657a13fb53617d9daec

Observation 4def9cc2-dc2c-448f-ae21-4702a0626c30 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis LLaMA: Open and Efficient Foundation Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T21:29:43.385693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:29:43.385693Z digest=sha256:41fdd9f7dc4237c89623e13327bdd325bb825267e8998fd7f01fd174e428c0c5

Observation 7648c2f8-f3da-4d9d-ae90-fab86882c09a · outbound

This paper cites Decoupled Weight Decay Regular- ization,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Decoupled Weight Decay Regular- ization,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.542313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.389703Z digest=sha256:b91028f50d67ea070c64f9d0ea65cf35577452244a01149e787c8c245078aac8

Observation 67b43322-bc93-41fb-b1cf-9886ae39087c · outbound

This paper cites emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.531861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.393629Z digest=sha256:3df0de7f5f7bb39ca119ba44b161b6d2495a932264f606c4e94b9ae6a331bafe

Observation 5f18c2de-8774-4623-9571-2f397f39f659 · outbound

This paper cites wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.521601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.397315Z digest=sha256:33118b44d6c58afdd326968f14a0f08c10917b36467ec2b384272054ac4f392c

Observation 9c39e78d-c66c-41b2-9c47-f32b48e68740 · outbound

This paper cites Sequence-to-Sequence Acoustic Modeling for V oice Conversion,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis Sequence-to-Sequence Acoustic Modeling for V oice Conversion,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.511085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.401154Z digest=sha256:4e1aa9e29d7a59ce878f3de24adc21e76908b541193688ab9b464e063f0a3069

Observation 4df5316b-c29f-4c04-9415-635dde855aaa · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models,.

JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis LoRA: Low-Rank Adaptation of Large Language Models,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:29:43.499432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T21:29:43.405193Z digest=sha256:442922669593fc84b1404089f51dcbbd0f6e943b75021b9a891b2d9a2bb5c9a1

Pith citing papers

No inbound Pith citation observations are available.