Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 45 inbound Pith citation observations for arXiv:2312.15185.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T20:48:43.642661Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T18:28:48.480697Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 8c1e784a-8609-4a47-b1bb-6dc4731fc21d · inbound
Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4f3719f0-2809-4359-831b-ac0361df2d2e · inbound
ParaLBench: A Large-Scale Benchmark for Computational Paralinguistics over Acoustic Foundation Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15e19581-d88f-41d5-8b30-29586fd03c7f · inbound
WavChat: A Survey of Spoken Dialogue Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 144
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7f9a7a9-556d-458a-bfbb-362be4585dec · inbound
Movie2Story: A framework for understanding videos and telling stories in the form of novel text emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3d606c3-af7f-405f-a4e5-839d6fb9decb · inbound
SemTalk: Holistic Co-speech Motion Generation with Frame-level Semantic Emphasis emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a9fb71a-b4f9-4d32-9de3-63d071f9e117 · inbound
EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cab6075e-e96e-4126-8053-974c7a93de7b · inbound
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a00d8ca-9701-4fe1-af19-c0862740addd · inbound
MoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animation emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49384579-9107-4709-b7a5-bcd2bcc0c23b · inbound
FleSpeech: Flexibly Controllable Speech Generation with Various Prompts emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfe2489b-1aaf-48eb-819b-defc8d520a24 · inbound
Leveraging Chain of Thought towards Empathetic Spoken Dialogue without Corresponding Question-Answering Data emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62909fd7-b90a-46aa-a123-24b808a9daba · inbound
Overview of the Amphion Toolkit (v0.2) emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75e6be26-a864-4409-b4b9-c61696ffcc7a · inbound
LUCY: Linguistic Understanding and Control Yielding Early Stage of Her emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 143e4c8e-f340-4307-b93c-822ef83fa9be · inbound
Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 974f1597-8794-4829-b9ec-2a0e2c86ae37 · inbound
Gender Bias in Instruction-Guided Speech Synthesis Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f619372-433f-4e44-baf3-81c64b1af81b · inbound
EASY: Emotion-aware Speaker Anonymization via Factorized Distillation emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa03a2b8-92fd-4ae8-afab-9bee14f50ef5 · inbound
EmoSign: A Multimodal Dataset for Understanding Emotions in American Sign Language emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a46fb90-8592-4e1b-9623-cba6f0a2de59 · inbound
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1283c4a1-509d-4616-a396-e1175b8febbd · inbound
Probing the Robustness Properties of Neural Speech Codecs emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a23ebd0-dc44-4dec-b950-692437c41977 · inbound
Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17b384e8-dc7f-4680-a281-53e03cb3b570 · inbound
Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c0e4eed-aa8b-4d57-84ec-3030fd4b73d1 · inbound
Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb7ebb49-ba2d-4e51-805d-2b1b66e5e02f · inbound
Optimizing Multilingual Text-To-Speech with Accents & Emotions emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 382fc309-fb3f-43a4-85a1-672827b718f9 · inbound
MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled Embedding emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f780c7fd-ff9e-48e9-abaf-946cd67323b8 · inbound
BoSS: Beyond-Semantic Speech emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 301254d5-bccf-44a0-9c29-c9adc82369e6 · inbound
Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd2dd173-740f-4bdb-b45a-89eb39ad134e · inbound
Listening to the Unspoken: Exploring "365" Aspects of Multimodal Interview Performance Assessment emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d904b568-df3c-4f04-9e7c-0eceb1425215 · inbound
MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f5dcc674-87c2-4a7e-a7b8-417df35ec02d · inbound
Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f720ed5-cb13-4286-845c-a4e5e3b0d86e · inbound
VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 875a69b9-3d43-4033-95b0-c1a2e6d7ba3b · inbound
Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f2f20dca-d5f0-4a8d-8723-6e5aaa377292 · inbound
Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 26e9a64d-ee20-421a-8ce6-207d9ac59a60 · inbound
Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1235f466-957a-4462-850b-41217ba13cd5 · inbound
ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4331fecc-4aaf-471d-a446-b8f915a6a3d1 · inbound
CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c34293b0-51c8-46a8-8778-04e736d509da · inbound
EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f2cdfda3-401b-417d-b4a1-06c2b5264529 · inbound
VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3400357e-5d39-49d9-a2bb-3ad06fbcff50 · inbound
EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fb0a5d89-b9d0-46f5-9587-71f8ac923026 · inbound
Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a1f077c2-ab30-4e47-8e9c-e03bb13a4806 · inbound
SpeechDx: A Multi-Task Benchmark for Clinical Speech AI emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e9feee97-ee97-4e78-a2f2-09a872113081 · inbound
UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0f87d94c-cb0a-4502-9e79-0d9d381219bf · inbound
FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 189
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e98932ef-6fc6-4774-b43e-c1653e5faa6c · inbound
X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e30aec3-6e10-456d-b113-a85825e87fb0 · inbound
AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 152
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c81d5b5-1e73-428e-941f-4a7c2c63a788 · inbound
Is Self-Pretraining really useful to improve diagnosis in medical Time Series? emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91568008-84fe-4be4-84d6-d88c38f2ae7d · inbound
Is Self-Pretraining really useful to improve diagnosis in medical Time Series? emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.