Pith. sign in

Paper Citation Record · LEDGER

emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 45 inbound Pith citation observations for arXiv:2312.15185.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.15185 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 45 of 45 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:48:43.642661Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T18:28:48.480697Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8c1e784a-8609-4a47-b1bb-6dc4731fc21d · inbound

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS cites this paper.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:33:25.612916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:0b070f8c43b5876e3a8b1a40d366ea16e36a52216dfa719cded9385d45dde7f9

Observation 4f3719f0-2809-4359-831b-ac0361df2d2e · inbound

ParaLBench: A Large-Scale Benchmark for Computational Paralinguistics over Acoustic Foundation Models cites this paper.

ParaLBench: A Large-Scale Benchmark for Computational Paralinguistics over Acoustic Foundation Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T20:48:43.642661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:48:43.642661Z digest=sha256:f88514d3e28b088154cc0950d3bc62cdc44434d1f79dd79a8650be959209d52d

Observation 15e19581-d88f-41d5-8b30-29586fd03c7f · inbound

WavChat: A Survey of Spoken Dialogue Models cites this paper.

WavChat: A Survey of Spoken Dialogue Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 144

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.752127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.752127Z digest=sha256:3320ca3dc687af85ece2f751f7129e57ad02803d423e9e333a9f8d627501351d

Observation e7f9a7a9-556d-458a-bfbb-362be4585dec · inbound

Movie2Story: A framework for understanding videos and telling stories in the form of novel text cites this paper.

Movie2Story: A framework for understanding videos and telling stories in the form of novel text emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T11:48:07.145680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:48:07.145680Z digest=sha256:65589ee39766c78ed989938474589e92740ec3bac82710437d2ab10e0c1ad205

Observation e3d606c3-af7f-405f-a4e5-839d6fb9decb · inbound

SemTalk: Holistic Co-speech Motion Generation with Frame-level Semantic Emphasis cites this paper.

SemTalk: Holistic Co-speech Motion Generation with Frame-level Semantic Emphasis emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T10:30:47.727145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:30:47.727145Z digest=sha256:6e043e1d455b390d1cdbe4208f728bec8ef2b5654293ecc5158ee596d0ca06cf

Observation 8a9fb71a-b4f9-4d32-9de3-63d071f9e117 · inbound

EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion cites this paper.

EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T23:28:04.187733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:28:04.187733Z digest=sha256:4490c77e43b356822c7296829ab5362f180ee624ff98e1fcb73fd23925143776

Observation cab6075e-e96e-4126-8053-974c7a93de7b · inbound

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios cites this paper.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.987198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.987198Z digest=sha256:5ecb3f080756e9eb1f5c29b0a59c4fc94f9abb66d49608fc52fae0f1b5b67c11

Observation 5a00d8ca-9701-4fe1-af19-c0862740addd · inbound

MoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animation cites this paper.

MoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animation emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:23:50.605657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:23:50.605657Z digest=sha256:2b3c2c661053c4722a90d53a6c70c54f34fc3f2d124ca96e2b60ab6d5105448d

Observation 49384579-9107-4709-b7a5-bcd2bcc0c23b · inbound

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts cites this paper.

FleSpeech: Flexibly Controllable Speech Generation with Various Prompts emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T21:32:17.086034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:32:17.086034Z digest=sha256:ababe1ad7e7ebb86f0fcf34fa94fe3084a67324204f9ee846eac6c6cdfb82e9c

Observation dfe2489b-1aaf-48eb-819b-defc8d520a24 · inbound

Leveraging Chain of Thought towards Empathetic Spoken Dialogue without Corresponding Question-Answering Data cites this paper.

Leveraging Chain of Thought towards Empathetic Spoken Dialogue without Corresponding Question-Answering Data emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T18:52:42.974249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:52:42.974249Z digest=sha256:29cbefa2b296775bb328bf2a22ca0e07af012b6573847dc8b4e9b23348130a34

Observation 62909fd7-b90a-46aa-a123-24b808a9daba · inbound

Overview of the Amphion Toolkit (v0.2) cites this paper.

Overview of the Amphion Toolkit (v0.2) emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.866566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.866566Z digest=sha256:4225e062f57cbbd2d1ed82e2264bb9284255dec51875a0bab21d5b2300c195f4

Observation 75e6be26-a864-4409-b4b9-c61696ffcc7a · inbound

LUCY: Linguistic Understanding and Control Yielding Early Stage of Her cites this paper.

LUCY: Linguistic Understanding and Control Yielding Early Stage of Her emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T13:34:25.681995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:34:25.681995Z digest=sha256:807630a3d1f7e36f853fdd1cf7ccebcd35f8e740148fd7a51484a376b936aade

Observation 143e4c8e-f340-4307-b93c-822ef83fa9be · inbound

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model cites this paper.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.182473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.182473Z digest=sha256:6bc182f3b9ae281b22af5a1d0adfadc6a44763936e3f25b2b4e931caeb95b977

Observation 974f1597-8794-4829-b9ec-2a0e2c86ae37 · inbound

Gender Bias in Instruction-Guided Speech Synthesis Models cites this paper.

Gender Bias in Instruction-Guided Speech Synthesis Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:11.621926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:11.621926Z digest=sha256:ea47c5b863bcce378a8ac75d2c16b9f44f0f789a918c6318cd0386dd26921339

Observation 4f619372-433f-4e44-baf3-81c64b1af81b · inbound

EASY: Emotion-aware Speaker Anonymization via Factorized Distillation cites this paper.

EASY: Emotion-aware Speaker Anonymization via Factorized Distillation emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:30:36.078119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:30:36.078119Z digest=sha256:c3cd953550eccaee294b8c1d52648b4f3e613078dba8513747dbeff696b88553

Observation fa03a2b8-92fd-4ae8-afab-9bee14f50ef5 · inbound

EmoSign: A Multimodal Dataset for Understanding Emotions in American Sign Language cites this paper.

EmoSign: A Multimodal Dataset for Understanding Emotions in American Sign Language emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:30:02.058069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:30:02.058069Z digest=sha256:be3a51e7b5c96918d1d13e497543da999214c222cccda112281fa7f540713278

Observation 2a46fb90-8592-4e1b-9623-cba6f0a2de59 · inbound

MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt cites this paper.

MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:39.181667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:39.181667Z digest=sha256:6eb9760376cd541c94c5b59ffd1d60eaf778cda6e77b275a873a0c8feea20202

Observation 1283c4a1-509d-4616-a396-e1175b8febbd · inbound

Probing the Robustness Properties of Neural Speech Codecs cites this paper.

Probing the Robustness Properties of Neural Speech Codecs emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:31:39.523254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:31:39.523254Z digest=sha256:f34b16078a478cd98e1edc60033d13a06af10f6f0e1ff18b7669aa07d2750376

Observation 3a23ebd0-dc44-4dec-b950-692437c41977 · inbound

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching cites this paper.

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:56.482687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:56.482687Z digest=sha256:46f4af80f14b8f6f3448b4b9cc32f2e3cc4ac21d6bc2664bbb5a32c9ac1b86f9

Observation 17b384e8-dc7f-4680-a281-53e03cb3b570 · inbound

Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions cites this paper.

Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:21:09.566555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:21:09.566555Z digest=sha256:bc442749cec166858ed6a65d8900f2cae4663ed32001dba98d9bc8548c0ddde6

Observation 2c0e4eed-aa8b-4d57-84ec-3030fd4b73d1 · inbound

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion cites this paper.

Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.353059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.353059Z digest=sha256:82196b44f52a88dea85dbb8c76ece361294e95ccf8eaf1c8bafe861356f66ff3

Observation eb7ebb49-ba2d-4e51-805d-2b1b66e5e02f · inbound

Optimizing Multilingual Text-To-Speech with Accents & Emotions cites this paper.

Optimizing Multilingual Text-To-Speech with Accents & Emotions emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:29.542314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:29.542314Z digest=sha256:ba95ace63d741d4f3f1c6c9fce7b9dc0db024344162245f2c67130470e3545fc

Observation 382fc309-fb3f-43a4-85a1-672827b718f9 · inbound

MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled Embedding cites this paper.

MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled Embedding emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:16:57.081692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:16:57.081692Z digest=sha256:42c6fe50ceb0fa375ff20d4fbdbc08d7512795855adce82989cc9379aced9df9

Observation f780c7fd-ff9e-48e9-abaf-946cd67323b8 · inbound

BoSS: Beyond-Semantic Speech cites this paper.

BoSS: Beyond-Semantic Speech emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:50:04.223838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:50:04.223838Z digest=sha256:46fb3907beef85b5e3fb6baaec85bb2eae1c8cced143feb0cf0d7040cfef91b4

Observation 301254d5-bccf-44a0-9c29-c9adc82369e6 · inbound

Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors cites this paper.

Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T11:50:25.414247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:50:25.414247Z digest=sha256:8fc87b6a2f96dc76632defb258cdb6f230b4482ee54b665d49bb624a14559d7d

Observation fd2dd173-740f-4bdb-b45a-89eb39ad134e · inbound

Listening to the Unspoken: Exploring "365" Aspects of Multimodal Interview Performance Assessment cites this paper.

Listening to the Unspoken: Exploring "365" Aspects of Multimodal Interview Performance Assessment emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T11:29:41.195435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:29:41.195435Z digest=sha256:8be474e10b3b2133d73fa28ee50a74d330b3a1c35b4ff99b247c0c92e0974e6e

Observation d904b568-df3c-4f04-9e7c-0eceb1425215 · inbound

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks cites this paper.

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:41:59.596971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-19T02:41:52.996457Z digest=sha256:2e1c4beb1b8302935111a8eb11fee1084ebe1ef377742b05baca93e923c8242c

Observation f5dcc674-87c2-4a7e-a7b8-417df35ec02d · inbound

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody cites this paper.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:43.524538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:43.524538Z digest=sha256:cff778f0e8cc2bcfd57cc1bb13d442eb26798573e9b2ffef2bc769a2a4c42696

Observation 4f720ed5-cb13-4286-845c-a4e5e3b0d86e · inbound

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents cites this paper.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.444581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.444581Z digest=sha256:a46360ec80ffbe6e7f25ceff4ad8727d2b4f3d0627af19e4cbd784f477e5c4d0

Observation 875a69b9-3d43-4033-95b0-c1a2e6d7ba3b · inbound

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models cites this paper.

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:46:03.634188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T07:42:43.077644Z digest=sha256:ec99483a769908dd026471278945015071501073e2c86387b0de4dc8e8a24a48

Observation f2f20dca-d5f0-4a8d-8723-6e5aaa377292 · inbound

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models cites this paper.

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:00:39.123236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T20:56:37.533183Z digest=sha256:92c72348ea04dc1cebb374a9297d87095721554bad0067f13f2a9ccb18171f0e

Observation 26e9a64d-ee20-421a-8ce6-207d9ac59a60 · inbound

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models cites this paper.

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T09:49:44.597311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:49:44.597311Z digest=sha256:0d9513b4643ebf3403d9988874e97e70b2ccda6e324eefeee5646a2289508e5f

Observation 1235f466-957a-4462-850b-41217ba13cd5 · inbound

ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining cites this paper.

ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T16:09:50.207908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:09:50.207908Z digest=sha256:126c74db5ac25f010e7156fbc43c75340fd35a8e777ecde864fd6b94d577aedc

Observation 4331fecc-4aaf-471d-a446-b8f915a6a3d1 · inbound

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation cites this paper.

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:16:00.004189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T17:15:28.918204Z digest=sha256:6af4770f7752fd1c1396368e714f224cc9fab9fccab93803ceb7af59708cd97c

Observation c34293b0-51c8-46a8-8778-04e736d509da · inbound

EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses cites this paper.

EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:01:26.002498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-07T13:13:49.855219Z digest=sha256:fb2ba2ed76dcb23dca1d40c6874e5fb224329ab02667072007f904b6afd3b717

Observation f2cdfda3-401b-417d-b4a1-06c2b5264529 · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:56.208024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:0c01b44e28a1e637dc47c9e57efc67b82178a613cfd2a8357b58e36ab601798b

Observation 3400357e-5d39-49d9-a2bb-3ad06fbcff50 · inbound

EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection cites this paper.

EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:48:04.313900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T05:47:45.259359Z digest=sha256:b9ddd0bdf8f6fa24f52edf973f9576fad660c5e89334bbac94f11c795379a771

Observation fb0a5d89-b9d0-46f5-9587-71f8ac923026 · inbound

Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models cites this paper.

Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T04:56:05.289172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T04:54:35.815868Z digest=sha256:79241eb049e6403f05f4567a458d521cf3df6a9c14d2d1d21b2da1fcc9169d95

Observation a1f077c2-ab30-4e47-8e9c-e03bb13a4806 · inbound

SpeechDx: A Multi-Task Benchmark for Clinical Speech AI cites this paper.

SpeechDx: A Multi-Task Benchmark for Clinical Speech AI emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:28:48.482501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T03:09:07.295810Z digest=sha256:7e1eeed3341808c38123124862f63e4c3401e5ab16ad6eeb57464621dbb855ce

Observation e9feee97-ee97-4e78-a2f2-09a872113081 · inbound

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling cites this paper.

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:45:46.191735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-07-01T04:05:27.343684Z digest=sha256:fc2e4a134d077b8103214b85068dad2f62448d107ffb9a1954c131bed688ab1b

Observation 0f87d94c-cb0a-4502-9e79-0d9d381219bf · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 189

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.155193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:66aef6a5a9183c5ee6e1d13b337917aa5124db362032473d7f62cd39ba2191ae

Observation e98932ef-6fc6-4774-b43e-c1653e5faa6c · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:11.459753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:11.459753Z digest=sha256:3d3a27762a61233f516a48cbdda5fb50c175bfb46413654fa37d76c076ea0028

Observation 3e30aec3-6e10-456d-b113-a85825e87fb0 · inbound

AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks cites this paper.

AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-08T11:50:21.349519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:50:21.349519Z digest=sha256:f070be5dc48a7731a334469ab9b78cf597c64604b00025203cf80b00864535d2

Observation 4c81d5b5-1e73-428e-941f-4a7c2c63a788 · inbound

Is Self-Pretraining really useful to improve diagnosis in medical Time Series? cites this paper.

Is Self-Pretraining really useful to improve diagnosis in medical Time Series? emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:38.489642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:38.489642Z digest=sha256:f4484993d93b617096de1144cb80bbeb84d620bbb9eaa0ce3366752c1e443a48

Observation 91568008-84fe-4be4-84d6-d88c38f2ae7d · inbound

Is Self-Pretraining really useful to improve diagnosis in medical Time Series? cites this paper.

Is Self-Pretraining really useful to improve diagnosis in medical Time Series? emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:22.726620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:22.726620Z digest=sha256:f8245834c639b1c27712cc94b6335b8c83d9a03cc255f7245555640ac1c32040