Pith. sign in

Paper Citation Record · LEDGER

emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2312.15185.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.15185 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:49:29.542314Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T18:28:48.480697Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8c1e784a-8609-4a47-b1bb-6dc4731fc21d · inbound

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS cites this paper.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-23T20:33:25.612916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:44fb673c436dbf1113aefddaa7c0e9e713e9b62f601c932be652e0103345606f

Observation eb7ebb49-ba2d-4e51-805d-2b1b66e5e02f · inbound

Optimizing Multilingual Text-To-Speech with Accents & Emotions cites this paper.

Optimizing Multilingual Text-To-Speech with Accents & Emotions emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:29.542314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:29.542314Z digest=sha256:c2e82b1d31d7f1a84ade0d580b3c5031f0258a62c38197aae172a90c6de7cbca

Observation 382fc309-fb3f-43a4-85a1-672827b718f9 · inbound

MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled Embedding cites this paper.

MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled Embedding emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:16:57.081692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:16:57.081692Z digest=sha256:0292a99f2906ac12207c29b76225051aff76728cfbc2a7c8da8c405640847703

Observation f780c7fd-ff9e-48e9-abaf-946cd67323b8 · inbound

BoSS: Beyond-Semantic Speech cites this paper.

BoSS: Beyond-Semantic Speech emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:50:04.223838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:50:04.223838Z digest=sha256:c8621c0a9551d10c7b250ef3f353bc087e04766683bcdccd0560724d6c4e667b

Observation 301254d5-bccf-44a0-9c29-c9adc82369e6 · inbound

Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors cites this paper.

Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T11:50:25.414247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:50:25.414247Z digest=sha256:7220c364facbc487aacfa5503d4cd0f01a454d7792a31632099bfcfdee0dd0bd

Observation fd2dd173-740f-4bdb-b45a-89eb39ad134e · inbound

Listening to the Unspoken: Exploring "365" Aspects of Multimodal Interview Performance Assessment cites this paper.

Listening to the Unspoken: Exploring "365" Aspects of Multimodal Interview Performance Assessment emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T11:29:41.195435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:29:41.195435Z digest=sha256:4192fac94e787b4ebebcd87a449874607ddc5201703d4392e5819acfcd295208

Observation d904b568-df3c-4f04-9e7c-0eceb1425215 · inbound

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks cites this paper.

MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:41:59.596971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T02:41:52.996457Z digest=sha256:a176ae142c49ca7833e27d7cb9763997a57f5121daef5809fd8288415ba5d603

Observation f5dcc674-87c2-4a7e-a7b8-417df35ec02d · inbound

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody cites this paper.

Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T22:34:43.524538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:34:43.524538Z digest=sha256:d74d339fe3b08bc6c1e2208a1f6c454eedd757af34e72605cbf38d30ecd31405

Observation 4f720ed5-cb13-4286-845c-a4e5e3b0d86e · inbound

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents cites this paper.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.444581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.444581Z digest=sha256:688f0065ceeba4faf38e32db7add174f56c113550ccb72ef21b58c305879a41f

Observation 875a69b9-3d43-4033-95b0-c1a2e6d7ba3b · inbound

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models cites this paper.

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:46:03.634188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T07:42:43.077644Z digest=sha256:ecd847bb02b6f09000c302826d76870cf94ccfb36831cf9f046468df3abd998d

Observation f2f20dca-d5f0-4a8d-8723-6e5aaa377292 · inbound

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models cites this paper.

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:00:39.123236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T20:56:37.533183Z digest=sha256:7413d3c837c86c6c02102577411148bff091f578fd06725e661df3230d1db670

Observation 26e9a64d-ee20-421a-8ce6-207d9ac59a60 · inbound

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models cites this paper.

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T09:49:44.597311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:49:44.597311Z digest=sha256:8b3182993d60748ede835e28626bc80d5bc33ff079dce326064a79c89c988dbb

Observation 1235f466-957a-4462-850b-41217ba13cd5 · inbound

ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining cites this paper.

ParaSpeechCLAP: A Dual-Encoder Speech-Text Model for Rich Stylistic Language-Audio Pretraining emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T16:09:50.207908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:09:50.207908Z digest=sha256:c67b3ecc3f45ec925f7c4f78d597d45cd96df19465954f3b50d30025ac7e4fa2

Observation 4331fecc-4aaf-471d-a446-b8f915a6a3d1 · inbound

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation cites this paper.

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:16:00.004189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:15:28.918204Z digest=sha256:aeb0bc5297aa85e8c98d8fca230653dd325b8da1d0308979b48d94ece244fa53

Observation c34293b0-51c8-46a8-8778-04e736d509da · inbound

EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses cites this paper.

EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:01:26.002498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T13:13:49.855219Z digest=sha256:d3837522d0c14d9b0d896f6a35eeb84e832cd4925d8e733f0d7bb449ed78d6c1

Observation f2cdfda3-401b-417d-b4a1-06c2b5264529 · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:56.208024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:48a0a361277b9388fcd979852daac582f9fc7f58df26296b235849e3e89dc9ae

Observation 3400357e-5d39-49d9-a2bb-3ad06fbcff50 · inbound

EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection cites this paper.

EMO-BOOST: Emotion-Augmented Audio-Visual Features for Improved Generalization in Deepfake Detection emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:48:04.313900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T05:47:45.259359Z digest=sha256:23594e8ea7675e524b886cac59b3aa96378180e10dc794795ef8b62272717439

Observation fb0a5d89-b9d0-46f5-9587-71f8ac923026 · inbound

Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models cites this paper.

Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-22T04:56:05.289172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T04:54:35.815868Z digest=sha256:512b3c3ec7ccd059cbaee79d9c366183848c7e3a5cde2c55cbefc1e56ca384c2

Observation a1f077c2-ab30-4e47-8e9c-e03bb13a4806 · inbound

SpeechDx: A Multi-Task Benchmark for Clinical Speech AI cites this paper.

SpeechDx: A Multi-Task Benchmark for Clinical Speech AI emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:28:48.482501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T03:09:07.295810Z digest=sha256:1fa2b9b4135725893c5f5984681f5a20be5fe975024de01d65a7a4fb759e1445

Observation e9feee97-ee97-4e78-a2f2-09a872113081 · inbound

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling cites this paper.

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:45:46.191735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T04:05:27.343684Z digest=sha256:fee577d83c01fd0781b63c3126cbadbbb8553c527ab1b247068380a3e5fa1f29

Observation 0f87d94c-cb0a-4502-9e79-0d9d381219bf · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 189

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.155193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:2ee8c2231b3ba6aeb4f5ff0bd56e70f64d825cff20f28f2eae2521fdf536a130

Observation e98932ef-6fc6-4774-b43e-c1653e5faa6c · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:11.459753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:11.459753Z digest=sha256:097d891568f0e7459366a5afd675ea506aef9e69a9aeca29b4796c0e9fe3a248