Pith. sign in

Paper Citation Record · LEDGER

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition

As of 15 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2412.05558.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.05558 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:38:01.992997Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4b7eceb0-f550-487e-a270-1ebc06d8669b · outbound

This paper cites Pattern recognition, 44(3):572–587 (2011).

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition Pattern recognition, 44(3):572–587 (2011)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:02.798188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:38:01.808303Z digest=sha256:1eb30ad0362d98f61b8783631cdc37547fe685cc925bab9ccadcd31f98804fb2

Observation 75c1d878-9e3e-4efc-b926-49b24acc88c6 · outbound

This paper cites In 2021 7th International Conference on Computer and Communications (ICCC), pages 514–518 (2021).

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition In 2021 7th International Conference on Computer and Communications (ICCC), pages 514–518 (2021)

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:02.779794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:38:01.824900Z digest=sha256:dd035a48dfc629ec7dc65c618c7c7820ac0cfd678e646e60e5786d77122d1c4e

Observation 891e4fb8-ae2d-4add-a0a3-bc01825ea43a · outbound

This paper cites In Third international conference on natural computation (ICNC 2007), volume 5, pages 809–813 (2007).

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition In Third international conference on natural computation (ICNC 2007), volume 5, pages 809–813 (2007)

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:02.764821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:38:01.830165Z digest=sha256:b586d34146ac2620bf5397a896f43d4b6a891afa2630adfe84c040e6a6ec86ed

Observation 90355c7f-f9aa-42e6-97bf-fb0fcd779427 · outbound

This paper cites In 2022 IEEE 8th World Forum on Internet of Things (WF-IoT), pages 1–6 (2022).

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition In 2022 IEEE 8th World Forum on Internet of Things (WF-IoT), pages 1–6 (2022)

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:02.731202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:38:01.840892Z digest=sha256:50035838b2b07045709b8f045cf9267eec7e57bf6385104eecf1d1a612a8f907

Observation 8f0d4748-a828-46ff-a306-34bb81aa4fdb · outbound

This paper cites In Applied Information Processing Systems: Proceedings of ICCET 2021, pages 83–92.

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition In Applied Information Processing Systems: Proceedings of ICCET 2021, pages 83–92

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:02.711651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:38:01.846676Z digest=sha256:bd70a597a8b0d1bfb8004e63f7147acc925ef285defcda40095b44000069389a

Observation 4258f317-d558-40b5-8bf0-51206ae62049 · outbound

This paper cites In 2017 seventh international conference on Affective Computing and Intelligent Interaction Workshops and Demos (ACIIW), pages 79–80 (2017).

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition In 2017 seventh international conference on Affective Computing and Intelligent Interaction Workshops and Demos (ACIIW), pages 79–80 (2017)

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:02.685887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:38:01.852707Z digest=sha256:ddc4935d10150aa21747f1531894dee2308ca2cdde6b72b539bbe056cf714019

Observation c81c9578-815a-42b1-a169-c0e68cb2537d · outbound

This paper cites In Proceedings of the AAAI conference on artificial intelligence, volume 30 (2016).

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition In Proceedings of the AAAI conference on artificial intelligence, volume 30 (2016)

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:02.657784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:38:01.859494Z digest=sha256:071e08b32d87f6648eb3b453724745eef047a47bfa3c9c686b7167e353bce980

Observation d22bed2b-83f6-4328-9592-841e4a64bdae · outbound

This paper cites Multimodal emotion recognitionusingdeeplearning.

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition Multimodal emotion recognitionusingdeeplearning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:02.636722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:38:01.866378Z digest=sha256:c049db487ac918977fd113a1acc9a5363bcf06cd916279ff4035e79f20062a5e

Observation 11d941ef-eedb-499e-b5e7-c41d451b3a85 · outbound

This paper cites In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 6269– 6273 (2021).

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 6269– 6273 (2021)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:02.619088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:38:01.872690Z digest=sha256:a0be677fd4316ca4aa36514cea57c6260cf7c3ef9c220a3ffc85e95343d4f6c4

Observation e592e0b5-83b1-4628-a472-b0c67f6711fb · outbound

This paper cites In AVSP 2001-International Conference on Auditory-Visual Speech Processing (2001).

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition In AVSP 2001-International Conference on Auditory-Visual Speech Processing (2001)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:02.583663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:38:01.880060Z digest=sha256:afaa9b3ece87f30ef089df1b46f3fda74f03c324d9b1b7ac488b9fe0f4fefc25

Observation 93ff3b66-8bd4-479b-a77d-f3712b9c5ae7 · outbound

This paper cites In Proceedings of the conference.

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition In Proceedings of the conference

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:02.567658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:38:01.888215Z digest=sha256:ad52074e10834795a9f0ea43afbf17720f21bf84325f6bbccc984dc6a25f7583

Observation 9b0bb16a-6534-4836-848a-bc5d017117c6 · outbound

This paper cites IEEE Transactions on Multimedia (2022).

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition IEEE Transactions on Multimedia (2022)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:02.528916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:38:01.894927Z digest=sha256:3a7c80a0880ffa2d362b083f637f7a2a8a998b7e0802dad0d2e7abf44b75991b

Observation 139cbc43-57bf-432a-bfbd-a019b089ad22 · outbound

This paper cites IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:3592–3603 (2021).

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition IEEE/ACM Transactions on Audio, Speech, and Language Processing, 29:3592–3603 (2021)

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:02.495930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:38:01.900223Z digest=sha256:ffacb0b7963dd5179a3354ecb012ea79df2bae630092463e6634acc31e11fdef

Observation 027aeed2-1e83-4d08-9ca8-778caace0dc1 · outbound

This paper cites In Proceedings of the 28th ACM international conference on multimedia, pages 1122–1131 (2020).

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition In Proceedings of the 28th ACM international conference on multimedia, pages 1122–1131 (2020)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:02.467416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:38:01.907033Z digest=sha256:a9ff8f8d0df89ae1b2b8ec4a092438104a92925ffe919c69283e8862daeb65d0

Observation 9c108565-c1c1-4f0c-9044-b4c0cd9116f3 · outbound

This paper cites DialogueTRM: Exploring the Intra- and Inter-Modal Emotional Behaviors in the Conversation.

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition DialogueTRM: Exploring the Intra- and Inter-Modal Emotional Behaviors in the Conversation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:01.914107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:38:01.914107Z digest=sha256:6fc7619196cab85d23318bd3d15432cdb9244107c1010d6bcb7a47fb2a120449

Observation 2d01b38e-ec0a-413d-ab04-70287b931bbf · outbound

This paper cites MMGCN: Multimodal Fusion via Deep Graph Convolution Network for Emotion Recognition in Conversation.

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition MMGCN: Multimodal Fusion via Deep Graph Convolution Network for Emotion Recognition in Conversation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:01.920850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:38:01.920850Z digest=sha256:b915b4e72cec894a4562a1df5da2c1d3bdee2932bba6f63558360ffc268fa572

Observation 1231d2dc-80e6-4105-a184-1b9cce8eaa8e · outbound

This paper cites In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 7037– 7041 (2022).

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 7037– 7041 (2022)

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:02.432840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:38:01.929176Z digest=sha256:511518b3ee972fccf305fc5036ba368c2c36c1a524306a703563e0227e271627

Observation f5f342c8-ee45-411c-b7e5-b6e75c0a8ebc · outbound

This paper cites In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4652–4661 (2022).

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4652–4661 (2022)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:02.413287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:38:01.935087Z digest=sha256:be467078a21ef1fe2c71979b3f0be03d29458b370779877591eb44155b10a133

Observation fc0c21ce-1eb2-46e7-9463-fa96279327bc · outbound

This paper cites Advances in neural information processing systems, 33:12449–12460 (2020).

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition Advances in neural information processing systems, 33:12449–12460 (2020)

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:02.380907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:38:01.941879Z digest=sha256:12b9fcdeea8c1f4ae1675753db586d33900b0467601075ed47247e1a6265378a

Observation 5fd2a026-115f-43ff-b3f0-3017e4ab8058 · outbound

This paper cites Centralized feature pyramid for object detection.

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition Centralized feature pyramid for object detection

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:02.338648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:38:01.946381Z digest=sha256:7b1cf9b9c670f8236e50b5805fffb38da35c6de53e79e6f92f3c9cdb00a04755

Observation 21ac3475-0b06-40ac-b452-8d20b1e91306 · outbound

This paper cites Language resources and evaluation, 42:335–359 (2008).

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition Language resources and evaluation, 42:335–359 (2008)

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:02.319049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:38:01.951740Z digest=sha256:b7cd50d026298ac092b3cd42bd1c6d11c237308a8e931493f336d295e3afe8c2

Observation 807482db-a36f-4fea-bd0d-bbce639f42f7 · outbound

This paper cites MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations.

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:01.958786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:38:01.958786Z digest=sha256:0cbc2cd19a1d418cbd9273650a2c920e74baa75c5e4c2d8037c1028e0109db35

Observation 5ec6c1a4-8b09-4d32-ab03-0300abee401d · outbound

This paper cites In Proceedings of the 28th International Conference on Computational Linguistics, pages 4190–4200 (2020).

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition In Proceedings of the 28th International Conference on Computational Linguistics, pages 4190–4200 (2020)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:02.289001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:38:01.965587Z digest=sha256:2c56eb31653012a91b996cdf742a680301a29a628bdf4fcc0b866e763220997f

Observation 8754cb24-90ab-4a5e-9963-46622dff684b · outbound

This paper cites In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 13789–13797 (2021).

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition In Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, pages 13789–13797 (2021)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:02.261810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:38:01.972177Z digest=sha256:8cd2669b7727edfc74f3b0b9ed22132899aece3d5364a39ffeee5c697502e863

Observation ce9f55e5-8624-468e-88f7-161c0aa4d89a · outbound

This paper cites COGMEN: COntextualized GNN based Multimodal Emotion recognitioN.

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition COGMEN: COntextualized GNN based Multimodal Emotion recognitioN

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:01.978552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:38:01.978552Z digest=sha256:bb2b4031480d144beef6cb2a4c74f8e322894050859856d2cf0e083df0910f39

Observation b9b28c35-92b2-43d9-b28a-dd8b96bc7a1a · outbound

This paper cites Neural Computing and Applica- tions, pages 1–14 (2023).

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition Neural Computing and Applica- tions, pages 1–14 (2023)

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:02.228561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T20:38:01.984168Z digest=sha256:5b95de2b5aaa69ad02dbcc1bd84435c3563afaa6f19a6656ded23b6b3a62f387

Observation 65dcd5cc-aecb-498f-8bf5-4a31b6e0ac83 · outbound

This paper cites UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition.

WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition UniMSE: Towards Unified Multimodal Sentiment Analysis and Emotion Recognition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:01.992997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:38:01.992997Z digest=sha256:7fee996f47ee986f94b95ba14949dd67bcc3c9dbd7965ab324044235662bd348

Pith citing papers

No inbound Pith citation observations are available.