Pith. sign in

Paper Citation Record · LEDGER

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025

As of 11 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2506.02088.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02088 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:40:11.848020Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:39:46.886300Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T11:40:12.056485Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact3
  • verified fuzzy23
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cbc91fd8-2696-431c-83e7-60c230cb87b5 · outbound

This paper cites Early SER relied on hand-crafted features but struggled with real- world generalization [2].

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Early SER relied on hand-crafted features but struggled with real- world generalization [2]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.900210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:39:46.828246Z digest=sha256:7ceae2f94a1d9f5cd2cb9e5230795decde8450cb2c3b625ffd9bfaa6ce00aeff

Observation b0fd2927-ab92-4ea7-b6d2-df576f3a353d · outbound

This paper cites Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:40:12.074820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:39:46.886300Z digest=sha256:49dfdf948d3aa4af199cf63e3dcbc220782c10e7d3e86f0f780113892ef00cd5

Observation d0e4715b-d7fe-41df-a432-fdb412559c61 · outbound

This paper cites The hidden states of the last layerLof the text encoder are denoted byZ L T (j)for positions j= 1,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 The hidden states of the last layerLof the text encoder are denoted byZ L T (j)for positions j= 1,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.725410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:39:46.937316Z digest=sha256:87fe0003b4c7e8b1a55c1222eab605cda102e6a002f565ef00eafa09e88ee1d4

Observation 568c6db1-c9bf-445c-b9e1-92b773217c62 · outbound

This paper cites an unresolved cited work.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:40:58.649776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:39:46.962061Z digest=sha256:b88601d98a4c135b234d4456b3bac884f079f1823d9b556a8b357ee18a3e260b

Observation 33001ac9-da6f-4328-9c06-f8c518f3b824 · outbound

This paper cites We report results for unimodal speech models, bi- modal fusion with text, prosodic and spectral feature integra- tion.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 We report results for unimodal speech models, bi- modal fusion with text, prosodic and spectral feature integra- tion

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.509032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:39:47.019513Z digest=sha256:3f8fcb40bc95f3fa6aa423acb05a65230bb7d3a104fa7ce54999da2fdc33ef25

Observation a15c6334-fba5-4b40-977e-75b1be96e4a9 · outbound

This paper cites Our evaluation of unimodal models demonstrated the strong performance of Whisper and XEUS, highlighting their robustness for SER in spontaneous speech.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Our evaluation of unimodal models demonstrated the strong performance of Whisper and XEUS, highlighting their robustness for SER in spontaneous speech

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.356842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:39:47.249227Z digest=sha256:5610d55e67d1a0cdb2a83440b3b5a1c257eb475a74c63f4b4d6a580124a2b7e9

Observation 8a356d0a-2d75-4065-b968-ab684529c6e7 · outbound

This paper cites We also thank the Artificial Intelligence Lab at Re- cod.ai, the Institute of Computing, University of Campinas.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 We also thank the Artificial Intelligence Lab at Re- cod.ai, the Institute of Computing, University of Campinas

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.217464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:39:47.380433Z digest=sha256:f961489c1ca44513bc38dde95c6217332aceb4390e9d575e54f18001950baa2a

Observation bd61b554-8a33-449d-b640-1d6fbccfcf00 · outbound

This paper cites Affective computing mit press,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Affective computing mit press,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.093220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:39:47.456692Z digest=sha256:0e9efce62cd79826fcdaf7665386e982869ed3bf9942b5dd6e8f430b7ed671f5

Observation 95edbc09-aad6-4d1d-ae6e-d62d26767b08 · outbound

This paper cites Iemocap: Interactive emotional dyadic motion capture database,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Iemocap: Interactive emotional dyadic motion capture database,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:47.566911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:47.566911Z digest=sha256:c30ac61afb97f5ea836dd4c32fb13f234b1b8413e83aea16c83906e54dee02a7

Observation 575e1dbe-4310-41de-a8e4-6c02731c8573 · outbound

This paper cites Every rating matters: Joint learning of subjective labels and individual annotators for speech emotion classification,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Every rating matters: Joint learning of subjective labels and individual annotators for speech emotion classification,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.052371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:39:48.879938Z digest=sha256:aeeb7f9db83943f94319a43f9a91da41793382a7349c3cfeb8fbaff6ffb88827

Observation 5a988f1c-8399-4005-9333-b47543d5203f · outbound

This paper cites Speech emotion recognition using self-supervised features,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Speech emotion recognition using self-supervised features,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:51.667321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:51.667321Z digest=sha256:a3b561e56a1168c7b5f08b11035a6be811253016b95f8995821137bfc660d1c3

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:57.908096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:39:51.789972Z digest=sha256:efa5d84025d4e41a9d8b07a2c6a3ab384351678249cf6903de2d27d1ee3a43b4

Observation f8ca3961-213f-4b2b-91ec-7514c4909d14 · outbound

This paper cites Speech emotion recognition with multi-task learning,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Speech emotion recognition with multi-task learning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:57.829289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:39:51.886757Z digest=sha256:777c06bf8cc1515b69416ce08a80a571823ab64afbef212263154d6efedd8515

Observation af2b8935-da7c-4af9-ab15-bbac4f956498 · outbound

This paper cites Improving speech emotion recogni- tion using self-supervised learning with domain-specific audiovi- sual tasks,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Improving speech emotion recogni- tion using self-supervised learning with domain-specific audiovi- sual tasks,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:57.755073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:39:51.999237Z digest=sha256:2ec21953fe03892c3301ffa3759670886a83278fbd96e358b497e2baf29ede00

Observation dc666eda-a686-4562-ac18-c5b4e6d3a2df · outbound

This paper cites Odyssey 2024-speech emotion recognition challenge: Dataset, baseline framework, and results,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Odyssey 2024-speech emotion recognition challenge: Dataset, baseline framework, and results,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:57.518830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:39:52.100807Z digest=sha256:d29125d36562411e90939e312dcf3c9d5fd20bd3717ef866f4d8a9519be6e2d0

Observation e4a34ba8-918f-4d92-a487-dfd354d892a4 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:52.179588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:52.179588Z digest=sha256:c249fbf50e5d7169b9c008de45cf88b97ee8f228f9f62c00398dcf4e07ada208

Observation cafa7854-99b7-47db-b70b-92195a3418cd · outbound

This paper cites Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:52.259720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:52.259720Z digest=sha256:97e20eb22cfd2301a4a1247b1c63043132954847ccd39bf9044af6c40e98dc58

Observation 82513e37-8794-4a55-986b-5d009f0a68c3 · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:52.329379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:52.329379Z digest=sha256:88523aba39a7533e82ec68c9b0dacc52d14e711ff7643badf5e3a6bc2f9371f4

Observation 814768dd-203d-4df1-bc62-9d272e49614e · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Robust speech recognition via large-scale weak supervision,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:52.399609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:52.399609Z digest=sha256:1f79e4a311ea0f1354ba4110cfd2ff58f7d5495e9435ff44fbe572d473d396c7

Observation 609fc84d-5d93-4fe4-a39b-a0eb1dd6b082 · outbound

This paper cites Towards robust speech representation learning for thousands of languages,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Towards robust speech representation learning for thousands of languages,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:52.465169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:52.465169Z digest=sha256:854249ef89ad68ebb708e08647cbcb9e70111cc3ef613b3d7f7c0d50c28e2b31

Observation fb018c13-febc-43f5-9629-7294bac6d95a · outbound

This paper cites A robustly optimized BERT pre-training approach with post-training,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 A robustly optimized BERT pre-training approach with post-training,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:55.952387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:39:52.554243Z digest=sha256:60818d32a76d28500f3950cfe9aa3ec9908784eca1ee81e8e35cd220394db959

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:40:12.025604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:39:52.777099Z digest=sha256:b0879a3a2968ed6bd7f926828907e1aa4655684599edd6d82697d7cf08cc8447

Observation 6c227393-77bb-4ef8-ad33-c932444ebe03 · outbound

This paper cites Enhancing cross-language multimodal emotion recognition with dual attention transform- ers,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Enhancing cross-language multimodal emotion recognition with dual attention transform- ers,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:54.883253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:39:53.019255Z digest=sha256:6e29ae3ad18663873c6f66aa659223549ee584f4440d17ca264f7bd2fa83f343

Observation 57f0aca7-3596-43b6-8b87-bf654988b716 · outbound

This paper cites Ced: Con- sistent ensemble distillation for audio tagging,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Ced: Con- sistent ensemble distillation for audio tagging,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:46.567879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:39:53.103227Z digest=sha256:e4dc91a534357dfa3e700ca841107457f91d4b57303d2d3207e2ae5c173f2857

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:40:11.978878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:39:53.213116Z digest=sha256:efedbee36e296ab471b94c2c4c3f3c4c808c89bacaa434c6aba76db80c6b9158

Observation adf8d24e-0b38-43d3-8eb2-5c3958ba2a54 · outbound

This paper cites Graph attention networks,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Graph attention networks,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:36.044619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:39:53.374984Z digest=sha256:0fe51085ac6c2405940b112e2df3c37be289a68fe4bf528eb285279b199d0eaf

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:07.102687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:07.102687Z digest=sha256:954e4b0dd73db4b5b72ec18039e78481afc695ea2999d0cc024fdf5e0fc3856f

Observation 2e2d95a2-2ebd-46ec-8686-ddc1d47f9b94 · outbound

This paper cites Espnet: End-to-end speech pro- cessing toolkit,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Espnet: End-to-end speech pro- cessing toolkit,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:35.916452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:40:08.043740Z digest=sha256:98c8475f1dd77db9dd18fd70031805f8e652f8c24983455a65474a08363140f8

Observation 571a2128-c921-4234-92c9-21618fd4f1e6 · outbound

This paper cites Less is more: Accu- rate speech recognition & translation without web-scale data,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Less is more: Accu- rate speech recognition & translation without web-scale data,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:09.312368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:09.312368Z digest=sha256:a70374d8d484f1672554a21dbf3031a3df0ce7064488b88c314ac5492dd37eed

Observation 738e5e1b-7056-46d6-ab15-de24cbf70da5 · outbound

This paper cites 1st place solution to odyssey emotion recognition chal- lenge task1: Tackling class imbalance problem,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 1st place solution to odyssey emotion recognition chal- lenge task1: Tackling class imbalance problem,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:55.870025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:40:09.642774Z digest=sha256:db464a6f636a49566fe4403fdd16fbfc8936245dcce4c86f2b313d2b37524ca5

Observation c2ef2ada-153e-4b7b-b0f4-8a58e4ec2f0f · outbound

This paper cites Fundamental frequency ex- traction in speech emotion recognition,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Fundamental frequency ex- traction in speech emotion recognition,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:35.694011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:40:09.779236Z digest=sha256:f20f49e97f35ecd32eb45bea1b00493f43c4867a829781258eae7fbf253cdf13

Observation b59c1520-ded0-4b57-84eb-320e3a08162b · outbound

This paper cites Autoregressive neural f0 model for statistical parametric speech synthesis,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Autoregressive neural f0 model for statistical parametric speech synthesis,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:31.875345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:40:09.869009Z digest=sha256:bf16eb6b4d335a34f7c4a9e0337b5ced316db3932db59921cc713aa7b12a43f1

Observation f9bb7dae-4b5b-4664-94b4-2ea372393358 · outbound

This paper cites Rmvpe: A robust model for vocal pitch estimation in polyphonic music,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Rmvpe: A robust model for vocal pitch estimation in polyphonic music,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:23.794768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:40:09.900326Z digest=sha256:e018c2e1494fd9af5c68f6ff681e5e3fe98ea79a91ab1e154391561f6ea4076e

Observation b5824415-23d1-4220-ac0e-c03afc0ea926 · outbound

This paper cites Enhancing skin can- cer diagnosis using swin transformer with hybrid shifted window- based multi-head self-attention and swiglu-based mlp,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Enhancing skin can- cer diagnosis using swin transformer with hybrid shifted window- based multi-head self-attention and swiglu-based mlp,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:23.219875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:40:09.918778Z digest=sha256:a0119e83d01c4afcceee2e24a65a950b172299db22953bafa2d8a80cf508933e

Observation 61f5b1ec-e191-4293-afb4-86cceefebbc0 · outbound

This paper cites Searching for activation functions,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Searching for activation functions,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:12.252580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:40:09.980654Z digest=sha256:b043878aa0247e660ae3be54f3a9bf3513260d1b45142836be599bb353a8ab47

Observation dec73175-a10b-4b58-8360-d52c414541d0 · outbound

This paper cites Decoupled weight de- cay regularization,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Decoupled weight de- cay regularization,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.029630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.029630Z digest=sha256:7de042d5ba99ad8888dc65583ba40f1401f33ddf7777eee54c8ad77cdbb182a3

Observation 072fc2c8-d40f-4cfd-8b25-d6b33dbf5e32 · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Panns: Large-scale pretrained audio neural networks for audio pattern recognition,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.634578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.634578Z digest=sha256:b23b3c1b841929f050c7091edd40fb44e7ef35fe9699ccb6f578f304f75eaa72

Observation b66fe9de-cb41-4a86-b269-35a9afcb68dc · outbound

This paper cites Focal loss for dense object detection,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Focal loss for dense object detection,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:12.140211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:40:11.768796Z digest=sha256:9768ed8630b9fbea8a1107110f1b4c44d48c0ab7c5bc16d94aebba2af912195e

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.848020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.848020Z digest=sha256:86dd15c13e7e70257c9c26c4ef2ccf8c7ba6701b7e7a9127190e93a0aaeb4045

Pith citing papers

Observation b0fd2927-ab92-4ea7-b6d2-df576f3a353d · inbound

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 cites this paper.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:40:12.074820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-07T11:39:46.886300Z digest=sha256:49dfdf948d3aa4af199cf63e3dcbc220782c10e7d3e86f0f780113892ef00cd5