Pith. sign in

Paper Citation Record · LEDGER

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025

As of 19 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2506.02088.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02088 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:40:11.848020Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:39:46.886300Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T11:40:12.056485Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact3
  • verified fuzzy23
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cbc91fd8-2696-431c-83e7-60c230cb87b5 · outbound

This paper cites Early SER relied on hand-crafted features but struggled with real- world generalization [2].

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Early SER relied on hand-crafted features but struggled with real- world generalization [2]

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.900210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:39:46.828246Z digest=sha256:b661ccef84798c3f3671c546c23afb9d2cd48f48dce440be6600cf286cd3c164

Observation b0fd2927-ab92-4ea7-b6d2-df576f3a353d · outbound

This paper cites Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:40:12.074820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:39:46.886300Z digest=sha256:eb15687f86f5b9197d039b3727d9f88b93b30f38ee82be26010b7db9552a5e4d

Observation d0e4715b-d7fe-41df-a432-fdb412559c61 · outbound

This paper cites The hidden states of the last layerLof the text encoder are denoted byZ L T (j)for positions j= 1,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 The hidden states of the last layerLof the text encoder are denoted byZ L T (j)for positions j= 1,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.725410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:39:46.937316Z digest=sha256:7602af7c6094ae6e456775ac036581e04f3c9d823ea43599c329e1fea79c160f

Observation 568c6db1-c9bf-445c-b9e1-92b773217c62 · outbound

This paper cites an unresolved cited work.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:40:58.649776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:39:46.962061Z digest=sha256:0782ce5bdb6776f94571e3831921f8313c74f07ea3be1ecff93f1070b5b8a3b4

Observation 33001ac9-da6f-4328-9c06-f8c518f3b824 · outbound

This paper cites We report results for unimodal speech models, bi- modal fusion with text, prosodic and spectral feature integra- tion.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 We report results for unimodal speech models, bi- modal fusion with text, prosodic and spectral feature integra- tion

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.509032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:39:47.019513Z digest=sha256:b3614f52f480c3d3c765b36b5d79b1b6553424f261d771a7fdbe836aed559bdb

Observation a15c6334-fba5-4b40-977e-75b1be96e4a9 · outbound

This paper cites Our evaluation of unimodal models demonstrated the strong performance of Whisper and XEUS, highlighting their robustness for SER in spontaneous speech.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Our evaluation of unimodal models demonstrated the strong performance of Whisper and XEUS, highlighting their robustness for SER in spontaneous speech

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.356842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:39:47.249227Z digest=sha256:1832520cd78c4967a8d0b7409c947ef9fe320487cf0fc0fc1ebeb707594e2014

Observation 8a356d0a-2d75-4065-b968-ab684529c6e7 · outbound

This paper cites We also thank the Artificial Intelligence Lab at Re- cod.ai, the Institute of Computing, University of Campinas.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 We also thank the Artificial Intelligence Lab at Re- cod.ai, the Institute of Computing, University of Campinas

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.217464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:39:47.380433Z digest=sha256:d5828a39dc9ebb59ffc0ba57fdc0f3cd6b693e5036fe46e46d9accd171bbbd3b

Observation bd61b554-8a33-449d-b640-1d6fbccfcf00 · outbound

This paper cites Affective computing mit press,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Affective computing mit press,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.093220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:39:47.456692Z digest=sha256:8520ebe1f8078bd016e6dbb260dad3943ecc8a34a60c8280914f9199850e7eb1

Observation 95edbc09-aad6-4d1d-ae6e-d62d26767b08 · outbound

This paper cites Iemocap: Interactive emotional dyadic motion capture database,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Iemocap: Interactive emotional dyadic motion capture database,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:47.566911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:47.566911Z digest=sha256:0d132a2595a06b0c6801f8e1dcd917653b8af590f2eeab41df708cb6ac79d1dc

Observation 575e1dbe-4310-41de-a8e4-6c02731c8573 · outbound

This paper cites Every rating matters: Joint learning of subjective labels and individual annotators for speech emotion classification,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Every rating matters: Joint learning of subjective labels and individual annotators for speech emotion classification,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:58.052371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:39:48.879938Z digest=sha256:f61bede58e7a372a4bfedd93846a08093b1109f5e00aa29a81cd260d77f60acb

Observation 5a988f1c-8399-4005-9333-b47543d5203f · outbound

This paper cites Speech emotion recognition using self-supervised features,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Speech emotion recognition using self-supervised features,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:51.667321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:51.667321Z digest=sha256:d6df89ccf9f49271aaeceed650d859c06b0b9491210ddfb53a892dd23f2d6926

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:57.908096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:39:51.789972Z digest=sha256:50d8b00db0e428446277e95377ac3815fbe0cb5f82b4e254f209d6f62a90d915

Observation f8ca3961-213f-4b2b-91ec-7514c4909d14 · outbound

This paper cites Speech emotion recognition with multi-task learning,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Speech emotion recognition with multi-task learning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:57.829289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:39:51.886757Z digest=sha256:138fc105a5cde5a34feca26c8b2ae4e7c906e257cb542b6d5b1982f63aeda255

Observation af2b8935-da7c-4af9-ab15-bbac4f956498 · outbound

This paper cites Improving speech emotion recogni- tion using self-supervised learning with domain-specific audiovi- sual tasks,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Improving speech emotion recogni- tion using self-supervised learning with domain-specific audiovi- sual tasks,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:57.755073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:39:51.999237Z digest=sha256:9d00bd7864334e8c6d38ab8088be8233c21886e2fbfa7fb9409d018bacfd7aa3

Observation dc666eda-a686-4562-ac18-c5b4e6d3a2df · outbound

This paper cites Odyssey 2024-speech emotion recognition challenge: Dataset, baseline framework, and results,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Odyssey 2024-speech emotion recognition challenge: Dataset, baseline framework, and results,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:57.518830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:39:52.100807Z digest=sha256:fbfef464f0efe0f2a34db7d95cdc20c5e3b564da0fa040129c2a78d12b3919d2

Observation e4a34ba8-918f-4d92-a487-dfd354d892a4 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:52.179588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:52.179588Z digest=sha256:9b6bacb83aa591fbb318e19f32dd47964fe126147f78aba2db03bad0f942a3aa

Observation cafa7854-99b7-47db-b70b-92195a3418cd · outbound

This paper cites Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:52.259720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:52.259720Z digest=sha256:7da1bc4a4f062a41bf6e77960ea198d3481867d0498e7086855d8cf63f8b68c2

Observation 82513e37-8794-4a55-986b-5d009f0a68c3 · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:52.329379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:52.329379Z digest=sha256:a010cb353f37178b8a6fdc3d3152713258b060ffd513e5d22885663460bc1ad8

Observation 814768dd-203d-4df1-bc62-9d272e49614e · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Robust speech recognition via large-scale weak supervision,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:52.399609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:52.399609Z digest=sha256:68dae0f063802b182e21fdf59ee98cca0f6827aadf807dceaedcd3cf808e6501

Observation 609fc84d-5d93-4fe4-a39b-a0eb1dd6b082 · outbound

This paper cites Towards robust speech representation learning for thousands of languages,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Towards robust speech representation learning for thousands of languages,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:52.465169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:52.465169Z digest=sha256:7631a9dc77baebc9b7f61a0c64000759f33a80981ec237af79f2300ada556041

Observation fb018c13-febc-43f5-9629-7294bac6d95a · outbound

This paper cites A robustly optimized BERT pre-training approach with post-training,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 A robustly optimized BERT pre-training approach with post-training,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:55.952387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:39:52.554243Z digest=sha256:104b050815b954d3bb85cfbd9172a4cf25a3607b896faf6144920e050fee22e4

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:40:12.025604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:39:52.777099Z digest=sha256:3e97b6be660dbd34ab7b45b002647469946e5d10507e3b646ed93ce847cd2697

Observation 6c227393-77bb-4ef8-ad33-c932444ebe03 · outbound

This paper cites Enhancing cross-language multimodal emotion recognition with dual attention transform- ers,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Enhancing cross-language multimodal emotion recognition with dual attention transform- ers,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:54.883253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:39:53.019255Z digest=sha256:59b1fa69b492d12b9af63aa6afd38f4578c398c7d30fae7348b865b3b99addfb

Observation 57f0aca7-3596-43b6-8b87-bf654988b716 · outbound

This paper cites Ced: Con- sistent ensemble distillation for audio tagging,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Ced: Con- sistent ensemble distillation for audio tagging,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:46.567879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:39:53.103227Z digest=sha256:3945bff9fed7aedc33a965d12085eb04b6349774178a570c93e083f6c668842f

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:40:11.978878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:39:53.213116Z digest=sha256:f1bddf4a7972dc0e12e7e87b3c1fe1acdb4c8f2c210312fd71f748c35b483a2d

Observation adf8d24e-0b38-43d3-8eb2-5c3958ba2a54 · outbound

This paper cites Graph attention networks,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Graph attention networks,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:36.044619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:39:53.374984Z digest=sha256:0d9c6ca67bfec803578945d9900323a5f3298f53785fcbafe999680185cddd3c

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:07.102687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:07.102687Z digest=sha256:d888ed308aef7cae84ce28eba8e77b36df17906b84c7682a0c96781172509501

Observation 2e2d95a2-2ebd-46ec-8686-ddc1d47f9b94 · outbound

This paper cites Espnet: End-to-end speech pro- cessing toolkit,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Espnet: End-to-end speech pro- cessing toolkit,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:35.916452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:40:08.043740Z digest=sha256:ecd532ac9afe5ac4f8704ad49484a70df4b006927d730040fd001e097f63c86d

Observation 571a2128-c921-4234-92c9-21618fd4f1e6 · outbound

This paper cites Less is more: Accu- rate speech recognition & translation without web-scale data,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Less is more: Accu- rate speech recognition & translation without web-scale data,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:09.312368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:09.312368Z digest=sha256:5223c4ffffecfd83208fa2b47196149ac47575df237feb04aa02f69771b33f07

Observation 738e5e1b-7056-46d6-ab15-de24cbf70da5 · outbound

This paper cites 1st place solution to odyssey emotion recognition chal- lenge task1: Tackling class imbalance problem,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 1st place solution to odyssey emotion recognition chal- lenge task1: Tackling class imbalance problem,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:55.870025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:40:09.642774Z digest=sha256:819227eb24655453cef0af286f304a7d4377d47043d76f7403807ac3d77c89fa

Observation c2ef2ada-153e-4b7b-b0f4-8a58e4ec2f0f · outbound

This paper cites Fundamental frequency ex- traction in speech emotion recognition,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Fundamental frequency ex- traction in speech emotion recognition,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:35.694011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:40:09.779236Z digest=sha256:b9858d1c855cc6b441654291f426212870799243de7b8fd8f2ae05a573ed4d7a

Observation b59c1520-ded0-4b57-84eb-320e3a08162b · outbound

This paper cites Autoregressive neural f0 model for statistical parametric speech synthesis,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Autoregressive neural f0 model for statistical parametric speech synthesis,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:31.875345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:40:09.869009Z digest=sha256:05d43f0bf2731685d4d3ca8d5bba087c175c2d29de3cd9732ce5492e852643cb

Observation f9bb7dae-4b5b-4664-94b4-2ea372393358 · outbound

This paper cites Rmvpe: A robust model for vocal pitch estimation in polyphonic music,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Rmvpe: A robust model for vocal pitch estimation in polyphonic music,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:23.794768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:40:09.900326Z digest=sha256:67c48981e0c43e4b3fee13323d615168c0f031ee1efc3df54b5c4acef252ca9d

Observation b5824415-23d1-4220-ac0e-c03afc0ea926 · outbound

This paper cites Enhancing skin can- cer diagnosis using swin transformer with hybrid shifted window- based multi-head self-attention and swiglu-based mlp,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Enhancing skin can- cer diagnosis using swin transformer with hybrid shifted window- based multi-head self-attention and swiglu-based mlp,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:23.219875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:40:09.918778Z digest=sha256:dbd62054ea71e434f64d4324fd74c8947937006ff9e96f61a79867e4d48901cb

Observation 61f5b1ec-e191-4293-afb4-86cceefebbc0 · outbound

This paper cites Searching for activation functions,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Searching for activation functions,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:12.252580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:40:09.980654Z digest=sha256:321904efac1e8a0a0980855ae0acae4dd31ac2944e5d51ae385f3d4ed3c835e0

Observation dec73175-a10b-4b58-8360-d52c414541d0 · outbound

This paper cites Decoupled weight de- cay regularization,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Decoupled weight de- cay regularization,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.029630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.029630Z digest=sha256:63324186da6a63315bf2c259a9e6eccda84643b43516bbd082e77f14d053121a

Observation 072fc2c8-d40f-4cfd-8b25-d6b33dbf5e32 · outbound

This paper cites Panns: Large-scale pretrained audio neural networks for audio pattern recognition,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Panns: Large-scale pretrained audio neural networks for audio pattern recognition,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.634578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.634578Z digest=sha256:26580a5f06283e98f9dcea4e7c216b3b26fd97ac71ea19effca04f3f8c0fe665

Observation b66fe9de-cb41-4a86-b269-35a9afcb68dc · outbound

This paper cites Focal loss for dense object detection,.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Focal loss for dense object detection,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:40:12.140211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:40:11.768796Z digest=sha256:272abf77ff0965e1d8793a2aec6edd8551f9ede364a0e701384a0f5e918cfede

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.848020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.848020Z digest=sha256:bebf79ededfd2e479a5f287dc7c8fa65f80b41dffaa79c4a6d8ab6e109b7f11c

Pith citing papers

Observation b0fd2927-ab92-4ea7-b6d2-df576f3a353d · inbound

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 cites this paper.

Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025 Enhancing Speech Emotion Recognition with Graph-Based Multimodal Fusion and Prosodic Features for the Speech Emotion Recognition in Naturalistic Conditions Challenge at Interspeech 2025

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:40:12.074820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T11:39:46.886300Z digest=sha256:eb15687f86f5b9197d039b3727d9f88b93b30f38ee82be26010b7db9552a5e4d