Pith. sign in

Paper Citation Record · LEDGER

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues

As of 9 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 1 inbound Pith citation observation for arXiv:2510.07355.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.07355 v2

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T11:06:50.337605Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:56:58.344964Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8b51d3a7-2e73-4cf8-bf6d-d8a2ba21af06 · outbound

This paper cites Emotional communication in speech and music: The role of melodic and rhythmic contrasts,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Emotional communication in speech and music: The role of melodic and rhythmic contrasts,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:43.632154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:43.632154Z digest=sha256:bfe1d3d6ee5a657a4308f87daf63f19eeefaff987450f1003f8d8d5dac472f31

Observation 6d4917ed-eef9-49dc-b21f-0e0d0879834c · outbound

This paper cites Effects of variation in emotional tone of voice on speech perception,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Effects of variation in emotional tone of voice on speech perception,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:43.675176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:43.675176Z digest=sha256:744287571d03da1104c6702bcdc4dd689e5f999ffcab7411b7fb0fb40a286416

Observation 56007e68-f4c7-4b26-8fd2-92161abff7f0 · outbound

This paper cites Analysis of emotion recognition using facial ex- pressions, speech and multimodal information,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Analysis of emotion recognition using facial ex- pressions, speech and multimodal information,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:43.765144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:43.765144Z digest=sha256:e32080dc0d852adc75bab14756350ca37a4c4a6b7322708816006b6dafbe3b57

Observation 50d76174-2266-43d9-a942-edcaf98c9b63 · outbound

This paper cites Language models are few-shot learners,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Language models are few-shot learners,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:43.901861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:43.901861Z digest=sha256:cdcfc882fe37910287abf44fb8791ce2e04d8daab379621e76bcd54accc389e3

Observation 55ce28e0-4813-4353-8197-81b50a8f88de · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:44.086529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:44.086529Z digest=sha256:2c782034e649d9f4bde762bc9e830a93dfb1c5344baf903ca4652e2f00b16875

Observation 1b8bbee8-5951-4f8c-a92d-3c152e36ae7c · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:44.257304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:44.257304Z digest=sha256:0594455e08c4a9c50bf4bdafc04c74337eedd27463c1acc713f520a1b2f60341

Observation 875ecdf8-0960-418d-b24a-1bc98129c738 · outbound

This paper cites Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:44.360493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:44.360493Z digest=sha256:a365978dabebff2d1c2dee16d412c0b11bac8490934bce468d8039215a557296

Observation 6d8c3881-632a-40ca-9478-ee325d43010f · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:44.532704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:44.532704Z digest=sha256:7e86d0d12ffc9e11f36a340df78044e79c8d09bdda7bf620915175166b18760c

Observation 88a46cbb-0276-4a0c-b631-cfcac641ba3a · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Moshi: a speech-text foundation model for real-time dialogue

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:44.649264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:44.649264Z digest=sha256:1ef828dce77dd44e5366c21361b3e5c768450f04922fe9eeddd58b198ef3c3ec

Observation d21cad1c-f787-4e84-98ac-87728cf598ea · outbound

This paper cites EMO-Reasoning: Benchmarking Emotional Rea- soning Capabilities in Spoken Dialogue Systems,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues EMO-Reasoning: Benchmarking Emotional Rea- soning Capabilities in Spoken Dialogue Systems,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:44.765981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:44.765981Z digest=sha256:0210bc699a1e7197b40bef03774e601f067611986851199fa7455a4e23ccd1e9

Observation 39279745-9fba-4474-9f3e-652bed7a1193 · outbound

This paper cites Textually pretrained speech language mod- els,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Textually pretrained speech language mod- els,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:44.936563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:44.936563Z digest=sha256:33fd1f453dbf5d64089ffeadd6de47bc2edd5d91da364d0b8126ba59c328306a

Observation f0f98697-da92-443a-b0aa-be673aa1316b · outbound

This paper cites GSQA: An End-to-End Model for Generative Spoken Question Answering.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues GSQA: An End-to-End Model for Generative Spoken Question Answering

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:45.071560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:45.071560Z digest=sha256:b8742b747965070b51256376d42892388973d056dbc83294e5c861c40f73f771

Observation da687337-5354-4610-8d33-81c741879d91 · outbound

This paper cites Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI Feedback.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI Feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:45.216359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:45.216359Z digest=sha256:6b8906748c1a884c174fe4529f2fef215089e704825269c8c5c4966e4c0e28bf

Observation 125e2bb9-743b-4c43-a38f-8fc9c0a22b07 · outbound

This paper cites Dynamic-superb: Towards a dynamic, collab- orative, and comprehensive instruction-tuning benchmark for speech,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Dynamic-superb: Towards a dynamic, collab- orative, and comprehensive instruction-tuning benchmark for speech,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:45.330196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:45.330196Z digest=sha256:d88bff3b7d09c8ee724faef15a8a51096e17fe247652d0adc265a5a59a90fc49

Observation 2e7e6682-6c59-46bc-a742-1b2da70970a0 · outbound

This paper cites Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Dynamic-SUPERB Phase-2: A Collaboratively Expanding Benchmark for Measuring the Capabilities of Spoken Language Models with 180 Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:45.452664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:45.452664Z digest=sha256:f8c7303c23c45d08d4ccb2976e1d207632aed3bd175277e96cd705827632ac21

Observation 5bcdd55d-b602-4683-9c78-67d02463c97b · outbound

This paper cites Sd-eval: A benchmark dataset for spoken dialogue understanding beyond words,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Sd-eval: A benchmark dataset for spoken dialogue understanding beyond words,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:45.617090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:45.617090Z digest=sha256:e7d56846254e3ac7fd1fb6e74d4554b7ad326302b62aa584d3b710e288f85e8d

Observation 294f0709-7083-4820-8032-f79d87d4813c · outbound

This paper cites Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:45.730320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:45.730320Z digest=sha256:127c9dccfa0c388fcc681e39c17cebb544dcdd0ddf595e3443e67f2f0e925f1a

Observation d79b70b3-246a-4063-ad4d-4050923f06a6 · outbound

This paper cites Can LLMs Understand the Implication of Emphasized Sentences in Dialogue?.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Can LLMs Understand the Implication of Emphasized Sentences in Dialogue?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:45.892058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:45.892058Z digest=sha256:9825b5d02a259859081134670d8e55ae2fdea61dbaf7acbb6e2aee7440aeccad

Observation 59b66f37-2dda-4843-acd9-3416ffaabd9e · outbound

This paper cites Paralinguistics-enhanced large language model- ing of spoken dialogue,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Paralinguistics-enhanced large language model- ing of spoken dialogue,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:46.069842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:46.069842Z digest=sha256:7178c6a2302a6a23987c73beb97ffcbe36b543aee72f2a078c91625ca05b8f1e

Observation 8cf06290-04af-4d4b-b500-0508e399abb4 · outbound

This paper cites Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:46.172940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:46.172940Z digest=sha256:c46fb097ae61a0afa7fd457e1b56275346fae4e12c8a0c9d0d4cc3e5074bc1e1

Observation 393e8571-d257-4c2c-85f1-4bf558dec92d · outbound

This paper cites Full-Duplex-Bench v1.5: Evaluating Overlap Handling for Full-Duplex Speech Models.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Full-Duplex-Bench v1.5: Evaluating Overlap Handling for Full-Duplex Speech Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:46.285556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:46.285556Z digest=sha256:190d5c07cc313ba78fb7e428f86cfe6caacc72c6d68b57462296524e7f13efa4

Observation 6a3bd3b4-2cb7-481c-9ffb-82fbbd7114b7 · outbound

This paper cites EmoNet-Face: An Expert-Annotated Benchmark for Synthetic Emotion Recognition.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues EmoNet-Face: An Expert-Annotated Benchmark for Synthetic Emotion Recognition

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:46.449195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:46.449195Z digest=sha256:e43cd4d8bcd7ba905fcbd226351bfa4615e7f0c03fcde67d4ea6297880295960

Observation 1952de1b-9e3c-48fb-83b2-c9ab0b44afeb · outbound

This paper cites Dfme: A new benchmark for dynamic facial micro- expression recognition,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Dfme: A new benchmark for dynamic facial micro- expression recognition,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:46.510825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:46.510825Z digest=sha256:836ad6e7c01aed69ddae0cf7b0d30a63f8bd85641939e10b3d1d4086f5c614de

Observation 650e59fa-b87e-4d33-996e-ec183875082e · outbound

This paper cites What comprises a good talking-head video generation?: A Survey and Benchmark.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues What comprises a good talking-head video generation?: A Survey and Benchmark

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:46.566015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:46.566015Z digest=sha256:a10323dd3d931117b6d24a462d2bcad13053d085d7d8813785a45cbe6abd63fc

Observation 8e28c4f6-9261-415c-b69f-acb4787ddc3d · outbound

This paper cites Subjective and objective quality-of-experience assessment for 3d talking heads,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Subjective and objective quality-of-experience assessment for 3d talking heads,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:46.658569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:46.658569Z digest=sha256:d46a83e72dda4956ff26615e780c49ddd205b0413881b02e2044e7e3c03dba90

Observation a6619faf-fe91-4fdc-91ca-efe844f7e6ac · outbound

This paper cites Av- data2vec: Self-supervised learning of audio-visual speech representa- tions with contextualized target representations,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Av- data2vec: Self-supervised learning of audio-visual speech representa- tions with contextualized target representations,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:46.766534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:46.766534Z digest=sha256:8e525a719c34e8b1e57d1c6cfb54f4f8d321d9322f16a47caebe4317be026989

Observation c84f1758-c514-4d74-adb5-c47ff4736a94 · outbound

This paper cites Jointly Learning From Unimodal and Multimodal-Rated Labels in Audio-Visual Emotion Recognition,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Jointly Learning From Unimodal and Multimodal-Rated Labels in Audio-Visual Emotion Recognition,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:46.881448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:46.881448Z digest=sha256:d776ffa3c7b507291aaa5ee0cf43d31d9cf7f0c2679bca5e46495fc7178ac4df

Observation e491b503-38be-46cd-ae0d-99b8fc7f6426 · outbound

This paper cites Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:47.055060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:47.055060Z digest=sha256:e66ff98dc64aba767e33f2ba06aef63c703b83fdf6e0be697ba2699797e38b4b

Observation 94344922-5168-4498-b185-ff6548bd7b27 · outbound

This paper cites Cross-modal incongruity aligning and collaborating for multi-modal sarcasm detection,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Cross-modal incongruity aligning and collaborating for multi-modal sarcasm detection,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:47.160427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:47.160427Z digest=sha256:a703d3aaf4baa8e139c9f80daf5549775204ba87f0d44fefd3bae1646a3ed394

Observation 324591fa-222c-4683-aab5-a9d0f6d31fa5 · outbound

This paper cites Let’s Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Let’s Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:47.327139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:47.327139Z digest=sha256:d97629d726076822f20cafb559a0b0d494b133c9af4f875f499d8c9c9d5b96bb

Observation 5741f955-ca5f-4795-9766-ca92451055e3 · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:47.467247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:47.467247Z digest=sha256:de45095b901be6b5af56b56eeb706148a2eaf75064680677a835c54d6394180f

Observation 8b2d7aaa-cd9f-4a57-8f0c-a142052a77cb · outbound

This paper cites Facial Expression Recognition with Adaptive Frame Rate based on Multiple Testing Correction,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Facial Expression Recognition with Adaptive Frame Rate based on Multiple Testing Correction,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:47.604554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:47.604554Z digest=sha256:1e8fd4a30d3f4d033f99d4438a749303095a58058622727886ad6791e2b5d0e0

Observation 19302319-6d60-438c-ba27-2ed0c0b0ef8d · outbound

This paper cites Introducing the RECOLA multimodal corpus of remote collaborative and affective interactions,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Introducing the RECOLA multimodal corpus of remote collaborative and affective interactions,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:47.778184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:47.778184Z digest=sha256:701a7d1ad38cf7c554d0c046eacf65cb43cf52b64314a32cdcef25170d0e68aa

Observation e9d32696-507a-47d3-a9ea-25741e192fb7 · outbound

This paper cites Avec 2016: Depression, mood, and emotion recognition workshop and challenge,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Avec 2016: Depression, mood, and emotion recognition workshop and challenge,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:47.905073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:47.905073Z digest=sha256:218f5bc2b2105f85c44b66a2f394a5d5fbf14db3dec08d39d5883bfaf8df8012

Observation 69f15386-e08b-434c-b053-4dfb0cccb9d9 · outbound

This paper cites Hicmae: Hierarchical contrastive masked autoencoder for self-supervised audio-visual emotion recognition,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Hicmae: Hierarchical contrastive masked autoencoder for self-supervised audio-visual emotion recognition,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:48.047900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:48.047900Z digest=sha256:4bf8cd3fe04ae15e28e7dc64406da22b789dd892874eac2799ee1444d8172f17

Observation e33dd677-042a-4cc7-9867-0beb00476dde · outbound

This paper cites Baichuan-Omni-1.5 Technical Report,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Baichuan-Omni-1.5 Technical Report,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:48.131813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:48.131813Z digest=sha256:90b1f49248eec6c3f555181af009b9ead1868d4b44d80b354b14e043280bee42

Observation 158e03a8-2b26-406f-8e84-9d8f6cbc1791 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:48.189332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:48.189332Z digest=sha256:279ca27cef9cb4146634261439508202868f2fc0b9d6931608c4fc33d39d7e85

Observation 7e3ad335-ff11-4fb1-966f-20765b90f8e2 · outbound

This paper cites Qwen2.5-Omni Technical Report.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Qwen2.5-Omni Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:48.258673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:48.258673Z digest=sha256:27a3c742d8ad4746cd3cf613730ae9dd830c58d8af26975a96eab3776c8ce999

Observation 2b570d25-4ff0-4d74-915c-2a9b94f73a29 · outbound

This paper cites GPT-4 Technical Report.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues GPT-4 Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:48.330522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:48.330522Z digest=sha256:f64fbd0fe9a445437907be1b735d2876bdfe0dafc1ca3d61ccfc3813d0aeb2b5

Observation 21296364-8cff-4b23-8a30-2db2bb9e81d5 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:48.378240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:48.378240Z digest=sha256:09688b73ff837dbedb650b0b4a83910078ec6b6fc90d52810f072c5b81fcf6d8

Observation d394c90c-79b4-45d5-9426-971267eb3242 · outbound

This paper cites AI-Face: A Million-Scale Demographically Annotated AI-Generated Face Dataset and Fairness Benchmark,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues AI-Face: A Million-Scale Demographically Annotated AI-Generated Face Dataset and Fairness Benchmark,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:48.434176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:48.434176Z digest=sha256:689f32dec59e1d781e202831b57cabeda04e32d8e1f559af58cc6ecfca16ee42

Observation 18ba8698-1e52-45ca-8930-12de44413092 · outbound

This paper cites DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:48.586938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:48.586938Z digest=sha256:d26f962b8d4694fb8fa549f5495b250909bcfde076180767a21a3ecbe0f91def

Observation 92494f21-2696-4f33-8c6e-89b5d4088fb3 · outbound

This paper cites Unconstrained dysfluency modeling for dys- fluent speech transcription and detection,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Unconstrained dysfluency modeling for dys- fluent speech transcription and detection,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:48.730556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:48.730556Z digest=sha256:b5a4c8a40450a7cd3fe4a5218e4dc205323671aca6150127243b45807f2dbc1e

Observation 981e0606-c161-4694-be0a-c9722b56dde7 · outbound

This paper cites Towards hierarchical spo- ken language disfluency modeling,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Towards hierarchical spo- ken language disfluency modeling,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:48.867492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:48.867492Z digest=sha256:aacc3035af0ee7f1aca006bbd3c00b09811f30d7219a272de8a873b31d45cd0d

Observation 2e6ba329-a502-4cd3-aafb-36ff0892a010 · outbound

This paper cites Ssdm: Scalable speech dysfluency model- ing,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Ssdm: Scalable speech dysfluency model- ing,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:49.002050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:49.002050Z digest=sha256:a8b842fd906949197cfd063ff82be495186c25bb7b644bbcb01303e8dacdc74c

Observation 9ef1c0e1-47e9-4c16-8d7c-5b26e31ef006 · outbound

This paper cites Auto- matic detection of articulatory-based disfluencies in primary progres- sive aphasia,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Auto- matic detection of articulatory-based disfluencies in primary progres- sive aphasia,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:49.173181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:49.173181Z digest=sha256:b5b43599c63aa6ab26c673408515131ce8ec4aa570b1be1fd2fa0ce250ff332f

Observation 23c766c1-283b-400f-b708-8f41dd377b11 · outbound

This paper cites Yolo-stutter: End-to-end region-wise speech dysfluency detection,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Yolo-stutter: End-to-end region-wise speech dysfluency detection,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:49.337643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:49.337643Z digest=sha256:9b0475086d4d43d038d17a2020e233c10bd25029fbcc243b64bb635aa9b3a791

Observation e6c095a5-c59c-4296-90a6-f4270545c875 · outbound

This paper cites Stutter-solver: End-to-end multi- lingual dysfluency detection,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Stutter-solver: End-to-end multi- lingual dysfluency detection,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:49.474841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:49.474841Z digest=sha256:d187381b74a3baba64f49f9b109461615a1409cedd05f9a4edf201ae41f5a2fa

Observation c0fd4617-7d21-4a44-9fa9-f104ec2ee1bf · outbound

This paper cites Time and tokens: Benchmarking end-to-end speech dys- fluency detection,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Time and tokens: Benchmarking end-to-end speech dys- fluency detection,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:49.590004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:49.590004Z digest=sha256:5946483380e6e7fd51523054d8cd20ef40bd7edf4058cd44e2371bcd71c2c7cf

Observation 5f5d49ad-c691-4f74-946f-1d5f811bd2bb · outbound

This paper cites Towards accurate phonetic error detection through phoneme similarity modeling,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Towards accurate phonetic error detection through phoneme similarity modeling,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:49.731493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:49.731493Z digest=sha256:bab26792c02f97a2b8b499bb89a22f122f65505478113c1f28d7bf8070956e37

Observation ee530a01-f84e-46d9-a9e6-15199c6855d0 · outbound

This paper cites Dysfluent wfst: A framework for zero-shot speech dysfluency transcription and detection,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Dysfluent wfst: A framework for zero-shot speech dysfluency transcription and detection,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:49.811259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:49.811259Z digest=sha256:ec91af7f7eee914f0e31b547c31d73e0583c7e2edab864a6ddb81f7d38e136f4

Observation dbea160c-b9e7-43e5-828d-747d8831208d · outbound

This paper cites Analysis and evaluation of synthetic data generation in speech dysfluency detec- tion,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Analysis and evaluation of synthetic data generation in speech dysfluency detec- tion,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:49.906913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:49.906913Z digest=sha256:b44cd4ff87915b105c66f269bb48c2cabd23c78745f93cd00aac1aa41ef54ec9

Observation 6fab458a-3d21-45c4-9b3e-731303dd4d5f · outbound

This paper cites LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:49.960449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:49.960449Z digest=sha256:92c208a4d5f4548395923411be2a4025f32039e171fbc610cb8a7ed75f3a4082

Observation 7d32a70c-a167-4f3d-a62f-1c516f3e9576 · outbound

This paper cites Seamless dysfluent speech text alignment for disordered speech analysis,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Seamless dysfluent speech text alignment for disordered speech analysis,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:50.075040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:50.075040Z digest=sha256:650cf27c63fe14375feee78d7e54ba208a626fd1ca8bd11aef2ac2e99d9488b0

Observation 7f7edb46-3e57-4b63-9836-0d198d68fec6 · outbound

This paper cites K-function: Joint pronunciation transcription and feedback for evaluating kids language function,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues K-function: Joint pronunciation transcription and feedback for evaluating kids language function,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:50.181760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:50.181760Z digest=sha256:6942f91d25cd385280ea62ab6a53d2162ee037bc020eb4e27146cab6cf4501d2

Observation 50249b44-c621-400b-b776-11c39a144355 · outbound

This paper cites Deep Neural Convolutive Matrix Factorization for Articulatory Representation Decomposition,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Deep Neural Convolutive Matrix Factorization for Articulatory Representation Decomposition,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:50.261052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:50.261052Z digest=sha256:e1b9ed28e7f3c72cd41f3be9c319d3a399324ccd654bcabdab8557b9e6b8fc8a

Observation b7dea22d-a56e-4a87-a507-4492bc2313e8 · outbound

This paper cites Articulatory representation learn- ing via joint factor analysis and neural matrix factorization,.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues Articulatory representation learn- ing via joint factor analysis and neural matrix factorization,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:50.337605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:50.337605Z digest=sha256:b0d3168b9af38ce3b6a7517b12c221db6069d759e745bc389c366a6b1887c663

Pith citing papers

Observation 4bbd6493-6c83-43e5-9f56-dba555af0d74 · inbound

S-MARC: Causal Streaming Reasoning for Full-Duplex Conversational Behavior Modeling cites this paper.

S-MARC: Causal Streaming Reasoning for Full-Duplex Conversational Behavior Modeling AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T00:56:58.344964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:56:58.344964Z digest=sha256:fac3b065f3b346f8bdc9d6341c86f6d246900a86ba5889247824dec8d52e2f6b