Pith. sign in

Paper Citation Record · LEDGER

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand?

As of 18 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2608.13239.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.13239 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T15:40:22.428261Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved18
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fcb3e689-f5cc-4edf-b386-61120ec4cc92 · outbound

This paper cites In: Proceedings of the 33rd ACM International Confer- ence on Multimedia.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? In: Proceedings of the 33rd ACM International Confer- ence on Multimedia

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.511709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.183478Z digest=sha256:be9b8b636ec5c944d08926a19a4161d945629f226bfe5391be223a05bd2fbb2a

Observation ebf1d7cf-2b22-48f8-97c5-989a2cedf960 · outbound

This paper cites In: Belgrave, D., Zhang, C., Lin, H., Pascanu, R., Koniusz, P., Ghassemi, M., Chen, N.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? In: Belgrave, D., Zhang, C., Lin, H., Pascanu, R., Koniusz, P., Ghassemi, M., Chen, N

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.491598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.189912Z digest=sha256:4ede8c0e8eff58445fb14a5f6124f1c3362e4777cd3784149b63cb07c9e6263d

Observation b8ad8321-5d4f-47d3-a567-5808432c5690 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Gemma 2: Improving Open Language Models at a Practical Size

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.196318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.196318Z digest=sha256:547b80873485bd21e0ab764bc2c4cc398202707ecd938cbf015b6cf787117009

Observation c167fe86-a646-46ec-8322-762a47c04df8 · outbound

This paper cites Nature645(8081), 633–638 (2025).https://doi.org/10.1038/s41586-025-09422-z,http://dx.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Nature645(8081), 633–638 (2025).https://doi.org/10.1038/s41586-025-09422-z,http://dx

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.203126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.203126Z digest=sha256:7a90a584e9098a8df4c470b209f1637dc39dcaaaa93bb2010e2c0a9c966c5f1d

Observation 71ab9d94-9509-4867-b6f8-2b12e5127677 · outbound

This paper cites pre- ferring shorter thinking chains for improved llm reasoning.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? pre- ferring shorter thinking chains for improved llm reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.209126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.209126Z digest=sha256:f9333b7a016e1d9b4ff39e991d956d408a2d4140aee31ed7957e5dc85da1b51c

Observation dce71d75-8e6a-4879-9609-24d2d8e8ea70 · outbound

This paper cites In: The Fourteenth Inter- national Conference on Learning Representations (2026),https://openreview.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? In: The Fourteenth Inter- national Conference on Learning Representations (2026),https://openreview

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.473292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.215354Z digest=sha256:6a3b85a0a15950a2f60e403f1f82478924fba416cd6fcaad3c8432a9916441b7

Observation 1431cff5-861d-4542-97df-30a83cda322a · outbound

This paper cites an unresolved cited work.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:40:23.456227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.221895Z digest=sha256:345280cdce0dcdcb674102fda4b8ad58098d66fc312af065d0d8b3aeb11d37ba

Observation b26da899-3e01-432e-bcb1-63c107e0ea75 · outbound

This paper cites In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2026).

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2026)

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.439702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.227435Z digest=sha256:1044610f9582f4e6295d2d738f6084642f52b5ce17eaa3331af847ae773b6b82

Observation 94effb2d-3ea3-45a5-bdb6-51465b73bae9 · outbound

This paper cites In: Proceedings of the AAAI Conference on Artificial Intelligence.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? In: Proceedings of the AAAI Conference on Artificial Intelligence

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.421959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.233184Z digest=sha256:dbe633d4a5d6f9a422b9f150e6a0fefb151d7f3f9d895f8e7de1f80ee50afe5b

Observation 1bac5ce9-d9ce-4f25-8c19-c0abeef21e1d · outbound

This paper cites Transactions on Machine Learning Re- search (TMLR) (2026) 16 K.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Transactions on Machine Learning Re- search (TMLR) (2026) 16 K

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.402857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.239412Z digest=sha256:c8b3aadc369f81658f2be24ebbb3c004cc171b49dbb1e5d5cae0fa3a27e34477

Observation 94d9d133-0221-4577-b26d-72cd84f0b53a · outbound

This paper cites In: The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track (2026), https://openreview.net/forum?id=SSF4qgsNYE.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? In: The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track (2026), https://openreview.net/forum?id=SSF4qgsNYE

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.384098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.245607Z digest=sha256:f0e0a4105859579c090c15a3ec3c23f7bc10d31fefc4017fb700d9f5eea0336d

Observation 331b1625-dbc1-4104-b370-44f3383ab879 · outbound

This paper cites an unresolved cited work.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.251999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.251999Z digest=sha256:f137e5ce891ddba240cd77183b7b6e64b9bdde3d72bd8ec54dc09470ad7b8001

Observation b76175d4-b330-4aa5-b3ed-76396b6e712d · outbound

This paper cites Explainable Multimodal Emotion Recognition.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Explainable Multimodal Emotion Recognition

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.258726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.258726Z digest=sha256:5906e2356a5b4a1a6ab928275a8e4a3be7ed48a31630eaf60bb6b9ff6c79d873

Observation d2008d3f-f1e2-4300-8549-b3ef134667e3 · outbound

This paper cites Advances in Neural Information Processing Systems36, 34892–34916 (2023).

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Advances in Neural Information Processing Systems36, 34892–34916 (2023)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.361120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.266135Z digest=sha256:21b8edcdc76a9e1fff99e2334853f524f9fccf548ec925825ec7f468a48d17b6

Observation 0c1bdeb2-c3ea-4cfd-bde9-562187734be3 · outbound

This paper cites The Llama 3 Herd of Models.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.271573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.271573Z digest=sha256:8d9fbd301e099a20e51fd822e1ac0f764be5ada726dae962cb54f91502b44996

Observation cc2c9d48-c62e-4373-9422-43e9d060b4b9 · outbound

This paper cites International Conference on Learning Representations (ICLR) (2026).

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? International Conference on Learning Representations (ICLR) (2026)

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.340132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.277997Z digest=sha256:e385cf6aed6453b7ee5a8780f42ad96e9893fe3b1f3793419e315a877b472cd4

Observation 48de3eb2-7ea5-465d-ad4e-fd162ee927a7 · outbound

This paper cites MODF-SIR: A Multi-agent Omni-modal Distilled Framework for Social Intelligence Reasoning.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? MODF-SIR: A Multi-agent Omni-modal Distilled Framework for Social Intelligence Reasoning

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-08-14T15:40:22.902794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.283540Z digest=sha256:f9a008309a548b5a64f4e9652a81dca8c0b9af6bfd16d4ca1a4d9af0e4287a8c

Observation 5bf02b1c-97f5-4650-bfbd-5f8c224913fc · outbound

This paper cites In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.322801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.289268Z digest=sha256:e4f8a6905f7b5dab0106a3c57b79d3ca38bdcccaa61ccd1518147aa5f915b237

Observation 05fc6fc0-db95-4e07-b562-4ee30c328158 · outbound

This paper cites an unresolved cited work.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:40:23.305487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.295407Z digest=sha256:5057a84f729d25a4bf0a15a6496ae1dd45dc61d720b7a23a8e4ef61349f8a170

Observation c46bf565-a18f-4c7b-8d69-717186081ca5 · outbound

This paper cites an unresolved cited work.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:40:23.287468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.301379Z digest=sha256:812f41c2368ec27b5f4d193aaf153f3b86eba1f2a713f956006d0fde52f088f9

Observation 6ead7a3d-e8ed-448d-a535-028650533434 · outbound

This paper cites In: Proceedings of the AAAI Conference on Artificial Intelligence.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? In: Proceedings of the AAAI Conference on Artificial Intelligence

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.264045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.308927Z digest=sha256:6568013f4b1d95470ffdbbcdc72d10676368901de5a8429609a97ae298f35982

Observation a5661164-765a-40f0-a528-c1ce20f39c39 · outbound

This paper cites Qwen3.5-Omni Technical Report.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Qwen3.5-Omni Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.314683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.314683Z digest=sha256:a2bf17cc6435409e7d7b619803e7b92b69897930578b0eef4e2b7651fad3662f

Observation b2fdcf59-6add-4887-8a55-43043127a5c1 · outbound

This paper cites Proceedings of the AAAI Conference on Arti- ficial Intelligence40(3), 2029–2037 (Mar 2026).https://doi.org/10.1609/aaai.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Proceedings of the AAAI Conference on Arti- ficial Intelligence40(3), 2029–2037 (Mar 2026).https://doi.org/10.1609/aaai

Reference 23

Resolution
malformed identifier
no resolver link, observed 2026-08-14T15:40:22.320829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.320829Z digest=sha256:7c768229df0ae20065ec0b105473e1a1c4bed7bfd4b0c15489ce401f18eb59f4

Observation 4de6f7d6-9041-438e-aa85-4b7b2131ab37 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.330193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.330193Z digest=sha256:1aa388b29f61eed1fb88ac1c72ed5aa0e9be6046f4d8bd5d4100181ab5d8625b

Observation ef2d15aa-22c3-406c-a6a3-99f58c4eed8f · outbound

This paper cites IEEE Transactions on Affective Computing pp.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? IEEE Transactions on Affective Computing pp

Reference 25

Resolution
metadata mismatch
raw_fallback, observed 2026-08-14T15:40:22.837582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.338534Z digest=sha256:e3c3f998da2bbb7dcd35b679763224fd628367ba2bd879df25a6cbed274a8b6b

Observation 864397c2-61c0-4e69-88c3-bd32fbf521c0 · outbound

This paper cites NIPS ’22, Curran Associates Inc., Red Hook, NY, USA (2022) Reasoning for Social AV-QA: Where Do We Stand? 17.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? NIPS ’22, Curran Associates Inc., Red Hook, NY, USA (2022) Reasoning for Social AV-QA: Where Do We Stand? 17

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.245065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.346441Z digest=sha256:0ea656a621653288a1c02d0f83dd62283ce2618496b0ebf2cacf460580be6f72

Observation 99d44657-f5ad-439d-8113-0272b25107f4 · outbound

This paper cites an unresolved cited work.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:40:23.224191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.354610Z digest=sha256:866b8b71d7af2d6565408707d44ea3f7f90cf8265da63a24c0b6667453b2f6cf

Observation 9a5be136-94c5-48f3-9271-805bdd1e91cc · outbound

This paper cites In: The Fourteenth International Conference on Learning Representations (2026),https://openreview.net/forum?id=KttCXdjj4w.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? In: The Fourteenth International Conference on Learning Representations (2026),https://openreview.net/forum?id=KttCXdjj4w

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.205764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.361475Z digest=sha256:db11fa0a2ff7d32718e1c50cdaae72b92bf49a5f1b850896b34e6f001ec559af

Observation 9dde63e2-6531-4509-8203-9b8cbd09f804 · outbound

This paper cites Qwen2.5-Omni Technical Report.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Qwen2.5-Omni Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.370138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.370138Z digest=sha256:48b803754d8dd0701029e4ca9b42b3499eb8f0d5b656a4c8bbc8fc4d1d53949e

Observation 311bef0b-dcd4-4b30-829b-aeffeb57682e · outbound

This paper cites HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.379357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.379357Z digest=sha256:75d78c4b07a6a2491039e8ca49caea88138c33db8dfc2706d2f13c0790da1bcb

Observation 146a99af-8849-4886-a608-74b881604adc · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.185468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.387022Z digest=sha256:649b148081cca37aba30e75195c10218442a39f7f8c61f520a92894c6535822b

Observation c4488e5e-0328-4d8b-b5fb-e61f1fcafda1 · outbound

This paper cites In: The Fourteenth International Confer- ence on Learning Representations (2026),https://openreview.net/forum?id= xindJJLSr1.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? In: The Fourteenth International Confer- ence on Learning Representations (2026),https://openreview.net/forum?id= xindJJLSr1

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.163274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.392688Z digest=sha256:ca22d9fd81a13f32ccb8d9ef561370d40755654e95364219af2602d7877d007f

Observation 7595718e-6042-458e-b0b3-44aa283e9948 · outbound

This paper cites In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.142770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.398106Z digest=sha256:590b6a4747fb6ee248b138eeb48578c994ae0a52727607bce7d682fc4f87d1cc

Observation 1351b028-c51b-4d92-a936-9235dfbebfbb · outbound

This paper cites R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.407728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.407728Z digest=sha256:ab6d88c4aa6a3aa86cfaef34bc2dd364e01338c1b9d1a55ae1c91dd85831b022

Observation 3aecfed7-7764-4b78-9b86-e7a58187b973 · outbound

This paper cites arXiv preprint arXiv:2512.09616 (2025).

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? arXiv preprint arXiv:2512.09616 (2025)

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.415414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.415414Z digest=sha256:023b1285817cc156b6ca4ff3b6bc03b67b0651b43be8bac6227d6533d259173c

Observation b47a2d16-4a8f-482e-ac4d-67e3d64bc694 · outbound

This paper cites Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.421903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.421903Z digest=sha256:6309eda493467ba81bc7425f6a73386b26cad5b793a375a026fded178a785439

Observation d9427a58-448b-418a-82a9-a38044e6c743 · outbound

This paper cites arXiv preprint arXiv:2505.17862 (2025) 18 K.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? arXiv preprint arXiv:2505.17862 (2025) 18 K

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.428261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.428261Z digest=sha256:8a6769387d5a35de51f1f57a86271e71f6e63ab93bbf6fc78fd82d4d556657b4

Pith citing papers

No inbound Pith citation observations are available.