Pith. sign in

Paper Citation Record · LEDGER

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand?

As of 17 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2608.13239.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.13239 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T15:40:22.428261Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved18
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fcb3e689-f5cc-4edf-b386-61120ec4cc92 · outbound

This paper cites In: Proceedings of the 33rd ACM International Confer- ence on Multimedia.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? In: Proceedings of the 33rd ACM International Confer- ence on Multimedia

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.511709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.183478Z digest=sha256:c4cfac3c91faff7118734315c077571d87b52d6d60854d9d92438d7ab7c28d9f

Observation ebf1d7cf-2b22-48f8-97c5-989a2cedf960 · outbound

This paper cites In: Belgrave, D., Zhang, C., Lin, H., Pascanu, R., Koniusz, P., Ghassemi, M., Chen, N.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? In: Belgrave, D., Zhang, C., Lin, H., Pascanu, R., Koniusz, P., Ghassemi, M., Chen, N

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.491598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.189912Z digest=sha256:909de767ee9f5ebabc86af999eb5e115e63e46b1a56439af65a1bc8a1284332c

Observation b8ad8321-5d4f-47d3-a567-5808432c5690 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Gemma 2: Improving Open Language Models at a Practical Size

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.196318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.196318Z digest=sha256:8756434ec589f651dd543ed18e213b8c428f2284a675797b1d543c34c905e936

Observation c167fe86-a646-46ec-8322-762a47c04df8 · outbound

This paper cites Nature645(8081), 633–638 (2025).https://doi.org/10.1038/s41586-025-09422-z,http://dx.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Nature645(8081), 633–638 (2025).https://doi.org/10.1038/s41586-025-09422-z,http://dx

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.203126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.203126Z digest=sha256:25311c84f031459b709dccf0e4cac128c36696b3cc3fe8bfe10debf3aa6e52ec

Observation 71ab9d94-9509-4867-b6f8-2b12e5127677 · outbound

This paper cites pre- ferring shorter thinking chains for improved llm reasoning.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? pre- ferring shorter thinking chains for improved llm reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.209126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.209126Z digest=sha256:25129fe01977709122405324fd8dd09c70285c948c0bfe13b89c32b9548e68a8

Observation dce71d75-8e6a-4879-9609-24d2d8e8ea70 · outbound

This paper cites In: The Fourteenth Inter- national Conference on Learning Representations (2026),https://openreview.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? In: The Fourteenth Inter- national Conference on Learning Representations (2026),https://openreview

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.473292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.215354Z digest=sha256:740d5cb4ded943c95a37d2d24dac9f6ec6d1a3937c0dcae8cc00a1764d959f2f

Observation 1431cff5-861d-4542-97df-30a83cda322a · outbound

This paper cites an unresolved cited work.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:40:23.456227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.221895Z digest=sha256:c5069d80f3b5140ffd4d60594e2c29f45637c8e0c5a0e0ce477dcfb349615e66

Observation b26da899-3e01-432e-bcb1-63c107e0ea75 · outbound

This paper cites In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2026).

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2026)

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.439702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.227435Z digest=sha256:2b2a77818cabbf2de19f43259e8b65eda737f7e3dbc47fe4d38c47c995e224e2

Observation 94effb2d-3ea3-45a5-bdb6-51465b73bae9 · outbound

This paper cites In: Proceedings of the AAAI Conference on Artificial Intelligence.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? In: Proceedings of the AAAI Conference on Artificial Intelligence

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.421959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.233184Z digest=sha256:32f4a735eae64dbb4416faf9bfe7a11cd2e4a74e2bcb4d36632050a70d1235f4

Observation 1bac5ce9-d9ce-4f25-8c19-c0abeef21e1d · outbound

This paper cites Transactions on Machine Learning Re- search (TMLR) (2026) 16 K.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Transactions on Machine Learning Re- search (TMLR) (2026) 16 K

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.402857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.239412Z digest=sha256:d1b297ac6bc63e0163929ed2ceceab8aacd59d314027f6fc9167cf06851eecdd

Observation 94d9d133-0221-4577-b26d-72cd84f0b53a · outbound

This paper cites In: The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track (2026), https://openreview.net/forum?id=SSF4qgsNYE.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? In: The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track (2026), https://openreview.net/forum?id=SSF4qgsNYE

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.384098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.245607Z digest=sha256:61bf72a84a4e70ff9c677911bd98547e649b6095fe569e968426c33d0b911992

Observation 331b1625-dbc1-4104-b370-44f3383ab879 · outbound

This paper cites an unresolved cited work.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.251999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.251999Z digest=sha256:2833dc3aff368ae6837d64db89efdf82be96a87e91a55b421c69bcafc7ff1691

Observation b76175d4-b330-4aa5-b3ed-76396b6e712d · outbound

This paper cites Explainable Multimodal Emotion Recognition.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Explainable Multimodal Emotion Recognition

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.258726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.258726Z digest=sha256:4603e28eb42c24919b32d9527dd9bccbd19cc838b73d0ee15c97d256cb6c7542

Observation d2008d3f-f1e2-4300-8549-b3ef134667e3 · outbound

This paper cites Advances in Neural Information Processing Systems36, 34892–34916 (2023).

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Advances in Neural Information Processing Systems36, 34892–34916 (2023)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.361120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.266135Z digest=sha256:9adbc70b10c06366caab80421dbb2b471b71771243572fdc143f8f973bd28676

Observation 0c1bdeb2-c3ea-4cfd-bde9-562187734be3 · outbound

This paper cites The Llama 3 Herd of Models.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.271573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.271573Z digest=sha256:8416c6bffd6b6022d791182ee77de56275c2962f0c109a87a58d26aa410868aa

Observation cc2c9d48-c62e-4373-9422-43e9d060b4b9 · outbound

This paper cites International Conference on Learning Representations (ICLR) (2026).

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? International Conference on Learning Representations (ICLR) (2026)

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.340132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.277997Z digest=sha256:15a3021af237956b8483d923f2e23038cca5f42c26d26e804df2957e33571808

Observation 48de3eb2-7ea5-465d-ad4e-fd162ee927a7 · outbound

This paper cites MODF-SIR: A Multi-agent Omni-modal Distilled Framework for Social Intelligence Reasoning.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? MODF-SIR: A Multi-agent Omni-modal Distilled Framework for Social Intelligence Reasoning

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-08-14T15:40:22.902794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.283540Z digest=sha256:01b37929c4bd784ca79ac7b2d5edf49271aefd1f3fca5ca828bc8feee75be620

Observation 5bf02b1c-97f5-4650-bfbd-5f8c224913fc · outbound

This paper cites In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? In: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.322801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.289268Z digest=sha256:efe7f94de7eaa142c8090211cf981f18620ec2bd10309cd7b26f7bc5f7dcafbc

Observation 05fc6fc0-db95-4e07-b562-4ee30c328158 · outbound

This paper cites an unresolved cited work.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:40:23.305487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.295407Z digest=sha256:e59f78129d236711594a514ea80f60379f574402b03a9456618ab500fe06f4d1

Observation c46bf565-a18f-4c7b-8d69-717186081ca5 · outbound

This paper cites an unresolved cited work.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:40:23.287468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.301379Z digest=sha256:bb29928feb0b317ef6b509ea1db4cbd88e622db90a0708f8450fc97c28babdbe

Observation 6ead7a3d-e8ed-448d-a535-028650533434 · outbound

This paper cites In: Proceedings of the AAAI Conference on Artificial Intelligence.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? In: Proceedings of the AAAI Conference on Artificial Intelligence

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.264045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.308927Z digest=sha256:54ee97614bd8674e764ccdd2bc110e81a8cbfb3d05aa2e3f1725b5fcb6c0a2f1

Observation a5661164-765a-40f0-a528-c1ce20f39c39 · outbound

This paper cites Qwen3.5-Omni Technical Report.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Qwen3.5-Omni Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.314683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.314683Z digest=sha256:e15f5b538ac99e6c4870ca316c7b68e3be4508f54e900c62615836e7b69b2bde

Observation b2fdcf59-6add-4887-8a55-43043127a5c1 · outbound

This paper cites Proceedings of the AAAI Conference on Arti- ficial Intelligence40(3), 2029–2037 (Mar 2026).https://doi.org/10.1609/aaai.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Proceedings of the AAAI Conference on Arti- ficial Intelligence40(3), 2029–2037 (Mar 2026).https://doi.org/10.1609/aaai

Reference 23

Resolution
malformed identifier
no resolver link, observed 2026-08-14T15:40:22.320829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.320829Z digest=sha256:a6cb9c80f916faacafa65a21cfcd7e9b7d07b28868e0a19dde8b9d7fceb62d01

Observation 4de6f7d6-9041-438e-aa85-4b7b2131ab37 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.330193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.330193Z digest=sha256:34d96d4d1e97e66dc7e8f207aa2230f719311053f2f64c065bb7f9df52b5905f

Observation ef2d15aa-22c3-406c-a6a3-99f58c4eed8f · outbound

This paper cites IEEE Transactions on Affective Computing pp.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? IEEE Transactions on Affective Computing pp

Reference 25

Resolution
metadata mismatch
raw_fallback, observed 2026-08-14T15:40:22.837582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.338534Z digest=sha256:78b2e92652b25314727038e0826f8f9dce62fb436e10b340e7fb076566a122dd

Observation 864397c2-61c0-4e69-88c3-bd32fbf521c0 · outbound

This paper cites NIPS ’22, Curran Associates Inc., Red Hook, NY, USA (2022) Reasoning for Social AV-QA: Where Do We Stand? 17.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? NIPS ’22, Curran Associates Inc., Red Hook, NY, USA (2022) Reasoning for Social AV-QA: Where Do We Stand? 17

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.245065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.346441Z digest=sha256:f0f5a538aaf61140460686b6c9a33349c475ad6d534908979a1df59851fedce2

Observation 99d44657-f5ad-439d-8113-0272b25107f4 · outbound

This paper cites an unresolved cited work.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-14T15:40:23.224191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.354610Z digest=sha256:eed42ef27d0ca749ff3bcda59a72edc22be91a56c6415d12112131994852f003

Observation 9a5be136-94c5-48f3-9271-805bdd1e91cc · outbound

This paper cites In: The Fourteenth International Conference on Learning Representations (2026),https://openreview.net/forum?id=KttCXdjj4w.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? In: The Fourteenth International Conference on Learning Representations (2026),https://openreview.net/forum?id=KttCXdjj4w

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.205764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.361475Z digest=sha256:29dc22cbbef5eaa5c4c1a0f3fc33900193e75bdc7a63f6fb495382356607998b

Observation 9dde63e2-6531-4509-8203-9b8cbd09f804 · outbound

This paper cites Qwen2.5-Omni Technical Report.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Qwen2.5-Omni Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.370138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.370138Z digest=sha256:3e221696d3b18d9e36b2550682aed491a83738d6ae6f8d726a07ab20be5e15c9

Observation 311bef0b-dcd4-4b30-829b-aeffeb57682e · outbound

This paper cites HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.379357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.379357Z digest=sha256:5bb7293ebcaaecb8c98fbae0bd32ff6f6eb5dfb22a5fc40efad75c3463093547

Observation 146a99af-8849-4886-a608-74b881604adc · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.185468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.387022Z digest=sha256:cb39a6a4f2ee37d158006e6cc72e0fbc654bfd104b55236d87edc9f9162ef2c4

Observation c4488e5e-0328-4d8b-b5fb-e61f1fcafda1 · outbound

This paper cites In: The Fourteenth International Confer- ence on Learning Representations (2026),https://openreview.net/forum?id= xindJJLSr1.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? In: The Fourteenth International Confer- ence on Learning Representations (2026),https://openreview.net/forum?id= xindJJLSr1

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.163274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.392688Z digest=sha256:63fc61d44d77bd6d2db86e73d42a76203bf3eb51be119eecf99882700ff01e1c

Observation 7595718e-6042-458e-b0b3-44aa283e9948 · outbound

This paper cites In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T15:40:23.142770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-14T15:40:22.398106Z digest=sha256:171d9f3b0ddfd124db692fa1086fb9119055ff3b0319575ef601ebf137f1fc6d

Observation 1351b028-c51b-4d92-a936-9235dfbebfbb · outbound

This paper cites R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.407728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.407728Z digest=sha256:a8a2e6a3951af38468aff470ebfedda5da2d87f674da2b7ce43ef56e6b23c2f7

Observation 3aecfed7-7764-4b78-9b86-e7a58187b973 · outbound

This paper cites arXiv preprint arXiv:2512.09616 (2025).

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? arXiv preprint arXiv:2512.09616 (2025)

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.415414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.415414Z digest=sha256:a377f524601004c98042ecb5bb14353c01d0e48c29ce9cc0d1a1d0cde431bf48

Observation b47a2d16-4a8f-482e-ac4d-67e3d64bc694 · outbound

This paper cites Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.421903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.421903Z digest=sha256:43c1e2c930b1de701f9277fa2fa1212650b6c2f458a9063240c7d0c37eb90f3d

Observation d9427a58-448b-418a-82a9-a38044e6c743 · outbound

This paper cites arXiv preprint arXiv:2505.17862 (2025) 18 K.

Reasoning for Social Audio-Visual Question Answering: Where Do We Stand? arXiv preprint arXiv:2505.17862 (2025) 18 K

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-14T15:40:22.428261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T15:40:22.428261Z digest=sha256:60bb113ccbcc3538f75dc0d595b8940095c9a7852dd3dced25cdd01d1711fbfa

Pith citing papers

No inbound Pith citation observations are available.