Pith. sign in

Paper Citation Record · LEDGER

Audio-Visual LLM for Video Understanding

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2312.06720.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.06720 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:47:17.716820Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:27:36.866746Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9c3355a9-bca8-4cc7-8895-0fa71688eca1 · inbound

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs cites this paper.

VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs Audio-Visual LLM for Video Understanding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T02:44:53.634793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T02:44:53.284345Z digest=sha256:1228a34722af28ed3812b54f6aabd97c03f46d61f6c80ccb940d4a28de0abb48

Observation b3c31b57-5776-4d2e-8ad0-f9eaa60f7215 · inbound

Do Language Models Understand Time? cites this paper.

Do Language Models Understand Time? Audio-Visual LLM for Video Understanding

Reference 146

Resolution
unresolved
no resolver link, observed 2026-08-11T12:47:17.716820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:47:17.716820Z digest=sha256:6eb6ba0aeb14e8638c600aa846c4b05fcbbb5224745d9e23bc635e9bd4926061

Observation f56fd9f8-008a-434e-98bf-c4f74ea804ec · inbound

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction cites this paper.

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction Audio-Visual LLM for Video Understanding

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:08:19.785427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T21:08:19.570050Z digest=sha256:ab86f498cb57b3edadc80e52b1480230984eb15a8c4bd6dea9e6ecb44ebae20e

Observation d6b5824c-d880-4145-867f-cfec6ee61f70 · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Audio-Visual LLM for Video Understanding

Reference 157

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:20:00.157256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:431aab3f3730660facb82e06e7f601fbbbc46fdc840f3dbc19e42dc681c7b972

Observation f94a6481-4a61-4a81-821d-9751a9d44c09 · inbound

Multimodal Large Language Models for Image, Text, and Speech Data Augmentation: A Survey cites this paper.

Multimodal Large Language Models for Image, Text, and Speech Data Augmentation: A Survey Audio-Visual LLM for Video Understanding

Reference 288

Resolution
unresolved
no resolver link, observed 2026-08-10T04:36:38.400017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:36:38.400017Z digest=sha256:12b0490a42913b54d260807bffaba219ef5a00aa1f75d9caaa25d8c41a2b2ab0

Observation bf0cfcfd-0c56-4fce-a101-93c74b334982 · inbound

Video-R1: Reinforcing Video Reasoning in MLLMs cites this paper.

Video-R1: Reinforcing Video Reasoning in MLLMs Audio-Visual LLM for Video Understanding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:43:00.428279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T09:43:00.208065Z digest=sha256:2ab00aa0e149b44b2e7b113c96e6395797c3cef79c999eb068a43649cbdfda46

Observation 9e79c0e2-fdf6-45ea-910e-0fcd9692b7ae · inbound

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language cites this paper.

RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language Audio-Visual LLM for Video Understanding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:14.093875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:19:14.093875Z digest=sha256:83d3273ba262cbb100da509853057658bd7b8869662e9c50c8cc86656687ee54

Observation 868cd186-217e-47e2-9345-540ad555343e · inbound

Reinforcing Video Reasoning with Focused Thinking cites this paper.

Reinforcing Video Reasoning with Focused Thinking Audio-Visual LLM for Video Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:12.360528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:12.360528Z digest=sha256:d3cf56ea4ee9d473dc459fb2884de0c952dc302c8c634783ce268421a901cde1

Observation 7782f6bc-cd96-4608-a112-e0cc68e74532 · inbound

Learning Sparsity for Effective and Efficient Music Performance Question Answering cites this paper.

Learning Sparsity for Effective and Efficient Music Performance Question Answering Audio-Visual LLM for Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:50:44.668965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:50:44.668965Z digest=sha256:36c32bcf15370fe085a8001e7e117724938336c675652dd25f39e7f808105ce7

Observation 0117793b-5de8-4a38-8bf8-8c9c07a7b2ee · inbound

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing cites this paper.

Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Audio-Visual LLM for Video Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:26:50.104456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:26:50.104456Z digest=sha256:7721cf2cc2081c8e48f7e981a92527c2b66ef7992546f6d236a1a528352eb0de

Observation 9b435fac-ba84-4bce-a64c-509e62ab856b · inbound

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought cites this paper.

Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Audio-Visual LLM for Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:05:50.794819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:05:50.794819Z digest=sha256:fb5feff914e431aad70e3d06f55f7165919f6ec52830dbb48a69a9e5fc3a5cf5

Observation 3c1bf8c1-3523-453a-9fb7-56d662b584e5 · inbound

"Before, I Asked My Mom, Now I Ask ChatGPT": Visual Privacy Management with Generative AI for Blind and Low-Vision People cites this paper.

"Before, I Asked My Mom, Now I Ask ChatGPT": Visual Privacy Management with Generative AI for Blind and Low-Vision People Audio-Visual LLM for Video Understanding

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:45.751588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:45.751588Z digest=sha256:30bc3f9c1b2f39360be3b25f5c081757c51ed200d2e713f74101a9758f204afe

Observation bc1ed0eb-b562-42ae-a323-31a5a96752f1 · inbound

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning cites this paper.

IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Audio-Visual LLM for Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:29.216841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:29.216841Z digest=sha256:f3aec34b20b9a6cc93ea5d72528baf9edcaab3a6d3d6aa0d410beff7dad3d606

Observation b4a0a645-4785-4ba4-a8bd-573947db35c1 · inbound

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts cites this paper.

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Audio-Visual LLM for Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T13:12:40.249511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:12:40.249511Z digest=sha256:fbe87d2768673784d975e63aa8db7605fa3f306bd84386acc60e1e7f0335e2d2

Observation 0a5b208f-05a0-45e2-9930-fc1c196af9fd · inbound

AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering cites this paper.

AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering Audio-Visual LLM for Video Understanding

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:30:55.403906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T05:26:07.722390Z digest=sha256:c07da604d35a0152adcb28b419073bf7233d0439ad2dce04c24d79cc038c0668

Observation 52e6c9c0-0ce5-47fc-96f6-0e6a2b0a3fb1 · inbound

AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering cites this paper.

AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering Audio-Visual LLM for Video Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:26.153084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:55:26.153084Z digest=sha256:91d515e78fcdfa573f9e7d1658a22423e96e9aadbd0d9d1be71fd5c2bb96725c

Observation e99330e2-0e58-43c7-a6f6-f913315cf8e1 · inbound

Audio-Visual Exchange-Aware Token Pruning for Efficient Audio-Visual Captioning cites this paper.

Audio-Visual Exchange-Aware Token Pruning for Efficient Audio-Visual Captioning Audio-Visual LLM for Video Understanding

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:27:36.868245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T13:53:53.520545Z digest=sha256:93d33c7cc987e9517fec8397bfebb001a33fb9f7affe85654514bd872e9efc99