Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2312.06720.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T04:36:38.400017Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T04:27:36.866746Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 9c3355a9-bca8-4cc7-8895-0fa71688eca1 · inbound
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs Audio-Visual LLM for Video Understanding
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f56fd9f8-008a-434e-98bf-c4f74ea804ec · inbound
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction Audio-Visual LLM for Video Understanding
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d6b5824c-d880-4145-867f-cfec6ee61f70 · inbound
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Audio-Visual LLM for Video Understanding
Reference 157
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f94a6481-4a61-4a81-821d-9751a9d44c09 · inbound
Multimodal Large Language Models for Image, Text, and Speech Data Augmentation: A Survey Audio-Visual LLM for Video Understanding
Reference 288
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf0cfcfd-0c56-4fce-a101-93c74b334982 · inbound
Video-R1: Reinforcing Video Reasoning in MLLMs Audio-Visual LLM for Video Understanding
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9e79c0e2-fdf6-45ea-910e-0fcd9692b7ae · inbound
RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language Audio-Visual LLM for Video Understanding
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 868cd186-217e-47e2-9345-540ad555343e · inbound
Reinforcing Video Reasoning with Focused Thinking Audio-Visual LLM for Video Understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7782f6bc-cd96-4608-a112-e0cc68e74532 · inbound
Learning Sparsity for Effective and Efficient Music Performance Question Answering Audio-Visual LLM for Video Understanding
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0117793b-5de8-4a38-8bf8-8c9c07a7b2ee · inbound
Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing Audio-Visual LLM for Video Understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b435fac-ba84-4bce-a64c-509e62ab856b · inbound
Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought Audio-Visual LLM for Video Understanding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c1bf8c1-3523-453a-9fb7-56d662b584e5 · inbound
"Before, I Asked My Mom, Now I Ask ChatGPT": Visual Privacy Management with Generative AI for Blind and Low-Vision People Audio-Visual LLM for Video Understanding
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc1ed0eb-b562-42ae-a323-31a5a96752f1 · inbound
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning Audio-Visual LLM for Video Understanding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4a0a645-4785-4ba4-a8bd-573947db35c1 · inbound
ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts Audio-Visual LLM for Video Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a5b208f-05a0-45e2-9930-fc1c196af9fd · inbound
AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering Audio-Visual LLM for Video Understanding
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 52e6c9c0-0ce5-47fc-96f6-0e6a2b0a3fb1 · inbound
AV-Master: Dual-Path Comprehensive Perception Makes Better Audio-Visual Question Answering Audio-Visual LLM for Video Understanding
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e99330e2-0e58-43c7-a6f6-f913315cf8e1 · inbound
Audio-Visual Exchange-Aware Token Pruning for Efficient Audio-Visual Captioning Audio-Visual LLM for Video Understanding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.