Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2404.03413.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:29:59.428434Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T06:39:37.386392Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation ccdaa848-0a74-47df-a5d9-781b71649ec2 · inbound
MLVU: Benchmarking Multi-task Long Video Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1c5b397b-649d-401b-a0f3-626d2b80f660 · inbound
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d75d13d1-16a3-4873-9478-01a542fae4cd · inbound
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7bccb503-6199-46e9-9652-c5ee814d49e5 · inbound
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d5a8ee7e-ac4b-4091-9af7-9c75d21190cd · inbound
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d2f6c73f-2372-44a3-b160-0537fb465047 · inbound
EmoSign: A Multimodal Dataset for Understanding Emotions in American Sign Language MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b4b6cbc-051e-458e-86a3-e7dab822e61a · inbound
Multi-modal brain encoding models for multi-modal stimuli MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b324e589-68a1-4714-8d08-3deb4559a04d · inbound
VUDG: A Dataset for Video Understanding Domain Generalization MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b24fe65a-5f5a-4cb0-ba0e-6cc9c46d922e · inbound
From Motion to Behavior: Hierarchical Modeling of Humanoid Generative Behavior Control MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90c069d5-78fa-4889-9b20-ecc8e2e66d6c · inbound
LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8be7d67e-4201-4abc-acb6-e92aa4708ba6 · inbound
Vision Generalist Model: A Survey MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95c4507e-4f56-42ef-bd11-4cb240c38ea9 · inbound
ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33c76992-50e8-44ba-9246-b3c7cc996789 · inbound
MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53afb53d-4905-43c0-bc65-1518f0ee4249 · inbound
Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c40a1ee-0707-4abd-8809-7b85169751e1 · inbound
LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9885207-34f9-4eef-b420-a4429e96f580 · inbound
Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance? MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8aba6b87-1732-4856-9e30-2136e5deb11d · inbound
SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 255a21a5-7f8b-4120-acc6-f40f65ea49f1 · inbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7a72c2ee-6597-44f7-9e06-bbaebd2e3e40 · inbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 158
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 660701bf-6df2-4b57-801e-7e7628952385 · inbound
One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e44ff72d-fe3d-44eb-90f9-764ea821fa25 · inbound
Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 850ea5ce-203f-4cd4-ac50-d4befa26a7f0 · inbound
LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 670b309a-0678-4391-96da-8b21ed8b59ad · inbound
HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 147
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b7bb3f1d-09d1-4fc2-bb2b-9f0d9867f14d · inbound
VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 20a41ddb-143e-4c2f-a16e-6dcf548dd89c · inbound
TimeThink: Reasoning with Time for Video LLMs MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5bbbc41-1e21-43ea-bae8-2169698fb5a4 · inbound
Sign Language Question Answering: A New Task, Benchmark, and Baseline for Sign Language Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 556337f2-1942-4cbc-93e3-483748e20b39 · inbound
ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.