Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:15:50.652098Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 6 inbound Pith citation observations for arXiv:2504.20091.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:15:50.652098Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:20:12.739322Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-17T20:20:11.800532Z
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation aa1e6425-295d-4105-908b-86faefb6c1c4 · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering ENTER: Event Based Interpretable Reasoning for VideoQA
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5551cb9e-2eff-4e80-b4dc-475bf1407573 · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering HourVideo: 1-Hour Video-Language Understanding
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e0ccdbc-dc61-4c8e-9ba9-47889b54753e · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Videollama 2: Advancing spatial-temporal modeling and audio understanding in video- llms, 2024
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c2dd7f3-ec03-476e-906b-b993d97fa18c · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Adaptive video relationship detection via graph matching and attention
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6a83fc47-71a2-43f2-9405-f98f961ab9e5 · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering EVA-02: A Visual Representation for Neon Genesis
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d32cc318-6bac-40e6-9653-2b7e2686b168 · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5d50635-7b9b-4544-8983-b4e1ef5550e8 · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Linvt: Empower your image- level large language model to understand videos, 2024
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 676aa8b4-368e-4b74-9eaf-265bf3b8d0cf · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Learning situation hyper-graphs for video question answering
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5836d3ee-1455-4a25-a072-05b14e1848f2 · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering An image grid can be worth a video: Zero-shot video question answering using a vlm
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation acb824ff-54de-4b32-ae30-a082cbfa9acc · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering VDMA: Video Question Answering with Dynamically Generated Multi-Agents
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3416796f-bb40-4e78-b057-3022daf27627 · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e40d79c-2319-4eef-bd6c-4144028489e6 · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering In- tentqa: Context-aware video intent reasoning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bc3043f9-0cf1-473e-be77-879fd79acd5a · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Videoinsta: Zero-shot long video understanding via informative spatial- temporal reasoning with llms
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 80967dbf-8a0f-4626-a629-ed2604852816 · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Video-llava: Learning united visual representation by alignment before projection
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f11da199-0d8f-475b-bc97-8fe9002f8627 · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Video-llava: Learning united visual representation by alignment before projection
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4d65ce78-8e5d-41de-bee1-00065871fba9 · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering MM-VID: Advancing Video Understanding with GPT-4V(ision)
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6707995e-12e0-4b8b-a491-f2475c8c986c · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Khan, and Fahad Shahbaz Khan
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 45917096-e0dd-4428-bb9f-d77a939518b0 · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Egoschema: A diagnostic benchmark for very long- form video language understanding
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fa273361-0a44-4342-80b3-1be3f7b0f28d · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Hello GPT-4o: OpenAI’s Multimodal GPT-4 Omni Announcement
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f791700f-3d40-4d0e-a349-76e78e40ad0b · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb6589cf-d9c5-43c5-b877-59a6d59dbd4c · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Ts-llava: Constructing visual tokens through thumbnail-and-sampling for training-free video large language models, 2024
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 97a32367-ab6e-45cc-b227-af4ff9e9ec34 · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Action scene graphs for long- form understanding of egocentric videos, 2023
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 84acb3fe-1c1f-4021-9b55-da674452dbdf · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Kim, Bilge Soran, Raghuraman Krishnamoor- thi, Mohamed Elhoseiny, and Vikas Chandra
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2c8a76bc-621d-4092-acb6-e445bb50f265 · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Moviechat: From dense token to sparse memory for long video under- standing
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7a3f4cf5-44cc-4270-b6f8-d474997add27 · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Videoagent: Long-form video un- derstanding with large language model as agent
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a3c7cb4f-0f8f-4371-84f3-041ca0843897 · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Tarsier: Recipes for training and evaluating large video description models, 2024
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bce10fed-84dc-4364-9e42-6b80d4908e68 · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08c727bb-4fa1-4e23-83e3-4214c8c997a3 · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89ee353c-6709-4e50-adad-69c4ec3deb23 · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Lifelongmem- ory: Leveraging llms for answering queries in long-form egocentric videos
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 35c81d8f-7884-47de-9fd4-8725fa37067a · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6caa70b-ae80-4def-aee5-4b9b0d1a4316 · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering NExT-QA: Next phase of question-answering to explaining temporal actions
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 058e9ed9-55b4-4611-9aa0-7e1994c72b8a · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5040b133-e85b-4888-86ab-2128b412cfc9 · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Self-chained image-language model for video localization and question answering
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0e88c148-a2b3-47b4-a6da-388ddad66e52 · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering A simple llm framework for long-range video question-answering
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 09e99bca-4720-465d-9119-c997c76926be · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Video-llama: An instruction-tuned audio-visual language model for video un- derstanding
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ac40e92a-6de6-4aa7-bede-adcf353767c4 · outbound
VideoMultiAgents: A Multi-Agent Framework for Video Question Answering Hcqa @ ego4d egoschema challenge 2024, 2024
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bc3b5004-dfc5-4e02-a6dc-af4e8fc96a92 · inbound
DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 VideoMultiAgents: A Multi-Agent Framework for Video Question Answering
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f87ec8e5-4f9f-4518-a73d-372fddc88871 · inbound
AVATAAR: Agentic Video Answering via Temporal Adaptive Alignment and Reasoning VideoMultiAgents: A Multi-Agent Framework for Video Question Answering
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b91c65b1-a8d0-4f39-9fc3-459a2d524396 · inbound
GLANCE: A Global-Local Coordination Multi-Agent Framework for Music-Grounded Non-Linear Video Editing VideoMultiAgents: A Multi-Agent Framework for Video Question Answering
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 40b269ad-0e41-43eb-ac0d-a2e4c554f35b · inbound
Scaling Video Understanding via Compact Latent Multi-Agent Collaboration VideoMultiAgents: A Multi-Agent Framework for Video Question Answering
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b6d20c69-d121-4d76-bb05-73b40e4804ad · inbound
Child-Oriented AIGC Video Risk Reviewing: A Benchmark and Knowledge-Supported Iterative Reasoning Framework VideoMultiAgents: A Multi-Agent Framework for Video Question Answering
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de8679bc-353a-45a9-b2ca-0d4f70d4bb3c · inbound
AgenticVAU: Multi-Agent Explore-Verify Reasoning for Video Anomaly Understanding VideoMultiAgents: A Multi-Agent Framework for Video Question Answering
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.