Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2403.10517.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T23:55:58.560208Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T06:39:37.670946Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 6924df9c-13db-4996-b4db-9902ed93ab8a · inbound
MLVU: Benchmarking Multi-task Long Video Understanding VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 01e938e7-3868-42ae-998d-6a3b0f3a7954 · inbound
Progress-Aware Video Frame Captioning VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 389f5ced-991c-420c-addf-8fc75f0483c4 · inbound
Towards Long Video Understanding via Fine-detailed Video Story Generation VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 106
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d6627f5-da75-4f28-999a-4b5e21dc0f04 · inbound
V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 127
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f10a286-f30b-4dcf-a860-88cc3fc88c7d · inbound
Apollo: An Exploration of Video Understanding in Large Multimodal Models VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab93c02c-3f04-47d6-b5e6-5f508bb496a5 · inbound
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d418d1fe-bf5c-4005-82bb-888a32e90f20 · inbound
Four Eyes Are Better Than Two: Harnessing the Collaborative Potential of Large Models via Differentiated Thinking and Complementary Ensembles VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de31d2bc-24b0-4290-928d-141577e8d402 · inbound
HCQA-1.5 @ Ego4D EgoSchema Challenge 2025 VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7246b089-de6e-49b3-9bd5-6dd77cde997e · inbound
Frame-Level Captions for Long Video Generation with Complex Multi Scenes VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0c07cc0-f991-4305-80a8-37b2a2c996c5 · inbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6441a79e-3fa6-448a-ac90-99b4a81b0d0b · inbound
TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a207b0b-a775-4415-aac5-094f9018a8c6 · inbound
Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0c89a3f-5ec9-4803-ab29-2c4119d62c42 · inbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e77ea6c7-fc4c-4ce1-9bb7-cd6e577804e5 · inbound
Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6195eb98-5b57-4afb-98cc-49ea563e4268 · inbound
NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video Understanding VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1154590-e902-4fc2-94e3-dfcfe8eeecbe · inbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81dfd892-5e54-4733-8e46-4d28e297f7ed · inbound
Towards Sparse Video Understanding and Reasoning VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f41f82df-f88f-4bcf-98de-0bb210df5f4f · inbound
Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ff412391-d975-4854-8d2f-14f05c1657cc · inbound
SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation bd15e77c-f160-475d-99b8-5cb61f3dc4b4 · inbound
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d3ae6187-b553-4978-a25b-e0b33a43737c · inbound
Parameter-Efficient Multi-View Proficiency Estimation: From Discriminative Classification to Generative Feedback VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f76a6309-9f2a-4bf7-b4e9-9c9b3984d5d0 · inbound
MARS: Technical Report for the CASTLE Challenge at EgoVis 2026 VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3e3da286-8c98-4c32-9f42-4c1775aaebef · inbound
UNIVID: Unified Vision-Language Model for Video Moderation VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fe7ae386-fb18-4eae-8738-9287df3bb9c2 · inbound
HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 181
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.