Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:25:33.395157Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 8 inbound Pith citation observations for arXiv:2504.15681.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:25:33.395157Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T00:03:32.341621Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T20:46:13.651443Z
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 75545c5d-9008-4c3c-bb3d-62c7e814ba34 · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing Gemini: A Family of Highly Capable Multimodal Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1b128ed-282b-4a18-be43-1223ecefd36d · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dcb4d13-fd59-4cc3-98b2-9741a3fb604d · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 908c11c5-134e-461f-9765-87df8cc8216b · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a55b34b9-08aa-4fbb-b82a-9c9a86b8f33a · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfdf50ae-f463-4b04-ac97-3d186445b46a · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing TALL: temporal activity localization via language query
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 214ad4d7-1925-4596-82c5-de054e189fc1 · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f40f9da-8876-4edc-b47d-0853101dd3e1 · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing Fullstop: Multi- lingual deep models for punctuation prediction
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation eca0cda8-904b-41e6-bd67-3ee360580d5f · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2f08ae5d-be3a-4036-8e2e-b5338d041836 · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing GPT-4o System Card
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 993668c6-c7f5-494d-b87e-68b41cfece4f · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing Mistral 7B
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 925c9294-5948-4462-937f-d5de7c897a1f · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing Dense-captioning events in videos
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 52eb6fbe-0c18-4296-879c-85f36cf4b945 · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing D-Attn: Decomposed Attention for Large Vision-and-Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4de6e23-b1fc-47d4-86dc-8339a6c0d366 · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing Berg, and Mohit Bansal
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 921dc973-69cd-4219-bec8-e91e21bac565 · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing LLaVA-OneVision: Easy Visual Task Transfer
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79db2b3c-7825-41f7-b54c-cabe13a8bc0e · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing Structured Context Transformer for Generic Event Boundary Detection
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef0b583b-9242-4a78-a9a7-563b4c57b1f7 · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5688aec5-1346-4d95-b5d5-46b36a739e0c · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation cc2472a3-33a7-4874-8443-8ee7660d6d0c · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing Visual instruction tuning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4212afa9-c0d8-412e-92ff-3a793e0328f8 · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing Decoupled weight decay regularization
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7e6fb3ee-1f7a-4dc7-9853-8e70650fb3ec · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing ZoomV: Temporal Zoom-in for Efficient Long Video Understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aced7a25-51c5-4384-9cf3-48c9aa8dc898 · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing Robust speech recognition via large-scale weak supervision
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 03452171-a7be-4592-a52e-f95ebc2b79f9 · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing CinePile: A Long Video Question Answering Dataset and Benchmark
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 175a2ce4-8450-4f99-8052-409ba80863a3 · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing Timechat: A time-sensitive multimodal large language model for long video understanding
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4feb0279-91c2-4237-ab43-02099d89dfe7 · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing Gemma 2: Improving Open Language Models at a Practical Size
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 550b9d51-acc6-4c81-ab00-9b987a757ac1 · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing Moviechat: From dense token to sparse memory for long video understanding
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation dbe16d8b-ee4a-4e05-ac19-985e8cd3a226 · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a0225ed-6fde-4a55-91fd-4881953bc472 · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing Gomez, Lukasz Kaiser, and Illia Polosukhin
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5edaf4e0-9884-47d4-bbe6-5eda2d6cf1f4 · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 654cb160-705f-458e-947d-6be61b1082fa · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6c0f6e6-1ee6-4bd0-85f4-c9cdd88e94e5 · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing LVBench: An Extreme Long Video Understanding Benchmark
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5303013-c728-4eb4-9848-818e6ea3fa1e · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing Chi, Quoc V
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e7b47912-3644-4c6b-96a0-0d3c51f5c5be · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing Longvideobench: A benchmark for long-context interleaved video-language understanding
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3774260e-0ffa-4758-8876-f4d5e768af80 · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing T*: Re-thinking Temporal Search for Long-Form Video Understanding
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21125145-8afe-4382-a0dd-731c852c6066 · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing Sigmoid loss for language image pre-training
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation cce8b1d4-b89a-4d8d-933d-e1a496a1b9a1 · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a27b6c80-bc29-42e8-9a38-94194eee2fb4 · outbound
Vidi: Large Multimodal Models for Video Understanding and Editing InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c784aa37-b1ec-47ec-b4c2-37215622e784 · inbound
Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection Vidi: Large Multimodal Models for Video Understanding and Editing
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1dcda03-a4ba-4173-8a0b-7819012ba321 · inbound
Timeripple: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space Vidi: Large Multimodal Models for Video Understanding and Editing
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60d8b032-c162-4b77-b4cf-53ec4b9aae5d · inbound
EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation Vidi: Large Multimodal Models for Video Understanding and Editing
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3895ec79-b6ea-41d3-bc29-d9331292d0d8 · inbound
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Vidi: Large Multimodal Models for Video Understanding and Editing
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 20247188-fc73-4b75-8c2c-6c629235a58c · inbound
StoryTR: Narrative-Centric Video Temporal Retrieval with Theory of Mind Reasoning Vidi: Large Multimodal Models for Video Understanding and Editing
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c6fcef02-2efa-4065-b867-a466e360b740 · inbound
VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning Vidi: Large Multimodal Models for Video Understanding and Editing
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8536bac6-8a77-49a4-be09-1afb22f38f08 · inbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Vidi: Large Multimodal Models for Video Understanding and Editing
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48bf1810-df25-4868-a715-14b523634f5d · inbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Vidi: Large Multimodal Models for Video Understanding and Editing
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.