Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T14:22:51.060258Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 26 inbound Pith citation observations for arXiv:2412.12075.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T14:22:51.060258Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:43:37.679253Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T00:49:17.542022Z
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5572d0dc-b182-4137-bdce-d2805dd1e64d · outbound
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding Gemini: A Family of Highly Capable Multimodal Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3cf6368-4e3f-428a-bd60-eaf980c70b67 · outbound
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation daa6b931-4dae-412c-ad40-03b093044f16 · outbound
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f63edbc4-30f1-4295-a97a-ba733a4ba18f · outbound
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09751473-8f29-499f-bcff-c6650dc119af · outbound
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3961459f-b42d-4c2b-904b-f6c61d894f30 · outbound
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding TVQA: Localized, Compositional Video Question Answering
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5d29f87-55f5-4b11-b0c7-73eeaccb2ed3 · outbound
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding VideoChat: Chat-Centric Video Understanding
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ceaf6c78-9c05-4091-8cc0-a2bd2c06ba25 · outbound
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding Improved Baselines with Visual Instruction Tuning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ad0e2d1-bf2a-4d42-8477-199a3c4c9914 · outbound
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d49752f-6432-41a9-b1a4-dcbdc4b942ad · outbound
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3de094b4-c346-49fb-978a-a29965673799 · outbound
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2af867e-7a80-4ec6-a3c4-aec5c6374b46 · outbound
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding CinePile: A Long Video Question Answering Dataset and Benchmark
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5820709-7195-4e36-a83e-f2ecfae8add5 · outbound
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9587f3f-bdc5-4272-8752-2b9c817eaaea · outbound
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1116ab4d-0408-4fce-8208-1804b28f0e7d · outbound
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 091b6b9b-11a4-4835-a72c-8ac4d093758c · outbound
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d64fd37-344d-4fb0-ba99-12c2337a4ef3 · outbound
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 312c4651-ca5f-49e3-aa59-2085c7554e85 · outbound
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding LLaVA-OneVision: Easy Visual Task Transfer
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d86ed82-9362-493e-a91e-2c0729ec98b0 · outbound
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13e58c84-453f-446d-aa6e-b6ad2e415696 · outbound
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4427d20-f72d-465b-945a-e268ff6e010c · outbound
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6578085-a73b-4727-b924-867d9c0aec1c · inbound
Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fedff7f-aaec-4ec1-8767-dea3b12c325d · inbound
MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e150a606-02fb-4a4e-8b6f-c23c42111155 · inbound
DisTime: Distribution-based Time Representation for Video Large Language Models CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02223deb-a7f6-4379-8ff1-a37f4262b12d · inbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e150c5ae-358f-4b63-a745-006588ae60ac · inbound
Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d426b64f-a31c-4848-bfb2-48061f7ec977 · inbound
Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa3c841a-826e-4063-b441-59d7aee0678b · inbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0c1fecd-0d5e-46e2-96fd-1bf79aa69b5e · inbound
OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19e47eb9-e53e-4ef6-b09f-014ff4bff5ee · inbound
A Survey on Video Temporal Grounding with Multimodal Large Language Model CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60dad2f5-2f7a-48c4-8107-13874d6f42b4 · inbound
EMCompress: Video-LLMs with Endomorphic Multimodal Compression CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5b5217f4-7191-472c-8ce7-66cc9f303bd9 · inbound
REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation df201b4e-fc59-42cd-b341-1f002533f6fa · inbound
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8cb99da1-f860-46d7-964b-44a6e0bcaba9 · inbound
LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1e641f19-8de6-4e96-990a-0d80cdbc04a6 · inbound
Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and Prospects CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 267929d3-e852-4c10-a1e2-0fc519398c0b · inbound
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e610b167-f3da-415c-8fe0-d640e963b60a · inbound
Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 420a45c4-ccb8-4b28-ac77-fb6637033bd0 · inbound
Towards Temporal Compositional Reasoning in Long-Form Sports Videos CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb6e4c7b-c7d6-46af-9ab0-a9de1463aa6a · inbound
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c8e94c6d-64a2-4df4-bab7-3e695f1db87e · inbound
VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6aa79895-1600-4daf-9aec-21d29475123a · inbound
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9c8f22a7-42e6-4db0-b5e9-7ccde862993c · inbound
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f043ac01-df25-4a2e-86ed-07be12eda13a · inbound
Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6d90e396-1fe9-4aa3-8d49-02c641377eef · inbound
Rethinking RAG in Long Videos: What to Retrieve and How to Use It? CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 432ebe94-bbaf-4073-84f0-3a6c78567a71 · inbound
Incentivizing Vision Language Models to Search for Long Video Question Answering CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a86fe1e0-42ac-4ff6-a65f-c6d62ca74575 · inbound
TimeThink: Reasoning with Time for Video LLMs CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89ae2c35-e156-4fa2-8fdf-0a06d1aefbe6 · inbound
SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.