Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T21:09:45.965576Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2607.16189.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T21:09:45.965576Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0bb08b37-1d7a-4026-bb59-b6966e225fc1 · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cd76cc4-84a5-4b04-a048-2c579d8112cf · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Revisiting the “video” in video-language understanding
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbaa2b13-bfd2-4873-bcee-8ca162105cd8 · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA VideoMiner: Iteratively grounding key frames of hour-long videos via tree-based group relative policy optimization.arXiv preprint arXiv:2510.06040, 2025
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e9ccef5-c63d-42d8-94b8-88ec37652e5b · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA CG-Bench: Clue-grounded question answering benchmark for long video understanding
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44046130-c663-4c9c-98a3-c56a08721efa · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA ShareGPT4Video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems (NeurIPS), 2024
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e31af0a0-7e39-488b-9e68-b986d5e74fa7 · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43099faf-9860-4b49-99d3-a85d55c1c44f · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA VideoZoomer: Reinforcement-learned temporal focusing for long video reasoning.arXiv preprint arXiv:2512.22315, 2025
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6c9e309-563c-4c60-8957-33fe452cd3ca · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Video-R1: Reinforcing Video Reasoning in MLLMs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be5e3cd1-0857-4f95-9390-550c22a657e4 · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6babdc4c-0950-4397-8730-7ec9fca0dbaf · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA ReVisionLLM: Recursive vision-language model for temporal grounding in hour-long videos
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df930178-2b3a-43b9-a0b3-2c4480914b84 · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Long movie clip classification with state-space video models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40fca317-5763-4c9f-bd62-79b95b4e733d · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Efficient movie scene detection using state-space transformers
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44921443-2857-4240-83d4-3e84c564af26 · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Video ReCap: Recursive captioning of hour-long videos
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3fb1495-cc4e-4279-873e-002b7ba41fa5 · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA BIMBA: Selective-scan compression for long-range video question answering
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a9c70b5-7d51-4f0b-a03b-b8e2fc64a5bd · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d96c88b-f6e2-46a0-8165-73e347507928 · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca4c39a9-0a3f-4708-8551-7ace2a98966f · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA LLaMA-VID: An image is worth 2 tokens in large language models.European Conference on Computer Vision (ECCV), 2024
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbe511cd-bbb2-40ce-8e78-b68d1ce62936 · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA TimeSearch-R: Adaptive temporal search for long-form video understanding via self-verification reinforcement learning.arXiv preprint arXiv:2511.05489, 2025
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84c09683-d885-4508-b4c4-a0212dd8815e · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62e4b502-7a54-4aea-8e2c-6dcf227a67f6 · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Learning transferable visual models from natural language supervision
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f14f7b0-4d50-4382-8e79-5db2ade9af63 · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Understanding long videos in one multimodal language model pass
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdeeffc3-46f4-4f02-b7e9-6321d1ef67df · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7a55875-f353-4b32-9ba8-3b1c1b284223 · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 751eb134-49f0-4ee9-953f-35fff8f7cfeb · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5d2c529-0028-4e4a-b1de-c4ff5247ef5c · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA MovieChat: From dense token to sparse memory for long video understanding.Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4955792f-f0e2-48f5-b1f6-b3a1c10e2f66 · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Adaptive keyframe sampling for long video understanding
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a7e6fa5-5466-454f-8284-109e92a63dce · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Alvarez, Lei Zhang, and Zhiding Yu
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d06acb4e-218a-443d-ad58-723e803b012f · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA LVBench: An Extreme Long Video Understanding Benchmark
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 659a6dfb-9521-4811-ab31-a4ca266faee6 · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4169a6b-2f5d-4b69-910a-9db16aedc76f · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca71d312-04bf-4640-9f7a-976d50d0a3c2 · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA VideoTree: Adaptive tree-based video representation for LLM reasoning on long videos
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2160f28-107f-4777-aa54-4fcf8fa980b3 · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47ada167-aec7-4386-aa2c-5396ba2e10b0 · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA LongVideoBench: A benchmark for long- context interleaved video-language understanding.Advances in Neural Information Processing Systems (NeurIPS), 2024
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ffb45fa-d594-474d-aa85-e13fd2a50c35 · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02b1174a-3beb-4864-aa68-2cf561590e83 · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Generative Frame Sampler for Long Video Understanding
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bb651f5-d098-4d35-8655-7cfa911c0247 · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA T*: Re-thinking Temporal Search for Long-Form Video Understanding
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c761285d-106e-48eb-a426-a80e90c012f3 · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA MomentSeeker: A benchmark for long-video moment retrieval.arXiv preprint arXiv:2502.12558, 2025
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d04a5fb-0e4a-4430-bed8-93c0705a9293 · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80010330-3aa0-499f-8133-04338642dd97 · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA A simple LLM framework for long-range video question-answering
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc8524ad-ed63-4605-8c9d-a10dec30233e · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA SiLVR: A Simple Language-based Video Reasoning Framework
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a0e3e4e-fe69-4804-b274-743ef21e3497 · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d57e3b34-6e00-4cb7-8c0b-ecef8bc297ba · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Long Context Transfer from Language to Vision
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efea4c79-7408-4801-a2dd-e4c8152affec · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA MLVU: Benchmarking Multi-task Long Video Understanding
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3752c4c4-f498-4b2e-8689-473e325b665f · outbound
Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA User: Current Segment [t s-te]:
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.