Pith. sign in

Paper Citation Record · LEDGER

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA

As of 9 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2607.16189.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.16189 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T21:09:45.965576Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved43
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0bb08b37-1d7a-4026-bb59-b6966e225fc1 · outbound

This paper cites Qwen3-VL Technical Report.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:41.763169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:41.763169Z digest=sha256:9b2d7ca82241cb58744bcab69a79710b80c458bc3db70aff4481c39d8dda19d0

Observation 2cd76cc4-84a5-4b04-a048-2c579d8112cf · outbound

This paper cites Revisiting the “video” in video-language understanding.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Revisiting the “video” in video-language understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:41.875217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:41.875217Z digest=sha256:df905341f916496b646c4be7dbb023d26ccfe1be7df281eb275232a663260ce7

Observation fbaa2b13-bfd2-4873-bcee-8ca162105cd8 · outbound

This paper cites VideoMiner: Iteratively grounding key frames of hour-long videos via tree-based group relative policy optimization.arXiv preprint arXiv:2510.06040, 2025.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA VideoMiner: Iteratively grounding key frames of hour-long videos via tree-based group relative policy optimization.arXiv preprint arXiv:2510.06040, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:41.954161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:41.954161Z digest=sha256:bebfc5cabb59d68b09425d312d7d9922dcedff3faaac9ca33d72465bde7e9933

Observation 2e9ccef5-c63d-42d8-94b8-88ec37652e5b · outbound

This paper cites CG-Bench: Clue-grounded question answering benchmark for long video understanding.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA CG-Bench: Clue-grounded question answering benchmark for long video understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:42.030375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:42.030375Z digest=sha256:a3907ca00d41c202cec7a70148933b81aec2d33e402649012cd1b7f8ada45d1a

Observation 44046130-c663-4c9c-98a3-c56a08721efa · outbound

This paper cites ShareGPT4Video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems (NeurIPS), 2024.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA ShareGPT4Video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems (NeurIPS), 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:42.195155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:42.195155Z digest=sha256:d702a76fe9b5a66e1a03f07ccdee6ec92949417e83e67cbcfc35f8044a38415a

Observation e31af0a0-7e39-488b-9e68-b986d5e74fa7 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:42.330014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:42.330014Z digest=sha256:d82133ecd3c2d434c2c5ca9155e7c71c529d5267d66426925ece5d602259734c

Observation 43099faf-9860-4b49-99d3-a85d55c1c44f · outbound

This paper cites VideoZoomer: Reinforcement-learned temporal focusing for long video reasoning.arXiv preprint arXiv:2512.22315, 2025.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA VideoZoomer: Reinforcement-learned temporal focusing for long video reasoning.arXiv preprint arXiv:2512.22315, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:42.446315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:42.446315Z digest=sha256:006f29d59e2781955a184ee2e69f7ecf63e882d737a872ef93c963637fb43350

Observation d6c9e309-563c-4c60-8957-33fe452cd3ca · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:42.508311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:42.508311Z digest=sha256:6062d87a6c0d663bfe6b236ae7ba1dd2652e899b5faa6d7845b53c02c772e5da

Observation be5e3cd1-0857-4f95-9390-550c22a657e4 · outbound

This paper cites an unresolved cited work.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:42.622277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:42.622277Z digest=sha256:8c5ff400d1d7244f3bd9102b1876ffa520b668b47b65f0d8ed065370db710faa

Observation 6babdc4c-0950-4397-8730-7ec9fca0dbaf · outbound

This paper cites ReVisionLLM: Recursive vision-language model for temporal grounding in hour-long videos.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA ReVisionLLM: Recursive vision-language model for temporal grounding in hour-long videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:42.728739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:42.728739Z digest=sha256:a9a16d557fbc1ff1467505979d7206306decda7f372ea3d0d2cd435d1f22bb9f

Observation df930178-2b3a-43b9-a0b3-2c4480914b84 · outbound

This paper cites Long movie clip classification with state-space video models.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Long movie clip classification with state-space video models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:42.910965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:42.910965Z digest=sha256:98e595a3a0ce852a81f70fd94c948d793d16cf6da52b411ab813a7d907c4e166

Observation 40fca317-5763-4c9f-bd62-79b95b4e733d · outbound

This paper cites Efficient movie scene detection using state-space transformers.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Efficient movie scene detection using state-space transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:43.043540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:43.043540Z digest=sha256:1ed7f61f1fb25b94ec3dccc638b009d5144864036e76f1a5477a15ce7a4e72a2

Observation 44921443-2857-4240-83d4-3e84c564af26 · outbound

This paper cites Video ReCap: Recursive captioning of hour-long videos.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Video ReCap: Recursive captioning of hour-long videos

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:43.128553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:43.128553Z digest=sha256:9996dd388bc307acaa0cc5cad20f70af658470c96a2b91dfa6040fa6073c0029

Observation e3fb1495-cc4e-4279-873e-002b7ba41fa5 · outbound

This paper cites BIMBA: Selective-scan compression for long-range video question answering.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA BIMBA: Selective-scan compression for long-range video question answering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:43.287488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:43.287488Z digest=sha256:ec788d450ccd3e370b9b0d2e773d0c1afe6ec20d5d21d5b8617a5cf753d5e08e

Observation 2a9c70b5-7d51-4f0b-a03b-b8e2fc64a5bd · outbound

This paper cites an unresolved cited work.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:43.382041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:43.382041Z digest=sha256:8816e18ab85d6983c56c919e6bb67c8e0505ef2ad7a025f55730a7362b880bed

Observation 2d96c88b-f6e2-46a0-8165-73e347507928 · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:43.463242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:43.463242Z digest=sha256:36536a54f3a630c7726ff51f0102c5eb0c694bf56aa88f9ceb73a996e84d56c7

Observation ca4c39a9-0a3f-4708-8551-7ace2a98966f · outbound

This paper cites LLaMA-VID: An image is worth 2 tokens in large language models.European Conference on Computer Vision (ECCV), 2024.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA LLaMA-VID: An image is worth 2 tokens in large language models.European Conference on Computer Vision (ECCV), 2024

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:43.571420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:43.571420Z digest=sha256:48ec857d5f1ff82d6488d041f725db55fe85a9a4164c44be7836cafcad31161d

Observation bbe511cd-bbb2-40ce-8e78-b68d1ce62936 · outbound

This paper cites TimeSearch-R: Adaptive temporal search for long-form video understanding via self-verification reinforcement learning.arXiv preprint arXiv:2511.05489, 2025.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA TimeSearch-R: Adaptive temporal search for long-form video understanding via self-verification reinforcement learning.arXiv preprint arXiv:2511.05489, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:43.682875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:43.682875Z digest=sha256:c13c0f5d695e8f18ac32b6dd2400962d8204b1f3bacda81270a07f4b44a578a9

Observation 84c09683-d885-4508-b4c4-a0212dd8815e · outbound

This paper cites an unresolved cited work.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:43.802551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:43.802551Z digest=sha256:c9f0b7963d6fb41f7b5868e540fc98d5a13863c6e827ed5938d035cb20e18377

Observation 62e4b502-7a54-4aea-8e2c-6dcf227a67f6 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Learning transferable visual models from natural language supervision

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:43.929348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:43.929348Z digest=sha256:320e5cd83cf676e5a5281b2706d0018b0e1c5a0faab2d37ab6fd900f37797202

Observation 3f14f7b0-4d50-4382-8e79-5db2ade9af63 · outbound

This paper cites Understanding long videos in one multimodal language model pass.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Understanding long videos in one multimodal language model pass

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:44.052763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:44.052763Z digest=sha256:09727365046a1da0b2322914dc984aaa0c9ec3cd052636587b94dd4413644245

Observation fdeeffc3-46f4-4f02-b7e9-6321d1ef67df · outbound

This paper cites an unresolved cited work.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:44.148009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:44.148009Z digest=sha256:5be1986f21e6f861f30ecf4e2bd3a04e0ce91b96f506401fa12dce972830ef6f

Observation b7a55875-f353-4b32-9ba8-3b1c1b284223 · outbound

This paper cites xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:44.272531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:44.272531Z digest=sha256:9226fb4d81cf99d5fe289e6dad8ef612c04fc5de66eb8769a712332182b75019

Observation 751eb134-49f0-4ee9-953f-35fff8f7cfeb · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:44.385817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:44.385817Z digest=sha256:d6911ab7a1da6d411712a8a47b46071b7989fab482ef78e8788ad79b8ba65474

Observation b5d2c529-0028-4e4a-b1de-c4ff5247ef5c · outbound

This paper cites MovieChat: From dense token to sparse memory for long video understanding.Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA MovieChat: From dense token to sparse memory for long video understanding.Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:44.512577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:44.512577Z digest=sha256:785d1f2a8df526f1e4ca6f5ec8283b2a4d02851c589793f9aa447806f9be03ed

Observation 4955792f-f0e2-48f5-b1f6-b3a1c10e2f66 · outbound

This paper cites Adaptive keyframe sampling for long video understanding.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Adaptive keyframe sampling for long video understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:44.636461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:44.636461Z digest=sha256:04b2726ed6899fb4a7cc848c3dcabcf3e3da77179803c10667a5e60e4b13367e

Observation 4a7e6fa5-5466-454f-8284-109e92a63dce · outbound

This paper cites Alvarez, Lei Zhang, and Zhiding Yu.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Alvarez, Lei Zhang, and Zhiding Yu

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:44.742647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:44.742647Z digest=sha256:34646ce38e706105b54e50721a09516a27c79234c8aae3fe1ca96cbb0a9734cc

Observation d06acb4e-218a-443d-ad58-723e803b012f · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA LVBench: An Extreme Long Video Understanding Benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:44.860609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:44.860609Z digest=sha256:beca28649a2dade35aa44b5023d537d6671d0fa2f90ba417359e94f9db8fe9f5

Observation 659a6dfb-9521-4811-ab31-a4ca266faee6 · outbound

This paper cites AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:44.952695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:44.952695Z digest=sha256:df978ece841fbad1c5eedb97fc7b76df04079efec30b308d49954bedbfe8a589

Observation c4169a6b-2f5d-4b69-910a-9db16aedc76f · outbound

This paper cites Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:45.070624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:45.070624Z digest=sha256:30135653295827001b47d7f4f71d8691b623a3d3c8bac3117b0949408f7cce2f

Observation ca71d312-04bf-4640-9f7a-976d50d0a3c2 · outbound

This paper cites VideoTree: Adaptive tree-based video representation for LLM reasoning on long videos.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA VideoTree: Adaptive tree-based video representation for LLM reasoning on long videos

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:45.141034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:45.141034Z digest=sha256:8fd595855ef66a83165b7955dffdeb87e6de1d43f4d09cae2626976d79dccda7

Observation b2160f28-107f-4777-aa54-4fcf8fa980b3 · outbound

This paper cites Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:45.205921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:45.205921Z digest=sha256:023f1d121ba10e86220742dc18a5fb0d777fda4efa549cca5f6f460e6a21135e

Observation 47ada167-aec7-4386-aa2c-5396ba2e10b0 · outbound

This paper cites LongVideoBench: A benchmark for long- context interleaved video-language understanding.Advances in Neural Information Processing Systems (NeurIPS), 2024.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA LongVideoBench: A benchmark for long- context interleaved video-language understanding.Advances in Neural Information Processing Systems (NeurIPS), 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:45.254499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:45.254499Z digest=sha256:6aaf3c3a68619ee37d4807ac3e216592208401bdae9c1d1c34405d2a9f213925

Observation 7ffb45fa-d594-474d-aa85-e13fd2a50c35 · outbound

This paper cites LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:45.294808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:45.294808Z digest=sha256:ae4639a700a70a2739b94ecd6fa99aadce18d039de1eb770b2f2b40ba1e1be81

Observation 02b1174a-3beb-4864-aa68-2cf561590e83 · outbound

This paper cites Generative Frame Sampler for Long Video Understanding.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Generative Frame Sampler for Long Video Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:45.384742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:45.384742Z digest=sha256:7dea550c3ff8b6c23b204862233a06bfea6cda549a9d3e55f65ba62fd426d223

Observation 3bb651f5-d098-4d35-8655-7cfa911c0247 · outbound

This paper cites T*: Re-thinking Temporal Search for Long-Form Video Understanding.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA T*: Re-thinking Temporal Search for Long-Form Video Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:45.454954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:45.454954Z digest=sha256:ebdca5cd4fa42b852babd408824643b1a52a6048fd0ad28c519733ef132f658e

Observation c761285d-106e-48eb-a426-a80e90c012f3 · outbound

This paper cites MomentSeeker: A benchmark for long-video moment retrieval.arXiv preprint arXiv:2502.12558, 2025.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA MomentSeeker: A benchmark for long-video moment retrieval.arXiv preprint arXiv:2502.12558, 2025

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:45.511098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:45.511098Z digest=sha256:8b0994a2b9d2284f9438bcc02fde5f6cc15c16378d64929f3266afa1aa9b4218

Observation 2d04a5fb-0e4a-4430-bed8-93c0705a9293 · outbound

This paper cites Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Video-o3: Native Interleaved Clue Seeking for Long Video Multi-Hop Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:45.593997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:45.593997Z digest=sha256:b948c36f0b7826593779cc1d2676fa3e25f2c11c67bef12ad40b06d631d8ff4a

Observation 80010330-3aa0-499f-8133-04338642dd97 · outbound

This paper cites A simple LLM framework for long-range video question-answering.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA A simple LLM framework for long-range video question-answering

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:45.641212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:45.641212Z digest=sha256:855cc322a19dbc1549ed9f144b09a7b8c7789f9a07b7556d550723cd58b0af05

Observation bc8524ad-ed63-4605-8c9d-a10dec30233e · outbound

This paper cites SiLVR: A Simple Language-based Video Reasoning Framework.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA SiLVR: A Simple Language-based Video Reasoning Framework

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:45.691759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:45.691759Z digest=sha256:a96e7e44fc6323998ec99d15552ea26f0c17c439c1a067485cb5a44516528e21

Observation 1a0e3e4e-fe69-4804-b274-743ef21e3497 · outbound

This paper cites Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:45.802892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:45.802892Z digest=sha256:7842c7a24d3df1c5ac9149aecae1d6da50d9a9f185cbb595c5b7168d0628b3e8

Observation d57e3b34-6e00-4cb7-8c0b-ecef8bc297ba · outbound

This paper cites Long Context Transfer from Language to Vision.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA Long Context Transfer from Language to Vision

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:45.855573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:45.855573Z digest=sha256:1d344f25cdb66bb54632f424470021abcead59e2f7c362e164f8f821520a91e1

Observation efea4c79-7408-4801-a2dd-e4c8152affec · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA MLVU: Benchmarking Multi-task Long Video Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T21:09:45.916527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:45.916527Z digest=sha256:00d60d06ec5937f4d33fc32c1d68264ec83c81ed2f015355bff2c5a3def40613

Observation 3752c4c4-f498-4b2e-8689-473e325b665f · outbound

This paper cites User: Current Segment [t s-te]:.

Searching Videos as Trees: Self-Correcting Agents for Grounded Long Video QA User: Current Segment [t s-te]:

Reference 44

Resolution
malformed identifier
no resolver link, observed 2026-08-01T21:09:45.965576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:09:45.965576Z digest=sha256:a875802a7f919e69615d3b4307050597b5d51949d57fae7a2704a2a29e6e522a

Pith citing papers

No inbound Pith citation observations are available.