Pith. sign in

Paper Citation Record · LEDGER

VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2405.19209.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.19209 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T21:21:45.795412Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T07:55:33.235804Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 66e6dcd5-d004-49f6-b368-7eaf9ffc69f5 · inbound

$\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation cites this paper.

$\infty$-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T21:21:45.795412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T21:21:45.795412Z digest=sha256:40c7e687e3820ac17f5d8ca64f9ca57ae3134e0f30839a5c214530976c3293c6

Observation 497456f7-0595-408d-ab9a-510bbceffe90 · inbound

AirVista-II: An Agentic System for Embodied UAVs Toward Dynamic Scene Semantic Understanding cites this paper.

AirVista-II: An Agentic System for Embodied UAVs Toward Dynamic Scene Semantic Understanding VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:55:33.239119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T07:51:33.652470Z digest=sha256:4252ea839975d6c7f8da50d3eb9f4d49424efcf7eb1a551809bdc574c45d8572

Observation 2fc4f736-9708-4578-9ce3-60917c7ecfe6 · inbound

Uneven Event Modeling for Partially Relevant Video Retrieval cites this paper.

Uneven Event Modeling for Partially Relevant Video Retrieval VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:13.027251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:13.027251Z digest=sha256:51778e7569bd0c8ffb766f96920ace1e1902d51cc4dc28d0f00ffe3cda6eb07a

Observation 5c494f6b-4435-49f1-9976-d8d62903bfc5 · inbound

TriPSS: A Tri-Modal Keyframe Extraction Framework Using Perceptual, Structural, and Semantic Representations cites this paper.

TriPSS: A Tri-Modal Keyframe Extraction Framework Using Perceptual, Structural, and Semantic Representations VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:41.249261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:41.249261Z digest=sha256:cec99b402b3eaf2c9c09a1f46e702e5568fc1a324884b04eb216b68d1461940c

Observation bae2896b-f6e0-4548-b8bc-34a2e87c3ace · inbound

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval cites this paper.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.825064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.825064Z digest=sha256:eabebbabed82de786b416e86778a45db3dc917b9f8b0976b31ba3371e8389be8

Observation bafa5580-7d0f-462f-a507-2c1989dff989 · inbound

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding cites this paper.

Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:56.990818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:56.990818Z digest=sha256:f8b4f751a9c35b8ef94e52f040a2fbd2081036f981bbeb04766619e402f60432

Observation 28f45b09-2bf9-437a-841d-7b2e7489b3af · inbound

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks cites this paper.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.428376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.428376Z digest=sha256:309d2a9b3e259617817e65bdd71464c213b02bbe92370e4d154558e398ac09ea

Observation 392bf377-9646-4ff1-87ec-15f8401b57b1 · inbound

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision cites this paper.

TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:54:24.495885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:54:24.495885Z digest=sha256:b7342fbc391177002d92edd850d002132877e8e8b39b91b4ee1d4f433c0a3204

Observation fd24e478-93f4-46ad-87fa-1ca87619db4e · inbound

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning cites this paper.

Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T00:34:36.343155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:34:36.343155Z digest=sha256:3e88ce4e146e988ea57f146109d3be0750e1c5d9dad74e354348be568f7ac974

Observation 283646a8-8b70-457c-b779-f6eb04f50ef7 · inbound

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 cites this paper.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:13.236841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:13.236841Z digest=sha256:7da6a9bd92da47a5a3fdb4afc57e325a7db7ce3e9a3b07574e440a38db000ce7

Observation f1cc9b77-687e-4840-91f8-69b982c35dc5 · inbound

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames cites this paper.

Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T21:06:26.787020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:06:26.787020Z digest=sha256:fd3c5a2c6bdc8c3672b22e4cf53264443c97bf2d9c7938e8d48b020d3d67a39e

Observation 5d65b882-b812-4eb2-9817-d7922d9a8c0b · inbound

Infinite Video Understanding cites this paper.

Infinite Video Understanding VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:10:14.223529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:10:14.223529Z digest=sha256:f6313190c7e67be09364397b8a5b56dd557afbac39b9e9c888a177b8aee2584a

Observation 315dcd01-b057-4bad-816e-2f20f409bed2 · inbound

LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering cites this paper.

LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:53:34.166522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:53:34.166522Z digest=sha256:298f5b5c230300e7f29688a7dd4f30b9c4a1d446f31a0072fc904609f5a87bdd

Observation 8aa244f2-782b-4304-88de-19c590604bd3 · inbound

ProPy: Building Interactive Prompt Pyramids upon CLIP for Partially Relevant Video Retrieval cites this paper.

ProPy: Building Interactive Prompt Pyramids upon CLIP for Partially Relevant Video Retrieval VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T16:05:18.756354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:05:18.756354Z digest=sha256:32dc9c73fc9da0f522e5bd4a06673c4e22b5b62a52fa2a9eaed37b105433c213

Observation d7f41027-5e6f-4fd2-b83f-dcac4abe59d1 · inbound

CAViAR: Critic-Augmented Video Agentic Reasoning cites this paper.

CAViAR: Critic-Augmented Video Agentic Reasoning VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T21:27:35.899881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T21:27:35.899881Z digest=sha256:7f8d55ee18d3e420b794d78a70605a1eedad5082a780472c2737a90b2b27d17b

Observation ecc2491c-45b2-4cab-bf74-e1e0057f03fa · inbound

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding cites this paper.

Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T18:22:31.365284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:22:31.365284Z digest=sha256:522f3e7053b4da740626536e9192e4a832a6b5ee8b04428fc7424e1d3ea43e63

Observation 5f1de1df-84e0-4e28-bce3-6e91b5a6de45 · inbound

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding cites this paper.

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:57:53.878322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T12:55:04.564442Z digest=sha256:eb88ecc167a601638d9b256e0c08e35d39ba9083db8bd592875d50c2867d7077

Observation adbb7cc1-15e2-4710-80ed-7eb32774ec8f · inbound

Towards Sparse Video Understanding and Reasoning cites this paper.

Towards Sparse Video Understanding and Reasoning VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T23:33:14.887050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:33:14.887050Z digest=sha256:bb4cc077341f93a67d33b94b6b3b885faa3229afd2a22c4d20a761e532b49b9b

Observation 23bca80f-f2ec-488a-9935-5b6b2802f143 · inbound

Progressive Video Condensation with MLLM Agent for Long-form Video Understanding cites this paper.

Progressive Video Condensation with MLLM Agent for Long-form Video Understanding VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:43:14.883302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T20:40:41.380829Z digest=sha256:40fe4e1c05d47295daac531c81c53ca71ef0c656f6736bdb165d2be290ad1ebf

Observation 7ffd72af-016d-45fa-8211-2631d70216db · inbound

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration cites this paper.

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:48.953290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T20:20:08.590407Z digest=sha256:aae8799f4d5560a632eaa9ea6ac0be83f1886d9b00c77f4b843e422bd9096e3e

Observation d0820b5f-edbd-4f18-a35e-41b19a32e8bc · inbound

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding cites this paper.

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:30:26.614514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:28:58.920442Z digest=sha256:82547332abae73e9962648ef19a6e77a0ec3cda3c7a2e2c14e3d9579c50cf603

Observation 01fe7a43-bdd3-4e21-a48d-92919b0efde8 · inbound

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models cites this paper.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:31.687909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:6af6b0ee536a752274d42f84f5aaa3eb5d26452b0db2ae115debf1b595f666be

Observation d6fa54ed-9f06-42b7-bc43-25a5d061f534 · inbound

PhysMRV: Physical Memory Retrieval and Verification for Physics Plausibility Reasoning cites this paper.

PhysMRV: Physical Memory Retrieval and Verification for Physics Plausibility Reasoning VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T13:38:16.395135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T13:38:16.395135Z digest=sha256:da8536716f3e6a6e842f9b5bac1dd37dd96a1efb42928566a29768a89fefbcc5