Pith. sign in

Paper Citation Record · LEDGER

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding

As of 21 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2507.02946.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02946 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:59:17.302898Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2df4ca94-0573-4ef3-8129-00b94da7f0f4 · outbound

This paper cites Qwen2.5-VL Technical Report.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:15.678922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:15.678922Z digest=sha256:9d244267bde52f91c1925d7103072c3460cfbe189561958e0b36681228def5a3

Observation 29312749-d544-4313-a39f-c76c0a07fece · outbound

This paper cites AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:15.758936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:15.758936Z digest=sha256:ad48e0555ab1d7eb0cc42dd817b4c734e253961895e8015a7269f21bd914e2ed

Observation b4e04e31-54a1-4856-ae77-3dd741cf5db7 · outbound

This paper cites Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:15.855545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:15.855545Z digest=sha256:9af41d772e8aa995cdf73348605744c97b175609e9f995182a58d54a21b5ef97

Observation f02e1bab-94e6-4fcd-92b6-0f79569b36ec · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.010886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.010886Z digest=sha256:a814fd44d776ab905b4082ecc1d4354381e022bbece4eb1f516655262feeff14

Observation 48c49d92-36e3-48a5-acaf-059faa0018a9 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.134001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.134001Z digest=sha256:6e3a86890bc0746c01439a5a3f68b216c7b7997375f33bf9d9699368fdcf0813

Observation 6048adb1-94be-4ff2-b950-086bcb07d394 · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Reasoning with Language Model is Planning with World Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.221161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.221161Z digest=sha256:81a9f7e06ed79750f641a4beb7f0b84ada4055baf0f7ab7e971989bd6a18a8d7

Observation 8faf21c6-a3bf-4ade-a893-cd6fe75dedf0 · outbound

This paper cites From Image to Video, what do we need in multimodal LLMs?.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding From Image to Video, what do we need in multimodal LLMs?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.278173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.278173Z digest=sha256:8433a08f574df7f057bc89ee0b151ba0e3b5836e12edde33ee0231d96ada46a5

Observation 110c9dea-d75d-4a06-9dd8-94baa733e875 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.444148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.444148Z digest=sha256:0b759a95eaecca3c478dbaedbc8d468ba1850452372ad8e2712d51e8af34f74e

Observation 08d20e2c-0330-46ff-9b67-fe558e1ebfba · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.488837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.488837Z digest=sha256:3e64aada296e75761075cba2237709cb3f2eb05202e2f66d60836a26649d2099

Observation a2ad6307-3313-41c9-887d-fdf025e90a0e · outbound

This paper cites PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.530036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.530036Z digest=sha256:0f56d14fc494c9f22bc870f6c5eab61e6afaa65b0f2cb8ad15d44eea6711e584

Observation b822c8c7-c8c3-4c25-8055-bba79cbf1794 · outbound

This paper cites Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.614884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.614884Z digest=sha256:13c69e939eb19843534c8daf194abb53baf1caf25b42fedf3a6bbc279a66d380

Observation 7d870a68-30b1-48c1-adc5-d6dd3459f69a · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.696016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.696016Z digest=sha256:87c8c1dabb386f253d4bf6c72f124c3978b8fbac1bac1f5a749f6d8314c32ce2

Observation ca388100-361a-4ccc-b685-43025ec203b4 · outbound

This paper cites ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.773771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.773771Z digest=sha256:c1580b30827b531317b0b4faac4b29a7dc5de03fe6e790a76fcd1f9b8c759b39

Observation a03ecc67-4006-4690-9bf6-bdae02f2ad6f · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.899236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.899236Z digest=sha256:ed42ca53bdf408f2feb104ccd6f9d5059d0d208a1ebffb1d7aea00b2a62ab26e

Observation 91492b46-e8be-4b35-8952-a6a972a07059 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.981202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.981202Z digest=sha256:08018c6e656e0ba130d645a3f29061e66012a621588d3fa0a0582af21ae8ae0f

Observation 587ba512-870c-4a04-bbaa-48b8069a308f · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:17.153103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:17.153103Z digest=sha256:5c0a54de5da82ee677b8b73dc98f97edc9797f19492d758aba41b9d701af6094

Observation 53a350c1-9421-43e5-a25b-1b0516cc88a1 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:17.302898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:17.302898Z digest=sha256:2c43b5532daf993241b772903b6039e07e1486384df93bc7aad6eb0e93888fcd

Observation ca2d00a3-b6ba-4815-aa27-272eeedd7da0 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 1985

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.364927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.364927Z digest=sha256:2b823b2e9d44b290767698f34d2306c41f4e0c60f979a3eaf700a46ba703d539

Observation a3101767-ba83-4287-90e5-94ed395b9614 · outbound

This paper cites Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:17.224605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:17.224605Z digest=sha256:e60ab22af5b35f215dff6dc78d78c254ed0dd89be85473cbd6617f5e10236a5a

Observation ec3ffded-ef9e-43d7-ad38-a67c8cf17179 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:17.058810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:17.058810Z digest=sha256:673265c1b86715ccda0da341e6f8a83e92481a608df7f4f05bc71a6db2ef1a49

Observation a3e1ddd9-6e69-43fd-9211-6e2a79ced978 · outbound

This paper cites GPT-4 Technical Report.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding GPT-4 Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:15.515921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:15.515921Z digest=sha256:8c86c7591879273db964589026f13e57090227bb8461609758d319fdfaae0930

Observation 07a3c23f-0848-4e87-9a93-13c1267286ba · outbound

This paper cites MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:15.950597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:15.950597Z digest=sha256:0afa7ee09fc4ed0ea0a160c7c0293783f41b4c1b78a96bdedf9aaba56cc514a5

Observation 0ebdbcd1-1622-47c7-8529-22f36dd1d1dc · outbound

This paper cites DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:15.578220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:15.578220Z digest=sha256:50ddeee6d269761f6b4c4c84df4ec5962d66493eb2434c8171ffa72d230336b6

Pith citing papers

No inbound Pith citation observations are available.