Pith. sign in

Paper Citation Record · LEDGER

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding

As of 9 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2507.02946.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02946 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:59:17.302898Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2df4ca94-0573-4ef3-8129-00b94da7f0f4 · outbound

This paper cites Qwen2.5-VL Technical Report.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:15.678922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:15.678922Z digest=sha256:57a7164ff497a6143a635315d2a09cf99af7579118600b11c85c26fd3ed7bace

Observation 29312749-d544-4313-a39f-c76c0a07fece · outbound

This paper cites AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:15.758936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:15.758936Z digest=sha256:f3d152d06c450dc29a4d4d3ed5858112d22296ed4b03183b8a891943cc0ea267

Observation b4e04e31-54a1-4856-ae77-3dd741cf5db7 · outbound

This paper cites Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:15.855545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:15.855545Z digest=sha256:4cad0d7adfc04c2d0893a8f94e5127a8fe9872d478864111637f67531f90ffd5

Observation f02e1bab-94e6-4fcd-92b6-0f79569b36ec · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.010886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.010886Z digest=sha256:c1ba2c4c48cb62039846e1f4d2963d0df37c16d220d190f890d1ad869b4e1201

Observation 48c49d92-36e3-48a5-acaf-059faa0018a9 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.134001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.134001Z digest=sha256:f048a49a74ec8e394e2e429ef1e4d05f5450063a0279ea122d7e02ea983cbc20

Observation 6048adb1-94be-4ff2-b950-086bcb07d394 · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Reasoning with Language Model is Planning with World Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.221161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.221161Z digest=sha256:582f45cfc309977bad1b9dc25a1d81ae4c72a0bae60e812515f2212d94c78a26

Observation 8faf21c6-a3bf-4ade-a893-cd6fe75dedf0 · outbound

This paper cites From Image to Video, what do we need in multimodal LLMs?.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding From Image to Video, what do we need in multimodal LLMs?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.278173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.278173Z digest=sha256:3b7786b41f860aff3ee7316588aeef7f754b6fbba137a483bbc5bf8e9d8413ad

Observation 110c9dea-d75d-4a06-9dd8-94baa733e875 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding VideoChat: Chat-Centric Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.444148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.444148Z digest=sha256:d091de409c4f63b9b17acce28cd068b4c446181a22fe1a609d1e449da6efe138

Observation 08d20e2c-0330-46ff-9b67-fe558e1ebfba · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.488837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.488837Z digest=sha256:37f9b54733ffd368086a7ded596e0e9a91c884069bca9c4bfa759bd511e791b2

Observation a2ad6307-3313-41c9-887d-fdf025e90a0e · outbound

This paper cites PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.530036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.530036Z digest=sha256:61dc029048d2785295228d3ddff0172774c3a1c0b63bebff55fb2f51e9907796

Observation b822c8c7-c8c3-4c25-8055-bba79cbf1794 · outbound

This paper cites Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.614884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.614884Z digest=sha256:5bfd8c7002820eb5b3ef1521b86298485f2bf8922294a4ed6ccfc8a70b486ed1

Observation 7d870a68-30b1-48c1-adc5-d6dd3459f69a · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.696016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.696016Z digest=sha256:d51f9547f3e18366a8fb9d2ebf9378c53e8dc80fccf738dc02f2c1736bc51e25

Observation ca388100-361a-4ccc-b685-43025ec203b4 · outbound

This paper cites ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.773771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.773771Z digest=sha256:b4db63ec52229c12a34e1ded87a266f01e41ead055f53ef620ab0392b8c55e5f

Observation a03ecc67-4006-4690-9bf6-bdae02f2ad6f · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.899236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.899236Z digest=sha256:b9f0b3fcffb8146fc2ad3211a6f495d3d366c015b8b8a03262ecea51a4438f1b

Observation 91492b46-e8be-4b35-8952-a6a972a07059 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.981202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.981202Z digest=sha256:9f9bb5f4abd9adf1d177b851b5a83c5a42d487e6271862150fe1b2d7a32163bd

Observation 587ba512-870c-4a04-bbaa-48b8069a308f · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:17.153103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:17.153103Z digest=sha256:460c7e68ba96b842ab47e2af1e2d262a207e41b3eb75d0635626818dbed551f4

Observation 53a350c1-9421-43e5-a25b-1b0516cc88a1 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:17.302898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:17.302898Z digest=sha256:93207bcf4d8a919f71cb8972016f90cde74ca43aa642c288b6dd702d925c9340

Observation ca2d00a3-b6ba-4815-aa27-272eeedd7da0 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 1985

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:16.364927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:16.364927Z digest=sha256:ce796963172090fd852f5bbb0a32b81759d6dbd98ef30b3dd7aa4b17159c0169

Observation a3101767-ba83-4287-90e5-94ed395b9614 · outbound

This paper cites Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:17.224605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:17.224605Z digest=sha256:b0680ed13451ae75fc2871e4b4a147f64c30c531656f42ca7901809587ecac04

Observation ec3ffded-ef9e-43d7-ad38-a67c8cf17179 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:17.058810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:17.058810Z digest=sha256:f5537597c49c51c3b5e01c6575ab668c89e29ed251a30fedb97a79a6eb52f9b3

Observation a3e1ddd9-6e69-43fd-9211-6e2a79ced978 · outbound

This paper cites GPT-4 Technical Report.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding GPT-4 Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:15.515921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:15.515921Z digest=sha256:cbf9fcdf4dbeaa0ded4bf2b1dc52c21de0c04fc11fac40d8bacfee0b10301908

Observation 07a3c23f-0848-4e87-9a93-13c1267286ba · outbound

This paper cites MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:15.950597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:15.950597Z digest=sha256:8e4c00066d6668c2e337697a5f391fde71f00c6b2eaf2b5a18b2737de65727ea

Observation 0ebdbcd1-1622-47c7-8529-22f36dd1d1dc · outbound

This paper cites DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models.

Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T21:59:15.578220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:59:15.578220Z digest=sha256:b5b10924b7128d1e9dc5c094840bf0922a48b093162e67821e3545c0468af15b

Pith citing papers

No inbound Pith citation observations are available.