Pith. sign in

Paper Citation Record · LEDGER

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025

As of 17 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2506.21891.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21891 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:20:13.504824Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c61625f3-b6a4-4f2a-888a-2195dd747f97 · outbound

This paper cites VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:56.536060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:56.536060Z digest=sha256:46d889d2b3ec2597a8c631777f5738ab1b54eac27385e499d5c19ca6f074c70d

Observation 2a8ba645-1d35-41e4-89f6-481e20ca63be · outbound

This paper cites How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:12.499956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:12.499956Z digest=sha256:b38291dce14cf7f744779b9c28012b05dd7f9a61a8b916e614d46cd59397f505

Observation 9cde9d80-fc8e-42e3-8db5-5d172aee904d · outbound

This paper cites VDMA: Video Question Answering with Dynamically Generated Multi-Agents.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 VDMA: Video Question Answering with Dynamically Generated Multi-Agents

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:20:13.720355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:20:12.669954Z digest=sha256:f199cbc1aba84cd6922906cb79efe003a7caea88c0823d650dc1aca4e4e277c2

Observation bc3b5004-dfc5-4e02-a6dc-af4e8fc96a92 · outbound

This paper cites VideoMultiAgents: A Multi-Agent Framework for Video Question Answering.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 VideoMultiAgents: A Multi-Agent Framework for Video Question Answering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:12.739322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:12.739322Z digest=sha256:1aacbf18113a456c7dd7dda9d94504a0b269276895022c3192d187c6fa601918

Observation fb16a04b-c044-4c56-a2d8-b371f7fb96a6 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:20:15.304002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:20:12.784643Z digest=sha256:8704d150bc1302e6e5ab3e7c0b8e871a4d02f06ab145094a64d0ddd5dd9f7b1e

Observation a087af98-aff9-463e-8a79-5b22ead6f354 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 Llama-vid: An image is worth 2 tokens in large language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:20:15.067116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:20:12.823501Z digest=sha256:0eb1012ad7e0b207d1fa4c4cff983a90a4aa750c07a3c765174067544ac9753a

Observation 7e5ea4b5-c1c8-4bc9-b052-589030d25ee9 · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projection.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 Video-llava: Learning united visual representation by alignment before projection

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:20:14.893626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:20:12.861049Z digest=sha256:f7b8687567df17252cb97477335d95fb5c7f6fcdc7f0f845d9a9b3da2a7e6270

Observation a4b88318-1e65-4027-b18d-84a66ec415d2 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:12.887325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:12.887325Z digest=sha256:32dc4c62f9a2772ff61f4b678c1ddc8e174f37499969a55255346a2487eedda3

Observation 2d71e20c-afd2-494e-9538-9cc555bf9bea · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:20:14.709409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:20:12.904269Z digest=sha256:156f6b2b6e71ac8b84411099aee653c4431a916617d4029d832b916944cacfe1

Observation d5b3873f-d0f6-4d18-a067-566e4192ea4c · outbound

This paper cites Introducing deep research.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 Introducing deep research

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:20:14.561944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:20:12.925037Z digest=sha256:96df4c0c61684a4fd5553b2c616c3c8e411e5e112f947f492a48bfddde1b61df

Observation c76d8af5-6ecf-405e-8499-77ad6f20ccc0 · outbound

This paper cites TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:12.971888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:12.971888Z digest=sha256:00b7325cf45133cc180a267650adef06f3411963df8eef554e09e7198afb6b95

Observation aaa0365b-6b29-41ec-be0a-3520fe7119fc · outbound

This paper cites Traveler: A modular multi-lmm agent framework for video question-answering.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 Traveler: A modular multi-lmm agent framework for video question-answering

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:20:14.338089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:20:13.023972Z digest=sha256:2ee0f65618d9951a67736b6ef06fa424e6d699ce65bae5a0900bc733cc2995f9

Observation 65570e96-d72a-4e44-9caa-adc8fba857e2 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 Videoagent: Long-form video understanding with large language model as agent

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:20:14.180597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:20:13.078536Z digest=sha256:1967a7a4ec2bd1d5203f7088dd9459c6b7aec087a503866d0ef8b2fb696ef96e

Observation a2fe6278-5c86-4b72-9064-55b7470acee5 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:13.123695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:13.123695Z digest=sha256:0c1e708b36189737b982548b3bdad6150cf5fa88e0598debc65a57022d958738

Observation 697143f2-d66c-4bd1-8ec6-42d50ff53d0d · outbound

This paper cites Lifelongmemory: Leveraging llms for answering queries in long-form egocentric videos.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 Lifelongmemory: Leveraging llms for answering queries in long-form egocentric videos

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:20:14.068571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:20:13.161396Z digest=sha256:8337652c570e945b6ee195d6c96e22e92362b6bebc04b1b724ae25419e3d4609

Observation 283646a8-8b70-457c-b779-f6eb04f50ef7 · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:13.236841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:13.236841Z digest=sha256:13cd4a7c242103630f058e8f7005a61386ce56a4b9e0f6ead4b21f7412751646

Observation ec97086f-91a3-4ef2-95fd-5df53212c9d5 · outbound

This paper cites Zero-Shot Video Question Answering via Frozen Bidirectional Language Models.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 Zero-Shot Video Question Answering via Frozen Bidirectional Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:13.284415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:13.284415Z digest=sha256:1494ee19d44d7ac0dc5fc7ac28d2f777a7f2f77b69cae5963e1332ae1712b58e

Observation b8a684b2-0020-46df-8aab-ea856fc2d9bb · outbound

This paper cites A simple llm framework for long-range video question-answering.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 A simple llm framework for long-range video question-answering

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:20:13.971016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:20:13.339297Z digest=sha256:314b98432e3838045010dd49aa47c5e49fb689abeb4a0b9d676480cecb8cf5ee

Observation 34503728-fbe3-4366-9fa0-4d0be677738a · outbound

This paper cites Hcqa @ ego4d egoschema challenge.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 Hcqa @ ego4d egoschema challenge

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:20:13.870767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T22:20:13.389497Z digest=sha256:e8017744105ab0419bd073d53b7a8701349a451012a026dddb107891b7bb4904

Observation 1bfc5510-31fb-46b2-a18e-27f3ea93819f · outbound

This paper cites VideoAgent2: Enhancing the LLM-Based Agent System for Long-Form Video Understanding by Uncertainty-Aware CoT.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 VideoAgent2: Enhancing the LLM-Based Agent System for Long-Form Video Understanding by Uncertainty-Aware CoT

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:13.504824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:13.504824Z digest=sha256:8fb5722aea812c178e13c3b0a6b91ae8e7f4faebe8665f78ff237f178fd72ea9

Observation 573b2a95-b937-4130-a765-1ad2530963c4 · outbound

This paper cites HCQA @ Ego4D EgoSchema Challenge 2024.

DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025 HCQA @ Ego4D EgoSchema Challenge 2024

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:13.460205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:13.460205Z digest=sha256:1279fb5e96bd06344901c6b8985155bb21b34b2e14fea0b571ab3b127030cbe4

Pith citing papers

No inbound Pith citation observations are available.