Pith. sign in

Paper Citation Record · LEDGER

VideoRAG: Retrieval-Augmented Generation over Video Corpus

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2501.05874.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.05874 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:52:02.359068Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 54eae12c-d336-4b32-99a9-95fd043322e0 · inbound

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding cites this paper.

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding VideoRAG: Retrieval-Augmented Generation over Video Corpus

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:26.672299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:26.672299Z digest=sha256:8c2fd0c1483d7f368e977ad8d1191647bdde8573ec3a19454a38c02a2630ab4c

Observation d46fa311-2b27-4b04-a6ab-e182a8e94d36 · inbound

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding cites this paper.

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding VideoRAG: Retrieval-Augmented Generation over Video Corpus

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:02.359068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:02.359068Z digest=sha256:6540e5378b53fdc36fb77b6c88c294498a23b50a617ad328dab5543b506f4b27

Observation 8590ccea-a560-4f54-a566-bd67cde45dde · inbound

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding cites this paper.

SceneRAG: Scene-level Retrieval-Augmented Generation for Video Understanding VideoRAG: Retrieval-Augmented Generation over Video Corpus

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:50.801689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:50.801689Z digest=sha256:6e9eeec582c117bb796535913f54508e09dbc37234c3ef32ed1540a9704bae44

Observation aeaa6e0c-f7df-432c-8d78-712798e27b10 · inbound

RAVID: Retrieval-Augmented Visual Detection: A Knowledge-Driven Approach for AI-Generated Image Identification cites this paper.

RAVID: Retrieval-Augmented Visual Detection: A Knowledge-Driven Approach for AI-Generated Image Identification VideoRAG: Retrieval-Augmented Generation over Video Corpus

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T01:04:08.928124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T01:04:08.928124Z digest=sha256:a0e2a21547fd4dfb6d6cf0ee8487da7b2a44c642e43ee33de56eefc2d2033aec

Observation 2d83ecfe-48f0-4669-aeb5-e8d9310fb367 · inbound

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation cites this paper.

Track and Caption Any Motion: Open-Vocabulary Spatiotemporal Captioning via Trajectory-Conditioned Generation VideoRAG: Retrieval-Augmented Generation over Video Corpus

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:20.433930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:20.433930Z digest=sha256:f525640b98c705a6dddd3a4fdcec89f0ff1c60f854e1bd548f87e36217cab2d9

Observation 8bc95e34-6faa-4aaf-8c14-fbb97c7c3366 · inbound

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs cites this paper.

Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs VideoRAG: Retrieval-Augmented Generation over Video Corpus

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T04:20:53.503874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:20:53.503874Z digest=sha256:1a5f4b5e01e344f770670dc54c6a137257e87c6857bc83d7cb367cdc787b76f1

Observation 745eaa08-5760-42de-b4b3-bbab4e0afe39 · inbound

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning cites this paper.

Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning VideoRAG: Retrieval-Augmented Generation over Video Corpus

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:30:52.533572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:46:16.975267Z digest=sha256:d932e652fa79236c1becb15f17c42a053c54389533bc6f3d27751a7b476a7259

Observation 6fba3a16-47c4-4041-843e-2a1cdccdcdf4 · inbound

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG cites this paper.

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG VideoRAG: Retrieval-Augmented Generation over Video Corpus

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:50.251652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:15:02.124035Z digest=sha256:b20b18f2224bfc8a5b120069bc27c86447239b2a29ba65b62068267850136110

Observation b2a4c83f-20a5-404b-ae7f-ef26dd1707f6 · inbound

UrbanClipAtlas: A Visual Analytics Framework for Event and Scene Retrieval in Urban Videos cites this paper.

UrbanClipAtlas: A Visual Analytics Framework for Event and Scene Retrieval in Urban Videos VideoRAG: Retrieval-Augmented Generation over Video Corpus

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:04:06.093362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T10:00:16.944922Z digest=sha256:1d40a3908ca8cc52b0a5d6910ac095e1be3270a8317f6d815e328a808a4d4362

Observation 54deaee5-99ef-4616-a056-ee59baee9b5a · inbound

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning cites this paper.

GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning VideoRAG: Retrieval-Augmented Generation over Video Corpus

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:56:47.367481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T06:33:32.090913Z digest=sha256:1daf09ce026ff2da1feede3feb069dda3fde0cf1fea934940c7ee85268ee6b44

Observation 7c2ee8c2-0367-4fbe-b531-b3d3b71668f7 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning VideoRAG: Retrieval-Augmented Generation over Video Corpus

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:39:37.602804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:85d07c19922cbb4b221ebaf174fa87b22dc911562b47c1f78c9d11204184071b