Pith. sign in

Paper Citation Record · LEDGER

MM-VID: Advancing Video Understanding with GPT-4V(ision)

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2310.19773.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.19773 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T15:04:36.969203Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:09:44.465524Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 06ed1e70-1367-4eb2-9204-f235df35d07e · inbound

VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos cites this paper.

VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos MM-VID: Advancing Video Understanding with GPT-4V(ision)

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T15:04:36.969203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:04:36.969203Z digest=sha256:7e55b02d6b10d4d0208ac0fd1153370cae57d69ad3d9049eff87cf3dc14b5371

Observation 495627fe-3d36-4200-a04e-5b3679e4326d · inbound

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models cites this paper.

MESH -- Understanding Videos Like Human: Measuring Hallucinations in Large Video Models MM-VID: Advancing Video Understanding with GPT-4V(ision)

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T20:32:56.628654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:32:56.628654Z digest=sha256:f60634eed196618f9ba40fd71a92fd00c55829f08e0da020322f8dc7277cf498

Observation ebbb2ed5-4fcf-489a-bc72-066c46ea492a · inbound

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG cites this paper.

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG MM-VID: Advancing Video Understanding with GPT-4V(ision)

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:50.209434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:15:02.124035Z digest=sha256:30d02f043d105929ccfa9c33465aad237f508ad23c9b6edc6458f489a839fb54

Observation a14c7dde-17b6-4d58-bd66-0a8d1c0f5471 · inbound

Making AI Drafts Count: A Quality Threshold in Audio Description Workflows cites this paper.

Making AI Drafts Count: A Quality Threshold in Audio Description Workflows MM-VID: Advancing Video Understanding with GPT-4V(ision)

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:26:07.983894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T16:07:12.813384Z digest=sha256:c2f1d231caf929b094bbe645dd43a99d74398774f0155f34dbeb0430a628bb8b

Observation a29d900f-fa37-4fe9-903d-ee3b9f3b9e64 · inbound

Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration cites this paper.

Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration MM-VID: Advancing Video Understanding with GPT-4V(ision)

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:18:18.351189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T13:15:54.413960Z digest=sha256:ee8cd5a9b0163d81dfd401ee93d5a8843a7558d37ef3110e1057adecbb8cdb22

Observation e48ed5ef-bf5f-40ea-851b-10f4389b881f · inbound

StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset cites this paper.

StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset MM-VID: Advancing Video Understanding with GPT-4V(ision)

Reference 89

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:26:57.177575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:05:47.810096Z digest=sha256:4f2c70d7f38066dfc01484786346fc8acfdf8676f130562b30df69e384f12115

Observation 0e7c4054-6bbd-4e1d-abc9-afb5768d06ba · inbound

READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations cites this paper.

READ More than What You See: Reinforcement Learning for Accurate and Coherent Audio Description Generations MM-VID: Advancing Video Understanding with GPT-4V(ision)

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:09:44.466891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T09:09:23.918802Z digest=sha256:b8585b0a8948b781d728066ed575a2f2c45cb923db191ae64cef1f2b96c1085b

Observation 2ab7d186-2e3e-43ec-8421-433eb97269ab · inbound

ReQuest: Rethinking-based Question-Aware Frame Selection for Long-Form Video QA cites this paper.

ReQuest: Rethinking-based Question-Aware Frame Selection for Long-Form Video QA MM-VID: Advancing Video Understanding with GPT-4V(ision)

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:48:39.319110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T16:45:20.139230Z digest=sha256:ba3aae1d744249c3bf31f2d1a967c47597f8e90d99ad719525218c30c0076700

Observation ebcd45b0-7ca6-4881-adda-8b9c5ec37d2d · inbound

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG cites this paper.

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG MM-VID: Advancing Video Understanding with GPT-4V(ision)

Reference 141

Resolution
unresolved
no resolver link, observed 2026-08-02T10:20:56.204492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:20:56.204492Z digest=sha256:d3ee2d964c49e83d9004b6eaad2a04d2208046bcbd05bd8e217b3ae926ff5430