Pith. sign in

Paper Citation Record · LEDGER

VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2406.11303.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.11303 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:43.788047Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:26:57.184700Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation df47b175-ed30-402b-9ad1-4104542abff3 · inbound

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models cites this paper.

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:45:28.342097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:44:31.546843Z digest=sha256:68b11868e959e398eb29ed283169c2e9d362c540bc650eeaf2bc1a6b01f64903

Observation 29aaa453-09ca-4625-84b3-d9457014c874 · inbound

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation cites this paper.

VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:43.788047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:43.788047Z digest=sha256:bf907ae40b44fce79e54452a8612170c69fffef02ac7e1c4efa89b8a97a12bb7

Observation fc66876d-c16b-4d4b-977f-7b397612597f · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:01.131871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:03:01.131871Z digest=sha256:320e75f97a4e31bb9be3a0ab2798d9ec61df8c29630d2237ae44e38ea688343c

Observation 7514bf63-6fd6-4d1f-91bb-5fac0532564c · inbound

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation cites this paper.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:03.738031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:03.738031Z digest=sha256:5262d96e9c221a7bda71c1bc2fada08028c0b0d426508857d5bb8c731024ac6a

Observation f4e62474-1e58-4261-b56c-11dbe3e6481f · inbound

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos cites this paper.

VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.866771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:45:46.866771Z digest=sha256:467cbe87d02b6cc6b989dbeed12a1553183c94b7e4e5d4ca02bed7ea9f8a58c9

Observation 6f5a593c-f7dc-4764-85f5-da4346e18514 · inbound

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding cites this paper.

ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:44.943565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:40:44.943565Z digest=sha256:9eff59753e9c296ca3eadd9447daeda574f3a54fa86dfaf2b44eb86f1eafb9b7

Observation 408e7f61-f69b-46cd-97ec-da0311486d57 · inbound

VUDG: A Dataset for Video Understanding Domain Generalization cites this paper.

VUDG: A Dataset for Video Understanding Domain Generalization VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:32:10.535033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:32:10.535033Z digest=sha256:28a5b3b174e37db7e2a0d7c71ce3b0769c0e11ffa17211d75dff5ee859216fbb

Observation f72b07e1-89ec-4e63-a23d-54c6a082e19c · inbound

SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning cites this paper.

SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:37:15.739001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T11:36:36.687324Z digest=sha256:ccd810ebd9a9bc7fe3f80b8594c06db300e4f4678996b7307909d1a3155c8179

Observation 72944cf2-74c2-4dd2-91fa-073fcf1672fa · inbound

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos cites this paper.

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:55.901902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:55.901902Z digest=sha256:962a0396eb87089f840e3e137f6ab285e980d541d131fea237f46438b07e9387

Observation fd3247c8-5278-4a34-86de-a9e8723a234c · inbound

SmartHome-Bench: A Comprehensive Benchmark for Video Anomaly Detection in Smart Homes Using Multi-Modal Large Language Models cites this paper.

SmartHome-Bench: A Comprehensive Benchmark for Video Anomaly Detection in Smart Homes Using Multi-Modal Large Language Models VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:51.290582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:51.290582Z digest=sha256:84958f5b1e3a2c39e3bb3696d9421f378f888ca8d0e0ebcb327fd56137ab6fc9

Observation 7d166b95-e89a-4822-9a73-8309863d4e2a · inbound

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering cites this paper.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:39.159644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:39.159644Z digest=sha256:3a5dcff4c680904b003b261d4b8a4455848cda3aecdf2b2ae5874a74e76a9538

Observation 4880c9fe-dc2e-4103-adad-2d91bd29d674 · inbound

NeMo: Needle in a Montage for Video-Language Understanding cites this paper.

NeMo: Needle in a Montage for Video-Language Understanding VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T13:54:15.574763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:54:15.574763Z digest=sha256:0f4133acaa3d4d55d8cbfade795a509ad8bd4170864a3dbd401271ecc11a63ac

Observation 4f18c017-fa11-49b1-9ff5-5a40e6d0cc64 · inbound

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark cites this paper.

Seeing the Scene Matters: Revealing Forgetting in Video Understanding Models with a Scene-Aware Long-Video Benchmark VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:04.406514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T22:05:07.326202Z digest=sha256:bed0cbcd8ee12f56a61b19a4a924298f5e68d881612bc87a1938d87541c31edc

Observation c02b375f-471e-495f-ba2d-dd8d1b12bdb0 · inbound

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding cites this paper.

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:03:08.051487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T16:02:53.887605Z digest=sha256:91e6c5f754b4399543ec7a33bba7d443c5d2f876d34bf1a4155f1821f516ee3e

Observation b7b3ffe2-8116-4a87-b31b-d0d2b8287e22 · inbound

VidMsg: A Benchmark for Implicit Message Inference in Short Videos cites this paper.

VidMsg: A Benchmark for Implicit Message Inference in Short Videos VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:30.085934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:25:06.594946Z digest=sha256:f8f9ae9b179642101fcfac3fec533e2ab7b4c63e31334d830d89450cec98ea71

Observation fa3a6b1c-d6de-40ef-bfcd-2f77df218b05 · inbound

StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset cites this paper.

StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:26:57.186187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T02:05:47.810096Z digest=sha256:370e965cd579adcb7fd122c6e24b719bb87f7677038648447f68a5faf4c53f33

Observation 7ad99eda-7280-48fb-a4c6-8582ee1d7c21 · inbound

Animation2Code: Evaluating Temporal Visual Reasoning in Video-to-Code Generation cites this paper.

Animation2Code: Evaluating Temporal Visual Reasoning in Video-to-Code Generation VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:05:49.750946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T00:53:52.724902Z digest=sha256:3a8499448a024b62554cae1a6e535b214a5104faa27a9fc25fa15bda06609a06

Observation e059e240-3680-4a48-8baa-dbd830e98b72 · inbound

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping cites this paper.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.981358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.981358Z digest=sha256:4a6ea9b96060b2f10da2116a1afc46f0f9b806492b245d5041f48fd0f418a717