Pith. sign in

Paper Citation Record · LEDGER

An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2403.18406.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.18406 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:31:16.929272Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T05:45:28.275913Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 02d99ae4-fa34-4897-85c4-fb20e8603cbb · inbound

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning cites this paper.

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:21:57.957298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T20:21:57.873354Z digest=sha256:61ffc4d1ae9efea8cd9774d1d2c2b56e1222e49a5829b8223ac0deca21fcdb60

Observation 1a86423a-5ea2-446c-a0fa-169954ba3fae · inbound

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models cites this paper.

MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:45:28.278643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T05:44:31.546843Z digest=sha256:cc950314c0494386ec81b8d3203ab21e2aad5b4b1fcb5907e7d4359139f1b301

Observation 1d5c33d3-2e3b-476c-bf38-ed8bb295810e · inbound

CoS: Chain-of-Shot Prompting for Long Video Understanding cites this paper.

CoS: Chain-of-Shot Prompting for Long Video Understanding An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T15:31:16.929272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:31:16.929272Z digest=sha256:79deb1544ab626c0408c86d2ce6f44a480e3c7e2184bb49bfac1796b233d43b5

Observation 1697f5ac-2c46-4d4c-be1a-f72ab08548b7 · inbound

RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models cites this paper.

RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:07.878207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:24:07.878207Z digest=sha256:9541666eb04739c3561d69cd3a2d7d8e19a80375364a995ebf66b33a0bc87ab6

Observation 40d3da12-d19b-4694-bdbe-5413a14da1da · inbound

LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering cites this paper.

LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:53:29.967745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:53:29.967745Z digest=sha256:06a052ebb2fc9fc96d03fcf731335abbc977108ecdeae197f21a1dc279a91b23

Observation 86d87ae0-451a-468d-ba72-f19c6f13db1c · inbound

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding cites this paper.

DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:34:44.410539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:34:44.410539Z digest=sha256:a6129d056bd59affaa15fbf36b83df28eee789a226ff26ac9aca7b628a5ff44b

Observation 2bea321f-524f-4942-b37b-b774c9ab31f1 · inbound

Progressive Video Condensation with MLLM Agent for Long-form Video Understanding cites this paper.

Progressive Video Condensation with MLLM Agent for Long-form Video Understanding An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:43:14.853772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T20:40:41.380829Z digest=sha256:618793ab407ef5af4a1005d71aa2c9a9bb5a3efbef9047b537bbe1243b0a5e4e

Observation 71b9b300-e04e-4249-bfa0-b5a8df08b993 · inbound

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models cites this paper.

Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:31.716021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T19:17:20.745425Z digest=sha256:1e628f2cf4eb932755e9eb73299b02f34feeb8fe17de9e19185c796ef16c6d7b

Observation 920c2d81-8894-4271-bb3e-3b84f6186c3c · inbound

Enhancing Visual Token Representations for Video Large Language Models via Training-Free Spatial-Temporal Pooling and Gridding cites this paper.

Enhancing Visual Token Representations for Video Large Language Models via Training-Free Spatial-Temporal Pooling and Gridding An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:21:10.868264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T06:16:32.972099Z digest=sha256:f5277866dc948e2f811793ca537340d4d77d50f7a1d542ff28407a7a2f59bab4