REVIEW 4 cited by
An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Stimulated by the sophisticated reasoning capabilities of recent Large Language Models (LLMs), a variety of strategies for bridging video modality have been devised. A prominent strategy involves Video Language Models (VideoLMs), which train a learnable interface with video data to connect advanced vision encoders with LLMs. Recently, an alternative strategy has surfaced, employing readily available foundation models, such as VideoLMs and LLMs, across multiple stages for modality bridging. In this study, we introduce a simple yet novel strategy where only a single Vision Language Model (VLM) is utilized. Our starting point is the plain insight that a video comprises a series of images, or frames, interwoven with temporal information. The essence of video comprehension lies in adeptly managing the temporal aspects along with the spatial details of each frame. Initially, we transform a video into a single composite image by arranging multiple frames in a grid layout. The resulting single image is termed as an image grid. This format, while maintaining the appearance of a solitary image, effectively retains temporal information within the grid structure. Therefore, the image grid approach enables direct application of a single high-performance VLM without necessitating any video-data training. Our extensive experimental analysis across ten zero-shot video question answering benchmarks, including five open-ended and five multiple-choice benchmarks, reveals that the proposed Image Grid Vision Language Model (IG-VLM) surpasses the existing methods in nine out of ten benchmarks.
Forward citations
Cited by 4 Pith papers
-
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding
DynImg represents a video snippet as a keyframe plus four resized neighboring frames as temporal prompts, with a 4D rotary position embedding, and reports improved video QA accuracy.
-
RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models
RTime-QA is a video-question benchmark where models choose between temporally opposite descriptions of the same event, and current AI models score far below humans.
-
LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering
LeAdQA improves video question answering by using LLM-rewritten causal queries to drive temporal grounding that selects relevant video segments for the answering model.
-
CoS: Chain-of-Shot Prompting for Long Video Understanding
A training-free method that uses an AI model's yes/no judgments on mosaic clips to build positive and negative shot sets, improving long-video question answering by a few points.
Discussion (0). Continue with ORCID to comment.