Pith. sign in

REVIEW 3 cited by

ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.07022 v1 pith:GPBZCCIC submitted 2023-11-13 cs.CL cs.AIcs.CV

classification cs.CLcs.AIcs.CV
keywords modelsvidlmsvilmabenchmarkcapabilitiesneedtestsassessment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With the ever-increasing popularity of pretrained Video-Language Models (VidLMs), there is a pressing need to develop robust evaluation methodologies that delve deeper into their visio-linguistic capabilities. To address this challenge, we present ViLMA (Video Language Model Assessment), a task-agnostic benchmark that places the assessment of fine-grained capabilities of these models on a firm footing. Task-based evaluations, while valuable, fail to capture the complexities and specific temporal aspects of moving images that VidLMs need to process. Through carefully curated counterfactuals, ViLMA offers a controlled evaluation suite that sheds light on the true potential of these models, as well as their performance gaps compared to human-level understanding. ViLMA also includes proficiency tests, which assess basic capabilities deemed essential to solving the main counterfactual tests. We show that current VidLMs' grounding abilities are no better than those of vision-language models which use static images. This is especially striking once the performance on proficiency tests is factored in. Our benchmark serves as a catalyst for future research on VidLMs, helping to highlight areas that still need to be explored.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Introduces GLIMPSE, a video-QA benchmark whose questions cannot be answered from single frames; best model GPT-o3 scores 66.43% vs 94.82% human accuracy.

  2. EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?

    cs.CV 2025-06 conditional novelty 6.0 of 10

    EOC-Bench evaluates MLLMs on egocentric object cognition across past, present, and future temporal dimensions, finding large gaps versus humans, especially in absolute time perception.

  3. Argus Inspection: Do Multimodal Large Language Models Possess the Eye of Panoptes?

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A new 1,430-item multimodal benchmark shows that leading multimodal LLMs rarely notice small visual traps needed for commonsense safety reasoning, with top scores near 0.46 on a scale whose maximum is about 0.97.

Pith tools