REVIEW 2 cited by
Verbs in Action: Improving verb understanding in video-language models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Understanding verbs is crucial to modelling how people and objects interact with each other and the environment through space and time. Recently, state-of-the-art video-language models based on CLIP have been shown to have limited verb understanding and to rely extensively on nouns, restricting their performance in real-world video applications that require action and temporal understanding. In this work, we improve verb understanding for CLIP-based video-language models by proposing a new Verb-Focused Contrastive (VFC) framework. This consists of two main components: (1) leveraging pretrained large language models (LLMs) to create hard negatives for cross-modal contrastive learning, together with a calibration strategy to balance the occurrence of concepts in positive and negative pairs; and (2) enforcing a fine-grained, verb phrase alignment loss. Our method achieves state-of-the-art results for zero-shot performance on three downstream tasks that focus on verb understanding: video-text matching, video question-answering and video classification. To the best of our knowledge, this is the first work which proposes a method to alleviate the verb understanding problem, and does not simply highlight it.
Forward citations
Cited by 2 Pith papers
-
CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image Generation
CatchPhrase improves audio-to-image generation by enriching weak class labels with LLM- and audio-caption-based prompts, filtering and retrieving the best prompt per clip, and training a mapping adapter with contrasti...
-
Causal Graphical Models for Vision-Language Compositional Understanding
Ordering word prediction by a dependency tree instead of left-to-right improves vision-language compositional understanding across five benchmarks.
Discussion (0). Continue with ORCID to comment.