Pith. sign in

REVIEW 2 cited by

Revealing Single Frame Bias for Video-and-Language Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.03428 v1 pith:6YQPEPUD submitted 2022-06-07 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords video-and-languageframesmodelmultipletasksbiasdatasetsexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Training an effective video-and-language model intuitively requires multiple frames as model inputs. However, it is unclear whether using multiple frames is beneficial to downstream tasks, and if yes, whether the performance gain is worth the drastically-increased computation and memory costs resulting from using more frames. In this work, we explore single-frame models for video-and-language learning. On a diverse set of video-and-language tasks (including text-to-video retrieval and video question answering), we show the surprising result that, with large-scale pre-training and a proper frame ensemble strategy at inference time, a single-frame trained model that does not consider temporal information can achieve better performance than existing methods that use multiple frames for training. This result reveals the existence of a strong "static appearance bias" in popular video-and-language datasets. Therefore, to allow for a more comprehensive evaluation of video-and-language models, we propose two new retrieval tasks based on existing fine-grained action recognition datasets that encourage temporal modeling. Our code is available at https://github.com/jayleicn/singularity

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    RTime-QA is a video-question benchmark where models choose between temporally opposite descriptions of the same event, and current AI models score far below humans.

  2. MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering

    cs.CV 2025-06

Pith tools