REVIEW 12 cited by
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Understanding fine-grained temporal dynamics is crucial for multimodal video comprehension and generation. Due to the lack of fine-grained temporal annotations, existing video benchmarks mostly resemble static image benchmarks and are incompetent at evaluating models for temporal understanding. In this paper, we introduce TemporalBench, a new benchmark dedicated to evaluating fine-grained temporal understanding in videos. TemporalBench consists of ~10K video question-answer pairs, derived from ~2K high-quality human annotations detailing the temporal dynamics in video clips. As a result, our benchmark provides a unique testbed for evaluating various temporal understanding and reasoning abilities such as action frequency, motion magnitude, event order, etc. Moreover, it enables evaluations on various tasks like both video question answering and captioning, both short and long video understanding, as well as different models such as multimodal video embedding models and text generation models. Results show that state-of-the-art models like GPT-4o achieve only 38.5% question answering accuracy on TemporalBench, demonstrating a significant gap (~30%) between humans and AI in temporal understanding. Furthermore, we notice a critical pitfall for multi-choice QA where LLMs can detect the subtle changes in negative captions and find a centralized description as a cue for its prediction, where we propose Multiple Binary Accuracy (MBA) to correct such bias. We hope that TemporalBench can foster research on improving models' temporal reasoning capabilities. Both dataset and evaluation code will be made available.
Forward citations
Cited by 12 Pith papers
-
The TIME Machine: On The Power of Motion for Efficient Perception
TIME is a motion-based embedding from point tracks, trained only on synthetic data via masked autoencoding, that matches state-of-the-art video model performance with up to 10,000x less training data.
-
From Segments to Scenes: Temporal Understanding for Agentic Autonomous Driving via Vision-Language Models
TAD, a 5,861-question benchmark, shows VLMs score far below humans on temporal understanding of driving videos, and an ego-trajectory text summary (TCogMap) substantially boosts their scores.
-
AdsQA: Towards Advertisement Video Understanding
AdsQA adds an ad-video question-answering benchmark and ReAd-R, a GRPO-trained model that beats 7B baselines but not larger closed models.
-
CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos
CausalStep introduces a stepwise video QA protocol and reports that top multimodal models (chain success rate 51%) remain far below human performance (79%) on explicit causal chains.
-
Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding
A new 1,000-video benchmark with adversarial question variants shows video LLMs remain far below human accuracy and robustness on short real-world videos.
-
GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?
Introduces GLIMPSE, a video-QA benchmark whose questions cannot be answered from single frames; best model GPT-o3 scores 66.43% vs 94.82% human accuracy.
-
ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
ProactiveVideoQA is a benchmark for proactive video question answering, and the proposed PAUC metric jointly scores response timing and content, claiming better alignment with human preferences.
-
Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times
Video-language models perform far below humans on a new quadrilingual benchmark that tests understanding of action completion and duration through grammatical aspect.
-
Fostering Video Reasoning via Next-Event Prediction
Next-event prediction, training video language models to caption unseen future frames, improves their scores on several temporal benchmarks while roughly preserving general video understanding.
-
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos
TUNA introduces a 1,000-video benchmark with dense temporal captions and 1,432 multiple-choice questions, and finds that current video LMMs are weakest at camera motion, action sequences, and multi-subject scenes.
-
RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models
RTime-QA is a video-question benchmark where models choose between temporally opposite descriptions of the same event, and current AI models score far below humans.
-
NeMo: Needle in a Montage for Video-Language Understanding
NeMoBench, an automatically generated benchmark with 31,378 QA pairs, shows that video LLMs struggle with temporal grounding of relevant clips hidden in long montages.
Discussion (0). Sign in to comment.