Pith. sign in

REVIEW 12 cited by

TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.10818 v2 pith:DEIB5PUN submitted 2024-10-14 cs.CV cs.AIcs.CLcs.LG

classification cs.CVcs.AIcs.CLcs.LG
keywords temporalvideomodelsunderstandingtemporalbenchfine-grainedevaluatingmultimodal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Understanding fine-grained temporal dynamics is crucial for multimodal video comprehension and generation. Due to the lack of fine-grained temporal annotations, existing video benchmarks mostly resemble static image benchmarks and are incompetent at evaluating models for temporal understanding. In this paper, we introduce TemporalBench, a new benchmark dedicated to evaluating fine-grained temporal understanding in videos. TemporalBench consists of ~10K video question-answer pairs, derived from ~2K high-quality human annotations detailing the temporal dynamics in video clips. As a result, our benchmark provides a unique testbed for evaluating various temporal understanding and reasoning abilities such as action frequency, motion magnitude, event order, etc. Moreover, it enables evaluations on various tasks like both video question answering and captioning, both short and long video understanding, as well as different models such as multimodal video embedding models and text generation models. Results show that state-of-the-art models like GPT-4o achieve only 38.5% question answering accuracy on TemporalBench, demonstrating a significant gap (~30%) between humans and AI in temporal understanding. Furthermore, we notice a critical pitfall for multi-choice QA where LLMs can detect the subtle changes in negative captions and find a centralized description as a cue for its prediction, where we propose Multiple Binary Accuracy (MBA) to correct such bias. We hope that TemporalBench can foster research on improving models' temporal reasoning capabilities. Both dataset and evaluation code will be made available.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The TIME Machine: On The Power of Motion for Efficient Perception

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    TIME is a motion-based embedding from point tracks, trained only on synthetic data via masked autoencoding, that matches state-of-the-art video model performance with up to 10,000x less training data.

  2. From Segments to Scenes: Temporal Understanding for Agentic Autonomous Driving via Vision-Language Models

    cs.CV 2025-12 conditional novelty 6.0 of 10

    TAD, a 5,861-question benchmark, shows VLMs score far below humans on temporal understanding of driving videos, and an ego-trajectory text summary (TCogMap) substantially boosts their scores.

  3. AdsQA: Towards Advertisement Video Understanding

    cs.CV 2025-09 conditional novelty 6.0 of 10

    AdsQA adds an ad-video question-answering benchmark and ReAd-R, a GRPO-trained model that beats 7B baselines but not larger closed models.

  4. CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos

    cs.CV 2025-07 conditional novelty 6.0 of 10

    CausalStep introduces a stepwise video QA protocol and reports that top multimodal models (chain success rate 51%) remain far below human performance (79%) on explicit causal chains.

  5. Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new 1,000-video benchmark with adversarial question variants shows video LLMs remain far below human accuracy and robustness on short real-world videos.

  6. GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Introduces GLIMPSE, a video-QA benchmark whose questions cannot be answered from single frames; best model GPT-o3 scores 66.43% vs 94.82% human accuracy.

  7. ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    ProactiveVideoQA is a benchmark for proactive video question answering, and the proposed PAUC metric jointly scores response timing and content, claiming better alignment with human preferences.

  8. Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Video-language models perform far below humans on a new quadrilingual benchmark that tests understanding of action completion and duration through grammatical aspect.

  9. Fostering Video Reasoning via Next-Event Prediction

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Next-event prediction, training video language models to caption unseen future frames, improves their scores on several temporal benchmarks while roughly preserving general video understanding.

  10. TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos

    cs.CV 2025-05 conditional novelty 6.0 of 10

    TUNA introduces a 1,000-video benchmark with dense temporal captions and 1,432 multiple-choice questions, and finds that current video LMMs are weakest at camera motion, action sequences, and multi-subject scenes.

  11. RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    RTime-QA is a video-question benchmark where models choose between temporally opposite descriptions of the same event, and current AI models score far below humans.

  12. NeMo: Needle in a Montage for Video-Language Understanding

    cs.CV 2025-09 conditional novelty 5.0 of 10

    NeMoBench, an automatically generated benchmark with 31,378 QA pairs, shows that video LLMs struggle with temporal grounding of relevant clips hidden in long montages.

Pith tools