Pith. sign in

REVIEW 8 cited by

FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.14935 v1 pith:FAY3QOSZ submitted 2025-03-19 cs.CV cs.AI

FAVOR-Bench: A Comprehensive Benchmark for Fine-Grained Video Motion Understanding

classification cs.CV cs.AI
keywords favor-benchmotionunderstandingvideocomprehensivefavor-trainfine-grainedmllms
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Multimodal Large Language Models (MLLMs) have shown remarkable capabilities in video content understanding but still struggle with fine-grained motion comprehension. To comprehensively assess the motion understanding ability of existing MLLMs, we introduce FAVOR-Bench, comprising 1,776 videos with structured manual annotations of various motions. Our benchmark includes both close-ended and open-ended tasks. For close-ended evaluation, we carefully design 8,184 multiple-choice question-answer pairs spanning six distinct sub-tasks. For open-ended evaluation, we develop both a novel cost-efficient LLM-free and a GPT-assisted caption assessment method, where the former can enhance benchmarking interpretability and reproducibility. Comprehensive experiments with 21 state-of-the-art MLLMs reveal significant limitations in their ability to comprehend and describe detailed temporal dynamics in video motions. To alleviate this limitation, we further build FAVOR-Train, a dataset consisting of 17,152 videos with fine-grained motion annotations. The results of finetuning Qwen2.5-VL on FAVOR-Train yield consistent improvements on motion-related tasks of TVBench, MotionBench and our FAVOR-Bench. Comprehensive assessment results demonstrate that the proposed FAVOR-Bench and FAVOR-Train provide valuable tools to the community for developing more powerful video understanding models. Project page: \href{https://favor-bench.github.io/}{https://favor-bench.github.io/}.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. NextMotionQA: Benchmarking and Judging Human Motion Understanding with Vision-Language Models

    cs.CV 2026-06 unverdicted novelty 7.0

    NextMotionQA benchmark reveals VLMs have critical gaps in fine-grained human motion understanding and align with experts on coarse judgment (κ=0.70) but not fine-grained (κ=0.10).

  2. Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events

    cs.CV 2026-06 conditional novelty 7.0

    Moment-Video benchmark shows top video MLLM achieves only 39.6% accuracy on momentary visual event tasks, with most open-source models below 25%.

  3. Can Multimodal Large Language Models Truly Understand Small Objects?

    cs.CV 2026-04 unverdicted novelty 7.0

    Current MLLMs show weak performance on small object understanding tasks, but fine-tuning with the new SOU-Train dataset measurably improves their capabilities.

  4. Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents

    cs.CV 2025-11 conditional novelty 7.0

    CFG-Bench uses 19,562 questions across four cognitive tiers to show that vision-language models are weak at fine-grained physical action understanding, and that SFT on its data improves scores on external embodied benchmarks.

  5. Natural Language Camera Movement Understanding

    cs.CV 2026-07 accept novelty 6.5

    A cinematographic taxonomy, atomic real+synthetic benchmark (ACaM), and targeted-augmentation SFT let an 8B VLM outperform Gemini-3.1-Pro by 10-11% on camera-movement recognition, yet a large human gap remains.

  6. MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models

    cs.CV 2026-06 unverdicted novelty 6.0

    MotionEnhancer distills motion priors from video diffusion models into VLMs via parameter-free attention alignment modules to improve motion-level video understanding.

  7. Exploring High-Order Self-Similarity for Video Understanding

    cs.CV 2026-04 unverdicted novelty 6.0

    The MOSS module learns and combines multi-order space-time self-similarity features to enhance temporal dynamics modeling in videos across action recognition, VQA, and robotic tasks.

  8. OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering

    cs.CV 2026-04 unverdicted novelty 5.0

    OmniJigsaw is a self-supervised proxy task that reconstructs shuffled audio-visual clips via joint integration, sample-level selection, and clip-level masking strategies, yielding gains on 15 video, audio, and reasoni...