Pith. sign in

REVIEW 15 cited by

CinePile: A Long Video Question Answering Dataset and Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.08813 v3 pith:MIRWZLLT submitted 2024-05-14 cs.CV cs.LGcs.MM

classification cs.CVcs.LGcs.MM
keywords datasetvideolong-formunderstandingbenchmarkcinepilecomprehensioncurrent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current datasets for long-form video understanding often fall short of providing genuine long-form comprehension challenges, as many tasks derived from these datasets can be successfully tackled by analyzing just one or a few random frames from a video. To address this issue, we present a novel dataset and benchmark, CinePile, specifically designed for authentic long-form video understanding. This paper details our innovative approach for creating a question-answer dataset, utilizing advanced LLMs with human-in-the-loop and building upon human-generated raw data. Our comprehensive dataset comprises 305,000 multiple-choice questions (MCQs), covering various visual and multimodal aspects, including temporal comprehension, understanding human-object interactions, and reasoning about events or actions within a scene. Additionally, we fine-tuned open-source Video-LLMs on the training split and evaluated both open-source and proprietary video-centric LLMs on the test split of our dataset. The findings indicate that although current models underperform compared to humans, fine-tuning these models can lead to significant improvements in their performance.

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ARGUS: Hallucination and Omission Evaluation in Video-LLMs

    cs.CV 2025-06 conditional novelty 7.0 of 10

    ARGUS measures hallucination and omission in free-form video captions using LLM-based entailment and temporal alignment, finding that even the best video-LLM still produces roughly 40% hallucinated content.

  2. Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding

    cs.CV 2025-06 conditional novelty 7.0 of 10

    MF2 evaluates long-movie understanding by asking models to classify fact/fib claim pairs; the best model trails humans by 23.5 points in pairwise accuracy.

  3. ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding

    cs.CV 2025-05 conditional novelty 7.0 of 10

    ScaleLong embeds four timescale question types into the same long videos, and evaluation of 23 MLLMs reveals a U-shaped accuracy curve across timescales.

  4. Vid-SME: Membership Inference Attacks against Large Video Understanding Models

    cs.CV 2025-05 reject novelty 7.0 of 10

    Vid-SME computes Sharma-Mittal entropy differences between natural and reversed video frame sequences to infer training membership in video understanding LLMs, but its effectiveness is confounded by member/non-member ...

  5. The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping

    cs.AI 2026-08 conditional novelty 6.0 of 10

    Trace-grounded parametric profiling of three synthetic counting tasks shows current video-language models only count reliably at low event counts and low rates, and final-answer accuracy masks poor timestamp-level eve...

  6. RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

    cs.RO 2026-07 conditional novelty 6.0 of 10

    RynnBrain 1.1 reports state-of-the-art embodied cognition and localization scores with a 122B-A10B model and improved real-robot VLA policies via joint multi-embodiment training.

  7. VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

    cs.CV 2026-07 conditional novelty 6.0 of 10

    An open 4B video MLLM with inflated-3D ViT tokenization and adaptive streaming perception outperforms comparable open models on general, long-video, and streaming benchmarks while using fewer visual tokens.

  8. VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos

    cs.CV 2025-06 conditional novelty 6.0 of 10

    VRBench is a benchmark of 960 long narrative videos with 8,243 human-written multi-step questions, plus a two-level evaluation of answer accuracy and reasoning quality for 31 large models.

  9. CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CausalVQA provides 793 paired real-video causal reasoning questions on which the best multimodal model scores 61.66% versus 84.78% for humans, with the largest gaps on anticipation and hypothetical questions.

  10. Ming-Omni: A Unified Multimodal Model for Perception and Generation

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A single model with modality-specific routing processes image, text, audio, and video inputs and generates text, speech, and images, with public benchmarks reported across all of these abilities.

  11. MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks

    cs.CV 2025-06 conditional novelty 6.0 of 10

    AVHaystacks is a new 3100-question benchmark for audio-visual QA across 500 videos, and the MAGNET multi-agent pipeline beats current baselines on it.

  12. ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A reinforcement-learned frame selection policy, trained with reward margins from a reference video-LLM, improves video QA accuracy of LLaVA-OV and InternVL3 across several benchmarks.

  13. Reading Between the Frames: Interpreting Implicit and Non-literal Meaning in Social Media Videos

    cs.CL 2026-08 conditional novelty 5.0 of 10

    A new video benchmark, DrivelHub+, shows that current video-language models can describe what happens in social media clips but largely fail to infer the implicit humour, irony, or cultural meaning.

  14. LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    LaVi encodes visual context into LayerNorm affine parameters, bypassing visual token concatenation, and reports LLaVA-comparable accuracy at a 94% FLOP reduction.

  15. MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering

    cs.CV 2025-06

Pith tools