Pith. sign in

REVIEW 9 cited by

Video Question Answering: Datasets, Algorithms and Challenges

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.01225 v2 pith:NVU3LNWU submitted 2022-03-02 cs.CV

classification cs.CV
keywords videoqaalgorithmsdatasetsvideoansweringchallengesdifferentlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Video Question Answering (VideoQA) aims to answer natural language questions according to the given videos. It has earned increasing attention with recent research trends in joint vision and language understanding. Yet, compared with ImageQA, VideoQA is largely underexplored and progresses slowly. Although different algorithms have continually been proposed and shown success on different VideoQA datasets, we find that there lacks a meaningful survey to categorize them, which seriously impedes its advancements. This paper thus provides a clear taxonomy and comprehensive analyses to VideoQA, focusing on the datasets, algorithms, and unique challenges. We then point out the research trend of studying beyond factoid QA to inference QA towards the cognition of video contents, Finally, we conclude some promising directions for future exploration.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos

    cs.CV 2025-06 conditional novelty 7.0 of 10

    VidEvent provides a large annotated video dataset for extracting structured event scripts and predicting future events, with baseline benchmarks.

  2. Mitigating Easy Option Bias in Multiple-Choice Question Answering

    cs.CV 2025-08 conditional novelty 6.0 of 10

    In six VQA benchmarks, models can often choose the correct option from image plus options alone, and the GroundAttack toolkit generates visually plausible hard negatives to remove this shortcut.

  3. AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning

    cs.CV 2025-08 reject novelty 6.0 of 10

    A new audio-visual reasoning benchmark and a two-part metric (factual consistency, core inference) claim to expose a gap between answer accuracy and reasoning quality in AV-LLMs.

  4. CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CausalVQA provides 793 paired real-video causal reasoning questions on which the best multimodal model scores 61.66% versus 84.78% for humans, with the largest gaps on anticipation and hypothetical questions.

  5. MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks

    cs.CV 2025-06 conditional novelty 6.0 of 10

    AVHaystacks is a new 3100-question benchmark for audio-visual QA across 500 videos, and the MAGNET multi-agent pipeline beats current baselines on it.

  6. Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Video-language models perform far below humans on a new quadrilingual benchmark that tests understanding of action completion and duration through grammatical aspect.

  7. Vid2Coach: Transforming How-To Videos into Task Assistants

    cs.HC 2025-05 conditional novelty 6.0 of 10

    Vid2Coach converts how-to videos into a real-time, wearable task assistant that helps blind and low vision people cook with fewer errors.

  8. VIBE: Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR

    cs.CV 2025-05 conditional novelty 6.0 of 10

    VIBE selects task-relevant video summaries by combining a grounding score (video-text alignment) and a utility score (task informativeness), improving human accuracy by up to 61% in user studies.

  9. Exploring the Application of Visual Question Answering (VQA) for Classroom Activity Monitoring

    cs.CV 2025-07 reject novelty 5.0 of 10

    Four open-source VQA models reach moderate accuracy on a new classroom video dataset, with yes/no questions easiest and counting/reasoning hardest.

Pith tools