REVIEW 9 cited by
Video Question Answering: Datasets, Algorithms and Challenges
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Video Question Answering (VideoQA) aims to answer natural language questions according to the given videos. It has earned increasing attention with recent research trends in joint vision and language understanding. Yet, compared with ImageQA, VideoQA is largely underexplored and progresses slowly. Although different algorithms have continually been proposed and shown success on different VideoQA datasets, we find that there lacks a meaningful survey to categorize them, which seriously impedes its advancements. This paper thus provides a clear taxonomy and comprehensive analyses to VideoQA, focusing on the datasets, algorithms, and unique challenges. We then point out the research trend of studying beyond factoid QA to inference QA towards the cognition of video contents, Finally, we conclude some promising directions for future exploration.
Forward citations
Cited by 9 Pith papers
-
VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos
VidEvent provides a large annotated video dataset for extracting structured event scripts and predicting future events, with baseline benchmarks.
-
Mitigating Easy Option Bias in Multiple-Choice Question Answering
In six VQA benchmarks, models can often choose the correct option from image plus options alone, and the GroundAttack toolkit generates visually plausible hard negatives to remove this shortcut.
-
AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
A new audio-visual reasoning benchmark and a two-part metric (factual consistency, core inference) claim to expose a gap between answer accuracy and reasoning quality in AV-LLMs.
-
CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models
CausalVQA provides 793 paired real-video causal reasoning questions on which the best multimodal model scores 61.66% versus 84.78% for humans, with the largest gaps on anticipation and hypothetical questions.
-
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks
AVHaystacks is a new 3100-question benchmark for audio-visual QA across 500 videos, and the MAGNET multi-agent pipeline beats current baselines on it.
-
Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times
Video-language models perform far below humans on a new quadrilingual benchmark that tests understanding of action completion and duration through grammatical aspect.
-
Vid2Coach: Transforming How-To Videos into Task Assistants
Vid2Coach converts how-to videos into a real-time, wearable task assistant that helps blind and low vision people cook with fewer errors.
-
VIBE: Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR
VIBE selects task-relevant video summaries by combining a grounding score (video-text alignment) and a utility score (task informativeness), improving human accuracy by up to 61% in user studies.
-
Exploring the Application of Visual Question Answering (VQA) for Classroom Activity Monitoring
Four open-source VQA models reach moderate accuracy on a new classroom video dataset, with yes/no questions easiest and counting/reasoning hardest.
Discussion (0). Sign in to comment.