REVIEW 6 cited by
Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
A short clip of video may contain progression of multiple events and an interesting story line. A human need to capture both the event in every shot and associate them together to understand the story behind it. In this work, we present a new multi-shot video understanding benchmark Shot2Story with detailed shot-level captions, comprehensive video summaries and question-answering pairs. To facilitate better semantic understanding of videos, we provide captions for both visual signals and human narrations. We design several distinct tasks including single-shot video captioning, multi-shot video summarization, and multi-shot video question answering. Preliminary experiments show some challenges to generate a long and comprehensive video summary for multi-shot videos. Nevertheless, the generated imperfect summaries can already achieve competitive performance on existing video understanding tasks such as video question-answering, promoting an under-explored setting of video understanding with detailed summaries.
Forward citations
Cited by 6 Pith papers
-
HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting
HIPPO-Video contributes 2,040 LLM-simulated watch-history and saliency-score pairs, and the HiPHer model uses these histories to beat generic and query-based baselines on the new benchmark.
-
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding
A 2B-parameter video-language model using an RWKV linear-RNN backbone and sorted token merging achieves competitive long-video QA accuracy with far lower memory cost than transformer-based models.
-
CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models
CausalVQA provides 793 paired real-video causal reasoning questions on which the best multimodal model scores 61.66% versus 84.78% for humans, with the largest gaps on anticipation and hypothetical questions.
-
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
GRPO post-training with structured thinking and dual think/caption rewards improves Qwen2-VL-7B video captioning over the base model and SFT on DREAM-1K, VDC, and CAREBENCH.
-
NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video Understanding
NoteIt converts instructional videos into interactive notes that preserve chapter and step structure and key visual and verbal information, and users significantly preferred it over a commercial baseline.
-
Comparing Learning Paradigms for Egocentric Video Summarization
A prompt-engineered GPT-4o (quality score 64.95) outperformed Shotluck Holmes (61.19) and TAC-SUM (58.43) on a 21-video egocentric summary evaluation, though all scores were modest.
Discussion (0). Sign in to comment.