Pith. sign in

REVIEW 13 cited by

STAR: A Benchmark for Situated Reasoning in Real-World Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.09711 v1 pith:DUAOWGTZ submitted 2024-05-15 cs.AI cs.CLcs.CV

classification cs.AIcs.CLcs.CV
keywords reasoningbenchmarkreal-worldsituatedvideossituationsituationsabstraction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reasoning in the real world is not divorced from situations. How to capture the present knowledge from surrounding situations and perform reasoning accordingly is crucial and challenging for machine intelligence. This paper introduces a new benchmark that evaluates the situated reasoning ability via situation abstraction and logic-grounded question answering for real-world videos, called Situated Reasoning in Real-World Videos (STAR Benchmark). This benchmark is built upon the real-world videos associated with human actions or interactions, which are naturally dynamic, compositional, and logical. The dataset includes four types of questions, including interaction, sequence, prediction, and feasibility. We represent the situations in real-world videos by hyper-graphs connecting extracted atomic entities and relations (e.g., actions, persons, objects, and relationships). Besides visual perception, situated reasoning also requires structured situation comprehension and logical reasoning. Questions and answers are procedurally generated. The answering logic of each question is represented by a functional program based on a situation hyper-graph. We compare various existing video reasoning models and find that they all struggle on this challenging situated reasoning task. We further propose a diagnostic neuro-symbolic model that can disentangle visual perception, situation abstraction, language understanding, and functional reasoning to understand the challenges of this benchmark.

Discussion (0). Sign in to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 22 citations worldwide. Full citation record

  1. ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding

    cs.CV 2025-05 conditional novelty 7.0 of 10

    ScaleLong embeds four timescale question types into the same long videos, and evaluation of 23 MLLMs reveals a U-shaped accuracy curve across timescales.

  2. Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding

    cs.CV 2026-01 conditional novelty 6.0 of 10

    A structured five-step reasoning template plus diverse-trajectory cold start and diversity-preserving two-stage RL lifts a 7B multimodal model to state-of-the-art multi-image reasoning on several benchmarks.

  3. VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

    cs.CV 2025-11 conditional novelty 6.0 of 10

    A 1,680-question video benchmark shows leading multimodal models lag humans by ~15 points on visual knowledge, and a See-Think-Answer RL-trained model narrows the gap.

  4. InterAct-Video: Reasoning-Rich Video QA for Urban Traffic

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new traffic-intersection VideoQA benchmark containing roughly 28,800 human-verified GPT-seeded QA pairs, with evaluations showing fine-tuning improves three video-language models.

  5. IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A new multi-shot video dataset and an instance-prompt video LLM report large gains, but the main benchmark is built by the same authors and the model is not released.

  6. VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos

    cs.CV 2025-06 conditional novelty 6.0 of 10

    VRBench is a benchmark of 960 long narrative videos with 8,243 human-written multi-step questions, plus a two-level evaluation of answer accuracy and reasoning quality for 31 large models.

  7. TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Weakly supervised vision-language model jointly generating open-ended video QA answers with temporal groundings, reporting SOTA on NExT-GQA, MSVD-QA, and ActivityNet-QA.

  8. MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks

    cs.CV 2025-06 conditional novelty 6.0 of 10

    AVHaystacks is a new 3100-question benchmark for audio-visual QA across 500 videos, and the MAGNET multi-agent pipeline beats current baselines on it.

  9. InterRVOS: Interaction-aware Referring Video Object Segmentation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    InterRVOS extends referring video object segmentation to segment both actor and target objects separately for interaction expressions, with a new dataset and MLLM-based method.

  10. Rethinking Causal Mask Attention for Vision-Language Inference

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Relaxing causal masking so image tokens can preview future image and text context during prefill improves several vision-language benchmarks, and pooling future attention into a single prefix token preserves most of the gain.

  11. IMBench: A Benchmark for Intuitive Robotic Manipulation

    cs.RO 2026-07 conditional novelty 5.0 of 10

    IMBench is a 35-task robosuite benchmark with a three-stage evaluation showing current VLMs and robot policies fail to convert physical reasoning into executable manipulation.

  12. Reinforcing Video Reasoning with Focused Thinking

    cs.CV 2025-05 reject novelty 5.0 of 10

    A GRPO variant with token-level KL weighting and partial-credit rewards improves video-QA on modified multi-answer benchmarks, but transfer to original single-answer benchmarks is not established.

  13. POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency

    cs.CV 2025-10 reject novelty 4.0 of 10

    POVQA reports large F1 gains on a new 239-example video QA dataset after rationale-based fine-tuning, but its own keyframe-only ablation matches the full pooling pipeline, and fine-tuning hurts zero-shot TVQA accuracy.

Pith tools