Pith. sign in

REVIEW 11 cited by

VideoAgent: Long-form Video Understanding with Large Language Model as Agent

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.10517 v1 pith:6DVJE5ZL submitted 2024-03-15 cs.CV cs.AIcs.CLcs.IR

classification cs.CVcs.AIcs.CLcs.IR
keywords long-formunderstandingvideomodelvideoagentagentagent-basedinformation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Long-form video understanding represents a significant challenge within computer vision, demanding a model capable of reasoning over long multi-modal sequences. Motivated by the human cognitive process for long-form video understanding, we emphasize interactive reasoning and planning over the ability to process lengthy visual inputs. We introduce a novel agent-based system, VideoAgent, that employs a large language model as a central agent to iteratively identify and compile crucial information to answer a question, with vision-language foundation models serving as tools to translate and retrieve visual information. Evaluated on the challenging EgoSchema and NExT-QA benchmarks, VideoAgent achieves 54.1% and 71.3% zero-shot accuracy with only 8.4 and 8.2 frames used on average. These results demonstrate superior effectiveness and efficiency of our method over the current state-of-the-art methods, highlighting the potential of agent-based approaches in advancing long-form video understanding.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EgoExo-Con: Exploring View-Invariant Video Temporal Understanding

    cs.CV 2025-10 conditional novelty 6.0 of 10

    Most Video-LLMs answer temporal questions far less consistently when the same event is shown from ego and exo views, and a GRPO variant with a reasoning-similarity reward partially closes the gap.

  2. Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities

    cs.LG 2025-07 conditional novelty 6.0 of 10

    By probing visual, projection, and response representations, the authors find that most VLM visual knowledge loss for recognition and counting occurs in the language decoder, while spatial understanding is lost in the...

  3. AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A 2B-parameter video-language model using an RWKV linear-RNN backbone and sorted token merging achieves competitive long-video QA accuracy with far lower memory cost than transformer-based models.

  4. Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames

    cs.LG 2025-07 conditional novelty 6.0 of 10

    TCoT uses a single VLM to select question-relevant video frames from segments, then answers from that curated context, improving video QA accuracy across four benchmarks and three VLMs.

  5. TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Weakly supervised vision-language model jointly generating open-ended video QA answers with temporal groundings, reporting SOTA on NExT-GQA, MSVD-QA, and ActivityNet-QA.

  6. MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks

    cs.CV 2025-06 conditional novelty 6.0 of 10

    AVHaystacks is a new 3100-question benchmark for audio-visual QA across 500 videos, and the MAGNET multi-agent pipeline beats current baselines on it.

  7. Towards Sparse Video Understanding and Reasoning

    cs.CV 2026-02 conditional novelty 5.0 of 10

    A video-QA agent that carries only a structured text summary between rounds beats dense-frame baselines on accuracy while using a handful of frames per video.

  8. NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video Understanding

    cs.HC 2025-08 conditional novelty 5.0 of 10

    NoteIt converts instructional videos into interactive notes that preserve chapter and step structure and key visual and verbal information, and users significantly preferred it over a commercial baseline.

  9. Frame-Level Captions for Long Video Generation with Complex Multi Scenes

    cs.CV 2025-05 conditional novelty 5.0 of 10

    Frame-level captions with per-frame cross-attention and parallel multi-window denoising reduce semantic confusion in long multi-scene video generation in the authors' internal evaluation.

  10. HCQA-1.5 @ Ego4D EgoSchema Challenge 2025

    cs.CV 2025-05 conditional novelty 4.0 of 10

    An ensemble of LLMs with confidence filtering and low-confidence re-reasoning reaches 77% accuracy on the EgoSchema benchmark, up from 75% for the prior HCQA system.

  11. Four Eyes Are Better Than Two: Harnessing the Collaborative Potential of Large Models via Differentiated Thinking and Complementary Ensembles

    cs.CV 2025-05 reject novelty 4.0 of 10

    A training-free ensemble of commercial VLMs with prompt and chain-of-thought engineering reaches 79% on EgoSchema, ranking 2nd in the CVPR 2025 challenge, but the ensemble weights are fit to the test labels.

Pith tools