Pith. sign in

REVIEW 3 cited by

Object-Shot Enhanced Grounding Network for Egocentric Video

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.04270 v1 pith:QCE4NE5K submitted 2025-05-07 cs.CV cs.AI

classification cs.CVcs.AI
keywords egocentricvideovideosgroundinginformationosgnetenhancedexocentric
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Egocentric video grounding is a crucial task for embodied intelligence applications, distinct from exocentric video moment localization. Existing methods primarily focus on the distributional differences between egocentric and exocentric videos but often neglect key characteristics of egocentric videos and the fine-grained information emphasized by question-type queries. To address these limitations, we propose OSGNet, an Object-Shot enhanced Grounding Network for egocentric video. Specifically, we extract object information from videos to enrich video representation, particularly for objects highlighted in the textual query but not directly captured in the video features. Additionally, we analyze the frequent shot movements inherent to egocentric videos, leveraging these features to extract the wearer's attention information, which enhances the model's ability to perform modality alignment. Experiments conducted on three datasets demonstrate that OSGNet achieves state-of-the-art performance, validating the effectiveness of our approach. Our code can be found at https://github.com/Yisen-Feng/OSGNet.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OSGNet @ Ego4D Episodic Memory Challenge 2025

    cs.CV 2025-06 conditional novelty 4.0 of 10

    OSGNet, an early-fusion grounding model, wins all three Ego4D Episodic Memory Challenge tracks by converting localization tasks into retrieval problems.

  2. Technical Report for Ego4D Long-Term Action Anticipation Challenge 2025

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A three-stage pipeline using the EgoVideo-V encoder, a verb-noun co-occurrence reranker, SAM2 hand-object features, and a fine-tuned Llama 2 model took first place in the Ego4D 2025 long-term action anticipation challenge.

  3. HCQA-1.5 @ Ego4D EgoSchema Challenge 2025

    cs.CV 2025-05 conditional novelty 4.0 of 10

    An ensemble of LLMs with confidence filtering and low-confidence re-reasoning reaches 77% accuracy on the EgoSchema benchmark, up from 75% for the prior HCQA system.

Pith tools