Pith. sign in

REVIEW 2 cited by

GroundNLQ @ Ego4D Natural Language Queries Challenge 2023

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.15255 v1 pith:Y6OARUHJ submitted 2023-06-27 cs.CV cs.CL

classification cs.CVcs.CL
keywords groundinggroundnlqmodelvideochallengeeffectiveego4degocentric
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this report, we present our champion solution for Ego4D Natural Language Queries (NLQ) Challenge in CVPR 2023. Essentially, to accurately ground in a video, an effective egocentric feature extractor and a powerful grounding model are required. Motivated by this, we leverage a two-stage pre-training strategy to train egocentric feature extractors and the grounding model on video narrations, and further fine-tune the model on annotated data. In addition, we introduce a novel grounding model GroundNLQ, which employs a multi-modal multi-scale grounding module for effective video and text fusion and various temporal intervals, especially for long videos. On the blind test set, GroundNLQ achieves 25.67 and 18.18 for R1@IoU=0.3 and R1@IoU=0.5, respectively, and surpasses all other teams by a noticeable margin. Our code will be released at\url{https://github.com/houzhijian/GroundNLQ}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GazeNLQ @ Ego4D Natural Language Queries Challenge 2025

    cs.CV 2025-06 conditional novelty 5.0 of 10

    GazeNLQ adds contrastively pretrained gaze embeddings to a GroundNLQ-style grounding model, reporting 27.82 R1@0.3 on the Ego4D NLQ test split only when ensembled with GroundVQA.

  2. OSGNet @ Ego4D Episodic Memory Challenge 2025

    cs.CV 2025-06 conditional novelty 4.0 of 10

    OSGNet, an early-fusion grounding model, wins all three Ego4D Episodic Memory Challenge tracks by converting localization tasks into retrieval problems.

Pith tools