Pith. sign in

REVIEW 2 cited by

Self-view Grounding Given a Narrated 360{\deg} Video

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1711.08664 v1 pith:V6BMWCBQ submitted 2017-11-23 cs.CV

classification cs.CV
keywords sentencegivennfovscandidategroundinglossnarratednfov
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Narrated 360{\deg} videos are typically provided in many touring scenarios to mimic real-world experience. However, previous work has shown that smart assistance (i.e., providing visual guidance) can significantly help users to follow the Normal Field of View (NFoV) corresponding to the narrative. In this project, we aim at automatically grounding the NFoVs of a 360{\deg} video given subtitles of the narrative (referred to as "NFoV-grounding"). We propose a novel Visual Grounding Model (VGM) to implicitly and efficiently predict the NFoVs given the video content and subtitles. Specifically, at each frame, we efficiently encode the panorama into feature map of candidate NFoVs using a Convolutional Neural Network (CNN) and the subtitles to the same hidden space using an RNN with Gated Recurrent Units (GRU). Then, we apply soft-attention on candidate NFoVs to trigger sentence decoder aiming to minimize the reconstruct loss between the generated and given sentence. Finally, we obtain the NFoV as the candidate NFoV with the maximum attention without any human supervision. To train VGM more robustly, we also generate a reverse sentence conditioning on one minus the soft-attention such that the attention focuses on candidate NFoVs less relevant to the given sentence. The negative log reconstruction loss of the reverse sentence (referred to as "irrelevant loss") is jointly minimized to encourage the reverse sentence to be different from the given sentence. To evaluate our method, we collect the first narrated 360{\deg} videos dataset and achieve state-of-the-art NFoV-grounding performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos

    cs.CV 2024-12 conditional novelty 6.0 of 10

    The authors train a view-switch predictor on pseudo-labeled web videos and show it transfers to selecting ego/exo views in new multi-view videos with limited labels.

  2. Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos

    cs.CV 2024-11 conditional novelty 6.0 of 10

    LangView uses per-view caption accuracy against view-agnostic narrations as pseudo-labels to train a view selector that outperforms heuristics and prior baselines on Ego-Exo4D and LEMMA.

Pith tools