Pith. sign in

REVIEW 2 cited by

Learning to Answer Visual Questions from Web Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.05019 v2 pith:VYHSGZL6 submitted 2022-05-10 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords datasetvideoqageneratevideosanswersdatasetsmanualquestion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent methods for visual question answering rely on large-scale annotated datasets. Manual annotation of questions and answers for videos, however, is tedious, expensive and prevents scalability. In this work, we propose to avoid manual annotation and generate a large-scale training dataset for video question answering making use of automatic cross-modal supervision. We leverage a question generation transformer trained on text data and use it to generate question-answer pairs from transcribed video narrations. Given narrated videos, we then automatically generate the HowToVQA69M dataset with 69M video-question-answer triplets. To handle the open vocabulary of diverse answers in this dataset, we propose a training procedure based on a contrastive loss between a video-question multi-modal transformer and an answer transformer. We introduce the zero-shot VideoQA task and the VideoQA feature probe evaluation setting and show excellent results, in particular for rare answers. Furthermore, our method achieves competitive results on MSRVTT-QA, ActivityNet-QA, MSVD-QA and How2QA datasets. We also show that our VideoQA dataset generation approach generalizes to another source of web video and text data. We use our method to generate the WebVidVQA3M dataset from the WebVid dataset, i.e., videos with alt-text annotations, and show its benefits for training VideoQA models. Finally, for a detailed evaluation we introduce iVQA, a new VideoQA dataset with reduced language bias and high-quality manual annotations. Code, datasets and trained models are available at https://antoyang.github.io/just-ask.html

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos

    cs.CV 2024-11 conditional novelty 6.0 of 10

    LangView uses per-view caption accuracy against view-agnostic narrations as pseudo-labels to train a view selector that outperforms heuristics and prior baselines on Ego-Exo4D and LEMMA.

  2. Hierarchical Banzhaf Interaction for General Video-Language Representation Learning

    cs.CV 2024-12 conditional novelty 4.0 of 10

    HBI V2 models video-text alignment as a cooperative game with Hierarchical Banzhaf Interaction plus single/cross-modal representation fusion, improving retrieval, QA, and captioning benchmarks.

Pith tools