REVIEW 1 cited by
Dense-Captioning Events in Videos: SYSU Submission to ActivityNet Challenge 2020
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Dense-Captioning Events in Videos: SYSU Submission to ActivityNet Challenge 2020
read the original abstract
This technical report presents a brief description of our submission to the dense video captioning task of ActivityNet Challenge 2020. Our approach follows a two-stage pipeline: first, we extract a set of temporal event proposals; then we propose a multi-event captioning model to capture the event-level temporal relationships and effectively fuse the multi-modal information. Our approach achieves a 9.28 METEOR score on the test set.
Forward citations
Cited by 1 Pith paper
-
Sali4Vid: Saliency-Aware Video Reweighting and Adaptive Caption Retrieval for Dense Video Captioning
Sali4Vid improves dense video captioning by reweighting video features with timestamp-derived sigmoid importance and adaptively retrieving captions per semantic segment, achieving new SOTA on YouCook2 and ViTT.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.