Pith. sign in

REVIEW 1 cited by

MDMMT-2: Multidomain Multimodal Transformer for Video Retrieval, One More Step Towards Generalization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.07086 v1 pith:N3HPCEOS submitted 2022-03-14 cs.CV

classification cs.CV
keywords differentknowledgepairsretrievaltrainingadditionallyallowsanalysis
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this work we present a new State-of-The-Art on the text-to-video retrieval task on MSR-VTT, LSMDC, MSVD, YouCook2 and TGIF obtained by a single model. Three different data sources are combined: weakly-supervised videos, crowd-labeled text-image pairs and text-video pairs. A careful analysis of available pre-trained networks helps to choose the best prior-knowledge ones. We introduce three-stage training procedure that provides high transfer knowledge efficiency and allows to use noisy datasets during training without prior knowledge degradation. Additionally, double positional encoding is used for better fusion of different modalities and a simple method for non-square inputs processing is suggested.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review

    cs.CV 2025-05 conditional novelty 3.0 of 10

    A structured review of 81 text-to-video retrieval papers that leverage auxiliary information, organized by a taxonomy and compared on standard benchmarks.

Pith tools