Pith. sign in

REVIEW 2 cited by

Support-set bottlenecks for video-text representation learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.02824 v2 pith:WXUUDIJI submitted 2020-10-06 cs.CV

classification cs.CV
keywords representationssampleslearningcontrastivemethodnoiseotherpairs
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The dominant paradigm for learning video-text representations -- noise contrastive learning -- increases the similarity of the representations of pairs of samples that are known to be related, such as text and video from the same sample, and pushes away the representations of all other pairs. We posit that this last behaviour is too strict, enforcing dissimilar representations even for samples that are semantically-related -- for example, visually similar videos or ones that share the same depicted action. In this paper, we propose a novel method that alleviates this by leveraging a generative model to naturally push these related samples together: each sample's caption must be reconstructed as a weighted combination of other support samples' visual representations. This simple idea ensures that representations are not overly-specialized to individual samples, are reusable across the dataset, and results in representations that explicitly encode semantics shared between samples, unlike noise contrastive learning. Our proposed method outperforms others by a large margin on MSR-VTT, VATEX and ActivityNet, and MSVD for video-to-text and text-to-video retrieval.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PHA-Net: Prototype-based Hierarchical Alignment Network for Text-Video Retrieval

    cs.IR 2026-08 conditional novelty 5.0 of 10

    PHA-Net inserts shared prototype tokens into a three-level text-video alignment model and reports higher aggregate retrieval scores than the HBI baseline on four benchmarks, though several gains are small and unverified.

  2. Distribution-Alignment Bridge for Uncertainty-Aware Text-to-Video Retrieval

    cs.CV 2026-07 conditional novelty 5.0 of 10

    DAB refines Gaussian-distributed text embeddings toward video distributions through a deterministic truncated bridge with a directional KL contrastive loss, claiming state-of-the-art recall and mean-rank on three vide...

Pith tools