Video self-distillation with a next-frame dense prediction objective improves a single-image ViT's downstream ADE20K segmentation mIoU from 35.0 to 36.4 and COCO mAP from 33.0 to 33.5 after pre-training on one 2-hour video.
Is space-time attention all you need for video understanding? In Int
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Video Self-Distillation for Single-Image Encoders: A Step Toward Physically Plausible Perception
Video self-distillation with a next-frame dense prediction objective improves a single-image ViT's downstream ADE20K segmentation mIoU from 35.0 to 36.4 and COCO mAP from 33.0 to 33.5 after pre-training on one 2-hour video.