Pith. sign in

REVIEW 6 cited by

Scaling 4D Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.15212 v2 pith:CE3YOYL2 submitted 2024-12-19 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords videolearningmodelsscalingself-supervisedtasksclassificationestimation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Scaling has not yet been convincingly demonstrated for pure self-supervised learning from video. However, prior work has focused evaluations on semantic-related tasks $\unicode{x2013}$ action classification, ImageNet classification, etc. In this paper we focus on evaluating self-supervised learning on non-semantic vision tasks that are more spatial (3D) and temporal (+1D = 4D), such as camera pose estimation, point and object tracking, and depth estimation. We show that by learning from very large video datasets, masked auto-encoding (MAE) with transformer video models actually scales, consistently improving performance on these 4D tasks, as model size increases from 20M all the way to the largest by far reported self-supervised video model $\unicode{x2013}$ 22B parameters. Rigorous apples-to-apples comparison with many recent image and video models demonstrates the benefits of scaling 4D representations. Pretrained models are available at https://github.com/google-deepmind/representations4d .

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unique Lives, Shared World: Learning from Single-Life Videos

    cs.CV 2025-12 conditional novelty 7.0 of 10

    Vision models trained independently on single egocentric lives converge to aligned geometric representations, and about 30 hours of one life matches 30 hours of diverse video for depth-estimation pretraining.

  2. Self-Supervised Learning of Structured Dynamics from Videos

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A two-token 'primary/residual' future-feature predictor separates camera from object motion, outpacing frozen-feature baselines and matching larger supervised models on several probes.

  3. SeeSE3: Emergence of 3D Space in Vision Features

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Self-supervised vision features, especially DINOv2, contain a subspace that a small trained adapter can map to 3D camera motion, enabling pose estimation and latent-space navigation without explicit 3D reconstruction.

  4. Gen4U: Unifying Video Generation and Understanding via Diffusion

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Frozen video diffusion models, probed at optimal depth and noise levels, produce representations competitive with discriminative encoders across semantic and geometric video tasks in a single forward pass.

  5. SciVid: Cross-Domain Evaluation of Video Models in Scientific Applications

    cs.CV 2025-07 conditional novelty 6.0 of 10

    General-purpose video foundation models, adapted with lightweight readout heads, reach state-of-the-art performance on three of five scientific video benchmarks.

  6. MoSiC: Optimal-Transport Motion Trajectory for Dense Self-Supervised Learning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    MoSiC clusters dense point tracks in videos and propagates the cluster assignments along the tracks, improving DINOv2's dense representations by 1 to 6 percent on segmentation and in-context benchmarks.

Pith tools