Pith. sign in

REVIEW 9 cited by

Intuitive physics understanding emerges from self-supervised pretraining on natural videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.11831 v1 pith:CXGD6G4V submitted 2025-02-17 cs.CV cs.AI

classification cs.CVcs.AI
keywords intuitivephysicsunderstandingmodelsspacetrainedvideoachieve
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We investigate the emergence of intuitive physics understanding in general-purpose deep neural network models trained to predict masked regions in natural videos. Leveraging the violation-of-expectation framework, we find that video prediction models trained to predict outcomes in a learned representation space demonstrate an understanding of various intuitive physics properties, such as object permanence and shape consistency. In contrast, video prediction in pixel space and multimodal large language models, which reason through text, achieve performance closer to chance. Our comparisons of these architectures reveal that jointly learning an abstract representation space while predicting missing parts of sensory input, akin to predictive coding, is sufficient to acquire an understanding of intuitive physics, and that even models trained on one week of unique video achieve above chance performance. This challenges the idea that core knowledge -- a set of innate systems to help understand the world -- needs to be hardwired to develop an understanding of intuitive physics.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Prediction targets, not inputs, decide which physical parameters a latent world model acquires; a certified-recoverable drag parameter stays unlearned under every deterministic prediction objective tested.

  2. Apple-$\pi$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

    cs.CV 2026-07 conditional novelty 7.0 of 10

    A law-grounded benchmark, Apple-PI, grades video models stage-by-stage on physics reasoning and finds they top out at 0.473, well short of reliable simulation.

  3. IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments

    cs.CV 2025-06 conditional novelty 6.0 of 10

    IntPhys 2 evaluates models on permanence, immutability, continuity, and solidity in synthetic videos, and finds most models near chance while humans near perfect.

  4. WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A video-based benchmark shows that frontier AI models lag far behind humans on high-level world modeling and long-horizon procedural planning.

  5. VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    VideoREPA adds a token-relation distillation loss that aligns a text-to-video diffusion model's internal features with VideoMAEv2, boosting physical commonsense scores on VideoPhy and VideoPhy2.

  6. PAVXploreRL: Physical-Action-Visual World Model Reinforcement Learning with Action Exploration

    cs.CV 2026-07 conditional novelty 5.0 of 10

    PAVXploreRL post-trains action-conditioned world models with VJEPA-2 latent rewards and perturbed 'OOD' actions, reporting a 5.6% average gain and lowered policy-overestimation bias.

  7. Back to the Features: DINO as a Foundation for Video World Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A world model trained in frozen DINOv2 latent space on 66M videos beats much larger pixel-space models on forecasting and physics benchmarks, and fine-tunes for planning.

  8. Pixels to Principles: Probing Intuitive Physics Understanding in Multimodal Language Models

    cs.CL 2025-07 reject novelty 5.0 of 10

    Even the newest multimodal LLMs score near chance on intuitive physics videos, and the paper's probing evidence that vision encoders hold the relevant information is confounded by scene identity.

  9. From 2D to 3D Cognition: A Brief Survey of General World Models

    cs.CV 2025-06 conditional novelty 3.0 of 10

    A survey proposing a two-pillar, three-capability framework that organizes recent AI world models by their transition from 2D visual prediction to 3D cognition.

Pith tools