REVIEW 9 cited by
Intuitive physics understanding emerges from self-supervised pretraining on natural videos
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We investigate the emergence of intuitive physics understanding in general-purpose deep neural network models trained to predict masked regions in natural videos. Leveraging the violation-of-expectation framework, we find that video prediction models trained to predict outcomes in a learned representation space demonstrate an understanding of various intuitive physics properties, such as object permanence and shape consistency. In contrast, video prediction in pixel space and multimodal large language models, which reason through text, achieve performance closer to chance. Our comparisons of these architectures reveal that jointly learning an abstract representation space while predicting missing parts of sensory input, akin to predictive coding, is sufficient to acquire an understanding of intuitive physics, and that even models trained on one week of unique video achieve above chance performance. This challenges the idea that core knowledge -- a set of innate systems to help understand the world -- needs to be hardwired to develop an understanding of intuitive physics.
Forward citations
Cited by 9 Pith papers
-
What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations
Prediction targets, not inputs, decide which physical parameters a latent world model acquires; a certified-recoverable drag parameter stays unlearned under every deterministic prediction objective tested.
-
Apple-$\pi$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
A law-grounded benchmark, Apple-PI, grades video models stage-by-stage on physics reasoning and finds they top out at 0.473, well short of reliable simulation.
-
IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments
IntPhys 2 evaluates models on permanence, immutability, continuity, and solidity in synthetic videos, and finds most models near chance while humans near perfect.
-
WorldPrediction: A Benchmark for High-level World Modeling and Long-horizon Procedural Planning
A video-based benchmark shows that frontier AI models lag far behind humans on high-level world modeling and long-horizon procedural planning.
-
VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models
VideoREPA adds a token-relation distillation loss that aligns a text-to-video diffusion model's internal features with VideoMAEv2, boosting physical commonsense scores on VideoPhy and VideoPhy2.
-
PAVXploreRL: Physical-Action-Visual World Model Reinforcement Learning with Action Exploration
PAVXploreRL post-trains action-conditioned world models with VJEPA-2 latent rewards and perturbed 'OOD' actions, reporting a 5.6% average gain and lowered policy-overestimation bias.
-
Back to the Features: DINO as a Foundation for Video World Models
A world model trained in frozen DINOv2 latent space on 66M videos beats much larger pixel-space models on forecasting and physics benchmarks, and fine-tunes for planning.
-
Pixels to Principles: Probing Intuitive Physics Understanding in Multimodal Language Models
Even the newest multimodal LLMs score near chance on intuitive physics videos, and the paper's probing evidence that vision encoders hold the relevant information is confounded by scene identity.
-
From 2D to 3D Cognition: A Brief Survey of General World Models
A survey proposing a two-pillar, three-capability framework that organizes recent AI world models by their transition from 2D visual prediction to 3D cognition.
Discussion (0). Continue with ORCID to comment.