Pith. sign in

REVIEW 7 cited by

IntPhys: A Framework and Benchmark for Visual Intuitive Physics Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1803.07616 v3 pith:XGLI2B44 submitted 2018-03-20 cs.AI cs.CV

classification cs.AIcs.CV
keywords physicssystemsintuitivepossiblebenchmarkphysicalpredictionreasoning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In order to reach human performance on complexvisual tasks, artificial systems need to incorporate a sig-nificant amount of understanding of the world in termsof macroscopic objects, movements, forces, etc. Inspiredby work on intuitive physics in infants, we propose anevaluation benchmark which diagnoses how much a givensystem understands about physics by testing whether itcan tell apart well matched videos of possible versusimpossible events constructed with a game engine. Thetest requires systems to compute a physical plausibilityscore over an entire video. It is free of bias and cantest a range of basic physical reasoning concepts. Wethen describe two Deep Neural Networks systems aimedat learning intuitive physics in an unsupervised way,using only physically possible videos. The systems aretrained with a future semantic mask prediction objectiveand tested on the possible versus impossible discrimi-nation task. The analysis of their results compared tohuman data gives novel insights in the potentials andlimitations of next frame prediction architectures.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 34 citations worldwide. Full citation record

  1. Apple-$\pi$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

    cs.CV 2026-07 conditional novelty 7.0 of 10

    A law-grounded benchmark, Apple-PI, grades video models stage-by-stage on physics reasoning and finds they top out at 0.473, well short of reliable simulation.

  2. VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

    cs.CV 2025-11 conditional novelty 6.0 of 10

    A 1,680-question video benchmark shows leading multimodal models lag humans by ~15 points on visual knowledge, and a See-Think-Answer RL-trained model narrows the gap.

  3. Do Vision-Language Models Have Internal World Models? Towards an Atomic Evaluation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new 23-dimension benchmark finds that even the best VLMs score near random on motion trajectory, temporal extension, and several prediction tasks, far below humans, suggesting weak internal world models.

  4. CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CausalVQA provides 793 paired real-video causal reasoning questions on which the best multimodal model scores 61.66% versus 84.78% for humans, with the largest gaps on anticipation and hypothetical questions.

  5. IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments

    cs.CV 2025-06 conditional novelty 6.0 of 10

    IntPhys 2 evaluates models on permanence, immutability, continuity, and solidity in synthetic videos, and finds most models near chance while humans near perfect.

  6. IMBench: A Benchmark for Intuitive Robotic Manipulation

    cs.RO 2026-07 conditional novelty 5.0 of 10

    IMBench is a 35-task robosuite benchmark with a three-stage evaluation showing current VLMs and robot policies fail to convert physical reasoning into executable manipulation.

  7. SlotPi: Physics-informed Object-centric Reasoning Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    SlotPi combines a learned Hamiltonian energy module with spatiotemporal attention to improve object-centric video prediction and visual question answering on several datasets.

Pith tools