REVIEW 13 cited by
VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large-scale video generative models, capable of creating realistic videos of diverse visual concepts, are strong candidates for general-purpose physical world simulators. However, their adherence to physical commonsense across real-world actions remains unclear (e.g., playing tennis, backflip). Existing benchmarks suffer from limitations such as limited size, lack of human evaluation, sim-to-real gaps, and absence of fine-grained physical rule analysis. To address this, we introduce VideoPhy-2, an action-centric dataset for evaluating physical commonsense in generated videos. We curate 200 diverse actions and detailed prompts for video synthesis from modern generative models. We perform human evaluation that assesses semantic adherence, physical commonsense, and grounding of physical rules in the generated videos. Our findings reveal major shortcomings, with even the best model achieving only 22% joint performance (i.e., high semantic and physical commonsense adherence) on the hard subset of VideoPhy-2. We find that the models particularly struggle with conservation laws like mass and momentum. Finally, we also train VideoPhy-AutoEval, an automatic evaluator for fast, reliable assessment on our dataset. Overall, VideoPhy-2 serves as a rigorous benchmark, exposing critical gaps in video generative models and guiding future research in physically-grounded video generation. The data and code is available at https://videophy2.github.io/.
Forward citations
Cited by 13 Pith papers
-
GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models
A new real-world-grounded benchmark shows that physics engines and video world models each fail differently, with video models often fitting the shape of a physical law while recovering wrong parameters.
-
Apple-$\pi$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
A law-grounded benchmark, Apple-PI, grades video models stage-by-stage on physics reasoning and finds they top out at 0.473, well short of reliable simulation.
-
MemoBench: Benchmarking World Modeling in Dynamically Changing Environments
None of ten tested video-generation models reliably remembers objects after occlusion in dynamic scenes; static-camera videos inflate consistency scores.
-
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System
Using executable Blender code as an intermediate simulation draft improves physical consistency in text-to-video generation, lifting OmniWeaving from 0.475 to 0.558 on PhyGenBench and from 52.18% to 77.88% on VBench-2.0.
-
Physics-Grounded Fluid Video Generation with a Simulation Dataset and Dual-Stream Optical-Flow Supervision
Adding a dual-stream optical-flow decoder and a simulation+real fluid dataset to a frozen video diffusion model improves the physical plausibility of generated pours and splashes.
-
Learning Explicit Physical Parameter Control and Benchmarking for Video Generation
Explicit instance-level physical parameter conditioning with routing attention improves physical-law consistency in image-to-video generation, as measured on the authors' new simulator-based benchmark.
-
Thinking in Video: Can Video Generators Really Reason About the Real World?
Video generators show a perception-prediction gap: they can generate plausible continuations while failing explicit visual reasoning tests.
-
Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos
A learned human-action manifold, combining 3D pose, 2D keypoints, appearance, and motion derivatives, scores generated videos by distance to real-action centroids and embedding smoothness, beating prior metrics on hum...
-
Human Preference-Aligned Concept Customization Benchmark via Decomposed Evaluation
D-GPTScore, which averages GPT-4o's per-aspect ratings of concept-customized images, correlates with human preference at 0.78 Pearson on the new CC-AlignBench, beating prior metrics.
-
Can Your Model Separate Yolks with a Water Bottle? Benchmarking Physical Commonsense Understanding in Video Generation Models
PhysVidBench evaluates text-to-video models with 383 PIQA-derived prompts and a caption-based QA pipeline, finding all tested models score below 40% on physical commonsense, with spatial and temporal reasoning the weakest.
-
"PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
PhyWorldBench evaluates 12 text-to-video models on 1,050 physics prompts; the best model passes both semantic adherence and physical commonsense checks in only 26.2% of videos.
-
VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models
VideoREPA adds a token-relation distillation loss that aligns a text-to-video diffusion model's internal features with VideoMAEv2, boosting physical commonsense scores on VideoPhy and VideoPhy2.
-
When Physical Preferences Meet Semantic Constraints: Physical and Semantic Direct Preference Optimization for Text-to-Video Generation
A semantic-conflict reweighting and staged-training version of DPO improves physical plausibility in text-to-video generation while partly preserving prompt semantics.
Discussion (0). Continue with ORCID to comment.