Pith. sign in

REVIEW 13 cited by

VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.06800 v1 pith:UGCAGKK2 submitted 2025-03-09 cs.CV

classification cs.CV
keywords physicalcommonsensevideomodelsvideophy-2adherenceevaluationgenerative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large-scale video generative models, capable of creating realistic videos of diverse visual concepts, are strong candidates for general-purpose physical world simulators. However, their adherence to physical commonsense across real-world actions remains unclear (e.g., playing tennis, backflip). Existing benchmarks suffer from limitations such as limited size, lack of human evaluation, sim-to-real gaps, and absence of fine-grained physical rule analysis. To address this, we introduce VideoPhy-2, an action-centric dataset for evaluating physical commonsense in generated videos. We curate 200 diverse actions and detailed prompts for video synthesis from modern generative models. We perform human evaluation that assesses semantic adherence, physical commonsense, and grounding of physical rules in the generated videos. Our findings reveal major shortcomings, with even the best model achieving only 22% joint performance (i.e., high semantic and physical commonsense adherence) on the hard subset of VideoPhy-2. We find that the models particularly struggle with conservation laws like mass and momentum. Finally, we also train VideoPhy-AutoEval, an automatic evaluator for fast, reliable assessment on our dataset. Overall, VideoPhy-2 serves as a rigorous benchmark, exposing critical gaps in video generative models and guiding future research in physically-grounded video generation. The data and code is available at https://videophy2.github.io/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models

    cs.AI 2026-08 conditional novelty 7.0 of 10

    A new real-world-grounded benchmark shows that physics engines and video world models each fail differently, with video models often fitting the shape of a physical law while recovering wrong parameters.

  2. Apple-$\pi$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

    cs.CV 2026-07 conditional novelty 7.0 of 10

    A law-grounded benchmark, Apple-PI, grades video models stage-by-stage on physics reasoning and finds they top out at 0.473, well short of reliable simulation.

  3. MemoBench: Benchmarking World Modeling in Dynamically Changing Environments

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    None of ten tested video-generation models reliably remembers objects after occlusion in dynamic scenes; static-camera videos inflate consistency scores.

  4. VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Using executable Blender code as an intermediate simulation draft improves physical consistency in text-to-video generation, lifting OmniWeaving from 0.475 to 0.558 on PhyGenBench and from 52.18% to 77.88% on VBench-2.0.

  5. Physics-Grounded Fluid Video Generation with a Simulation Dataset and Dual-Stream Optical-Flow Supervision

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Adding a dual-stream optical-flow decoder and a simulation+real fluid dataset to a frozen video diffusion model improves the physical plausibility of generated pours and splashes.

  6. Learning Explicit Physical Parameter Control and Benchmarking for Video Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Explicit instance-level physical parameter conditioning with routing attention improves physical-law consistency in image-to-video generation, as measured on the authors' new simulator-based benchmark.

  7. Thinking in Video: Can Video Generators Really Reason About the Real World?

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Video generators show a perception-prediction gap: they can generate plausible continuations while failing explicit visual reasoning tests.

  8. Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos

    cs.CV 2025-12 conditional novelty 6.0 of 10

    A learned human-action manifold, combining 3D pose, 2D keypoints, appearance, and motion derivatives, scores generated videos by distance to real-action centroids and embedding smoothness, beating prior metrics on hum...

  9. Human Preference-Aligned Concept Customization Benchmark via Decomposed Evaluation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    D-GPTScore, which averages GPT-4o's per-aspect ratings of concept-customized images, correlates with human preference at 0.78 Pearson on the new CC-AlignBench, beating prior metrics.

  10. Can Your Model Separate Yolks with a Water Bottle? Benchmarking Physical Commonsense Understanding in Video Generation Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    PhysVidBench evaluates text-to-video models with 383 PIQA-derived prompts and a caption-based QA pipeline, finding all tested models score below 40% on physical commonsense, with spatial and temporal reasoning the weakest.

  11. "PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    PhyWorldBench evaluates 12 text-to-video models on 1,050 physics prompts; the best model passes both semantic adherence and physical commonsense checks in only 26.2% of videos.

  12. VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    VideoREPA adds a token-relation distillation loss that aligns a text-to-video diffusion model's internal features with VideoMAEv2, boosting physical commonsense scores on VideoPhy and VideoPhy2.

  13. When Physical Preferences Meet Semantic Constraints: Physical and Semantic Direct Preference Optimization for Text-to-Video Generation

    cs.CV 2026-07 conditional novelty 4.0 of 10

    A semantic-conflict reweighting and staged-training version of DPO improves physical plausibility in text-to-video generation while partly preserving prompt semantics.

Pith tools