REVIEW 7 cited by
PISA Experiments: Exploring Physics Post-Training for Video Diffusion Models by Watching Stuff Drop
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large-scale pre-trained video generation models excel in content creation but are not reliable as physically accurate world simulators out of the box. This work studies the process of post-training these models for accurate world modeling through the lens of the simple, yet fundamental, physics task of modeling object freefall. We show state-of-the-art video generation models struggle with this basic task, despite their visually impressive outputs. To remedy this problem, we find that fine-tuning on a relatively small amount of simulated videos is effective in inducing the dropping behavior in the model, and we can further improve results through a novel reward modeling procedure we introduce. Our study also reveals key limitations of post-training in generalization and distribution modeling. Additionally, we release a benchmark for this task that may serve as a useful diagnostic tool for tracking physical accuracy in large-scale video generative model development.
Forward citations
Cited by 7 Pith papers
-
GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models
A new real-world-grounded benchmark shows that physics engines and video world models each fail differently, with video models often fitting the shape of a physical law while recovering wrong parameters.
-
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System
Using executable Blender code as an intermediate simulation draft improves physical consistency in text-to-video generation, lifting OmniWeaving from 0.475 to 0.558 on PhyGenBench and from 52.18% to 77.88% on VBench-2.0.
-
Learning Explicit Physical Parameter Control and Benchmarking for Video Generation
Explicit instance-level physical parameter conditioning with routing attention improves physical-law consistency in image-to-video generation, as measured on the authors' new simulator-based benchmark.
-
VDAWorld: World Modelling via VLM-Directed Abstraction and Simulation
A vision-language model writes a simulation program—grounded by segmentation and 3D tools—that predicts physically plausible futures from an image and caption, outperforming video generators on modified benchmarks.
-
Epipolar Geometry Improves Video Generation Models
Ranking generated videos by their epipolar (Sampson) error and fine-tuning Wan2.1 with Flow-DPO cuts epipolar error 31% and raises human-rated 3D consistency from 54% to 72%.
-
From Black Box to Transparency: Enhancing Automated Interpreting Assessment with Explainable AI in College Classrooms
An explainable model using BLEURT, CometKiwi, pause features, and Chinese phraseological diversity predicts human-rated quality dimensions in English-Chinese consecutive interpreting, with SHAP identifying the stronge...
-
Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation
A hierarchical direct preference optimization with four alignment levels plus automated data selection improves physical plausibility of text-to-video models.
Discussion (0). Continue with ORCID to comment.