Pith. sign in

REVIEW 34 cited by

PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.16411 v2 pith:L5JB2337 submitted 2025-01-27 cs.CV cs.AIcs.CLcs.LGcs.RO

PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding

classification cs.CV cs.AIcs.CLcs.LGcs.RO
keywords physicalunderstandingvlmsworldmodelsphysbenchagentsembodied
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Understanding the physical world is a fundamental challenge in embodied AI, critical for enabling agents to perform complex tasks and operate safely in real-world environments. While Vision-Language Models (VLMs) have shown great promise in reasoning and task planning for embodied agents, their ability to comprehend physical phenomena remains extremely limited. To close this gap, we introduce PhysBench, a comprehensive benchmark designed to evaluate VLMs' physical world understanding capability across a diverse set of tasks. PhysBench contains 10,002 entries of interleaved video-image-text data, categorized into four major domains: physical object properties, physical object relationships, physical scene understanding, and physics-based dynamics, further divided into 19 subclasses and 8 distinct capability dimensions. Our extensive experiments, conducted on 75 representative VLMs, reveal that while these models excel in common-sense reasoning, they struggle with understanding the physical world -- likely due to the absence of physical knowledge in their training data and the lack of embedded physical priors. To tackle the shortfall, we introduce PhysAgent, a novel framework that combines the generalization strengths of VLMs with the specialized expertise of vision models, significantly enhancing VLMs' physical understanding across a variety of tasks, including an 18.4\% improvement on GPT-4o. Furthermore, our results demonstrate that enhancing VLMs' physical world understanding capabilities can help embodied agents such as MOKA. We believe that PhysBench and PhysAgent offer valuable insights and contribute to bridging the gap between VLMs and physical world understanding.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 34 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Apple-$\pi$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

    cs.CV 2026-07 conditional novelty 7.0

    A law-grounded benchmark, Apple-PI, grades video models stage-by-stage on physics reasoning and finds they top out at 0.473, well short of reliable simulation.

  2. RetroHolmes: When Semantic Plausibility Fails Retrospective Physical Process Reasoning

    cs.MM 2026-07 conditional novelty 7.0

    Current VLMs fail retrospective physical reachability and causal reconstruction on RetroHolmes; a simple analysis-by-synthesis loop with video simulation reduces bias and belief-conflict sensitivity.

  3. Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks

    cs.AI 2026-07 conditional novelty 7.0

    Frontier agentic AI models consistently fail at computational imaging tasks requiring physics-aware inversion, producing visually plausible but physically incorrect outputs.

  4. ChronoPhyBench: Do MLLMs Truly Understand the World or Merely Exploit Language Priors?

    cs.CV 2026-06 unverdicted novelty 7.0

    ChronoPhyBench is a new benchmark and dataset for chronological physical dynamics reasoning that combines video-conditioned next-state prediction with VQA to reduce language bias in MLLM evaluation.

  5. Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them

    cs.CV 2026-06 unverdicted novelty 7.0

    PhaseLock extracts motion priors from 2-step inference and enforces them via Latent Delta Guidance to raise physical consistency scores by 6.2 points on average in image-to-video diffusion models.

  6. Benchmarking Single-Factor Physical Video-to-Audio Generation

    cs.CV 2026-05 unverdicted novelty 7.0

    FlatSounds benchmark shows state-of-the-art V2A models rely more on text captions than visual input for physical and semantic accuracy, with captions improving correctness but degrading temporal alignment.

  7. World Models in Words: Auditing Physical State-Transition Commitments in Vision-Language Models

    cs.CL 2026-05 unverdicted novelty 7.0

    WMW audits VLMs by requiring typed physical state-transition traces and using a verifier to detect inconsistencies missed by answer-only evaluation, with TraceBank as a released resource of synthetic scenarios.

  8. ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop

    cs.CV 2026-05 unverdicted novelty 7.0

    ESI-Bench shows active exploration outperforms passive observation in multimodal LLMs on spatial tasks but reveals failures from poor action choices and overconfident belief commitment unlike humans.

  9. Tracing the Arrow of Time: Diagnosing Temporal Information Flow in Video-LLMs

    cs.CV 2026-05 unverdicted novelty 7.0

    Temporal information in Video-LLMs is encoded well by video-centric encoders but disrupted by standard projectors; time-preserved MLPs plus AoT supervision yield 98.1% accuracy on arrow-of-time and gains on other temp...

  10. Grounding Video Reasoning in Physical Signals

    cs.CV 2026-04 unverdicted novelty 7.0

    A new benchmark converts video clips into shared grounded event records and tests models across physics, semantic, and control prompts under original, shuffled, ablated, and masked conditions, finding selective robust...

  11. OptiVerse: A Comprehensive Benchmark towards Optimization Problem Solving

    cs.CL 2026-04 unverdicted novelty 7.0

    OptiVerse is a new benchmark spanning neglected optimization domains that shows LLMs suffer sharp accuracy drops on hard problems due to modeling and logic errors, with a Dual-View Auditor Agent proposed to improve pe...

  12. SCP: Spatial Causal Prediction in Video

    cs.CV 2026-03 unverdicted novelty 7.0

    SCP defines a new benchmark task for predicting spatial causal outcomes beyond direct observation and shows that 23 leading models lag far behind humans on it.

  13. V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions

    cs.CV 2025-12 conditional novelty 7.0

    V-REX shows that VLMs' multi-step visual exploration can be measured separately as planning (choosing sub-questions) and following (answering them), and that planning is the key bottleneck.

  14. Abstract 3D Perception for Spatial Intelligence in Vision-Language Models

    cs.CV 2025-11 unverdicted novelty 7.0

    SandboxVLM enhances VLMs' spatial intelligence by encoding 3D geometry with abstract bounding boxes in a four-stage zero-shot pipeline, yielding an 8.3% improvement on SAT Real benchmark.

  15. RigPI: Dynamic Parameter Identification of Rigid Body via VLM-Seeded Differentiable Simulation

    cs.RO 2026-06 unverdicted novelty 6.0

    RigPI combines VLM semantic priors with two-stage gradient optimization in differentiable simulation to identify inertial and frictional parameters of rigid bodies from robot-object interactions.

  16. Causal Scaffolding for Physical Reasoning: A Benchmark for Causally-Informed Physical World Understanding in VLMs

    cs.DB 2026-06 unverdicted novelty 6.0

    Introduces CausalPhys benchmark with causal graphs and CRFT fine-tuning to improve VLMs' causal physical reasoning accuracy and interpretability.

  17. $\Delta$ynamics: Language-Based Representation for Inferring Rigid-Body Dynamics From Videos

    cs.CV 2026-05 unverdicted novelty 6.0

    A vision-language framework generates text-based rigid-body scene configurations from videos using motion reasoning and optical flow, reporting 0.30 IoU on CLEVRER (7x over baselines) and transfer to 235 real videos.

  18. ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop

    cs.CV 2026-05 unverdicted novelty 6.0

    ESI-Bench is a new benchmark for embodied spatial intelligence with 10 task categories on OmniGibson that requires agents to actively explore via perception, locomotion, and manipulation, revealing that MLLMs suffer f...

  19. GeoWorld-VLM: Geometry from World Models for Vision-Language Models

    cs.CV 2026-05 unverdicted novelty 6.0

    GeoWorld-VLM aligns VLM image features with intermediate representations from camera-conditioned world models via fine-tuning only the encoder and projector, yielding ~4% gains on What'sUp and VSR spatial benchmarks a...

  20. Quantitative Video World Model Evaluation for Geometric-Consistency

    cs.CV 2026-05 unverdicted novelty 6.0

    PDI-Bench computes 3D projective residuals from segmented and tracked points to quantify geometric inconsistency in AI-generated videos.

  21. From Priors to Perception: Grounding Video-LLMs in Physical Reality

    cs.CV 2026-05 unverdicted novelty 6.0

    Video-LLMs fail physical reasoning due to semantic prior dominance rather than perception deficits; a new programmatic adversarial curriculum and visual-anchored reasoning chain enable substantial gains via standard L...

  22. Dual-Cluster Memory Agent: Resolving Multi-Paradigm Ambiguity in Optimization Problem Solving

    cs.CL 2026-04 unverdicted novelty 6.0

    DCM-Agent improves LLM performance on multi-paradigm optimization problems by 11-21% via dual-cluster memory construction and dynamic inference guidance.

  23. Multimodal Language Models Cannot Spot Spatial Inconsistencies

    cs.CV 2026-04 unverdicted novelty 6.0

    Multimodal LLMs significantly underperform humans at spotting objects that break 3D consistency in multi-view image pairs.

  24. CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates

    cs.CV 2025-12 conditional novelty 6.0

    VLMs perform near chance at detecting and correcting an erroneous step in visual sequential planning, and incrementally updating scene graphs step-by-step (SGI) yields small, partially inconsistent gains.

  25. SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios

    cs.CV 2025-11 conditional novelty 6.0

    SWITCH introduces a 193-video benchmark of tangible control-interface interactions and shows that frontier LMMMs struggle with fine-grained grounding and outcome verification.

  26. OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping

    cs.CV 2026-07 unverdicted novelty 5.0

    OmniView-Space framework with MPSM, tool-guided reasoning, and distillation achieves SOTA on spatial reasoning benchmarks for MLLMs while reducing external geometry dependencies.

  27. RigPI: Dynamic Parameter Identification of Rigid Body via VLM-Seeded Differentiable Simulation

    cs.RO 2026-06 unverdicted novelty 5.0

    RigPI combines VLM initialization with two-stage gradient-based optimization in differentiable simulation to estimate dynamic parameters of rigid bodies from real robot interactions.

  28. Physically Viable World Models: A Case for Query-Conditioned Embodied AI

    cs.AI 2026-05 unverdicted novelty 5.0

    Embodied AI requires query-conditioned world models that select the simplest physical abstraction sufficient to answer intervention queries.

  29. GeoWorld-VLM: Geometry from World Models for Vision-Language Models

    cs.CV 2026-05 unverdicted novelty 5.0

    GeoWorld-VLM distills geometric structure from camera-conditioned world models into VLMs by aligning visual features, improving spatial reasoning by about 4% on What'sUp and VSR benchmarks across two architectures whi...

  30. PhysBrain 1.0 Technical Report

    cs.RO 2026-05 unverdicted novelty 5.0

    PhysBrain 1.0 extracts scene elements, spatial dynamics, actions and depth relations from human egocentric video to create QA supervision for VLMs, then transfers the resulting physical priors to VLA policies via capa...

  31. Dual-Cluster Memory Agent: Resolving Multi-Paradigm Ambiguity in Optimization Problem Solving

    cs.CL 2026-04 unverdicted novelty 5.0

    DCM-Agent improves LLM optimization solving by 11–21% on seven benchmarks via dual-cluster memory of Approaches, Checklists, and Pitfalls plus adaptive path switching.

  32. Agentic Physical AI toward a Domain-Specific Foundation Model for Nuclear Reactor Control

    cs.AI 2025-12 unverdicted novelty 5.0

    A compact language model trained on scaled synthetic nuclear reactor control data exhibits variance collapse and emergent concentration on a single actuation strategy driven by physical execution success.

  33. MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models

    cs.CV 2025-11 unverdicted novelty 5.0

    MASS adds spatiotemporal motion signals and 3D grounding to VLMs and releases MASS-Bench, yielding physics-reasoning performance within 2% of Gemini-2.5-Flash after reinforcement fine-tuning.

  34. Evidence of a Cognitive Shift in AI Education: How Students Are Rethinking Human Intelligence?

    cs.CY 2026-04 unverdicted novelty 4.0

    Longitudinal poll data from 471 students in AI courses shows a shift toward preferring human intelligence, reaching 65% in technical courses and 90% in design courses by 2026.