Pith. sign in

REVIEW 7 cited by

PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.12054 v2 pith:ESCC4IKR submitted 2025-02-17 cs.AI

classification cs.AI
keywords physicsreasoningphysics-basedphysreasonbenchmarkcomprehensivehardmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models demonstrate remarkable capabilities across various domains, especially mathematics and logic reasoning. However, current evaluations overlook physics-based reasoning - a complex task requiring physics theorems and constraints. We present PhysReason, a 1,200-problem benchmark comprising knowledge-based (25%) and reasoning-based (75%) problems, where the latter are divided into three difficulty levels (easy, medium, hard). Notably, problems require an average of 8.1 solution steps, with hard requiring 15.6, reflecting the complexity of physics-based reasoning. We propose the Physics Solution Auto Scoring Framework, incorporating efficient answer-level and comprehensive step-level evaluations. Top-performing models like Deepseek-R1, Gemini-2.0-Flash-Thinking, and o3-mini-high achieve less than 60% on answer-level evaluation, with performance dropping from knowledge questions (75.11%) to hard problems (31.95%). Through step-level evaluation, we identified four key bottlenecks: Physics Theorem Application, Physics Process Understanding, Calculation, and Physics Condition Analysis. These findings position PhysReason as a novel and comprehensive benchmark for evaluating physics-based reasoning capabilities in large language models. Our code and data will be published at https:/dxzxy12138.github.io/PhysReason.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Agentic Exploration of Physics Models

    cs.AI 2025-09 conditional novelty 7.0 of 10

    A general-purpose LLM agent can discover physics models, including ODEs and spin Hamiltonians, by autonomously choosing experiments and fitting hypotheses to numeric data.

  2. SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopy

    cs.AI 2026-02 reject novelty 6.0 of 10

    A new benchmark of 2,703 automatically generated multimodal questions for scanning probe microscopy, plus a modified F1 metric that penalizes over-selection and labels model 'personalities'.

  3. Mitigating Easy Option Bias in Multiple-Choice Question Answering

    cs.CV 2025-08 conditional novelty 6.0 of 10

    In six VQA benchmarks, models can often choose the correct option from image plus options alone, and the GroundAttack toolkit generates visually plausible hard negatives to remove this shortcut.

  4. PhySense: Principle-Based Physics Reasoning Benchmarking for Large Language Models

    cs.LG 2025-05 reject novelty 6.0 of 10

    A new benchmark of 380 principle-based physics problems shows that state-of-the-art LLMs struggle to apply symmetry, conservation, and dimensional-analysis shortcuts, achieving under 50 percent average accuracy with h...

  5. Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new multi-image benchmark and graph-based scoring method show that MLLMs produce weak, poorly-grounded chains of thought on visual physics tasks, and that RL post-training can degrade spatial reasoning.

  6. PhyX: Does Your Model Have the "Wits" for Physical Reasoning?

    cs.AI 2025-05 conditional novelty 6.0 of 10

    PhyX is a new 3,000-question visual physics benchmark; the best AI model tested scores 45.8 percent, well below the 75.6 to 78.9 percent of a small human student sample.

  7. Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges

    cs.AI 2025-07 reject novelty 1.0 of 10

    A broad survey of LLM alignment that catalogs objectives, benchmarks, SFT/RLHF/DPO methods, and safety challenges, without contributing new experimental or theoretical results.

Pith tools