Pith. sign in

REVIEW 3 major objections 6 minor 29 references

Towards a Large Physics Benchmark

T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read A physics benchmark can score LLM understanding and creativity using expert difficulty and surprise ratings.

desk verdict A thoughtful framework proposal for a living physics-LLM benchmark, but the Type 3 surprise score conflates performance with the paper's own definition of creativity, and the pilot is far too small to validate the claims. read the letter →

arxiv 2507.21695 v1 pith:MSOI4GL7 submitted 2025-07-29 physics.data-an cs.AIhep-phphysics.comp-phphysics.hist-ph

classification physics.data-ancs.AIhep-phphysics.comp-phphysics.hist-ph
keywords largelanguagemodelsphysicsbenchmarkscientificunderstandingcreativityexpertscoringmethodologyparticleclassification
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes a benchmark for evaluating large language models on fundamental physics that aims to measure not just whether an answer is correct but how much scientific understanding and creativity the model displayed. Its core idea is to have expert physicists rate each question on difficulty — taken as a proxy for the epistemic depth required to answer it — and rate the known correct answer on surprise, which together with correctness stands in for scientific creativity. Questions come in three formats: multiple choice, derivations with a unique mathematical answer, and open-ended coding challenges that maximize a scalar score, so the framework covers recall, formal reasoning, and authentic problem solving. The authors also propose a 'living benchmark' in which physicists submit questions that are peer-reviewed by at least three experts and recorded with version histories, with only a small public subset released at each iteration. Preliminary results show all four tested models answering the sample multiple-choice and derivation questions correctly and generating working classifiers for the four-top-quark signal that approach but do not yet match specialized physics models.

What carries the argument

Equations (1)–(3) of the paper carry the whole argument. They turn the philosophical concepts into arithmetic: difficulty ratings $d_i$ and surprise ratings $s_i$ (each on a 1–5 expert scale) are summed only over correctly answered questions, so correctness acts as the 'value' condition and filters out wrong answers; the normalization constants $\alpha$ and $\beta$ are inversely proportional to the number of questions of each type; and Type 3 challenges replace the binary correctness flag with a step-function discretization of the achieved scalar score $c$. The same equations also define the final aggregates $D_F$ and $S_F$, the single numbers the benchmark would use to rank models. Underneath the arithmetic sits the assumption that difficulty of the question tracks the epistemic depth needed to answer it and that surprise of the known correct answer tracks the creativity a model must have exercised to reproduce it.

What would settle it

Take a set of Type 2 derivation questions whose expert-rated surprise is high, give structurally identical variants that use different variables, phrasing, and order of steps, and record each model's cut-off date; if a model solves the original questions but fails the unseen variants, the aggregated surprise score $S_F$ is measuring memorized retrieval rather than recreated surprising reasoning.

Watch

Extended reading notes

Core claim

The paper's central claim is that expert-elicited scores for difficulty and surprise, combined with a binary correctness value, give a workable operational measure of an LLM's scientific understanding and creativity in physics. For Type 1 and Type 2 questions the understanding score is the difficulty-weighted average over correctly answered questions, $D_{1,2} = \alpha_{1,2}\sum_i c_i d_i$, and the creativity score is the corresponding surprise-weighted average $S_{1,2} = \beta_{1,2}\sum_i c_i s_i$, with $c_i=1$ when the answer is correct; for Type 3 coding challenges the continuous metric $c\in(0,1)$ is mapped through a step function onto 0–5 difficulty and surprise ratings before aggregation. The two aggregates $D_F$ and $S_F$ are then each an equal-weight average over the three question types. The working assumption is that consistently solving a question whose correct answer experts find surprising means the model has recreated at least part of that surprising reasoning, so its output is itself a creative product.

Load-bearing premise

The whole creativity score rests on assuming that a model that produces a surprising correct answer has actually recreated the surprising reasoning instead of recalling it, guessing, or pattern-matching; the paper notes this is doubtful for multiple-choice questions but does not test it.

Editorial extensions

If this is right

  • A model could be tracked over time by two numbers — an understanding score and a creativity score — instead of accuracy on a fixed question set, letting the community monitor progress across model versions.
  • The community-driven, peer-reviewed submission pipeline is designed to keep the question set growing and refreshed, which would make benchmark saturation and training-data contamination harder than with static datasets.
  • Type 3 coding challenges produce a single metric comparable across models and against specialized physics software, so the same framework can ask whether a general-purpose LLM is approaching domain-specific classifiers on real analysis tasks.
  • If the operationalization holds, the benchmark supplies a concrete test of the philosophy-of-science definitions it builds on: 'understanding' and 'creativity' become things you can measure in a standardized way.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An immediate test of the paper's central proxy would be to check whether models that solve expert-rated-surprising questions can also solve structurally rephrased versions with the same solution method; if transfer fails, the surprise credit is largely memorization rather than recreated reasoning.
  • The equal weighting of the three question types is provisional, as the paper states explicitly, so the aggregates $D_F$ and $S_F$ will need versioned weights if later evidence shows, say, Type 3 performance is more indicative of research capability; otherwise cross-version comparisons would be compromised.
  • The same difficulty-and-surprise machinery could generalize to other sciences, or to human test-takers, but only if expert ratings are calibrated across fields; nothing in the paper establishes that.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This manuscript proposes a framework for a living, community-built benchmark to evaluate large language models on scientific understanding and creativity in fundamental physics. It defines three question types (multiple-choice, analytically unique, and open-ended coding challenges) and scores each by expert ratings of difficulty and surprise, with correctness serving as the value weight. Equations (1)-(3) aggregate these scores into final difficulty and surprise scores D_F and S_F. The paper presents two multiple-choice, two analytical, and one coding example (the FOURTOPS event classification task), along with preliminary results from several LLMs. The contribution is framed as a proposal and design study rather than a fully realized benchmark dataset.

Significance. If the proposed scoring scheme were valid, it would fill a real gap: most existing physics benchmarks measure answer correctness, not conceptual depth or original reasoning. The paper has genuine strengths: a transparent versioned-release plan, a triple-expert review pipeline, a detailed and reproducible sandboxed code harness for the FOURTOPS challenge, and an explicit attempt to ground the metrics in philosophical definitions of understanding and creativity. The treatment of I-surprise relative to a model's training-data cutoff is also thoughtful. However, the paper currently does not deliver a working benchmark: the dataset consists of only five examples, no aggregate D_F or S_F is computed for any model, and the central mappings from difficulty to understanding and from surprise to creativity are assumed rather than validated. Most importantly, the Type 3 surprise score in Eq. (2) measures the scalar performance metric rather than departure from pre-established generative principles, which contradicts the manuscript's own definition of surprise in Section 3.2. The headline creativity score S_F is therefore not measuring the intended construct as it stands.

major comments (3)
  1. [§3.4, Eq. (2)] The Type 3 surprise score s_i is defined purely as a step function of the continuous scalar c (e.g., AUC). The manuscript's own definition in §3.2 is that surprise measures how poorly a product can be explained by pre-established generative principles, and it states explicitly that following a known step plan is not creative. With Eq. (2), a model that reaches AUC 0.85 by applying a standard gradient-boosted tree and a model that reaches the same AUC with a genuinely new architecture receive the same s_i; conversely, a novel approach that scores slightly lower receives a lower s_i. The score is therefore a performance threshold, not a surprise rating. Since S_F in Eq. (3) averages S1, S2, and S3, the final creativity score is invalid as currently formulated. The limitation section (§7) does not mention this mismatch. I recommend either having experts rate the generated code's surprise directly (for example, its deviation from known methods in the prompt), or removing Type 3 from S_F and reporting the three type scores separately.
  2. [§3.4, Eq. (1)] The validity of S1 and S2 rests on the assumption that a correct answer to a question whose correct answer is expert-rated as surprising implies that the model "internally recreated part of the surprising reasoning." The authors acknowledge this is weaker for Type 1 because guessing is possible, but no evidence is provided for either type. A model could produce the correct answer through memorization, paraphrase of textbook material, or shallow pattern matching, especially for the standard examples in §5. This is load-bearing because S_F is presented as a measure of scientific creativity. The paper should include contamination checks (e.g., paraphrased variants, questions published after the model's cutoff, held-out question sets) and should compare the proxy scores against human creativity judgments for the same answers.
  3. [§3.4 and §4] The operationalization of understanding as expert-rated difficulty and creativity as expert-rated surprise is not empirically supported. The manuscript asserts that with enough evaluators, subjective differences in rating will average out, but it reports no inter-annotator agreement, no pilot study, and no evidence that expert difficulty correlates with epistemic depth as opposed to computational length, obscurity, or arithmetic load. Since D_F and S_F are defined directly from these ratings through Eqs. (1)-(3), the entire benchmark's construct validity depends on this step. A minimal validation would be a reliability study (e.g., intraclass correlation or Cohen's kappa across the proposed triple-review protocol) and a comparison of difficulty ratings with independent measures such as human success rates or time-to-solution.
minor comments (6)
  1. [§6] The paper never computes D_F or S_F for any model, even though Eqs. (1)-(3) define them; adding a worked example with one model would make the aggregation concrete and testable.
  2. [§3.4, Eq. (2)] The normalization weights α3 and β3 in Eq. (2) are never defined; the text only defines α1,2 and β1,2 in Eq. (1). Please specify whether they are also inversely proportional to K or chosen in some other way.
  3. [§5.1] The accepted answer to Type 1 Q2, which cites "mathematical problems (e.g. divergences)" as the motivation for the Higgs mechanism, is not the standard argument; the standard motivation is the violation of gauge invariance and unitarity when mass terms are added by hand. As an expert-validated question, this should be corrected or rephrased.
  4. [§5.3.1] Table 4 reports single AUC values with no uncertainty or number of runs; since LLM code generation is stochastic, repeated runs would be needed to support any comparison between models.
  5. [§5.2] The statement that all four models answered both Type 1 and Type 2 questions correctly is best labeled as a sanity check rather than a benchmark result, since the examples are standard textbook problems and the sample is very small.
  6. [Table 1] There are typographical issues in Table 1 (e.g., "Ours (This W ork)" and a broken "SciF act" entry); please proofread the table and the surrounding text.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the scoring pipeline is an explicit operationalization whose proxy assumptions are acknowledged, and no fitted parameter is relabeled as a prediction.

full rationale

The paper does not claim to derive its understanding/creativity scores from something else; it explicitly constructs them. Equation (1) defines D and S for Types 1/2 as expert-rated difficulty and surprise weighted by correctness, Equation (2) defines Type 3 difficulty/surprise as step functions of the scalar AUC metric, and Equation (3) is a stated aggregation rule. These equations are presented as definitions of the benchmark scores, not as empirical predictions of understanding or creativity from hidden inputs. The text repeatedly labels difficulty and surprise as 'proxies' and as an 'operational' measure, and it explicitly acknowledges the key assumption that solving a surprising question implies recreating surprising reasoning ('This assumption is more reasonable for type 2 questions than for type 1 questions'). The type 3 mapping of surprise to AUC is a stipulated measurement choice that may create a construct-validity gap—performance is not the same as generative novelty—but it is not circular because the paper does not hide the mapping or claim to measure the code's surprise directly; it states that 'Evaluating the surprise of each generated answer directly is unfeasible.' Self-citations [2] and [3] are contextual (question-type grounding and the large-physics-models proposal) rather than load-bearing for the equations, and the philosophical definitions of surprise and creativity are drawn from external sources (Boden, de Regt, Gaut). No uniqueness theorem is imported, no earlier result is invoked to forbid alternatives, and no fitted parameter is later renamed as a prediction. The acknowledged limitations—duplicate-problem detection and difficulty of human-model comparison—affect benchmark quality and validity, not the circularity of the derivation.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central claim does not rest on fitted parameters or invented physical entities. It rests on assumptions about the validity of expert ratings as proxies for understanding and creativity, and on unspecified threshold parameters in the scoring equations. The ledger reflects those assumptions.

free parameters (3)
  • Type 3 difficulty and surprise threshold breakpoints t_d, t_s = not specified
    Equation 2 maps continuous performance c to discrete difficulty and surprise scores using expert-chosen intervals; the intervals are free parameters of the framework and no values are given.
  • Aggregation weights alpha1,2, beta1,2 and equal type weights = equal weights, one third per question type
    Equation 3 assumes equal weights across question types, a modeling choice not derived from evidence.
  • Expert scores di and si = integer ratings from 1 to 5
    The core metrics are subjective expert ratings; their calibration is a free input to the benchmark.
assumptions (4)
  • domain assumption Difficulty ratings by experts are a valid proxy for the degree of scientific understanding required.
    Section 3.1 asserts that a higher difficulty score indicates more understanding is required, without independent validation.
  • domain assumption An LLM that correctly solves a question whose answer is rated surprising has internally recreated part of the surprising reasoning.
    Section 3.4 states this explicitly as an assumption; it is the bridge from answer correctness to creativity.
  • domain assumption I-novelty relative to an LLM's training data is a meaningful and checkable notion of individual novelty.
    Section 3.2 defines I-novelty this way, but determining what is in an LLM's training data is not operationalized.
  • domain assumption Correctness can be reliably verified for open-ended Type 2 answers, for example via Mathematica or numerical routines.
    Section 6 says verification software will be used, but no procedure, tolerance, or accepted solution format is specified.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Towards a Large Physics Benchmark." pith.science (2026). https://pith.science/paper/MSOI4GL7

@misc{pith2026250721695,
  author       = {Pith},
  title        = {Pith review of: Towards a Large Physics Benchmark},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MSOI4GL7}},
  note         = {Machine review of arXiv:2507.21695}
}
read the original abstract

We introduce a benchmark framework developed by and for the scientific community to evaluate, monitor and steer large language model development in fundamental physics. Building on philosophical concepts of scientific understanding and creativity, we develop a scoring system in which each question is scored by an expert for its correctness, difficulty, and surprise. The questions are of three forms: (i) multiple-choice questions for conceptual understanding, (ii) analytical problems requiring mathematical derivation, and (iii) openended tasks requiring complex problem solving. Our current dataset contains diverse set of examples, including a machine learning challenge to classify high-energy physics events, such as the four top quark signal. To ensure continued relevance, we propose a living benchmark, where physicists contribute questions, for instance alongside new publications. We invite contributions via: http://www.physicsbenchmarks.org/. We hope that this benchmark will enable a targeted AI development that can make a meaningful contribution to fundamental physics research.

Figures

Figures reproduced from arXiv: 2507.21695 by the authors.

Figure 1
Figure 1. Framework for Physics Creativity and Reasoning Benchmark [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Framework for Physics Creativity and Reasoning Benchmark: (a) Question Generation [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Training and validation accuracy versus epoch for six evaluated models of the example [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: ROC Curves and AUC scores for the example question of the FOURTOPS challenge. [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

29 extracted references · 21 canonical work pages

  1. [1]

    The Leaderboard Illusion

    Shivalika Singh et al. “The Leaderboard Illusion”. In: (2025). arXiv: 2504.20879 [cs.LG]

  2. [2]

    Towards a Benchmark for Scientific Understanding in Humans and Machines

    Kristian G. Barman et al. “Towards a Benchmark for Scientific Understanding in Humans and Machines”. In: Minds and Machines 34.1 (2024), p. 6. doi: 10.1007/s11023- 024- 09657-1. arXiv: 2304.10327 [cs.AI]

  3. [3]

    Large Physics Models: Towards a collaborative approach with Large Language Models and Foundation Models

    Kristian G. Barman et al. “Large Physics Models: Towards a collaborative approach with Large Language Models and Foundation Models”. In: (Jan. 2025). arXiv: 2501 . 05382 [physics.data-an]

  4. [4]

    SciQAG: A Framework for Auto-Generated Science Question Answering Dataset with Fine-grained Evaluation

    Yuwei Wan et al. “SciQAG: A Framework for Auto-Generated Science Question Answering Dataset with Fine-grained Evaluation”. In: (2024). arXiv: 2405.09939 [cs.CL]

  5. [5]

    GPQA: A Graduate-Level Google-Proof Q&A Benchmark

    David Rein et al. “GPQA: A Graduate-Level Google-Proof Q&A Benchmark”. In: Proceed- ings of the First Conference on Language Modeling . 2024. arXiv: 2311.12022 [cs.AI]

  6. [6]

    Scieval: A multi-level large language model evaluation benchmark for scientific research

    Liangtai Sun et al. “Scieval: A multi-level large language model evaluation benchmark for scientific research”. In: Proceedings of the AAAI Conference on Artificial Intelligence . Vol. 38. 17. 2024, pp. 19053–19061

  7. [7]

    Fact or Fiction: Verifying Scientific Claims

    David Wadden et al. “Fact or Fiction: Verifying Scientific Claims”. In: (2020). arXiv: 2004. 14974 [cs.CL]

  8. [8]

    BRIGHT: A Realistic and Challenging Benchmark for Reasoning- Intensive Retrieval

    Hongjin Su et al. “BRIGHT: A Realistic and Challenging Benchmark for Reasoning- Intensive Retrieval”. In: (2024). arXiv: 2407.12883 [cs.IR]

Show all 29 references
  1. [9]

    Evaluating and Enhancing Large Language Models for Novelty Assessment in Scholarly Publications

    Ethan Lin, Zhiyuan Peng, and Yi Fang. “Evaluating and Enhancing Large Language Models for Novelty Assessment in Scholarly Publications”. In: (2024). arXiv:2409.16605 [cs.CL]

  2. [10]

    Humanity’s Last Exam

    Long Phan et al. “Humanity’s Last Exam”. In: (2025). arXiv: 2501.14249 [cs.AI]

  3. [11]

    Theoretical Physics Benchmark (TPBench)—A Dataset and Study of AI Reasoning Capabilities in Theoretical Physics

    Daniel J. H. Chung et al. “Theoretical Physics Benchmark (TPBench)—A Dataset and Study of AI Reasoning Capabilities in Theoretical Physics”. In: (2025). arXiv: 2502.15815 [physics.comp-ph]

  4. [12]

    Henk W. de Regt. Understanding Scientific Understanding . Oxford University Press, 2017

  5. [13]

    Is Understanding a Species of Knowledge?

    Stephen R. Grimm. “Is Understanding a Species of Knowledge?” In: The British Journal for the Philosophy of Science 57.3 (2006), pp. 515–535

  6. [14]

    Understanding

    Stephen R. Grimm. “Understanding”. In: The Routledge Companion to Epistemology . Ed. by Sven Bernecker and Duncan Pritchard. Routledge, 2010, pp. 84–94. 13

  7. [15]

    External Representations and Scientific Under- standing

    Jaakko Kuorikoski and Petri Ylikoski. “External Representations and Scientific Under- standing”. In: Synthese 192 (2015), pp. 3817–3837

  8. [16]

    Attributing creativity

    Dustin Stokes Elliot Samuel Paul. “Attributing creativity”. In: Creativity and philosophy . Ed. by Gaut and Kieran. Feb. 2018, pp. 193–209. isbn: 9781351199797. doi: 10.4324/ 9781351199797-6

  9. [17]

    The creative mind: Myths and mechanisms

    Margaret A Boden. The creative mind: Myths and mechanisms . Routledge, 2004

  10. [19]

    The philosophy of creativity

    Berys Gaut. “The philosophy of creativity”. In: Philosophy Compass 5.12 (2010), pp. 1034– 1046

  11. [20]

    The value of creativity

    Berys Gaut. “The value of creativity”. In: Creativity and philosophy . Ed. by Gaut and Kieran. Feb. 2018, pp. 124–140. isbn: 9781351199797. doi: 10.4324/9781351199797-6

  12. [21]

    Explanations of Creativity

    David Novitz. “Explanations of Creativity”. In: The Creation of Art: New Essays in Philo- sophical Aesthetics. Ed. by Berys Gaut and Paisley Livingston. Cambridge: Cambridge UP, 2003, pp. 174–91

  13. [22]

    Creativity and Constraint

    David Novitz. “Creativity and Constraint”. In: Australasian Journal of Philosophy 77.1 (1999), pp. 67–82. doi: 10.1080/00048409912348811

  14. [23]

    Values in Science

    Ernan McMullin. “Values in Science”. In: PSA: Proceedings of the Biennial Meeting of the Philosophy of Science Association 1982 (1982), pp. 3–28. issn: 0270-8647. JSTOR: 192409. (Visited on 12/22/2024)

  15. [24]

    Attention to the strengths of physical interactions: Transformer and graph-based event classification for particle physics experiments

    Luc Builtjes et al. “Attention to the strengths of physical interactions: Transformer and graph-based event classification for particle physics experiments”. In: (Nov. 2022). arXiv: 2211.05143 [hep-ph]. 8 Acknowledgments The work of Caron, De Regt, Barman was supported by an I...

  16. [26]

    Do not add any formatting, such as markdown, to the response

  17. [27]

    >” comment, in the code template, with the required code

    Replace each ”# < LLM: ... >” comment, in the code template, with the required code. No placeholder should remain

  18. [28]

    Before finalizing your answer, double-check that your code runs without errors and meets all requirements (all functions implemented, correct tensor shapes, etc.)

  19. [29]

    To prevent dimensional mismatches make sure to annotate tensor sizes as comments

  20. [30]

    verbose =

    IMPORTANT: Remember, your first, and most important priority is to produce (syntacti- 19 cally) correct code. Prioritise what you can implement reliably above all else. Then prioritise maximising the metric. 9.2.2 Example Response: Google Gemini (23/06) AUC: 0.847 1 # 0. - - -...

  21. [512]

    EPOCHS

    , 207 c o l l a t e _ f n = collate , 208 l o a d e r _ c l s = l o a d e r _ c l s ) 209 210 # 2. Build model 211 f i r s t _ b a t c h = next ( iter ( t r a i n _ l o a d e r ) ) 212 e x a m p l e _ s a m p l e = f i r s t _ b a t c h [0] 213 model = m a k e _ m o d e l ( e ...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.