Pith. sign in

REVIEW 4 major objections 5 minor 34 references

Superstudent intelligence in thermodynamics

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read The reasoning model o3, given a 90-minute thermodynamics exam zero-shot, solved all problems correctly and scored above all 90 students, at the level of the best scores seen in more than 10,000 similar exams since 1985.

desk verdict Real exam, real result, but the diagram handling is murky and the abstract overstates the case; plausible and worth reviewing. read the letter →

arxiv 2506.09822 v1 pith:EHG2ZOGX submitted 2025-06-11 cs.CE cs.AI

classification cs.CEcs.AI
keywords largelanguagemodelso3reasoningmodelthermodynamicsexamzero-shotpromptingengineeringeducationassessmentmachineintelligence
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This short report claims that the reasoning model o3, used zero-shot (off the shelf, with no exam-specific training), outscored all 90 students who took a real university exam in engineering thermodynamics, with a score in the range of the best results the instructor has seen across more than 10,000 comparable exams since 1985. The exam is deliberately built from new, carefully crafted problems that are meant to defeat pattern learning, and the class results were typical: a 58 percent failure rate and exactly one A grade. The authors therefore read o3's performance as evidence that machines now excel at a task long treated as proof of exceptional human intellectual ability, and they discuss what that means for engineering work, for the education and examination of future engineers, and for what it means to understand thermodynamics.

What carries the argument

The measuring instrument is the exam itself, and the argument stands on its design premise: every problem is new and carefully crafted so that pattern matching from training data cannot solve it, and the grading rewards the analysis — writing and combining the relevant equations of the first and second law — far more than the numerical result, so a top grade is attainable without any numerical calculation. Against that instrument the authors put a strict comparison protocol: three zero-shot runs of o3, each prompted only with a brief instruction to act as an expert and solve the exam, with the three problems submitted one at a time, and each response assessed exactly as the students' papers were, against the published reference solution. The historical record of more than 10,000 similar exams graded by the same instructor since 1985 supplies the human score distribution that makes the model's single score interpretable.

What would settle it

Have the three o3 transcripts blind-graded by independent examiners with the graphical tasks strictly enforced — diagrams actually drawn and scored exactly as they are for students — and compare the result with the class distribution; if the score drops below the top student's, the reported margin is a format artifact rather than equivalent problem-solving. A second check would run the identical zero-shot protocol on a different exam version from the same course and test whether o3 again lands in the historical best-score range.

Watch

Extended reading notes

Core claim

The central claim is that o3, prompted as a thermodynamics expert with no special training, solved a real 90-minute Thermodynamics I exam better than every one of the 90 students who took it. The exam contained three problems — two gas chambers separated by a movable piston, a combined heat-and-power plant with compressor, turbine and heat pump, and a Rankine cycle whose steam table had to be completed — including graphical tasks in which students had to draw $p,V$ and $T,s$ diagrams. Three independent runs of o3 were graded with the same rubric and the same reference solution applied to the students, and the runs differed hardly at all. o3 effectively solved the first two problems, losing only a few points for graphical presentation in problem 2, and made one major error in problem 3 on the caloric properties of a constant-density fluid; its overall score placed it above the entire class and in the range of the best scores recorded since 1985. The authors call this a turning point: a grade A in this exam had been treated as proof of genuine understanding, and a machine has now reached it.

Load-bearing premise

The comparison assumes that the exam's diagrams reached o3 in a form equivalent to what the students saw and that its graphical answers were graded under the same rubric as the students' handwritten sheets, since the paper never states how the $p,V$ and $T,s$ tasks were presented to the model or how its graphical output was assessed.

Editorial extensions

If this is right

  • A commercial reasoning model, used off the shelf with no fine-tuning, can outperform an entire class on a demanding conceptual exam designed specifically to defeat pattern recognition.
  • The exam-design premise that carefully crafted new problems separate genuine understanding from memorized patterns no longer separates humans from machines.
  • Engineering education will shift toward prompting, questioning, and judging the trustworthiness of machine results, with fundamental understanding becoming more important than routine calculation.
  • Engineering practice will move toward mixed human-AI teams, raising open questions about task assignment, verification of results, responsibility, and legal accountability.
  • University examinations that certify independent problem-solving will need to be rethought, since students will know a machine can outperform them on the current format.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension the paper does not attempt: run the same zero-shot protocol on several earlier exam versions whose score distributions are already on record, which would show whether o3's top-of-distribution score is stable or depends on the particular exam's content and on how its diagrams are encoded.
  • The authors leave open whether o3 reached its score by applying general thermodynamic principles or by drawing on the vast number of solved examples in its training data; that question is testable with problems deliberately unlike any textbook example.
  • An implication the authors state only implicitly: if this exam can no longer certify human understanding, the natural response in assessment is to extend the rubric to what the model cannot yet do — such as defending a solution under questioning, choosing the analysis approach, or judging whether an answer is physically plausible in an unfamiliar situation.
  • The authors themselves caution that extrapolating from exam problems to real engineering tasks is bold, since LLMs are better suited to the former than the latter; that caveat limits the paper's practical consequences for engineering practice more than for education and assessment.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The manuscript reports that OpenAI's o3 model, prompted zero-shot in German, solved a 90-minute first-course engineering thermodynamics exam taken by 90 students, and that its graded score exceeded all student scores, falling in the range of the best scores the authors report having seen in more than 10,000 similar exams since 1985. The exam includes graphical input (a process sketch) and graphical output requirements (p,V and T,s diagrams). The authors report three o3 runs with similar results, discuss the one major mistake in Problem 3 and minor graphical issues, and then use the outcome to argue that machines exhibiting such exam performance should be called intelligent and that engineering education and practice must change accordingly.

Significance. If the central measurement is sound, this is a striking result: a commercial reasoning LLM outperforming every student in a genuine, previously unseen university exam with substantial graphical components would indeed be a noteworthy benchmark for machine competence in engineering thermodynamics. The study has clear strengths: it uses an authentic external exam rather than LLM-tailored problems, reports three independent zero-shot runs, provides the exam and model solution in the Supplementary Information, and builds on the authors' earlier published LLM baseline on simpler problems. The comparison against a real cohort of 90 students with a high failure rate gives the claim concrete educational relevance. However, the manuscript as written does not make the measurement fully auditable, and the abstract's wording exceeds what Section 2 reports.

major comments (4)
  1. [Abstract and Section 2] The Abstract states that o3 'solved all problems correctly' and 'outperformed all students,' but Section 2 explicitly reports that 'the only major mistake o3 made was in Problem 3' and that points were lost for minor graphical issues. These statements cannot both be true under the same grading standard. The authors should either correct the Abstract to say that o3 obtained the highest overall score despite one major mistake and minor graphical losses, or provide a point-by-point reconciliation of the claimed 'all problems correctly' with the reported losses.
  2. [Section 2; Problems 1c, 2a, 2c and BHKW sketch] The central comparison rests on the claim that o3's answers were assessed 'exactly the same way' as the students' handwritten exams, but the manuscript never specifies how the exam's graphical content was presented to o3 or how o3's graphical outputs were produced and graded. The exam requires a p,V diagram (Problem 1c, 4 points), a T,s diagram with state points (Problem 2a, 3 points), marking heat in the T,s diagram (Problem 2c, 4.5 points), and the process refers to a sketch ('siehe Skizze') that is presumably part of the problem statement. Without a description of whether the sketch was transmitted as an image, whether o3 responded with image or text/ASCII diagrams, and which rubric was applied to those outputs, the equivalence of the assessment conditions is not established and the outperformance claim is not fully supported.
  3. [Supplementary Information and Section 2] The manuscript does not include o3's answer sheets or a problem-by-problem point breakdown of the three runs, even though it does include the exam and a model solution. Without the raw o3 answers or a grading transcript showing how each part of each problem was scored, a reader cannot audit the claim that the same rubric was applied or verify where points were lost. The authors should add the o3 outputs (or a detailed transcript) and the per-part scoring for all three runs.
  4. [Section 3 and Figure 2] The claim that the overall score was 'in the range of the best scores we have seen in more than 10,000 similar exams since 1985' is presented without supporting data. The manuscript gives no distribution of historical best scores, percentile values, or a definition of 'similar exams,' so this comparative assertion cannot be checked. If the historical comparison is to remain in the Abstract and Discussion, the authors need to provide the relevant historical statistics or weaken the claim to what the provided student cohort shows.
minor comments (5)
  1. [Table 1] The column header 'GTP-4o' should read 'GPT-4o'.
  2. [Author affiliations] Affiliation 4 contains a typo: 'Kaierslautern' should be 'Kaiserslautern'.
  3. [Figure captions] The captions of Figures 1 and 2 contain garbled '/uni00000013/...' strings that appear to be PDF text-extraction artifacts; the figures should be regenerated so the captions are clean and readable.
  4. [Section 2] The claim that the language of the exam 'hardly' influences LLM performance is stated without data in this manuscript; a citation to the previous study where this was tested would make the claim auditable.
  5. [Section 2] The phrase 'the prompt was chosen based on some preliminary tests' should be accompanied by the tested alternatives or at least a description of what was varied, so that the zero-shot claim is not undermined by undisclosed prompt selection.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the central claim is an external empirical measurement of o3 against a fixed exam and a fixed student cohort, not a derivation from fitted inputs.

full rationale

The paper's central claim is empirical and self-contained: a fixed thermodynamics exam was given to 90 students and to OpenAI's o3 in zero-shot mode, and the answers were assessed by the same rubric ("we can directly compare the performance of the o3 model with that of our students"). There is no fitted parameter, no equation derived from an input, and no prediction that reduces by construction to a training target or to the exam itself. The self-citations, principally Loubet et al. (2025) for earlier LLM test scores and Vollmer et al. (2024) for a knowledge-based thermodynamics system, are contextual and do not carry the load-bearing comparison; the o3 result is reported directly from the exam runs, with three independent scores shown in Figures 1 and 2. The internal inconsistency between the abstract's "solved all problems correctly" and Section 2's "only major mistake o3 made was in Problem 3" is a reporting inconsistency, not circular reasoning. The unstated details of how graphical exam content was supplied to o3 and how its graphical answers were graded are validity or evidence concerns about comparability, not a case of the derivation reducing to its own inputs. Accordingly, no circular step is identified and the score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper introduces no free parameters and no new physical entities. The central claim depends on evaluation assumptions: identical scoring, no training contamination, and the validity of the exam as a measure of understanding. These are domain assumptions about the measurement, not fitted parameters.

assumptions (4)
  • domain assumption The exam grading rubric is applied identically to o3 and student answers, and the manual grading of o3's outputs is unbiased.
    The entire comparison rests on this premise; the Abstract says answers were assessed exactly the same way but no raw o3 outputs or inter-rater checks are provided.
  • domain assumption The exam problems cannot be solved by pattern learning, so a high score implies understanding of thermodynamic principles.
    The paper asserts this in Section 1 (pattern learning does not help in the exam) and Section 3 (cannot be solved by relying on pattern learning), and uses it to conclude that the machine understands thermodynamics.
  • domain assumption o3 did not have access to the specific exam content during training.
    The exam is from 07.04.2025 and o3 was released a few days earlier (Section 2), making contamination unlikely, but the paper does not address it explicitly.
  • domain assumption The German language of the exam and the English system prompt do not materially affect o3's performance.
    The paper states this is known from earlier work (Section 2) but does not validate it for this exam.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Superstudent intelligence in thermodynamics." pith.science (2026). https://pith.science/paper/EHG2ZOGX

@misc{pith2026250609822,
  author       = {Pith},
  title        = {Pith review of: Superstudent intelligence in thermodynamics},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EHG2ZOGX}},
  note         = {Machine review of arXiv:2506.09822}
}
read the original abstract

In this short note, we report and analyze a striking event: OpenAI's large language model o3 has outwitted all students in a university exam on thermodynamics. The thermodynamics exam is a difficult hurdle for most students, where they must show that they have mastered the fundamentals of this important topic. Consequently, the failure rates are very high, A-grades are rare - and they are considered proof of the students' exceptional intellectual abilities. This is because pattern learning does not help in the exam. The problems can only be solved by knowledgeably and creatively combining principles of thermodynamics. We have given our latest thermodynamics exam not only to the students but also to OpenAI's most powerful reasoning model, o3, and have assessed the answers of o3 exactly the same way as those of the students. In zero-shot mode, the model o3 solved all problems correctly, better than all students who took the exam; its overall score was in the range of the best scores we have seen in more than 10,000 similar exams since 1985. This is a turning point: machines now excel in complex tasks, usually taken as proof of human intellectual capabilities. We discuss the consequences this has for the work of engineers and the education of future engineers.

Figures

Figures reproduced from arXiv: 2506.09822 by the authors.

Figure 1
Figure 1. Distribution of points earned by o3 (three runs) and [PITH_FULL_IMAGE:figures/full_fig_p008_1.png] view at source ↗
Figure 2
Figure 2. Distribution of points earned by o3 (three runs) and [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

34 extracted references · 22 canonical work pages

  1. [1]

    Using Large Language Models for Solving Thermodynamic Problems

    Loubet, R.; Zittlau, P.; Vollmer, L.; Hoffmann, M.; Fellenz, S.; Jirasek, F.; Leitte, H.; Hasse, H. Using Large Language Models for Solving Thermodynamic Problems. 2025; http://arxiv.org/pdf/2502.05195v1

  2. [2]

    2025; https://platform.openai.com/docs/models/o3

    OpenAI OpenAI Platform. 2025; https://platform.openai.com/docs/models/o3

  3. [3]

    Evaluating GPT- and Reasoning-based Large Language Models on Physics Olympiad Problems: Surpassing Human Performance and Implications for Educational Assessment

    Tschisgale, P.; Maus, H.; Kieser, F.; Kroehs, B.; Petersen, S.; Wulff, P. Evaluating GPT- and Reasoning-based Large Language Models on Physics Olympiad Problems: Surpassing Human Performance and Implications for Educational Assessment. 2025; https://arxiv.org/abs/2505.09438

  4. [4]

    UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models

    Xu, X.; Xu, Q.; Xiao, T.; Chen, T.; Yan, Y.; Zhang, J.; Diao, S.; Yang, C.; Wang, Y. UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models. 2025; https://arxiv.org/abs/2502.00334

  5. [5]

    Combining Machine Learning with Physical Knowledge in Thermodynamic Modeling of Fluid Mixtures

    Jirasek, F.; Hasse, H. Combining Machine Learning with Physical Knowledge in Thermodynamic Modeling of Fluid Mixtures. Annual review of chemical and biomolecular engineering 2023, 14, 31--51

  6. [6]

    HANNA: hard-constraint neural network for consistent activity coefficient prediction

    Specht, T.; Nagda, M.; Fellenz, S.; Mandt, S.; Hasse, H.; Jirasek, F. HANNA: hard-constraint neural network for consistent activity coefficient prediction. Chemical science 2024, 15, 19777--19786

  7. [7]

    When physics meets machine learning: a survey of physics-informed machine learning

    Meng, C.; Griesemer, S.; Cao, D.; Seo, S.; Liu, Y. When physics meets machine learning: a survey of physics-informed machine learning. Machine Learning for Computational Science and Engineering 2025, 1

  8. [8]

    Wu, Y.; Sicard, B.; Gadsden, S. A. Physics-informed machine learning: A comprehensive review on applications in anomaly detection and condition monitoring. Expert Systems with Applications 2024, 255, 124678

Show all 34 references
  1. [9]

    Integrating knowledge-driven and data-driven approaches to modeling

    Todorovski, L.; D z eroski, S. Integrating knowledge-driven and data-driven approaches to modeling. Ecological Modelling 2006, 194, 3--13

  2. [10]

    Knowledge-Driven versus Data-Driven Logics

    Dubois, D.; H \'a jek, P.; Prade, H. Knowledge-Driven versus Data-Driven Logics. Journal of Logic, Language and Information 2000, 9, 65--89

  3. [11]

    KnowTD--An Actionable Knowledge Representation System for Thermodynamics

    Vollmer, L.; Fellenz, S.; Jirasek, F.; Leitte, H.; Hasse, H. KnowTD--An Actionable Knowledge Representation System for Thermodynamics. Journal of chemical information and modeling 2024, 64, 5878--5887

  4. [12]

    Turing, A. M. I.---Computing Machinery and Intelligence. Mind 1950, LIX, 433--460

  5. [13]

    Universal Intelligence: A Definition of Machine Intelligence

    Legg, S.; Hutter, M. Universal Intelligence: A Definition of Machine Intelligence. Minds and Machines 2007, 17, 391--444

  6. [14]

    On the Measure of Intelligence

    Chollet, F. On the Measure of Intelligence. 2019; http://arxiv.org/pdf/1911.01547v2

  7. [15]

    French, R. M. The Turing Test: the first 50 years. Trends in Cognitive Sciences 2000, 4, 115--122

  8. [16]

    T.; Li, Y.; Lundberg, S.; Nori, H.; Palangi, H.; Ribeiro, M

    Bubeck, S.; Chandrasekaran, V.; Eldan, R.; Gehrke, J.; Horvitz, E.; Kamar, E.; Lee, P.; Lee, Y. T.; Li, Y.; Lundberg, S.; Nori, H.; Palangi, H.; Ribeiro, M. T.; Zhang, Y. Sparks of Artificial General Intelligence: Early experiments with GPT-4. 2023; https://arxiv.org/abs/2303.12712

  9. [17]

    GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models

    Eloundou, T.; Manning, S.; Mishkin, P.; Rock, D. GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models. 2023; https://arxiv.org/abs/2303.10130

  10. [18]

    Kasneci, E. et al. ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences 2023, 103, 102274

  11. [19]

    Messeri, L.; Crockett, M. J. Artificial intelligence and illusions of understanding in scientific research. Nature 2024, 627, 49--58

  12. [20]

    Binz, M. et al. How should the advancement of large language models affect the practice of science? Proceedings of the National Academy of Sciences of the United States of America 2025, 122, e2401227121

  13. [21]

    Smuha, N. A. Regulation 2024/1689 of the Eur. Parl. & Council of June 13, 2024 (Eu Artificial Intelligence Act). International Legal Materials 2025, 1--148

  14. [22]

    S.; Wen, Q

    Chu, Z.; Wang, S.; Xie, J.; Zhu, T.; Yan, Y.; Ye, J.; Zhong, A.; Hu, X.; Liang, J.; Yu, P. S.; Wen, Q. LLM Agents for Education: Advances and Applications. 2025; https://arxiv.org/abs/2503.11733

  15. [23]

    F.; Gulati, R.; Jasperson, B

    Shojaei, M. F.; Gulati, R.; Jasperson, B. A.; Wang, S.; Cimolato, S.; Cao, D.; Neiswanger, W.; Garikipati, K. AI-University: An LLM-based platform for instructional alignment to scientific classrooms. 2025; https://arxiv.org/abs/2504.08846

  16. [24]

    Using Artificial Intelligence Case Studies in a Thermodynamics Course

    Supan, K. Using Artificial Intelligence Case Studies in a Thermodynamics Course. 2024 ASEE Annual Conference & Exposition Proceedings. 23.06.2024 - 12.07.2024

  17. [25]

    D.; Stoyanovich, J

    Fuligni, C.; Figaredo, D. D.; Stoyanovich, J. Would You Want an AI Tutor? Understanding Stakeholder Perceptions of LLM-based Chatbots in the Classroom. http://arxiv.org/pdf/2503.02885v1

  18. [26]

    Designing Safe and Relevant Generative Chats for Math Learning in Intelligent Tutoring Systems

    Levonian, Z.; Henkel, O.; Li, C.; Postle, M.-E. Designing Safe and Relevant Generative Chats for Math Learning in Intelligent Tutoring Systems. Journal of Educational Data Mining 2025, 17, 66--97

  19. [27]

    On the application of Large Language Models for language teaching and assessment technology

    Caines, A.; Benedetto, L.; Taslimipoor, S.; Davis, C.; Gao, Y.; Andersen, O.; Yuan, Z.; Elliott, M.; Moore, R.; Bryant, C.; Rei, M.; Yannakoudakis, H.; Mullooly, A.; Nicholls, D.; Buttery, P. On the application of Large Language Models for language teaching and assessment tech...

  20. [28]

    Effective and Scalable Math Support: Evidence on the Impact of an AI- Tutor on Math Achievement in Ghana

    Henkel, O.; Horne-Robinson, H.; Kozhakhmetova, N.; Lee, A. Effective and Scalable Math Support: Evidence on the Impact of an AI- Tutor on Math Achievement in Ghana. http://arxiv.org/pdf/2402.09809v2

  21. [29]

    T.; Yamashita, M.; Prihar, E.; Heffernan, N.; Wu, X.; Graff, B.; Lee, D

    Shen, J. T.; Yamashita, M.; Prihar, E.; Heffernan, N.; Wu, X.; Graff, B.; Lee, D. MathBERT: A Pre-trained Language Model for General NLP Tasks in Mathematics Education. http://arxiv.org/pdf/2106.07340v5

  22. [30]

    Student Interaction with NewtBot: An LLM-as-tutor Chatbot for Secondary Physics Education

    Lieb, A.; Goel, T. Student Interaction with NewtBot: An LLM-as-tutor Chatbot for Secondary Physics Education. Extended Abstracts of the CHI Conference on Human Factors in Computing Systems. New York, NY, USA, 2024; pp 1--8

  23. [31]

    P.; Sachan, M

    Vanzo, A.; Chowdhury, S. P.; Sachan, M. GPT-4 as a Homework Tutor can Improve Student Engagement and Learning Outcomes. 2024; https://arxiv.org/abs/2409.15981

  24. [32]

    Toward AI grading of student problem solutions in introductory physics: A feasibility study

    Kortemeyer, G. Toward AI grading of student problem solutions in introductory physics: A feasibility study. Physical Review Physics Education Research 2023, 19

  25. [33]

    Grading assistance for a handwritten thermodynamics exam using artificial intelligence: An exploratory study

    Kortemeyer, G.; N \"o hl, J.; Onishchuk, D. Grading assistance for a handwritten thermodynamics exam using artificial intelligence: An exploratory study. Physical Review Physics Education Research 2024, 20

  26. [34]

    intelligent

    Flod \'e n, J. Grading exams using large language models: A comparison between human and AI grading of exams in higher education using ChatGPT. British Educational Research Journal 2025, 51, 201--224 mcitethebibliography 3_main.tex0000664000000000000000000006656715022313577011...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.