REVIEW 4 major objections 5 minor 34 references
Superstudent intelligence in thermodynamics
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The reasoning model o3, given a 90-minute thermodynamics exam zero-shot, solved all problems correctly and scored above all 90 students, at the level of the best scores seen in more than 10,000 similar exams since 1985.
desk verdict Real exam, real result, but the diagram handling is murky and the abstract overstates the case; plausible and worth reviewing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The measuring instrument is the exam itself, and the argument stands on its design premise: every problem is new and carefully crafted so that pattern matching from training data cannot solve it, and the grading rewards the analysis — writing and combining the relevant equations of the first and second law — far more than the numerical result, so a top grade is attainable without any numerical calculation. Against that instrument the authors put a strict comparison protocol: three zero-shot runs of o3, each prompted only with a brief instruction to act as an expert and solve the exam, with the three problems submitted one at a time, and each response assessed exactly as the students' papers were, against the published reference solution. The historical record of more than 10,000 similar exams graded by the same instructor since 1985 supplies the human score distribution that makes the model's single score interpretable.
What would settle it
Have the three o3 transcripts blind-graded by independent examiners with the graphical tasks strictly enforced — diagrams actually drawn and scored exactly as they are for students — and compare the result with the class distribution; if the score drops below the top student's, the reported margin is a format artifact rather than equivalent problem-solving. A second check would run the identical zero-shot protocol on a different exam version from the same course and test whether o3 again lands in the historical best-score range.
Extended reading notes
Core claim
The central claim is that o3, prompted as a thermodynamics expert with no special training, solved a real 90-minute Thermodynamics I exam better than every one of the 90 students who took it. The exam contained three problems — two gas chambers separated by a movable piston, a combined heat-and-power plant with compressor, turbine and heat pump, and a Rankine cycle whose steam table had to be completed — including graphical tasks in which students had to draw $p,V$ and $T,s$ diagrams. Three independent runs of o3 were graded with the same rubric and the same reference solution applied to the students, and the runs differed hardly at all. o3 effectively solved the first two problems, losing only a few points for graphical presentation in problem 2, and made one major error in problem 3 on the caloric properties of a constant-density fluid; its overall score placed it above the entire class and in the range of the best scores recorded since 1985. The authors call this a turning point: a grade A in this exam had been treated as proof of genuine understanding, and a machine has now reached it.
Load-bearing premise
The comparison assumes that the exam's diagrams reached o3 in a form equivalent to what the students saw and that its graphical answers were graded under the same rubric as the students' handwritten sheets, since the paper never states how the $p,V$ and $T,s$ tasks were presented to the model or how its graphical output was assessed.
Editorial extensions
If this is right
- A commercial reasoning model, used off the shelf with no fine-tuning, can outperform an entire class on a demanding conceptual exam designed specifically to defeat pattern recognition.
- The exam-design premise that carefully crafted new problems separate genuine understanding from memorized patterns no longer separates humans from machines.
- Engineering education will shift toward prompting, questioning, and judging the trustworthiness of machine results, with fundamental understanding becoming more important than routine calculation.
- Engineering practice will move toward mixed human-AI teams, raising open questions about task assignment, verification of results, responsibility, and legal accountability.
- University examinations that certify independent problem-solving will need to be rethought, since students will know a machine can outperform them on the current format.
Reading between the lines
- A natural extension the paper does not attempt: run the same zero-shot protocol on several earlier exam versions whose score distributions are already on record, which would show whether o3's top-of-distribution score is stable or depends on the particular exam's content and on how its diagrams are encoded.
- The authors leave open whether o3 reached its score by applying general thermodynamic principles or by drawing on the vast number of solved examples in its training data; that question is testable with problems deliberately unlike any textbook example.
- An implication the authors state only implicitly: if this exam can no longer certify human understanding, the natural response in assessment is to extend the rubric to what the model cannot yet do — such as defending a solution under questioning, choosing the analysis approach, or judging whether an answer is physically plausible in an unfamiliar situation.
- The authors themselves caution that extrapolating from exam problems to real engineering tasks is bold, since LLMs are better suited to the former than the latter; that caveat limits the paper's practical consequences for engineering practice more than for education and assessment.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript reports that OpenAI's o3 model, prompted zero-shot in German, solved a 90-minute first-course engineering thermodynamics exam taken by 90 students, and that its graded score exceeded all student scores, falling in the range of the best scores the authors report having seen in more than 10,000 similar exams since 1985. The exam includes graphical input (a process sketch) and graphical output requirements (p,V and T,s diagrams). The authors report three o3 runs with similar results, discuss the one major mistake in Problem 3 and minor graphical issues, and then use the outcome to argue that machines exhibiting such exam performance should be called intelligent and that engineering education and practice must change accordingly.
Significance. If the central measurement is sound, this is a striking result: a commercial reasoning LLM outperforming every student in a genuine, previously unseen university exam with substantial graphical components would indeed be a noteworthy benchmark for machine competence in engineering thermodynamics. The study has clear strengths: it uses an authentic external exam rather than LLM-tailored problems, reports three independent zero-shot runs, provides the exam and model solution in the Supplementary Information, and builds on the authors' earlier published LLM baseline on simpler problems. The comparison against a real cohort of 90 students with a high failure rate gives the claim concrete educational relevance. However, the manuscript as written does not make the measurement fully auditable, and the abstract's wording exceeds what Section 2 reports.
major comments (4)
- [Abstract and Section 2] The Abstract states that o3 'solved all problems correctly' and 'outperformed all students,' but Section 2 explicitly reports that 'the only major mistake o3 made was in Problem 3' and that points were lost for minor graphical issues. These statements cannot both be true under the same grading standard. The authors should either correct the Abstract to say that o3 obtained the highest overall score despite one major mistake and minor graphical losses, or provide a point-by-point reconciliation of the claimed 'all problems correctly' with the reported losses.
- [Section 2; Problems 1c, 2a, 2c and BHKW sketch] The central comparison rests on the claim that o3's answers were assessed 'exactly the same way' as the students' handwritten exams, but the manuscript never specifies how the exam's graphical content was presented to o3 or how o3's graphical outputs were produced and graded. The exam requires a p,V diagram (Problem 1c, 4 points), a T,s diagram with state points (Problem 2a, 3 points), marking heat in the T,s diagram (Problem 2c, 4.5 points), and the process refers to a sketch ('siehe Skizze') that is presumably part of the problem statement. Without a description of whether the sketch was transmitted as an image, whether o3 responded with image or text/ASCII diagrams, and which rubric was applied to those outputs, the equivalence of the assessment conditions is not established and the outperformance claim is not fully supported.
- [Supplementary Information and Section 2] The manuscript does not include o3's answer sheets or a problem-by-problem point breakdown of the three runs, even though it does include the exam and a model solution. Without the raw o3 answers or a grading transcript showing how each part of each problem was scored, a reader cannot audit the claim that the same rubric was applied or verify where points were lost. The authors should add the o3 outputs (or a detailed transcript) and the per-part scoring for all three runs.
- [Section 3 and Figure 2] The claim that the overall score was 'in the range of the best scores we have seen in more than 10,000 similar exams since 1985' is presented without supporting data. The manuscript gives no distribution of historical best scores, percentile values, or a definition of 'similar exams,' so this comparative assertion cannot be checked. If the historical comparison is to remain in the Abstract and Discussion, the authors need to provide the relevant historical statistics or weaken the claim to what the provided student cohort shows.
minor comments (5)
- [Table 1] The column header 'GTP-4o' should read 'GPT-4o'.
- [Author affiliations] Affiliation 4 contains a typo: 'Kaierslautern' should be 'Kaiserslautern'.
- [Figure captions] The captions of Figures 1 and 2 contain garbled '/uni00000013/...' strings that appear to be PDF text-extraction artifacts; the figures should be regenerated so the captions are clean and readable.
- [Section 2] The claim that the language of the exam 'hardly' influences LLM performance is stated without data in this manuscript; a citation to the previous study where this was tested would make the claim auditable.
- [Section 2] The phrase 'the prompt was chosen based on some preliminary tests' should be accompanied by the tested alternatives or at least a description of what was varied, so that the zero-shot claim is not undermined by undisclosed prompt selection.
Circularity Check
No circularity: the central claim is an external empirical measurement of o3 against a fixed exam and a fixed student cohort, not a derivation from fitted inputs.
full rationale
The paper's central claim is empirical and self-contained: a fixed thermodynamics exam was given to 90 students and to OpenAI's o3 in zero-shot mode, and the answers were assessed by the same rubric ("we can directly compare the performance of the o3 model with that of our students"). There is no fitted parameter, no equation derived from an input, and no prediction that reduces by construction to a training target or to the exam itself. The self-citations, principally Loubet et al. (2025) for earlier LLM test scores and Vollmer et al. (2024) for a knowledge-based thermodynamics system, are contextual and do not carry the load-bearing comparison; the o3 result is reported directly from the exam runs, with three independent scores shown in Figures 1 and 2. The internal inconsistency between the abstract's "solved all problems correctly" and Section 2's "only major mistake o3 made was in Problem 3" is a reporting inconsistency, not circular reasoning. The unstated details of how graphical exam content was supplied to o3 and how its graphical answers were graded are validity or evidence concerns about comparability, not a case of the derivation reducing to its own inputs. Accordingly, no circular step is identified and the score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption The exam grading rubric is applied identically to o3 and student answers, and the manual grading of o3's outputs is unbiased.
- domain assumption The exam problems cannot be solved by pattern learning, so a high score implies understanding of thermodynamic principles.
- domain assumption o3 did not have access to the specific exam content during training.
- domain assumption The German language of the exam and the English system prompt do not materially affect o3's performance.
Cite this review
Pith. "Pith review of Superstudent intelligence in thermodynamics." pith.science (2026). https://pith.science/paper/EHG2ZOGX
@misc{pith2026250609822,
author = {Pith},
title = {Pith review of: Superstudent intelligence in thermodynamics},
year = {2026},
howpublished = {\url{https://pith.science/paper/EHG2ZOGX}},
note = {Machine review of arXiv:2506.09822}
}
read the original abstract
In this short note, we report and analyze a striking event: OpenAI's large language model o3 has outwitted all students in a university exam on thermodynamics. The thermodynamics exam is a difficult hurdle for most students, where they must show that they have mastered the fundamentals of this important topic. Consequently, the failure rates are very high, A-grades are rare - and they are considered proof of the students' exceptional intellectual abilities. This is because pattern learning does not help in the exam. The problems can only be solved by knowledgeably and creatively combining principles of thermodynamics. We have given our latest thermodynamics exam not only to the students but also to OpenAI's most powerful reasoning model, o3, and have assessed the answers of o3 exactly the same way as those of the students. In zero-shot mode, the model o3 solved all problems correctly, better than all students who took the exam; its overall score was in the range of the best scores we have seen in more than 10,000 similar exams since 1985. This is a turning point: machines now excel in complex tasks, usually taken as proof of human intellectual capabilities. We discuss the consequences this has for the work of engineers and the education of future engineers.
Figures
Reference graph
Works this paper leans on
-
[1]
Using Large Language Models for Solving Thermodynamic Problems
Loubet, R.; Zittlau, P.; Vollmer, L.; Hoffmann, M.; Fellenz, S.; Jirasek, F.; Leitte, H.; Hasse, H. Using Large Language Models for Solving Thermodynamic Problems. 2025; http://arxiv.org/pdf/2502.05195v1
arXiv 2025
-
[2]
2025; https://platform.openai.com/docs/models/o3
OpenAI OpenAI Platform. 2025; https://platform.openai.com/docs/models/o3
work page 2025
-
[3]
Tschisgale, P.; Maus, H.; Kieser, F.; Kroehs, B.; Petersen, S.; Wulff, P. Evaluating GPT- and Reasoning-based Large Language Models on Physics Olympiad Problems: Surpassing Human Performance and Implications for Educational Assessment. 2025; https://arxiv.org/abs/2505.09438
work page Pith review arXiv 2025
-
[4]
UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models
Xu, X.; Xu, Q.; Xiao, T.; Chen, T.; Yan, Y.; Zhang, J.; Diao, S.; Yang, C.; Wang, Y. UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language Models. 2025; https://arxiv.org/abs/2502.00334
arXiv 2025
-
[5]
Combining Machine Learning with Physical Knowledge in Thermodynamic Modeling of Fluid Mixtures
Jirasek, F.; Hasse, H. Combining Machine Learning with Physical Knowledge in Thermodynamic Modeling of Fluid Mixtures. Annual review of chemical and biomolecular engineering 2023, 14, 31--51
work page 2023
-
[6]
HANNA: hard-constraint neural network for consistent activity coefficient prediction
Specht, T.; Nagda, M.; Fellenz, S.; Mandt, S.; Hasse, H.; Jirasek, F. HANNA: hard-constraint neural network for consistent activity coefficient prediction. Chemical science 2024, 15, 19777--19786
work page 2024
-
[7]
When physics meets machine learning: a survey of physics-informed machine learning
Meng, C.; Griesemer, S.; Cao, D.; Seo, S.; Liu, Y. When physics meets machine learning: a survey of physics-informed machine learning. Machine Learning for Computational Science and Engineering 2025, 1
work page 2025
-
[8]
Wu, Y.; Sicard, B.; Gadsden, S. A. Physics-informed machine learning: A comprehensive review on applications in anomaly detection and condition monitoring. Expert Systems with Applications 2024, 255, 124678
work page 2024
Show all 34 references
-
[9]
Integrating knowledge-driven and data-driven approaches to modeling
Todorovski, L.; D z eroski, S. Integrating knowledge-driven and data-driven approaches to modeling. Ecological Modelling 2006, 194, 3--13
2006
-
[10]
Knowledge-Driven versus Data-Driven Logics
Dubois, D.; H \'a jek, P.; Prade, H. Knowledge-Driven versus Data-Driven Logics. Journal of Logic, Language and Information 2000, 9, 65--89
2000
-
[11]
KnowTD--An Actionable Knowledge Representation System for Thermodynamics
Vollmer, L.; Fellenz, S.; Jirasek, F.; Leitte, H.; Hasse, H. KnowTD--An Actionable Knowledge Representation System for Thermodynamics. Journal of chemical information and modeling 2024, 64, 5878--5887
2024
-
[12]
Turing, A. M. I.---Computing Machinery and Intelligence. Mind 1950, LIX, 433--460
1950
-
[13]
Universal Intelligence: A Definition of Machine Intelligence
Legg, S.; Hutter, M. Universal Intelligence: A Definition of Machine Intelligence. Minds and Machines 2007, 17, 391--444
2007
-
[14]
On the Measure of Intelligence
Chollet, F. On the Measure of Intelligence. 2019; http://arxiv.org/pdf/1911.01547v2
2019 arXiv
-
[15]
French, R. M. The Turing Test: the first 50 years. Trends in Cognitive Sciences 2000, 4, 115--122
2000
-
[16]
T.; Li, Y.; Lundberg, S.; Nori, H.; Palangi, H.; Ribeiro, M
Bubeck, S.; Chandrasekaran, V.; Eldan, R.; Gehrke, J.; Horvitz, E.; Kamar, E.; Lee, P.; Lee, Y. T.; Li, Y.; Lundberg, S.; Nori, H.; Palangi, H.; Ribeiro, M. T.; Zhang, Y. Sparks of Artificial General Intelligence: Early experiments with GPT-4. 2023; https://arxiv.org/abs/2303.12712
2023 arXiv
-
[17]
GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models
Eloundou, T.; Manning, S.; Mishkin, P.; Rock, D. GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models. 2023; https://arxiv.org/abs/2303.10130
2023 arXiv
-
[18]
Kasneci, E. et al. ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences 2023, 103, 102274
2023
-
[19]
Messeri, L.; Crockett, M. J. Artificial intelligence and illusions of understanding in scientific research. Nature 2024, 627, 49--58
2024
-
[20]
Binz, M. et al. How should the advancement of large language models affect the practice of science? Proceedings of the National Academy of Sciences of the United States of America 2025, 122, e2401227121
2025
-
[21]
Smuha, N. A. Regulation 2024/1689 of the Eur. Parl. & Council of June 13, 2024 (Eu Artificial Intelligence Act). International Legal Materials 2025, 1--148
2024
-
[22]
S.; Wen, Q
Chu, Z.; Wang, S.; Xie, J.; Zhu, T.; Yan, Y.; Ye, J.; Zhong, A.; Hu, X.; Liang, J.; Yu, P. S.; Wen, Q. LLM Agents for Education: Advances and Applications. 2025; https://arxiv.org/abs/2503.11733
2025
-
[23]
F.; Gulati, R.; Jasperson, B
Shojaei, M. F.; Gulati, R.; Jasperson, B. A.; Wang, S.; Cimolato, S.; Cao, D.; Neiswanger, W.; Garikipati, K. AI-University: An LLM-based platform for instructional alignment to scientific classrooms. 2025; https://arxiv.org/abs/2504.08846
2025 arXiv
-
[24]
Using Artificial Intelligence Case Studies in a Thermodynamics Course
Supan, K. Using Artificial Intelligence Case Studies in a Thermodynamics Course. 2024 ASEE Annual Conference & Exposition Proceedings. 23.06.2024 - 12.07.2024
2024
-
[25]
D.; Stoyanovich, J
Fuligni, C.; Figaredo, D. D.; Stoyanovich, J. Would You Want an AI Tutor? Understanding Stakeholder Perceptions of LLM-based Chatbots in the Classroom. http://arxiv.org/pdf/2503.02885v1
-
[26]
Designing Safe and Relevant Generative Chats for Math Learning in Intelligent Tutoring Systems
Levonian, Z.; Henkel, O.; Li, C.; Postle, M.-E. Designing Safe and Relevant Generative Chats for Math Learning in Intelligent Tutoring Systems. Journal of Educational Data Mining 2025, 17, 66--97
2025
-
[27]
On the application of Large Language Models for language teaching and assessment technology
Caines, A.; Benedetto, L.; Taslimipoor, S.; Davis, C.; Gao, Y.; Andersen, O.; Yuan, Z.; Elliott, M.; Moore, R.; Bryant, C.; Rei, M.; Yannakoudakis, H.; Mullooly, A.; Nicholls, D.; Buttery, P. On the application of Large Language Models for language teaching and assessment tech...
2023 arXiv
-
[28]
Effective and Scalable Math Support: Evidence on the Impact of an AI- Tutor on Math Achievement in Ghana
Henkel, O.; Horne-Robinson, H.; Kozhakhmetova, N.; Lee, A. Effective and Scalable Math Support: Evidence on the Impact of an AI- Tutor on Math Achievement in Ghana. http://arxiv.org/pdf/2402.09809v2
-
[29]
T.; Yamashita, M.; Prihar, E.; Heffernan, N.; Wu, X.; Graff, B.; Lee, D
Shen, J. T.; Yamashita, M.; Prihar, E.; Heffernan, N.; Wu, X.; Graff, B.; Lee, D. MathBERT: A Pre-trained Language Model for General NLP Tasks in Mathematics Education. http://arxiv.org/pdf/2106.07340v5
-
[30]
Student Interaction with NewtBot: An LLM-as-tutor Chatbot for Secondary Physics Education
Lieb, A.; Goel, T. Student Interaction with NewtBot: An LLM-as-tutor Chatbot for Secondary Physics Education. Extended Abstracts of the CHI Conference on Human Factors in Computing Systems. New York, NY, USA, 2024; pp 1--8
2024
-
[31]
P.; Sachan, M
Vanzo, A.; Chowdhury, S. P.; Sachan, M. GPT-4 as a Homework Tutor can Improve Student Engagement and Learning Outcomes. 2024; https://arxiv.org/abs/2409.15981
2024 arXiv
-
[32]
Toward AI grading of student problem solutions in introductory physics: A feasibility study
Kortemeyer, G. Toward AI grading of student problem solutions in introductory physics: A feasibility study. Physical Review Physics Education Research 2023, 19
2023
-
[33]
Grading assistance for a handwritten thermodynamics exam using artificial intelligence: An exploratory study
Kortemeyer, G.; N \"o hl, J.; Onishchuk, D. Grading assistance for a handwritten thermodynamics exam using artificial intelligence: An exploratory study. Physical Review Physics Education Research 2024, 20
2024
-
[34]
intelligent
Flod \'e n, J. Grading exams using large language models: A comparison between human and AI grading of exams in higher education using ChatGPT. British Educational Research Journal 2025, 51, 201--224 mcitethebibliography 3_main.tex0000664000000000000000000006656715022313577011...
2025
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.