Pith. sign in

REVIEW 5 major objections 4 minor 43 references

From Canonical to Complex: Benchmarking LLM Capabilities in Undergraduate Thermodynamics

T0 review · 5 major / 4 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read Current LLMs fall short of the 95% accuracy needed for unsupervised thermodynamics tutoring, topping out at 82%.

desk verdict Useful new benchmark with a credible text-vs-diagram gap, but the tutoring-suitability claim rests on an unvalidated 95% bar and a missing human baseline; needs revision, not desk rejection. read the letter →

arxiv 2508.21452 v1 pith:TIGRJQ2N submitted 2025-08-29 physics.ed-ph cs.CLphysics.chem-ph

classification physics.ed-phcs.CLphysics.chem-ph
keywords largelanguagemodelsthermodynamicseducationbenchmarkingpromptengineeringdiagram-basedreasoningreversibilityentropyeducationalmeasurement
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces UTQA, a 50-question single-choice benchmark in undergraduate thermodynamics, and reports that the leading 2025-era language models score at most 82% overall, below a 95% accuracy bar the authors take as the reliability threshold for unsupervised tutoring. The shortfall is not uniform: text-only items are answered far better than questions requiring interpretation of p–V, T–S, and similar diagrams, where many models fall to near-chance levels. The benchmark isolates what the models lack: reasoning about finite-rate, irreversible processes and binding visual diagram features to thermodynamic meaning. The authors conclude that current LLMs are not yet suitable as unsupervised tutors in this domain, while noting that accuracy alone is not sufficient for tutoring.

What carries the argument

The central object is UTQA, a 50-item benchmark with 33 text-only and 17 diagram-based single-choice questions. Its design deliberately isolates single constructs (state vs path functions, q/w sign conventions, quasistatic vs non-quasistatic, entropy bookkeeping) and uses distractors that correspond to known student misconceptions, so accuracy measures principle-grounded reasoning rather than surface cueing. The benchmark's power comes from the contrast between items with canonical structure, solvable by one state-function identity, and items requiring multi-constraint integration or diagram-to-thermodynamics binding.

What would settle it

A concrete falsifier: run UTQA on a future or untested model with the same protocol; if it scores above 95% on both text and diagram subsets, the central suitability claim is directly refuted. A more targeted check: take items the models miss and present the same diagrams with axis labels added or rotated; if accuracy jumps sharply, the bottleneck is label-pairing rather than conceptual diagram binding, weakening the paper's interpretation.

Watch

Extended reading notes

Core claim

The central claim is that no leading 2025-era LLM reached the authors' 95% competence threshold on UTQA, a 50-item instrument covering ideal-gas processes, reversibility, and diagram interpretation. The best model achieved 82%; text-only items averaged 67% across 19 models, while diagram-based items averaged 32%. Error analysis locates the bottleneck in two specific capabilities: recognizing when an ideal-gas process is finite-rate or irreversible and therefore resisting quasistatic templates, and mapping perceptual features of diagrams (signed areas, path orientation, leg types, cycle constraints) to thermodynamic quantities. The authors argue this pattern shows fluent recall of canonical r

Load-bearing premise

The claim depends on treating 95% accuracy on UTQA's 50 items as the right bar for unsupervised tutoring; if the items overstate difficulty or a lower accuracy is acceptable with safeguards, the suitability conclusion could change even though the raw scores stand.

Editorial extensions

If this is right

  • No model currently reaches the 95% threshold, so unsupervised LLM tutoring in undergraduate thermodynamics is not supported by this evidence.
  • Diagram-based items are the main bottleneck; even models that parse axes and curves fail to bind them to thermodynamic meaning.
  • Prompt phrasing, including chain-of-thought and persona variants, shifts accuracy by a small amount but does not close the gap; elimination-style prompts hurt.
  • Clause count, a simple proxy for linguistic complexity, shows no significant correlation with accuracy over the observed 1–20 clause range.
  • Improvements in multimodal binding and finite-rate process reasoning are the most likely path to crossing the threshold.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • This suggests a testable extension: adding axis labels or verbal descriptions of diagrams to exam prompts could disproportionately improve scores if the binding deficit is genuine.
  • The 95% threshold, imported from tutoring-effectiveness research, may be conservative for hybrid human-AI tutoring; a supervised setting with model uncertainty flags might be viable below that bar.
  • UTQA's misconception-based distractors could double as a diagnostic instrument for categorizing LLM errors, such as regime misclassification versus entropy bookkeeping versus area orientation.
  • If diagram-binding is the bottleneck, progress in visually-grounded reasoning could be tracked directly on this benchmark before general benchmark improvements appear.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 4 minor

Summary. The paper introduces UTQA, a 50-item single-choice undergraduate thermodynamics benchmark (33 text-only, 17 diagram-based), and uses it to evaluate 19 commercial LLMs plus 17 prompting variants on gpt-4o. The headline finding is that no 2025-era model reaches the authors' 95% accuracy threshold for unsupervised tutoring; the best model (gpt-o3) scores 82% overall. Text-only accuracy is substantially higher than diagram accuracy (67% vs. 32% averaged across models), and the authors attribute the gap to failures in finite-rate/irreversible reasoning and in mapping visual diagram features to thermodynamic meaning. The paper concludes that current LLMs are not yet reliable enough for unsupervised undergraduate thermodynamics tutoring, while noting that accuracy is necessary but not sufficient for tutoring quality.

Significance. If the results hold, UTQA is a useful, publicly released benchmark for a domain that is underrepresented in existing LLM evaluation suites. The paper has clear strengths: the benchmark is downloadable, full prompts and solutions are promised in the SI, items were expert-validated, controlled linguistic degradations are systematically designed, and the text-only vs. diagram gap is a striking and actionable finding. The strongest raw result—no leading model exceeds 82% on a 50-item expert-constructed test—appears robust in direction, and the evidence that diagram comprehension is the main bottleneck is plausible. However, the paper's suitability conclusion depends on an absolute accuracy threshold whose calibration is not established, and several supporting claims (human equivalence, finite-rate difficulty concentration) lack direct measurement. The benchmark itself is a contribution; the interpretive framing needs substantial reinforcement.

major comments (5)
  1. [Conclusions / Fig. 8] The central claim—that no LLM is suitable for unsupervised tutoring because none reaches 95% on UTQA—requires an anchor showing that UTQA items are a fair measure of competent human tutor knowledge. No human-expert baseline is reported: no data on how thermodynamics faculty or PhD-level experts perform on the same 50 items, no per-item expert agreement, and no demonstration that the 17 diagram items are decidable from the image resolution used. If expert accuracy on UTQA is itself below 95% (due to item ambiguity or deliberate emphasis on edge cases), the conclusion would reflect instrument difficulty rather than an LLM-specific deficit. The paper should add a human-expert baseline or explicitly reframe the claim as 'below an aspirational threshold' rather than 'unsuitable for tutoring.'
  2. [Figure 3 caption / Results] Immediately before the Fig. 3 caption, the manuscript contains the leftover text 'OLD STUFF before problem with question set spottet' (sic). This is internal evidence that some displayed prompt results may have been generated before a known problem with the question set was identified. The paper must clarify whether the prompt-comparison results in Figs. 2–4, and the gpt-4o points in Figs. 5 and 8, were obtained with the final, released UTQA item set. If any displayed results predate a question-set fix, those numbers cannot be interpreted as evaluating the released benchmark.
  3. [Methods / Figs. 5–6] The run-to-run scatter σ≈0.05 is estimated from 51 gpt-4o prompt-variation batches (Methods, Fig. 2a), but it is then applied as a universal per-condition uncertainty to all 19 cross-model comparisons, including the 17-item diagram subset. For a 17-item binary test, the binomial standard error at p≈0.5 is roughly 0.12, so several adjacent model differences in Fig. 6 are not significant under a test-specific error. The paper should report run counts and standard errors for each model and subset rather than a single pooled σ, or explicitly justify why the pooled estimate is transferable.
  4. [Results: Common strengths and weaknesses / Conclusions] The claim that the performance gap 'concentrates in finite-rate/irreversible scenarios' is not supported by a defined item subset or statistical comparison. The text gives illustrative examples but no per-category accuracies, item IDs, or a significance test for the category. Similarly, the statement that the strongest text-only models 'approach the level of a well-prepared graduate tutor' is a human-equivalence claim with no human data behind it. Define the finite-rate/irreversible item subset and report its accuracy separately, or soften these claims to what the data directly show.
  5. [Introduction / Conclusions] The 95% competence threshold is imported from VanLehn (2011), which is a review of tutoring effectiveness expressed in effect sizes, not a calibration of accuracy thresholds for an MCQ test. The paper does not justify translating that literature into a 95% accuracy cut on this specific 50-item instrument. Because the threshold is decisive for the suitability conclusion, its calibration needs direct support—for example, human expert scores, an analysis of the consequences of an 18% vs. 5% error rate in tutoring, or a learning-outcome criterion. Without this, the suitability conclusion is an interpretive assertion, even though the raw accuracy numbers remain informative.
minor comments (4)
  1. [Figure 3] The caption states the figure shows '33 text-only questions,' but the panel labels mention 'no image interpretation' and 'image interpretation problems,' and the caption contains the typo 'spottet.' Please align labels and text and remove the leftover note.
  2. [Figures 5–6] The diagram-based results appear in both Fig. 5(b) and Fig. 6 with inconsistent panel labels ('image interpretation problems' vs. 'diagram based items') and model ordering. Consolidate or clearly distinguish the two figures to avoid duplication confusion.
  3. [Bibliography] There are minor typographical issues in the reference list, e.g., 'Langauge' in ref. 31 and 'arXiv preprint' duplicated in ref. 26. Please proofread.
  4. [Abstract/Conclusions] The abstract says 'the best LLMs achieved 82% accuracy,' which matches gpt-o3 in Fig. 8, but the Results section also emphasizes DeepSeek R1 on text-only items. Make the exact aggregation basis (overall vs. text-only) consistent wherever numbers are quoted.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: UTQA results are empirical measurements with an external threshold.

full rationale

The paper's central outputs—model accuracy scores, prompt comparisons, and linguistic degradation effects—are measurements on a fixed, expert-validated 50-item set. No load-bearing step reduces to its own inputs: the 95% competence threshold is an external criterion imported from VanLehn (ref. 17), not fitted to the model data; the item answer keys were determined by expert validation rather than derived from model outputs; and the prompting/degradation experiments compare fixed prompts on fixed items. There are no self-citations used as premises, nor any uniqueness theorem or ansatz smuggled in from the authors' prior work. Two limitations are flagged in the manuscript but are not circularity: (1) the Fig. 3 caption contains the stray note 'OLD STUFF before problem with question set spottet', indicating that some displayed prompt results may have been generated before a known problem with the question set was identified—this undermines the integrity of those specific figure values but does not make the argument circular; (2) no human-expert baseline is reported on the same items, so the absolute suitability conclusion is unanchored to item quality—this is an external-validity concern, not a circular derivation. The headline 'not suitable for unsupervised tutoring' depends on an imported threshold, but the threshold is not a function of the measurements themselves; the raw accuracies stand as independent observations.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The central claim rests on the validity of the benchmark and the transferability of the 95% threshold; no physics derivation is involved. The main additions beyond the data are hand-chosen thresholds and assumptions about item validity, contamination, and comparability of model interfaces.

free parameters (1)
  • 95% competence threshold = 95% accuracy
    Imported from VanLehn's tutoring-effectiveness review, not derived from UTQA data. It is the criterion for the paper's unsuitable-for-unsupervised-tutoring conclusion; changing it changes the headline claim.
assumptions (4)
  • domain assumption Accuracy on UTQA's 50 single-choice expert-validated items measures the targeted thermodynamic reasoning constructs.
    The paper follows educational measurement guidance but reports no human baseline, item difficulty, or discrimination indices; the construct-validity mapping is assumed. See Methods, Benchmark design.
  • domain assumption The 95% accuracy threshold is an appropriate minimum bar for unsupervised tutoring.
    Cited to VanLehn, but the transfer from tutoring effectiveness to a multiple-choice accuracy score is not established. The paper itself calls it provisional. See Conclusions.
  • domain assumption UTQA items were not seen by the tested models during training.
    The benchmark is newly released, but the English set was adapted from a previously reviewed German pool; no contamination check (for example n-gram overlap) is reported. See Methods, Question Set Design.
  • ad hoc to paper Cross-model web-CLI evaluations are comparable to each other and to the gpt-4o API evaluation, with a common sigma of about 0.05.
    Platform-specific system prompts, sampling, and image preprocessing are uncontrolled; the sigma estimate comes from gpt-4o prompt-run residuals and is applied to all models and item subsets, including the 17-item diagram set where binomial uncertainty is much larger. See Methods and Figure 5 caption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of From Canonical to Complex: Benchmarking LLM Capabilities in Undergraduate Thermodynamics." pith.science (2026). https://pith.science/paper/TIGRJQ2N

@misc{pith2026250821452,
  author       = {Pith},
  title        = {Pith review of: From Canonical to Complex: Benchmarking LLM Capabilities in Undergraduate Thermodynamics},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TIGRJQ2N}},
  note         = {Machine review of arXiv:2508.21452}
}
read the original abstract

Large language models (LLMs) are increasingly considered as tutoring aids in science education. Yet their readiness for unsupervised use in undergraduate instruction remains uncertain, as reliable teaching requires more than fluent recall: it demands consistent, principle-grounded reasoning. Thermodynamics, with its compact laws and subtle distinctions between state and path functions, reversibility, and entropy, provides an ideal testbed for evaluating such capabilities. Here we present UTQA, a 50-item undergraduate thermodynamics question answering benchmark, covering ideal-gas processes, reversibility, and diagram interpretation. No leading 2025-era model exceeded our 95\% competence threshold: the best LLMs achieved 82\% accuracy, with text-only items performing better than image reasoning tasks, which often fell to chance levels. Prompt phrasing and syntactic complexity showed modest to little correlation with performance. The gap concentrates in finite-rate/irreversible scenarios and in binding visual features to thermodynamic meaning, indicating that current LLMs are not yet suitable for unsupervised tutoring in this domain.

Figures

Figures reproduced from arXiv: 2508.21452 by the authors.

Figure 1
Figure 1. Representative figure from a bench￾mark item: four p–V diagrams depicting re￾versible state changes from which the case of largest pressure–volume work performed by the system must be identified. 3 [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. a) Distribution of deviations from mean accuracies across 51 batch runs, corre￾sponding to an overall spread of σ = 0.05. b) Comparative accuracies of 17 prompting strate￾gies; observed variation exceeds σ3 ≈ 0.03. By contrast, elimination-style prompts under￾perform on this benchmark. Several factors likely contribute to these dif￾ferences. First, longer reasoning chains intro￾duce more opportunities for error, and… view at source ↗
Figure 3
Figure 3. Accuracy scores for 17 prompting strategies using gpt-4o on 33 text-only ques￾tions. High- and low-performing groups are outlined in red. Values are three-run means; error bars indicate σ3 ≈ 0.03. Effect of Linguistic Degradation We tested robustness to degraded wording un￾der a fixed baseline prompt (Please answer the following single-choice question.), altering only the stem and options. Controlled edits targeted … view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Accuracy for reference, degraded, and deficient versions of 33 text-only questions across three linguistic dimensions: clarity & accuracy, spelling & punctuation, and techni￾cal terminology. Results shown for the basic prompt; means over ten runs with error bars indica…
Figure 5
Figure 5. Figure 5: Comparative accuracies of different LLMs on 33 text-only interpretation questions. Standard deviations of individual accuracy val￾ues are estimated at σ ≈ 0.05. Items involving diagram interpre￾tation Diagrams are central to thermodynamics in￾struction and problem solv…
Figure 7
Figure 7. Figure 7: Average model accuracy vs. number of clauses for the 33 text-only items. Numerals indicate item identifiers; the dashed line marks the random-guessing baseline of 0.25. Shaded bands show 95.4% confidence intervals. text-only problems, the strongest recent models approa…
Figure 8
Figure 8. Figure 8: Omni-model comparison: overall ac￾curacies of all tested LLMs on the complete 50- question benchmark (aggregate of text-only and diagram-related items). nition: when dissipation, feasibility bounds, or non-quasistatic driving matter, accuracies drop significantly. (ii)…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

43 extracted references · 33 canonical work pages

  1. [1]

    E.; Salakhutdinov, R

    Hinton, G. E.; Salakhutdinov, R. R. Reducing the dimensionality of data with neural networks. Science 2006, 313, 504--507

  2. [2]

    Learning deep architectures for AI

    Bengio, Y. Learning deep architectures for AI . Foundations and Trends in Machine Learning 2009, 2, 1--127

  3. [3]

    Deep learning

    LeCun, Y.; Bengio, Y.; Hinton, G. Deep learning. Nature 2015, 521, 436--444

  4. [4]

    N.; Kaiser, .; Polosukhin, I

    Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, .; Polosukhin, I. Attention is All You Need. Advances in Neural Information Processing Systems (NeurIPS). 2017

  5. [5]

    Brown, T. B. et al. Language Models are Few-Shot Learners. Advances in Neural Information Processing Systems 2020, 33, 1877--1901

  6. [6]

    2023; https://arxiv.org/abs/2303.08774

    OpenAI GPT-4 Technical Report. 2023; https://arxiv.org/abs/2303.08774

  7. [7]

    Thermodynamics and Chemistry: A Non-mathematical Treatise for Chemists and Students of Chemistry; J

    Duhem, P. Thermodynamics and Chemistry: A Non-mathematical Treatise for Chemists and Students of Chemistry; J. Wiley & sons, 1903

  8. [8]

    B.; Zenzen, M

    Hollinger, H. B.; Zenzen, M. J. Thermodynamic Irreversibility: I. What Is It? Journal of Chemical Education 1991, 68, 31--34

Show all 43 references
  1. [9]

    Unearthing a Buried Memory: Duhem’s Third Way to Thermodynamics

    Bordoni, S. Unearthing a Buried Memory: Duhem’s Third Way to Thermodynamics. Part 1. Centaurus 2012, 54, 124--147

  2. [10]

    Callen, H. B. Thermodynamics and an Introduction to Thermostatistics, 2nd ed.; Wiley: New York, 1985

  3. [11]

    Physical Chemistry, 11th ed.; Oxford University Press: Oxford, 2018

    Atkins, P.; de Paula, J.; Keeler, J. Physical Chemistry, 11th ed.; Oxford University Press: Oxford, 2018

  4. [12]

    Bain, K.; Towns, M. H. A review of research on the teaching and learning of thermodynamics at the university level. Chemistry Education Research and Practice 2014, 15, 320--335

  5. [13]

    Y.; Dirani, J.; Michael, J.; Bowman, S

    Rein, D.; Li Hou, B.; Cooper Stickland, A.; Petty, J.; Pang, R. Y.; Dirani, J.; Michael, J.; Bowman, S. R. GPQA: A Graduate Level Google-Proof Q&A Benchmark. arXiv preprint 2023, arXiv:2305.10408

  6. [14]

    Phan, L. e. a. Humanity's Last Exam: A Benchmark for Evaluating AI on the Benchmarks of Human Civilization. https://arxiv.org/abs/2501.14249 2025,

  7. [15]

    R.; Zhang, S.; Sun, Y.; Wang, W

    Wang, X.; Hu, Z.; Lu, P.; Zhu, Y.; Zhang, J.; Subramaniam, S.; Loomba, A. R.; Zhang, S.; Sun, Y.; Wang, W. SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models. Proceedings of the Forty-First International Conference on Machine Learn...

  8. [16]

    In Albert Einstein: Philosopher--Scientist; Schilpp, P

    Einstein, A. In Albert Einstein: Philosopher--Scientist; Schilpp, P. A., Ed.; Library of Living Philosophers: Evanston, IL, 1949

  9. [17]

    The Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring Systems

    VanLehn, K. The Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring Systems. Educational Psychologist 2011, 46, 197--221

  10. [18]

    https://learn.microsoft.com/en-us/azure/ai-foundry/openai/how-to/prompt-caching, OpenAI documentation (accessed 2025)

    OpenAI Prompt caching. https://learn.microsoft.com/en-us/azure/ai-foundry/openai/how-to/prompt-caching, OpenAI documentation (accessed 2025)

  11. [19]

    Holistic Evaluation of Language Models

    Liang, P.; Bommasani, R.; Lee, T.; Tsipras, D.; Soylu, D.; Yasunaga, M.; Zhang, Y.; Narayanan, D.; Wu, Y.; Kumar, A. Holistic Evaluation of Language Models. arXiv preprint arXiv:2305.14233 2023,

  12. [20]

    B.; Doan, K

    Shojaee, P.; Nguyen, N.-H.; Meidani, K.; Farimani, A. B.; Doan, K. D.; Reddy, C. K. LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language Models. 2025; https://arxiv.org/abs/2504.10415

  13. [21]

    Cognitive load theory, learning difficulty, and instructional design

    Sweller, J. Cognitive load theory, learning difficulty, and instructional design. Learning and Instruction 1994, 4, 295--312

  14. [22]

    M.; Rodriguez, M

    Haladyna, T. M.; Rodriguez, M. C. Developing and Validating Test Items, 3rd ed.; Routledge: New York, 2013

  15. [23]

    Validity of psychological assessment: Validation of inferences from persons' responses and performances as scientific inquiry into score meaning

    Messick, S. Validity of psychological assessment: Validation of inferences from persons' responses and performances as scientific inquiry into score meaning. American Psychologist 1995, 50, 741--749

  16. [24]

    Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm

    Reynolds, L.; McDonell, K. Prompt Programming for Large Language Models: Beyond the Few-Shot Paradigm. Proceedings of the 3rd Workshop on Natural Language Processing for Programming (NLP4Prog). 2021; pp 15--22

  17. [25]

    Chatterjee, A.; Renduchintala, H. S. V. N. S. K.; Bhatia, S.; Chakraborty, T. POSIX: A Prompt Sensitivity Index For Large Language Models. Findings of the Association for Computational Linguistics: EMNLP 2024. 2024; pp 14550--14565

  18. [26]

    X.; Hasan, S

    He, J.; Rungta, M.; Koleczek, D.; Sekhon, A.; Wang, F. X.; Hasan, S. Does Prompt Formatting Have Any Impact on LLM Performance? arXiv preprint arXiv:2411.10541 2024, arXiv preprint

  19. [27]

    Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

    Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E.; Le, Q.; Zhou, D. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. Advances in Neural Information Processing Systems. 2022; pp 24824--24837

  20. [28]

    Tree of Thoughts: Deliberate Problem Solving with Large Language Models

    Yao, S.; Yu, D.; Zhao, J.; Shafran, I.; Griffiths, T.; Cao, Y.; Narasimhan, K. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. Advances in Neural Information Processing Systems. 2023; pp 11809--11822

  21. [29]

    Graph of Thoughts: Solving Elaborate Problems with Large Language Models

    Besta, M.; Blach, N.; Kubicek, A.; Gerstenberger, R.; Podstawski, M.; Gianinazzi, L.; Gajda, J.; Lehmann, T.; Niewiadomski, H.; Nyczyk, P.; Hoefler, T. Graph of Thoughts: Solving Elaborate Problems with Large Language Models. Proceedings of the AAAI Conference on Artificial In...

  22. [30]

    L ogi C o T : Logical Chain-of-Thought Instruction Tuning

    Liu, H.; Teng, Z.; Cui, L.; Zhang, C.; Zhou, Q.; Zhang, Y. L ogi C o T : Logical Chain-of-Thought Instruction Tuning. Findings of the Association for Computational Linguistics: EMNLP 2023. Singapore, 2023; pp 2908--2921

  23. [31]

    Chain-of-Symbol Prompting Elicits Planning in Large Langauge Models

    Hu, H.; Lu, H.; Zhang, H.; Song, Y.-Z.; Lam, W.; Zhang, Y. Chain-of-Symbol Prompting Elicits Planning in Large Langauge Models. arXiv preprint arXiv:2305.09692. 2023

  24. [32]

    Large Language Models Understand and Can be Enhanced by Emotional Stimuli

    Li, C.; Wang, J.; Zhang, Y.; Zhu, K.; Hou, W.; Lian, J.; Luo, F.; Yang, Q.; Xie, X. Large Language Models Understand and Can be Enhanced by Emotional Stimuli. 2023; https://arxiv.org/abs/2307.11760

  25. [33]

    Selective Prompting Tuning for Personalized Conversations with LLMs

    Huang, Q.; Liu, X.; Ko, T.; Wu, B.; Wang, W.; Zhang, Y.; Tang, L. Selective Prompting Tuning for Personalized Conversations with LLMs. 2024; https://arxiv.org/abs/2406.18187

  26. [34]

    Measuring Faithfulness in Chain-of-Thought Reasoning

    Lanham, J.; et al. Measuring Faithfulness in Chain-of-Thought Reasoning. 2023; https://arxiv.org/abs/2307.13702

  27. [35]

    ASCoT: An Adaptive Self-Correction Chain-of-Thought Method for Late-Stage Fragility in LLMs

    Zhang, D.; Yang, N.; Zhu, J.; Yang, J.; Xin, M.; Tian, B. ASCoT: An Adaptive Self-Correction Chain-of-Thought Method for Late-Stage Fragility in LLMs. 2025; https://arxiv.org/abs/2508.05282

  28. [36]

    Large Language Models are Zero-Shot Reasoners

    Kojima, T.; et al. Large Language Models are Zero-Shot Reasoners. 2022; https://arxiv.org/abs/2205.11916

  29. [37]

    E.; Uccelli, P

    Snow, C. E.; Uccelli, P. In The Cambridge Handbook of Literacy; Olson, D. R., Torrance, N., Eds.; Cambridge University Press: Cambridge, 2009; pp 112--133

  30. [38]

    Measuring Massive Multitask Language Understanding

    Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Tang, D.; Song, D.; Steinhardt, J. Measuring Massive Multitask Language Understanding. Proceedings of the International Conference on Learning Representations (ICLR). 2021

  31. [39]

    H.; Simon, H

    Larkin, J. H.; Simon, H. A. Why a Diagram is (Sometimes) Worth Ten Thousand Words. Cognitive Science 1987, 11, 65--100

  32. [40]

    Chi, M. T. H.; Feltovich, P. J.; Glaser, R. Categorization and Representation of Physics Problems by Experts and Novices. Cognitive Science 1981, 5, 121--152

  33. [41]

    Construction and interference in learning from multiple representation

    Schnotz, W.; Bannert, M. Construction and interference in learning from multiple representation. Learning and Instruction 2003, 13, 141--156

  34. [42]

    C.; McNamara, D

    Graesser, A. C.; McNamara, D. S.; Louwerse, M. M.; Cai, Z. Coh-Metrix: Analysis of text on cohesion and language. Behavior Research Methods, Instruments, & Computers 2004, 36, 193--202

  35. [43]

    urzburg, 97074 W\

    Gibson, E. Linguistic complexity: Locality of syntactic dependencies. Cognition 1998, 68, 1--76 mcitethebibliography JChemEd_LLM_Benchmark.tex0000664000000000000000000010305215054271360014221 0ustar rootroot [journal=jceda8,manuscript=article,layout=twocolumn] achemso [utf8] i...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.