REVIEW 3 major objections 4 minor 113 references
Assessing AI in Introductory Physics Problem Solving
T0 review · 3 major / 4 minor · reviewed 2026-08-02 · deepseek-v4-flash
Pith's one-line read The paper reports that the o4-mini model solves about 90% of standard introductory physics problems, but accuracy drops from 96% on text-only problems to 79% when images must be interpreted, and declines with problem difficulty.
desk verdict A useful large benchmark of o4-mini on Halliday & Resnick, with plausible modality and difficulty effects, but sloppy reporting (effort contradiction, odds/probability mix-up) and a grading-validation gap limit the current credibility of the exact percentages. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The evaluation pipeline: 1,203 problem-answer pairs extracted from the textbook, formatted in LaTeX with unified units, images as PNG screenshots; each problem run five times through o4-mini; solutions scored right/wrong by a large language model (GPT-5) against the textbook answer key, with a 600-item human audit of the grader. Statistical analysis uses logistic regression and generalized estimating equations to assess the difficulty-accuracy relationship.
What would settle it
Take a random sample of, say, 300 problems from the dataset, have two independent human physics instructors score the o4-mini solutions blind to the AI grader's scores, and compare the resulting accuracy estimates to the paper's 0.90, 0.96, and 0.79 figures; a substantial disagreement would overturn the central claim.
Extended reading notes
Core claim
The model o4-mini attains an overall accuracy of 0.90±0.03 on 1,203 odd-numbered problems from a standard introductory physics textbook, with a marked modality gap (0.96 text-only vs 0.79 text+image) and a monotonic decline across difficulty levels (0.94 easy, 0.88 medium, 0.84 hard). Logistic regression and generalized estimating equations confirm the difficulty effect, with odds of a correct response reduced by 51% for medium and 64% for hard relative to easy. The model's output effort roughly doubles for medium and hard problems. These results indicate that a current reasoning model can solve most standard introductory physics problems, but performance remains strongly constrained by visu
Load-bearing premise
Every accuracy figure depends on the AI grader being reliable: only 600 of 6,015 scores were spot-checked, and the small error rate's direction is unknown, so if the grader systematically credits flawed solutions the 90%, 96%, and 79% numbers are too high.
Editorial extensions
If this is right
- Students using such models on text-only homework can expect mostly correct solutions; image-based problems are much less reliable.
- The difficulty gradient implies AI could serve as an objective difficulty classifier for problem banks, as the paper itself suggests.
- Accuracy is stable across mechanics, electromagnetism, quantum theory, and other topics, so the model's weakness is not tied to specific content.
- The modality gap suggests that improving multimodal grounding is the next barrier for AI physics problem solving.
- Teachers and students should treat the high overall accuracy with caution because it masks systematic weaknesses on visual and hard problems.
Reading between the lines
- One implication the paper leaves implicit is that the proposed 'computational difficulty measure' is only as meaningful as the textbook's own difficulty labels; a natural test is to compare model accuracy against student success rates on the same problems.
- Because only odd-numbered problems with provided answers were used, the 90% figure excludes open-ended, drawing, or explanation problems; the result likely overstates the model's ability on the full range of textbook tasks.
- The direction of the 14 grader errors in the 600-item audit is not reported; if most were cases where GPT-5 marked an incorrect solution as correct, the true accuracy could be meaningfully lower than 90%.
- Re-running this exact 1,203-problem set on successor reasoning models would yield a direct longitudinal measure of whether the modality gap narrows over time.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript evaluates OpenAI's o4-mini on N=1203 odd-numbered end-of-chapter problems from Halliday and Resnick's Fundamentals of Physics, with each problem solved five times (6,015 runs). The authors report overall accuracy 0.90±0.03, text-only accuracy 0.96±0.03 versus 0.79±0.04 for problems requiring images, and accuracy declining from 0.94 (easy) to 0.88 (medium) to 0.84 (hard). They test the difficulty trend with logistic regression and generalized estimating equations, and also report model effort in output tokens. The conclusions are that current reasoning LLMs can solve most standard introductory problems but remain materially weaker on multimodal and harder problems.
Significance. If valid, this is a useful, large-scale benchmark of a reasoning model on a canonical introductory textbook: the design includes repeated runs per problem, explicit model checkpoint and prompting details, an external answer-key ground truth, and clustered regression for the difficulty analysis. The reported 17-point text-versus-image gap and the difficulty gradient are plausible and of practical interest to PER and to AI-in-education researchers. However, the central accuracy estimates depend on LLM grading whose validation is underreported, and the paper contains unresolved arithmetic inconsistencies in the dataset counts. Because the scoring disagreement rate is comparable to the reported difficulty steps, the grading issue must be resolved before the quantitative claims can be accepted as stated.
major comments (3)
- [Section II (AI-based scoring)] All headline figures (0.90 overall, 0.96 text, 0.79 image, 0.94/0.88/0.84 by difficulty) are produced by GPT-5 grading. The 600-run validation is not described (who re-scored, using what rubric?), is not stratified by modality or difficulty, and only reports 14 disagreements without a confusion matrix or error direction. Absorbing 0.023 as a 'systematic error' into the error bars does not correct a possible unidirectional bias. Since the modality gap (0.17) and adjacent difficulty steps (0.06, 0.04) are of the same scale as the raw disagreement rate, the authors must report false-positive and false-negative rates by condition and either correct the point estimates or provide a quantitative bound on grading bias.
- [Section II, Fig. 1 and Tables I, II, VI] The dataset arithmetic is inconsistent. Fig. 1 gives 1,225 problem–answer pairs and then N=1,203; summing the two volume tables yields 1,205 problems (410+380 text, 208+207 image-based) with difficulty totals 507/608/90, while Table VI reports 507/607/89. Additionally, the text says 411 images were extracted, whereas Fig. 1 counts 411 figures + 12 tables = 423 PNG images, with Nimg=415. Because every accuracy and regression uses these denominators, the exclusion rules need to be stated explicitly and the tables corrected accordingly.
- [Appendix B and Conclusion] The text interprets odds ratios as probability reductions: 'the probability of solving a Medium-level problem is smaller ... by 0.51' and by 0.27 for Hard. These are reductions in odds, not in probability: an odds ratio of 0.49 does not imply a 0.51 decrease in probability. This overstates the effect sizes. The authors should restate the results in terms of odds ratios with confidence intervals, or convert to predicted probabilities at a specified baseline.
minor comments (4)
- [Section III and Conclusion] Figure 4 discussion says the output size 'decreases correspondingly' as difficulty increases, while the Conclusion says the effort 'almost doubled' for Medium and Hard problems. These statements contradict each other and should be reconciled.
- [Tables III–IV captions] The captions do not specify whether the bottom-row averages are unweighted means of chapter accuracies or pooled accuracies over all problems in the volume. Please clarify.
- [Reference [29]] The author string 'D. E. Trowbridge1981' appears malformed; it should be 'D. E. Trowbridge and L. C. McDermott' (or similar).
- [Overall] No data or code availability statement is provided. Given the novelty of using an LLM grader at scale, the problem-level scores and scoring prompts should be released to make the quantitative claims checkable.
Circularity Check
No significant circularity: accuracies are measured against an external textbook answer key, and no prediction is constructed from a fitted input.
full rationale
The paper is an empirical benchmark, not a derivation. The central claims—overall accuracy 0.90, text-only 0.96 vs. image-based 0.79, and difficulty gradient—are defined by Equation (2) as averages of binary scores assigned against Halliday & Resnick's textbook answers. The textbook answer key is external to the model and to the authors, so the accuracy numbers are not fitted parameters renamed as predictions. The difficulty analysis regresses model scores on the textbook authors' difficulty labels; the predictor is independent of the outcome, so the regression is a measurement of a correlation, not a self-defined prediction. The only potentially load-bearing auxiliary element is the use of GPT-5 as grader, but grading is anchored to the external answer key and a 600-sample validation check with 14 disagreements is reported; this converts the concern into a grader-reliability / measurement-validity issue, not a circularity of construction. There are no self-citations by the authors, no self-referential uniqueness theorem, and no ansatz smuggled in via citation. The paper's own limitation statement about assuming the answer key has no errors is an explicit assumption, not a circular step. Accordingly, no circular step can be exhibited by quoting equations that reduce to their inputs, and the appropriate score is 0.
Assumptions & free parameters
free parameters (3)
- β1 (Medium vs Easy difficulty coefficient) =
-0.7072
- β2 (Hard vs Easy difficulty coefficient) =
-1.0294
- Systematic scoring error =
0.023
assumptions (5)
- domain assumption Textbook answer key is error-free and serves as ground truth.
- ad hoc to paper GPT-5 can reliably grade o4-mini solutions against the answer key, with error rate fixed at 0.023.
- domain assumption Difficulty labels from Halliday & Resnick (Easy, Medium, Hard) are meaningful ordinal categories.
- standard math In tests 1–2 difficulty is equally spaced on the logit scale (1,2,3).
- domain assumption The textbook's odd-numbered problems with answers are representative of the curriculum.
Cite this review
Pith. "Pith review of Assessing AI in Introductory Physics Problem Solving." pith.science (2026). https://pith.science/paper/K2FBRLGI
@misc{pith2026260714303,
author = {Pith},
title = {Pith review of: Assessing AI in Introductory Physics Problem Solving},
year = {2026},
howpublished = {\url{https://pith.science/paper/K2FBRLGI}},
note = {Machine review of arXiv:2607.14303}
}
read the original abstract
Reasoning or inference-scaling models are the new generation of Large Language Models (LLMs) capable of complex problem solving. To investigate their problem-solving capability in physics, we evaluated model o4-mini by OpenAI on solving traditional, end-of-chapter problems from Halliday and Resnick's "Fundamentals of Physics," spanning core topics in the undergraduate physics curriculum. Performance was analyzed across modality and problem difficulty. The model solved the problems with overall accuracy of about 90%, but performance depended strongly on representation: accuracy was much higher on text-only problems (96%) than on problems requiring coordinated interpretation of text and images (79%). Accuracy also declined significantly as the problem difficulty increased from low to medium to high. These results show that state-of-the-art LLMs can solve much of the standard introductory physics problems, but that their performance remains uneven and constrained by problem modality and problem difficulty.
Figures
Reference graph
Works this paper leans on
-
[1]
What is the problem-solving capability of AI models on standard topics in the intro- ductory physics curriculum?
-
[2]
How does this capability depend on modality—language and vision?
-
[3]
o4-mini” by OpenAI was selected due to its design and affordability. Together with a larger model “o3,
How does this capability depend on different levels of difficulty in the problems solved? An additional property of a model when solving analytical problems is the amount of “effort” it exerts while engaging in this process. This quantity may be defined as the number of tokensproduced by the model in generating the full solution to a given problem. It pro...
-
[4]
Measurement 13 2 8 5 2
-
[5]
Motion Along a Straight Line 25 10 15 16 4
-
[6]
Motion in Two and Three Dimensions 33 8 13 22 6
-
[7]
Force and Motion–I 18 16 11 20 3
-
[8]
Force and Motion–II 15 14 10 16 3
Show all 113 references
-
[9]
Kinetic Energy and Work 17 9 9 15 2
-
[10]
Potential Energy and Conservation of Energy 12 18 9 15 6
-
[11]
Center of Mass and Linear Momentum 22 18 13 23 4
-
[12]
Rotation 25 9 14 16 4
-
[13]
Rolling, Torque, and Angular Momentum 18 17 16 14 5
-
[14]
Equilibrium and Elasticity 6 18 10 11 3
-
[15]
Gravitation 26 8 18 13 3
-
[16]
Oscillations 20 12 17 12 3
-
[17]
Waves–I 23 7 13 15 2
-
[18]
Waves–II 29 6 16 15 4
-
[19]
Temperature, Heat, and the First Law of Thermodynamics 23 10 14 17 2
-
[20]
The Kinetic Theory of Gases 26 6 13 17 2
-
[21]
The additional errors coming from run variation within each chapter problem were propagated, adding to the total statistical error in our chapter ac- curacy results
Entropy and the Second Law of Thermodynamics 16 6 8 11 3 410 208 258 299 61 culated by SEMσ chap/√Ni. The additional errors coming from run variation within each chapter problem were propagated, adding to the total statistical error in our chapter ac- curacy results. Finally, ...
-
[22]
Coulomb’s Law 11 8 6 10 3
-
[23]
Electric Fields 15 13 10 17 1
-
[24]
Gauss’ Law 16 12 12 13 3
-
[25]
Electric Potential 22 12 12 19 3
-
[26]
Capacitance 15 7 11 16 1
-
[27]
Current and Resistance 24 3 10 16 1
-
[28]
Circuits 13 20 12 18 3
-
[29]
Magnetic Fields 23 9 17 14 1
-
[30]
Magnetic Fields Due to Currents 12 20 13 16 3
-
[31]
Induction and Inductance 19 19 18 17 3
-
[32]
Electromagnetic Oscillations and Alternating Current 20 10 15 15
-
[33]
Maxwell’s Equations; Magnetism of Matter 16 9 9 15 1
-
[34]
Electromagnetic Waves 19 14 15 17 1
-
[35]
Interference 13 8 12 18 1
-
[36]
Diffraction 29 6 21 14
-
[37]
Relativity 26 3 11 17 1
-
[38]
Photons and Matter Waves 33 2 10 23 2
-
[39]
More About Matter Waves 17 6 9 14
-
[40]
RESUL TS Tables III and IV show the scores achieved by model o4-mini on all problems from 40 chapters from Hallidayet al.as graded by model GPT-5
All About Atoms 25 4 18 11 380 207 249 309 29 III. RESUL TS Tables III and IV show the scores achieved by model o4-mini on all problems from 40 chapters from Hallidayet al.as graded by model GPT-5. The column Text indicates the model accuracy on text-only problems, while colum...
-
[41]
Measurement 0.97(0.07) 1.00(0.00) 0.97(0.07)
-
[42]
Motion Along a Straight Line 1.00(0.00) 0.46(0.25) 0.85(0.10)
-
[43]
Vectors 0.98(0.06) 0.48(0.32) 0.86(0.12)
-
[44]
Motion in Two and Three Dimensions 0.98(0.05) 0.75(0.19) 0.94(0.07)
-
[45]
Force and Motion–I 0.94(0.08) 0.82(0.17) 0.89(0.11)
-
[46]
Force and Motion–II 1.00(0.00) 0.94(0.10) 0.97(0.06)
-
[47]
Kinetic Energy and Work 0.99(0.05) 0.78(0.20) 0.92(0.09)
-
[48]
Potential Energy and CoE 1.00(0.00) 0.84(0.16) 0.91(0.11)
-
[49]
Center of Mass and Linear Momentum 1.00(0.00) 0.82(0.14) 0.92(0.08)
-
[50]
Rotation 0.94(0.08) 0.71(0.20) 0.88(0.09)
-
[51]
Rolling, Torque, and AM 0.91(0.11) 0.75(0.15) 0.83(0.11)
-
[52]
Equilibrium and Elasticity 0.93(0.13) 0.59(0.23) 0.68(0.19)
-
[53]
Gravitation 0.99(0.04) 1.00(0.00) 0.99(0.03)
-
[54]
Fluids 0.98(0.05) 0.87(0.16) 0.95(0.07)
-
[55]
Oscillations 0.98(0.06) 0.97(0.08) 0.98(0.06)
-
[56]
Waves–I 0.93(0.09) 0.17(0.19) 0.74(0.12)
-
[57]
Waves–II 0.90(0.11) 1.00(0.00) 0.92(0.10)
-
[58]
Temperature, Heat, and the FLT 0.97(0.06) 0.84(0.18) 0.93(0.08)
-
[59]
The Kinetic Theory of Gases 0.93(0.10) 0.83(0.19) 0.91(0.10)
-
[60]
Entropy and the SLT 0.94(0.09) 0.93(0.13) 0.94(0.09) Average:0.96(0.05) 0.79(0.10) 0.90(0.06) 12 TABLE IV. Vol. 2 scores by chapter and modality. Some chapter titles were abbreviated for visual purposes. (AC: Alternating Current, MoM: Magnetism of Matter) Chapter T ext T ext +...
-
[61]
Coulomb’s Law 0.96(0.08) 0.95(0.11) 0.96(0.08)
-
[62]
Electric Fields 0.95(0.09) 0.82(0.20) 0.89(0.13)
-
[63]
Gauss’ Law 0.94(0.10) 0.80(0.17) 0.88(0.11)
-
[64]
Electric Potential 0.97(0.07) 0.92(0.11) 0.95(0.07)
-
[65]
Capacitance 1.00(0.00) 0.78(0.20) 0.90(0.11)
-
[66]
Current and Resistance 1.00(0.00) 1.00(0.00) 1.00(0.00)
-
[67]
Circuits 1.00(0.00) 0.88(0.12) 0.93(0.08)
-
[68]
Magnetic Fields 0.94(0.09) 0.87(0.17) 0.92(0.10)
-
[69]
Magnetic Fields Due to Currents 0.97(0.08) 0.58(0.22) 0.72(0.17)
-
[70]
Induction and Inductance 0.95(0.08) 0.89(0.12) 0.92(0.08)
-
[71]
Electromagnetic Oscillations and AC 1.00(0.00) 0.96(0.09) 0.99(0.04)
-
[72]
Maxwell’s Equations; MoM 0.90(0.11) 0.76(0.22) 0.85(0.12)
-
[73]
Electromagnetic Waves 1.00(0.00) 0.73(0.20) 0.88(0.10)
-
[74]
Images 1.00(0.00) 0.67(0.27) 0.89(0.11)
-
[75]
Interference 0.97(0.07) 0.86(0.13) 0.90(0.10)
-
[76]
Diffraction 0.90(0.10) 0.33(0.23) 0.81(0.11)
-
[77]
Relativity 0.95(0.07) 1.00(0.00) 0.96(0.06)
-
[78]
Photons and Matter Waves 0.84(0.12) 0.80(0.22) 0.84(0.13)
-
[79]
More About Matter Waves 0.98(0.06) 0.83(0.19) 0.94(0.08)
-
[80]
it appears that grouping 17 the items in terms of test objectives does not provide novel meaningful insights into the strengths and weaknesses of individual chatbots
All About Atoms 0.93(0.09) 0.90(0.18) 0.92(0.09) Average:0.95(0.05) 0.81(0.10) 0.90(0.06) 13 TABLE V. Accuracy results by physics topics. #T opic ChaptersN problems Accuracy 1 Mechanics 1–14 434 0.90(0.06) 2 Oscillations and Waves 15–17 96 0.89(0.06) 3 Thermodynamics and Kinet...
2000
-
[81]
Merriam-Webster, artificial intelligence (2026)
2026
-
[82]
OpenAI, GPT-4 Technical Report (2024), arXiv: 2303.08774
2024 arXiv
-
[83]
Polverini and B
G. Polverini and B. Gregorcic, Performance of ChatGPT on the test of understanding graphs in kinematics, Phys. Rev. Phys. Educ. Res.20, 10.1103/PhysRevPhysEducRes.20.010109 (2024)
2024 doi
-
[84]
Polverini, J
G. Polverini, J. Melin, E. ¨Onerud, and B. Gregorcic, Performance of ChatGPT on tasks involv- ing physics visual representations: The case of the brief electricity and magnetism assessment, Phys. Rev. Phys. Educ. Res.21, 10.1103/PhysRevPhysEducRes.21.010154 (2025)
2025 doi
-
[85]
Hestenes, M
D. Hestenes, M. Wells, and G. Swackhamer, Force concept inventory, The Physics Teacher 30, 10.1119/1.2343497 (1992)
1992 doi
-
[86]
R. R. Hake, Interactive-engagement versus traditional methods: A six-thousand-student sur- vey of mechanics test data for introductory physics courses, American Journal of Physics66, 10.1119/1.18809 (1998)
1998 doi
-
[87]
Aldazharova, G
S. Aldazharova, G. Issayeva, S. Maxutov, and N. Balta, Assessing AI’s problem solving in physics: Analyzing reasoning, false positives and negatives through the force concept inven- tory, Contemporary Educational Technology16, 10.30935/cedtech/15592 (2024)
2024 doi
-
[88]
R. J. Beichner, Testing student interpretation of kinematics graphs, American Journal of Physics62, 10.1119/1.17449 (1994)
1994 doi
-
[89]
Zavala, S
G. Zavala, S. Tejeda, P. Barniol, and R. J. Beichner, Modifying the test of understand- ing graphs in kinematics, Phys. Rev. Phys. Educ. Res.13, 10.1103/PhysRevPhysEdu- cRes.13.020111 (2017)
2017 doi
-
[90]
Polverini and B
G. Polverini and B. Gregorcic, Evaluating vision-capable chatbots in interpreting kinematics graphs: a comparative study of free and subscription-based models, Frontiers in Education9 (2024)
2024
-
[91]
Polverini and B
G. Polverini and B. Gregorcic, Multimodal large language models and physics visual tasks: Comparative analysis of performance and costs, European Journal of Physics46, 23 10.1088/1361-6404/ae03f8 (2025)
2025 doi
-
[92]
K. D. Wang, E. Burkholder, C. Wieman, S. Salehi, and N. Haber, Examining the potential and pitfalls of ChatGPT in science and engineering problem-solving, Frontiers in Education 8, 10.3389/feduc.2023.1330486 (2024)
2023
-
[93]
Kieser and P
F. Kieser and P. Wulff, Using Large Language Models to Probe Cognitive Constructs, Aug- ment Data, and Design Instructional Materials, inMachine Learning in Educational Sciences: Approaches, Applications and Advances, edited by M. S. Khine (Springer Nature Singapore,
-
[94]
J. C. Dunlap, R. Sissons, and R. Widenhorn, Descending an inclined plane with a large language model, Phys. Rev. Phys. Educ. Res.21, 10.1103/PhysRevPhysEducRes.21.010153 (2025)
2025 doi
-
[95]
Tschisgale, H
P. Tschisgale, H. Maus, F. Kieser, B. Kroehs, S. Petersen, and P. Wulff, Evaluating GPT- and reasoning-based large language models on Physics Olympiad problems: Surpassing human performance and implications for educational assessment, Phys. Rev. Phys. Educ. Res.21, 10.1103/6fm...
2025 doi
-
[96]
Kortemeyer, M
G. Kortemeyer, M. Babayeva, G. Polverini, R. Widenhorn, and B. Gregorcic, Multilingual performance of a multimodal artificial intelligence system on multisubject physics concept inventories, Phys. Rev. Phys. Educ. Res.21, 10.1103/98hg-rkrf (2025)
2025 doi
-
[97]
Horchani, ChatGPT’s problem-solving abilities in context-rich and traditional physics prob- lems, Physics Education60, 10.1088/1361-6552/adb473 (2025)
R. Horchani, ChatGPT’s problem-solving abilities in context-rich and traditional physics prob- lems, Physics Education60, 10.1088/1361-6552/adb473 (2025)
2025 doi
-
[98]
F. Yu, H. Wan, Q. Cheng, Y. Zhang, J. Chen, F. Han, Y. Wu, J. Yao, R. Hu, N. Ding, Y. Cheng, T. Chen, L. Bai, D. Zhou, Y. Luo, G. Cui, and P. Ye, HiPhO: How Far Are (M)LLMs from Humans in the Latest High School Physics Olympiad Benchmark? (2025), arXiv:2509.07894 [cs.AI]
2025
-
[99]
D. J. H. Chung, Z. Gao, Y. Kvasiuk, T. Li, M. M¨ unchmeyer, M. Rudolph, F. Sala, and S. C. Tadepalli, Theoretical physics benchmark (TPBench): A dataset and study of AI reasoning capabilities in theoretical physics, Machine Learning: Science and Technology6, 10.1088/2632- 2153...
2025 doi
-
[100]
K. Feng, Y. Zhao, Y. Liu, T. Yang, C. Zhao, J. Sous, and A. Cohan, PHYSICS: Benchmarking Foundation Models on University-Level Physics Problem Solving (2025), arXiv:2503.21821 [physics.ed-ph]. 24
2025 arXiv
-
[101]
Zheng, Q
S. Zheng, Q. Cheng, J. Yao, M. Wu, N. Ding, Y. Cheng, S. Hu, L. Bai, D. Zhou, G. Cui, and P. Ye, Scaling Physical Reasoning with the PHYSICS Dataset, arXiv:2506.00022 (2025), [Accepted by NeurIPS 2025]
2025
-
[102]
OpenAI, Introducing OpenAI o3 and o4-mini,https://openai.com/index/ introducing-o3-and-o4-mini/(2025)
2025
-
[103]
OpenAI, Introducing GPT-4.1 in the API,https://openai.com/index/gpt-4-1/(2025)
2025
-
[104]
OpenAI, Introducing GPT-5,https://openai.com/index/introducing-gpt-5/(2025)
2025
-
[105]
Halliday, R
D. Halliday, R. Resnick, and J. Walker,Fundamentals of Physics, 12th ed. (John Wiley & Sons, Inc., 2022)
2022
-
[106]
Wright, siunitx - A comprehensive (SI) units package (2025)
J. Wright, siunitx - A comprehensive (SI) units package (2025)
2025
-
[107]
Liang and S
K.-Y. Liang and S. L. Zeger, Longitudinal data analysis using generalized linear models, Biometrika73, 13 (1986)
1986
-
[108]
D. E. Trowbridge and L. C. McDermott, Investigation of student understanding of the concept of velocity in one dimension, American Journal of Physics48, 10.1119/1.12298 (1980)
1980 doi
-
[109]
D. E. Trowbridge1981 and L. C. McDermott, Investigation of student understanding of the concept of acceleration in one dimension, American Journal of Physics49, 10.1119/1.12525 (1981)
1981 doi
-
[110]
W. K. Adams and C. E. Wieman, Analyzing the many skills involved in solving complex physics problems, American Journal of Physics83, 10.1119/1.4913923 (2015)
2015 doi
-
[111]
J. L. Docktor, J. Dornfeld, E. Frodermann, K. Heller, L. Hsu, K. A. Jackson, A. Mason, Q. X. Ryan, and J. Yang, Assessing student written problem solutions: A problem-solving rubric with application to introductory physics, Phys. Rev. Phys. Educ. Res.12, 10.1103/Phys- RevPhysE...
2016 doi
-
[112]
A. M. Price, C. J. Kim, E. W. Burkholder, A. V. Fritz, and C. E. Wieman, A Detailed Characterization of the Expert Problem-Solving Process in Science and Engineering: Guidance for Teaching and Assessment, CBE Life Sci. Educ.20, 10.1187/cbe.20-12-0276 (2021)
2021 doi
-
[113]
A. Elby, A. Conte, E. Sohr, and J. Radoff, How do students use online homework solutions? (2025), AAPT Summer Meeting [Conference Presentation]. 25
2025
Reviewed August 2, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.