REVIEW 4 major objections 6 minor 1 cited by
Performance of ChatGPT on tasks involving physics visual representations: the case of the Brief Electricity and Magnetism Assessment
T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read ChatGPT-4o answers 67% of BEMA items correctly when scored by meaning, beating the 53.4% student average, but fails on spatial and visual tasks such as the right-hand rule.
desk verdict Solid, honest evaluation of ChatGPT-4/o on BEMA with a useful failure taxonomy; the headline claim of outperforming students overreaches because it compares meaning-coded chatbot answers with normally scored student answers. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is a two-pass scoring scheme. Each response is first coded by the letter the chatbot states, then recoded by the meaning of the answer's text; comparing the two separates errors in reading the answer-option layout from errors in physics reasoning. For the 14 low-scoring items, the authors further code each chain-of-thought explanation for three difficulty types—visual interpretation, physics-law statement, and spatial coordination—and for the seven right-hand-rule items they track whether the rule is stated, stated correctly, and applied correctly. The Brief Electricity and Magnetism Assessment itself, with 31 items and 30 visual representations, provides the stimulus set that makes this analysis possible.
What would settle it
Run ChatGPT-4o on BEMA items with answer-option labels randomly permuted across trials; if the letter-meaning mismatch rate stays high and the errors track particular letter positions rather than particular images, then the mismatch is driven by answer-letter bias rather than visual layout.
Extended reading notes
Core claim
On its own terms, the paper claims that ChatGPT-4o has crossed a threshold: it now outperforms an average university student on a standard conceptual electromagnetism inventory, and it clearly improves on its predecessor's ability to interpret physics diagrams. Yet this overall competence coexists with a distinctive failure mode. When answers are rescored by the meaning of what the model writes rather than the letter it states, ChatGPT-4o's performance rises from 59.8% to 67.0% and ChatGPT-4's from 50.2% to 60.7%, with letter-meaning mismatches occurring in 10.8% of ChatGPT-4o's responses and 26.8% of ChatGPT-4's; the authors attribute most of this gap to the chatbot misreading the spatial arrangement of answer-option lists, while acknowledging that answer-letter bias cannot be excluded. On the 14 items with below-average performance, qualitative coding of the model's chain-of-thought responses finds vision errors in 32% of responses, spatial-coordination errors in 60% on most of those items, and incorrect physics-law statements in 14%. The right-hand rule appears in seven of the weak items, and there the model's average is 35%; it usually states the rule correctly in most cases but misapplies it in 41% of the responses that invoke it.
Load-bearing premise
The paper treats the gap between letter-coded and meaning-coded answers as evidence of visual misinterpretation of answer-option layouts, but the authors acknowledge they cannot rule out a built-in model bias toward particular answer letters, which could also explain the gap.
Editorial extensions
If this is right
- Educators should not rely on ChatGPT-4o as a primary tutor for electricity and magnetism topics that hinge on spatial reasoning, because it frequently states correct rules and then applies them to the wrong geometry.
- The letter-meaning mismatch implies that a chatbot can reason correctly and still pick the wrong multiple-choice option when answer choices are arranged visually, which test designers can exploit to build chatbot-resistant items.
- ChatGPT-4o's error profile is sufficiently different from typical student errors that it is limited as a realistic model of a student for generating synthetic misconception data or for practicing Socratic teaching.
- Right-hand-rule tasks are a dependable weak spot, with an average of 35% correct across seven items, making them good candidates for targeted assessment design and for benchmarking future multimodal models.
Reading between the lines
- A randomized experiment that shuffles answer-option labels across trials would settle whether the letter-meaning gap is visual-layout misreading or letter-position bias; the paper's own caveat about answer-index bias leaves this open.
- The spatial-coordination failures suggest a general limitation of today's multimodal language models: they do not seem to hold a persistent geometric model of a scene, so other three-dimensional physics tasks beyond electromagnetism should show similar breakdowns.
- The same three-category coding scheme could be applied to other concept inventories, such as the Force Concept Inventory, to map which representation types each model generation handles and to compare models released over time.
- If the letter-meaning gap is truly visual layout rather than content misunderstanding, then simply reformatting multiple-choice answer options into a single column could raise a chatbot's effective score without any model improvement.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript evaluates ChatGPT-4 and ChatGPT-4o on the Brief Electricity and Magnetism Assessment (BEMA), a 31-item multiple-choice conceptual inventory rich in visual representations such as vector fields, circuit diagrams, and graphs. The authors submitted screenshots of the test items to the two chatbots, ran repeated trials (60 for ChatGPT-4, 30 for ChatGPT-4o), scored responses both by the stated answer letter and by the content of the reasoning ('meaning'), and compared performance with a published student dataset (N = 12,214). They also qualitatively coded 420 ChatGPT-4o responses on 14 low-scoring items into three difficulty categories (visual interpretation, physics-law errors, and spatial coordination/application) and conducted a focused analysis of right-hand-rule tasks. The central quantitative findings are that ChatGPT-4o scores 59.8% by letter and 67.0% by meaning, while ChatGPT-4 scores 50.2% and 60.7%, respectively; the qualitative analysis identifies persistent difficulties with visual and spatial reasoning.
Significance. If the results hold, the study is a useful contribution to the emerging literature on large multimodal models in physics education. Its strengths include repeated sampling with reported standard errors, an open data repository, a clearly described qualitative coding scheme with 89% intercoder agreement, and a focused analysis of right-hand-rule tasks that yields specific, actionable findings for educators. The identification of three distinct difficulty types and their relative prevalence is valuable for designing chatbot-resistant assessments and for understanding LMM limitations. However, the headline comparison to the student sample needs to be re-framed and statistically grounded, and the vision-specific interpretation of the letter-meaning mismatch requires qualification.
major comments (4)
- [IV.B and Abstract] The claim that ChatGPT-4o outperforms the student sample conflates two different scoring rules. The 67.0% figure is a meaning-coded score: the chatbot's reasoning content is matched to answer options regardless of the letter it states, and this re-coding changes 10.8% of ChatGPT-4o's responses (Secs. III.C and IV.A). The student scores from Wheatley et al. [14] are standard BEMA option scores, and the manuscript does not report whether those student scores were obtained with the basic or advanced scoring key, whereas the chatbot was scored only with the basic key (Sec. III.C footnote). The letter-coded chatbot score (59.8%) is the only like-for-like comparison, and it is not significance-tested, leaving the reported 'exceeds the average score for students' unsupported by any inferential statistic. Please either compare like-for-like with a significance test that respects the 31-item structure, or explicitly state that the 67.0% is a nonstandard 'reasoning-coded' score that cannot be directly benchmarked against published student averages.
- [IV.A and Abstract] The abstract's statement that ChatGPT-4o 'demonstrates improvements in ... vision interpretation ability' over ChatGPT-4 is not uniquely supported by the letter-vs-meaning mismatch data. The reduction in mismatch from 26.8% to 10.8% could reflect either improved reading of answer-option layouts or reduced answer-index bias, and the manuscript itself says in Sec. IV.A that 'our analysis cannot exclude other possible explanations' including the bias proposed by Zheng et al. [91]. Since the vision-specific interpretation is used to frame part of the analysis and the abstract, the claim should be softened or accompanied by a control analysis (e.g., shuffling answer letters or separating spatially arranged versus simple option lists).
- [IV.B and IV.A.3] The model-to-model and model-to-student comparisons are presented as point estimates without uncertainty quantification. The chatbot scores have reported item-level standard errors, yet the differences (e.g., 59.8% vs 53.4% or 67.0% vs 60.7%) are not accompanied by confidence intervals or significance tests. Because the two chatbots were run with different numbers of iterations (60 vs 30) and the student dataset is much larger, a formal comparison that respects the 31-item structure (e.g., a paired test on item scores with appropriate variance) is needed to support the performance claims.
- [III.C] The meaning-coding procedure, which is central to the main outperformance figure, is not checked for reliability. The qualitative difficulty coding in Sec. III.D reports 89% intercoder agreement, but no such check is reported for the meaning coding, even though the authors state that 10.8% of ChatGPT-4o responses (and 26.8% of ChatGPT-4 responses) are affected by the letter-meaning mismatch. A second coder's independent application of the meaning-coding rule to a subset of responses should be reported, or the coding should be described as an interpretive judgment whose reliability is unknown.
minor comments (6)
- [III.B] There is a typo in the sentence about ChatGPT-4o iterations: 'the largest uncertainties for item perfoemance' should be 'performance'.
- [VII] 'ChatGPT4o' appears without a hyphen in 'ChatGPT4o's (59.8%)'; standardize the model name throughout.
- [VII] The letter-coded ChatGPT-4o score (59.8%) is described as 'similar' to the student average (53.4%); a 6.4-percentage-point gap is not trivially similar, so specify whether this means 'not statistically different' or 'of the same order of magnitude'.
- [IV.A.3] The statement that 'ChatGPT-4o outperforms ChatGPT-4 on nearly half of the items' is vague; report the actual number of items (16 out of 31) as done in Sec. IV.A.2.
- [V.C] The sentence 'This type of difficulty was present in answers to all of the survey items, except items 8, 9, and 10...' refers to the 14 analyzed items, not the full 31-item survey; rephrase to avoid ambiguity.
- [Figure 1] The response text in the right panel is small; consider enlarging it or adding a transcript in the caption for readability.
Circularity Check
No circular reasoning found; the evaluation is anchored to external BEMA answer keys and a published student dataset, with caveats about scoring comparability acknowledged rather than hidden.
full rationale
I walked the paper's derivation chain and found no step in which a claimed prediction is equivalent to an input by construction, no fitted parameter renamed as a prediction, and no load-bearing argument that reduces to an unverified self-citation. The quantitative results (50.2%, 59.8%, 60.7%, 67.0%, and the 53.4% student comparison) are all generated by applying externally defined scoring rules to externally sourced materials: the BEMA test, its basic scoring key, and the published student sample of Wheatley et al. [14]. The two coding schemes, by letter and by meaning, are explicitly defined and independently applied to the same chatbot responses; the letter-meaning mismatch is an observed empirical quantity, not a definitional identity. The paper openly acknowledges in Sec. IV.A and Sec. VI.A that the mismatch could also reflect answer-index bias cited from Zheng et al. [91], which weakens a causal interpretation but does not make the measured performance circular. The qualitative difficulty categories (vision, physics-law errors, spatial coordination) were developed from the 420 analyzed responses with reported intercoder agreement, not imported from the authors' prior work; prior self-citations [12,39,40] are used for methodological continuity and comparison, not as the source of the present findings. No uniqueness theorem or ansatz is smuggled in via citation. The only substantive concern, that the headline outperformance compares meaning-coded chatbot scores with conventionally scored student scores, is a validity and fairness issue about like-for-like comparison, not circularity, because neither quantity is derived from the other or from a fitted parameter. Overall, the central claims are self-contained against external benchmarks, so the circularity score is 0.
Assumptions & free parameters
free parameters (1)
- Qualitative inclusion threshold =
67%
assumptions (3)
- domain assumption BEMA item scores, with the basic scoring key, are comparable to the student average from Wheatley et al. [14].
- domain assumption The chain-of-thought text generated by ChatGPT-4o reflects the model's actual processing, so coding it reveals underlying difficulties.
- domain assumption Repeated submissions to the web interface in separate chat windows are independent Bernoulli trials.
Cite this review
Pith. "Pith review of Performance of ChatGPT on tasks involving physics visual representations: the case of the Brief Electricity and Magnetism Assessment." pith.science (2026). https://pith.science/paper/W4VABBGH
@misc{pith2026241210019,
author = {Pith},
title = {Pith review of: Performance of ChatGPT on tasks involving physics visual representations: the case of the Brief Electricity and Magnetism Assessment},
year = {2026},
howpublished = {\url{https://pith.science/paper/W4VABBGH}},
note = {Machine review of arXiv:2412.10019}
}
read the original abstract
Artificial intelligence-based chatbots are increasingly influencing physics education due to their ability to interpret and respond to textual and visual inputs. This study evaluates the performance of two large multimodal model-based chatbots, ChatGPT-4 and ChatGPT-4o on the Brief Electricity and Magnetism Assessment (BEMA), a conceptual physics inventory rich in visual representations such as vector fields, circuit diagrams, and graphs. Quantitative analysis shows that ChatGPT-4o outperforms both ChatGPT-4 and a large sample of university students, and demonstrates improvements in ChatGPT-4o's vision interpretation ability over its predecessor ChatGPT-4. However, qualitative analysis of ChatGPT-4o's responses reveals persistent challenges. We identified three types of difficulties in the chatbot's responses to tasks on BEMA: (1) difficulties with visual interpretation, (2) difficulties in providing correct physics laws or rules, and (3) difficulties with spatial coordination and application of physics representations. Spatial reasoning tasks, particularly those requiring the use of the right-hand rule, proved especially problematic. These findings highlight that the most broadly used large multimodal model-based chatbot, ChatGPT-4o, still exhibits significant difficulties in engaging with physics tasks involving visual representations. While the chatbot shows potential for educational applications, including personalized tutoring and accessibility support for students who are blind or have low vision, its limitations necessitate caution. On the other hand, our findings can also be leveraged to design assessments that are difficult for chatbots to solve.
Figures
Forward citations
Cited by 1 Pith paper
-
Multilingual Performance of a Multimodal Artificial Intelligence System on Multisubject Physics Concept Inventories
GPT-4o averaged 71% on English physics concept inventories, outperformed average post-instruction undergraduates in most subjects but not laboratory skills, and scored far worse on image-dependent items and in non-Wes...
Reference graph
Works this paper leans on
-
[1]
The difference in performance of ChatGPT-4 when using the two coding systems When coding ChatGPT-4’s responses by letter, the chatbot scored an average of 50.2% on the test. Passing from coding by letter to coding by meaning, its performance improved by 10.5 percentage points, indicating its ability to extract and process relevant information from graphic...
-
[2]
The difference in performance of ChatGPT-4o when using the two coding systems When coding ChatGPT-4o’s responses by letter, the chatbot scored an average of 59.8% on the test. Switching to coding by meaning resulted in a 7.2 percentage point improvement, bringing the average performance to 67.0%. Figure 3 presents the scores for individual items and the o...
-
[3]
ChatGPT-4o as a model of a student As we lack corresponding student reasoning data, we are unable to systematically compare students’ and chatbots’ approaches to solving different tasks on BEMA. However, we can see that many errors made by the chatbot would be quite atypical for human students, based on our experience as physics instructors. One striking ...
-
[10]
S. Shetye, An Evaluation of Khanmigo, a Generative AI Tool, as a Computer-Assisted Language Learning App, Studies in Applied Linguistics and TESOL 24, 1 (2024). [11] S. Steinert, K. E. Avila, J. Kuhn, and S. Küchemann, Using GPT-4 as a guide during inquiry-based learning, The Physics Teacher 62, 618 (2024). [12] G. Polverini and B. Gregorcic, Performance ...
arXiv 2024
-
[35]
K. A. Pimbblet and L. Morrell, Can ChatGPT pass a physics degree? Making a case for reformation of assessment of undergraduate degrees, Eur. J. Phys. (2024). [36] K. D. Wang, E. Burkholder, C. Wieman, S. Salehi, and N. Haber, Examining the potential and pitfalls of ChatGPT in science and engineering problem-solving, Front. Educ. 8, (2024). [37] T. Kumar a...
arXiv 2024
-
[60]
Z. Chen and T. Wan, Using Large Language Models to Assign Partial Credit to Students’ Explanations of Problem-Solving Process: Grade at Human Level Accuracy with Grading Confidence Index and Personalized Student-Facing Feedback, arXiv:2412.06910. [61] J. R. Aguilar-Mejía, S. Tejeda, C. V. Ramirez-Lopez, and C. L. Garay-Rondero, Design and Use of a Chatbot...
work page Pith review arXiv 2010
-
[83]
T. S. Volkwyn, J. Airey, B. Gregorcic, F. Heijkensköld, and C. Linder, Physics Students Learning about Abstract Mathematical Tools When Engaging with “Invisible” Phenomena, in 2017 Physics Education Research Conference Proceedings (American Association of Physics Teachers, Cincinnati, OH, 2018), pp. 408–411. [84] M. B. Kustusch, Assessing the impact of re...
work page Pith review arXiv 2016
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.