Pith. sign in

REVIEW 4 major objections 6 minor 1 cited by

Performance of ChatGPT on tasks involving physics visual representations: the case of the Brief Electricity and Magnetism Assessment

T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read ChatGPT-4o answers 67% of BEMA items correctly when scored by meaning, beating the 53.4% student average, but fails on spatial and visual tasks such as the right-hand rule.

desk verdict Solid, honest evaluation of ChatGPT-4/o on BEMA with a useful failure taxonomy; the headline claim of outperforming students overreaches because it compares meaning-coded chatbot answers with normally scored student answers. read the letter →

arxiv 2412.10019 v3 pith:W4VABBGH submitted 2024-12-13 physics.ed-ph

classification physics.ed-ph PACS 01.40.Fk
keywords ChatGPT-4omultimodallargelanguagemodelsBriefElectricityandMagnetismAssessmentphysicsvisualrepresentationsright-handrulespatialreasoningconceptinventorychain-of-thought
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

ChatGPT-4o solves 67.0% of the Brief Electricity and Magnetism Assessment when its answers are coded by meaning, beating the 53.4% average of a 12,214-student sample; ChatGPT-4 scores 60.7% under the same coding. The paper's main point, however, is not the overall score but the pattern beneath it. On the 14 items where ChatGPT-4o scores below the test average, its failed responses exhibit three kinds of difficulties: misreading the visual representations, stating incorrect physics laws, and failing to coordinate correctly stated laws with the spatial layout of the problem. Right-hand-rule tasks are especially weak, averaging 35% correct. The authors conclude that the chatbot is not yet reliable as a tutor or accessibility tool for visually rich physics, and that such spatial tasks can be used to design assessments that are hard for chatbots to answer.

What carries the argument

The central mechanism is a two-pass scoring scheme. Each response is first coded by the letter the chatbot states, then recoded by the meaning of the answer's text; comparing the two separates errors in reading the answer-option layout from errors in physics reasoning. For the 14 low-scoring items, the authors further code each chain-of-thought explanation for three difficulty types—visual interpretation, physics-law statement, and spatial coordination—and for the seven right-hand-rule items they track whether the rule is stated, stated correctly, and applied correctly. The Brief Electricity and Magnetism Assessment itself, with 31 items and 30 visual representations, provides the stimulus set that makes this analysis possible.

What would settle it

Run ChatGPT-4o on BEMA items with answer-option labels randomly permuted across trials; if the letter-meaning mismatch rate stays high and the errors track particular letter positions rather than particular images, then the mismatch is driven by answer-letter bias rather than visual layout.

Watch

Extended reading notes

Core claim

On its own terms, the paper claims that ChatGPT-4o has crossed a threshold: it now outperforms an average university student on a standard conceptual electromagnetism inventory, and it clearly improves on its predecessor's ability to interpret physics diagrams. Yet this overall competence coexists with a distinctive failure mode. When answers are rescored by the meaning of what the model writes rather than the letter it states, ChatGPT-4o's performance rises from 59.8% to 67.0% and ChatGPT-4's from 50.2% to 60.7%, with letter-meaning mismatches occurring in 10.8% of ChatGPT-4o's responses and 26.8% of ChatGPT-4's; the authors attribute most of this gap to the chatbot misreading the spatial arrangement of answer-option lists, while acknowledging that answer-letter bias cannot be excluded. On the 14 items with below-average performance, qualitative coding of the model's chain-of-thought responses finds vision errors in 32% of responses, spatial-coordination errors in 60% on most of those items, and incorrect physics-law statements in 14%. The right-hand rule appears in seven of the weak items, and there the model's average is 35%; it usually states the rule correctly in most cases but misapplies it in 41% of the responses that invoke it.

Load-bearing premise

The paper treats the gap between letter-coded and meaning-coded answers as evidence of visual misinterpretation of answer-option layouts, but the authors acknowledge they cannot rule out a built-in model bias toward particular answer letters, which could also explain the gap.

Editorial extensions

If this is right

  • Educators should not rely on ChatGPT-4o as a primary tutor for electricity and magnetism topics that hinge on spatial reasoning, because it frequently states correct rules and then applies them to the wrong geometry.
  • The letter-meaning mismatch implies that a chatbot can reason correctly and still pick the wrong multiple-choice option when answer choices are arranged visually, which test designers can exploit to build chatbot-resistant items.
  • ChatGPT-4o's error profile is sufficiently different from typical student errors that it is limited as a realistic model of a student for generating synthetic misconception data or for practicing Socratic teaching.
  • Right-hand-rule tasks are a dependable weak spot, with an average of 35% correct across seven items, making them good candidates for targeted assessment design and for benchmarking future multimodal models.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A randomized experiment that shuffles answer-option labels across trials would settle whether the letter-meaning gap is visual-layout misreading or letter-position bias; the paper's own caveat about answer-index bias leaves this open.
  • The spatial-coordination failures suggest a general limitation of today's multimodal language models: they do not seem to hold a persistent geometric model of a scene, so other three-dimensional physics tasks beyond electromagnetism should show similar breakdowns.
  • The same three-category coding scheme could be applied to other concept inventories, such as the Force Concept Inventory, to map which representation types each model generation handles and to compare models released over time.
  • If the letter-meaning gap is truly visual layout rather than content misunderstanding, then simply reformatting multiple-choice answer options into a single column could raise a chatbot's effective score without any model improvement.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The manuscript evaluates ChatGPT-4 and ChatGPT-4o on the Brief Electricity and Magnetism Assessment (BEMA), a 31-item multiple-choice conceptual inventory rich in visual representations such as vector fields, circuit diagrams, and graphs. The authors submitted screenshots of the test items to the two chatbots, ran repeated trials (60 for ChatGPT-4, 30 for ChatGPT-4o), scored responses both by the stated answer letter and by the content of the reasoning ('meaning'), and compared performance with a published student dataset (N = 12,214). They also qualitatively coded 420 ChatGPT-4o responses on 14 low-scoring items into three difficulty categories (visual interpretation, physics-law errors, and spatial coordination/application) and conducted a focused analysis of right-hand-rule tasks. The central quantitative findings are that ChatGPT-4o scores 59.8% by letter and 67.0% by meaning, while ChatGPT-4 scores 50.2% and 60.7%, respectively; the qualitative analysis identifies persistent difficulties with visual and spatial reasoning.

Significance. If the results hold, the study is a useful contribution to the emerging literature on large multimodal models in physics education. Its strengths include repeated sampling with reported standard errors, an open data repository, a clearly described qualitative coding scheme with 89% intercoder agreement, and a focused analysis of right-hand-rule tasks that yields specific, actionable findings for educators. The identification of three distinct difficulty types and their relative prevalence is valuable for designing chatbot-resistant assessments and for understanding LMM limitations. However, the headline comparison to the student sample needs to be re-framed and statistically grounded, and the vision-specific interpretation of the letter-meaning mismatch requires qualification.

major comments (4)
  1. [IV.B and Abstract] The claim that ChatGPT-4o outperforms the student sample conflates two different scoring rules. The 67.0% figure is a meaning-coded score: the chatbot's reasoning content is matched to answer options regardless of the letter it states, and this re-coding changes 10.8% of ChatGPT-4o's responses (Secs. III.C and IV.A). The student scores from Wheatley et al. [14] are standard BEMA option scores, and the manuscript does not report whether those student scores were obtained with the basic or advanced scoring key, whereas the chatbot was scored only with the basic key (Sec. III.C footnote). The letter-coded chatbot score (59.8%) is the only like-for-like comparison, and it is not significance-tested, leaving the reported 'exceeds the average score for students' unsupported by any inferential statistic. Please either compare like-for-like with a significance test that respects the 31-item structure, or explicitly state that the 67.0% is a nonstandard 'reasoning-coded' score that cannot be directly benchmarked against published student averages.
  2. [IV.A and Abstract] The abstract's statement that ChatGPT-4o 'demonstrates improvements in ... vision interpretation ability' over ChatGPT-4 is not uniquely supported by the letter-vs-meaning mismatch data. The reduction in mismatch from 26.8% to 10.8% could reflect either improved reading of answer-option layouts or reduced answer-index bias, and the manuscript itself says in Sec. IV.A that 'our analysis cannot exclude other possible explanations' including the bias proposed by Zheng et al. [91]. Since the vision-specific interpretation is used to frame part of the analysis and the abstract, the claim should be softened or accompanied by a control analysis (e.g., shuffling answer letters or separating spatially arranged versus simple option lists).
  3. [IV.B and IV.A.3] The model-to-model and model-to-student comparisons are presented as point estimates without uncertainty quantification. The chatbot scores have reported item-level standard errors, yet the differences (e.g., 59.8% vs 53.4% or 67.0% vs 60.7%) are not accompanied by confidence intervals or significance tests. Because the two chatbots were run with different numbers of iterations (60 vs 30) and the student dataset is much larger, a formal comparison that respects the 31-item structure (e.g., a paired test on item scores with appropriate variance) is needed to support the performance claims.
  4. [III.C] The meaning-coding procedure, which is central to the main outperformance figure, is not checked for reliability. The qualitative difficulty coding in Sec. III.D reports 89% intercoder agreement, but no such check is reported for the meaning coding, even though the authors state that 10.8% of ChatGPT-4o responses (and 26.8% of ChatGPT-4 responses) are affected by the letter-meaning mismatch. A second coder's independent application of the meaning-coding rule to a subset of responses should be reported, or the coding should be described as an interpretive judgment whose reliability is unknown.
minor comments (6)
  1. [III.B] There is a typo in the sentence about ChatGPT-4o iterations: 'the largest uncertainties for item perfoemance' should be 'performance'.
  2. [VII] 'ChatGPT4o' appears without a hyphen in 'ChatGPT4o's (59.8%)'; standardize the model name throughout.
  3. [VII] The letter-coded ChatGPT-4o score (59.8%) is described as 'similar' to the student average (53.4%); a 6.4-percentage-point gap is not trivially similar, so specify whether this means 'not statistically different' or 'of the same order of magnitude'.
  4. [IV.A.3] The statement that 'ChatGPT-4o outperforms ChatGPT-4 on nearly half of the items' is vague; report the actual number of items (16 out of 31) as done in Sec. IV.A.2.
  5. [V.C] The sentence 'This type of difficulty was present in answers to all of the survey items, except items 8, 9, and 10...' refers to the 14 analyzed items, not the full 31-item survey; rephrase to avoid ambiguity.
  6. [Figure 1] The response text in the right panel is small; consider enlarging it or adding a transcript in the caption for readability.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular reasoning found; the evaluation is anchored to external BEMA answer keys and a published student dataset, with caveats about scoring comparability acknowledged rather than hidden.

full rationale

I walked the paper's derivation chain and found no step in which a claimed prediction is equivalent to an input by construction, no fitted parameter renamed as a prediction, and no load-bearing argument that reduces to an unverified self-citation. The quantitative results (50.2%, 59.8%, 60.7%, 67.0%, and the 53.4% student comparison) are all generated by applying externally defined scoring rules to externally sourced materials: the BEMA test, its basic scoring key, and the published student sample of Wheatley et al. [14]. The two coding schemes, by letter and by meaning, are explicitly defined and independently applied to the same chatbot responses; the letter-meaning mismatch is an observed empirical quantity, not a definitional identity. The paper openly acknowledges in Sec. IV.A and Sec. VI.A that the mismatch could also reflect answer-index bias cited from Zheng et al. [91], which weakens a causal interpretation but does not make the measured performance circular. The qualitative difficulty categories (vision, physics-law errors, spatial coordination) were developed from the 420 analyzed responses with reported intercoder agreement, not imported from the authors' prior work; prior self-citations [12,39,40] are used for methodological continuity and comparison, not as the source of the present findings. No uniqueness theorem or ansatz is smuggled in via citation. The only substantive concern, that the headline outperformance compares meaning-coded chatbot scores with conventionally scored student scores, is a validity and fairness issue about like-for-like comparison, not circularity, because neither quantity is derived from the other or from a fitted parameter. Overall, the central claims are self-contained against external benchmarks, so the circularity score is 0.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

The paper's empirical claims rest on comparability of the external student benchmark, on the validity of chain-of-thought text as evidence of internal processing, and on independence of repeated chatbot runs. There are no fitted physics parameters or invented entities; the only hand-chosen quantity is the 67% cutoff for qualitative analysis.

free parameters (1)
  • Qualitative inclusion threshold = 67%
    Hand-chosen cutoff used in Sec. III.D to select below-average items for qualitative analysis. It affects which items enter the difficulty taxonomy and the reported prevalence of each error type.
assumptions (3)
  • domain assumption BEMA item scores, with the basic scoring key, are comparable to the student average from Wheatley et al. [14].
    The comparison in Sec. IV.B assumes the student score of 53.4% was obtained under equivalent scoring and administration; the paper does not verify this.
  • domain assumption The chain-of-thought text generated by ChatGPT-4o reflects the model's actual processing, so coding it reveals underlying difficulties.
    Sec. II.C and V treat CoT output as a window into the chatbot's abilities, despite undisclosed image-processing internals.
  • domain assumption Repeated submissions to the web interface in separate chat windows are independent Bernoulli trials.
    Sec. III.B treats each of 60 or 30 responses as independent; model updates or session-level dependencies could correlate responses.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Performance of ChatGPT on tasks involving physics visual representations: the case of the Brief Electricity and Magnetism Assessment." pith.science (2026). https://pith.science/paper/W4VABBGH

@misc{pith2026241210019,
  author       = {Pith},
  title        = {Pith review of: Performance of ChatGPT on tasks involving physics visual representations: the case of the Brief Electricity and Magnetism Assessment},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/W4VABBGH}},
  note         = {Machine review of arXiv:2412.10019}
}
read the original abstract

Artificial intelligence-based chatbots are increasingly influencing physics education due to their ability to interpret and respond to textual and visual inputs. This study evaluates the performance of two large multimodal model-based chatbots, ChatGPT-4 and ChatGPT-4o on the Brief Electricity and Magnetism Assessment (BEMA), a conceptual physics inventory rich in visual representations such as vector fields, circuit diagrams, and graphs. Quantitative analysis shows that ChatGPT-4o outperforms both ChatGPT-4 and a large sample of university students, and demonstrates improvements in ChatGPT-4o's vision interpretation ability over its predecessor ChatGPT-4. However, qualitative analysis of ChatGPT-4o's responses reveals persistent challenges. We identified three types of difficulties in the chatbot's responses to tasks on BEMA: (1) difficulties with visual interpretation, (2) difficulties in providing correct physics laws or rules, and (3) difficulties with spatial coordination and application of physics representations. Spatial reasoning tasks, particularly those requiring the use of the right-hand rule, proved especially problematic. These findings highlight that the most broadly used large multimodal model-based chatbot, ChatGPT-4o, still exhibits significant difficulties in engaging with physics tasks involving visual representations. While the chatbot shows potential for educational applications, including personalized tutoring and accessibility support for students who are blind or have low vision, its limitations necessitate caution. On the other hand, our findings can also be leveraged to design assessments that are difficult for chatbots to solve.

Figures

Figures reproduced from arXiv: 2412.10019 by the authors.

Figure 5
Figure 5. Test performance of ChatGPT-4o (green) and a sample of students [14](purple). Test items are ranked from highest to lowest score for students (bottom) and ChatGPT-4o (top). V. QUALITATIVE ANALYSIS In the previous section, we presented the findings of the quantitative analysis of the performance of ChatGPT-4 and 4o on BEMA. This section focuses on answering the RQ2 through a qualitative analysis of a subset of ChatGP… view at source ↗
Figure 9
Figure 9. Shared image for items 4 and 5 (left) and the relevant [PITH_FULL_IMAGE:figures/full_fig_p017_9.png] view at source ↗
Figure 10
Figure 10. Image from item 25 (above) and the relevant part of ChatGPT- [PITH_FULL_IMAGE:figures/full_fig_p017_10.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multilingual Performance of a Multimodal Artificial Intelligence System on Multisubject Physics Concept Inventories

    physics.ed-ph 2025-01 conditional novelty 6.0 of 10

    GPT-4o averaged 71% on English physics concept inventories, outperformed average post-instruction undergraduates in most subjects but not laboratory skills, and scored far worse on image-dependent items and in non-Wes...

Reference graph

Works this paper leans on

7 extracted references · 5 canonical work pages · cited by 1 Pith paper

  1. [1]

    ChatGPT-4_meaning,

    The difference in performance of ChatGPT-4 when using the two coding systems When coding ChatGPT-4’s responses by letter, the chatbot scored an average of 50.2% on the test. Passing from coding by letter to coding by meaning, its performance improved by 10.5 percentage points, indicating its ability to extract and process relevant information from graphic...

  2. [2]

    mind’s eye,

    The difference in performance of ChatGPT-4o when using the two coding systems When coding ChatGPT-4o’s responses by letter, the chatbot scored an average of 59.8% on the test. Switching to coding by meaning resulted in a 7.2 percentage point improvement, bringing the average performance to 67.0%. Figure 3 presents the scores for individual items and the o...

  3. [3]

    However, we can see that many errors made by the chatbot would be quite atypical for human students, based on our experience as physics instructors

    ChatGPT-4o as a model of a student As we lack corresponding student reasoning data, we are unable to systematically compare students’ and chatbots’ approaches to solving different tasks on BEMA. However, we can see that many errors made by the chatbot would be quite atypical for human students, based on our experience as physics instructors. One striking ...

  4. [10]

    Shetye, An Evaluation of Khanmigo, a Generative AI Tool, as a Computer-Assisted Language Learning App, Studies in Applied Linguistics and TESOL 24, 1 (2024)

    S. Shetye, An Evaluation of Khanmigo, a Generative AI Tool, as a Computer-Assisted Language Learning App, Studies in Applied Linguistics and TESOL 24, 1 (2024). [11] S. Steinert, K. E. Avila, J. Kuhn, and S. Küchemann, Using GPT-4 as a guide during inquiry-based learning, The Physics Teacher 62, 618 (2024). [12] G. Polverini and B. Gregorcic, Performance ...

  5. [35]

    K. A. Pimbblet and L. Morrell, Can ChatGPT pass a physics degree? Making a case for reformation of assessment of undergraduate degrees, Eur. J. Phys. (2024). [36] K. D. Wang, E. Burkholder, C. Wieman, S. Salehi, and N. Haber, Examining the potential and pitfalls of ChatGPT in science and engineering problem-solving, Front. Educ. 8, (2024). [37] T. Kumar a...

  6. [60]

    Using Large Language Models to Assign Partial Credit to Students' Explanations of Problem-Solving Process: Grade at Human Level Accuracy with Grading Confidence Index and Personalized Student-facing Feedback

    Z. Chen and T. Wan, Using Large Language Models to Assign Partial Credit to Students’ Explanations of Problem-Solving Process: Grade at Human Level Accuracy with Grading Confidence Index and Personalized Student-Facing Feedback, arXiv:2412.06910. [61] J. R. Aguilar-Mejía, S. Tejeda, C. V. Ramirez-Lopez, and C. L. Garay-Rondero, Design and Use of a Chatbot...

  7. [83]

    Taking the Next Step with Generative Artificial Intelligence: The Transformative Role of Multimodal Large Language Models in Science Education

    T. S. Volkwyn, J. Airey, B. Gregorcic, F. Heijkensköld, and C. Linder, Physics Students Learning about Abstract Mathematical Tools When Engaging with “Invisible” Phenomena, in 2017 Physics Education Research Conference Proceedings (American Association of Physics Teachers, Cincinnati, OH, 2018), pp. 408–411. [84] M. B. Kustusch, Assessing the impact of re...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.