REVIEW 3 major objections 4 minor 39 references
VArsity: Can Large Language Models Keep Power Engineering Students in Phase?
T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read In two offerings of a power systems course, students found the errors in ChatGPT GPT-4 outputs much easier to spot than the subtler errors in ChatGPT o1 outputs, with 75% versus 36.66% success.
desk verdict Minimax fix is correct and the dataset is new, but the o1-vs-GPT-4 comparison is confounded and the paper needs a major revision before it should appear. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is a power factor correction problem with a fixed inductor and a switchable capacitor, where the source's reactive power is Qs = Qd + $V^{2}$/Xl with the switch open and Qs = Qd + $V^{2}$/Xl - $V^{2}$/Xc with the switch closed. Because the real power Ps is fixed at 50 W, the smallest power factor is determined by the largest magnitude of |Qs| over the demand range, so maximizing the worst-case power factor is equivalent to minimizing the worst-case |Qs|. The argument turns on the min-max geometry of |Qs| as a function of Qd: the best solution places the points of unity power factor at the quarter points of the interval [-36,60] VAr, where the resulting four line segments have equal maximum height, giving a worst-case |Qs| of 24 VAr.
What would settle it
A controlled A/B study in which the same cohort critiques both a GPT-4 and an o1 solution to the same power factor problem, with version order randomized and grading blinded, would settle whether the 75% versus 36.66% gap reflects model capabilities or the different problems; the technical claim could also be checked by a grid search showing whether any fixed inductor and switchable capacitor design beats a 24 VAr worst-case mismatch.
Extended reading notes
Core claim
The central claim is that, in a power engineering course, students found errors from ChatGPT o1 much harder to identify than errors from ChatGPT GPT-4: 36.66% of the Spring 2025 class identified all three o1 errors and 23.33% identified none, compared with 75% of the Fall 2023 class finding all or nearly all GPT-4 errors. The paper also demonstrates that the o1 solution is genuinely suboptimal: anchoring unity power factor at the extreme reactive demands gives a worst-case |Qs| of 48 VAr, while the optimal placement of the unity-power-factor points at the quarter points of the demand range gives 24 VAr and a worst-case power factor of 50/$\sqrt$($50^{2}$+$24^{2}$)=0.9015. The authors argue that the similarity between LLM errors and student misconceptions makes it hard to tell copied work from genuine misunderstanding, and that educators need assessment strategies that track rapidly changing LLM capabilities.
Load-bearing premise
The load-bearing premise is that the difference in error-detection rates between the two semesters is attributable to the switch from GPT-4 to o1, even though the semesters differed in problem, circuit values, number and obviousness of errors, student cohort, and grading procedures.
Editorial extensions
If this is right
- If the reported rates generalize, course assessments built on older-model LLM outputs may overestimate students' ability to catch errors; recent reasoning-model outputs need stronger scaffolding.
- The o1 solution's flaw is not cosmetic: at the extremes it reaches unity power factor but in the middle of the demand range the worst-case mismatch is twice the optimum.
- Error-identification exercises can separate surface-level reading from genuine understanding, since about a quarter of the Spring 2025 class detected none of the o1 errors while a smaller share failed all corrections.
- LLM error profiles are changing quickly enough that an assessment built on one model version cannot be assumed stable across semesters.
- The similarity between LLM errors and student errors implies that simple 'was this copied?' judgments are unreliable in power engineering coursework.
Reading between the lines
- One extension beyond the paper: the o1 error pattern, anchoring compensation at the endpoints of the demand interval, could be used as a diagnostic for the same misconception in students, with a quick check at the quarter points revealing the mistake.
- If the resemblance between LLM errors and student errors is general, LLM outputs could serve as automatic generators of plausible incorrect answers for power engineering assessments.
- The reported difficulty gap may partly stem from the Spring 2025 task asking students to find three subtle conceptual errors without a checklist; an intervention study could test whether a hint such as 'check the midpoint and the quarters' raises detection rates.
- A practical implication for instructors is to re-benchmark LLM outputs at each model release, since the error profile that made GPT-4 easy to critique did not carry over to o1.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports an educational case study from ECE 4320 at Georgia Tech in which students were asked to identify, explain, and correct errors in ChatGPT-generated power factor correction solutions. In Spring 2025, ChatGPT o1 produced a solution with three errors; the paper shows that the o1 solution is genuinely suboptimal: anchoring unity power factor at the endpoints of the reactive demand range gives worst-case |Qs| = 48 VAr, whereas the minimax placement at the quarter points gives worst-case |Qs| = 24 VAr and a worst-case power factor of 0.9015. The paper reports that 36.66% of 30 students identified all three o1 errors and about a quarter identified none, and contrasts this with 75% of 28 Fall 2023 students who found all or nearly all errors in a GPT-4 solution. Appendix C reports that ChatGPT o3 solves the power factor problems essentially correctly. The paper concludes that students find o1 errors harder to identify and discusses implications for LLM use in power engineering education.
Significance. The paper's technical contribution is solid: the minimax argument in Section IV.B is a parameter-free derivation, the corrected reactances (Xl = 833.3 Ω, Xc = -208.3 Ω) follow from the stated circuit relations, and the result is independently confirmed by the o3 grid search in Appendix C. The Spring 2025 student performance data, taken as a descriptive case study, is a useful data point for the community. However, the paper's headline comparative claim—that students found o1 errors much more difficult than GPT-4 errors—is not supported by the evidence, because the two semesters differ in problem, circuit scale, number of errors, error salience, cohort, and grading rubric. The retrospective design is acknowledged, but the abstract and conclusion nonetheless assert the comparative claim. If the comparison is reframed as two separate descriptive case studies, the paper would be a modest but useful contribution; in its current form the central empirical claim overreaches.
major comments (3)
- [Abstract and Section V.C] The central comparative claim that students found o1 errors much more difficult to identify than GPT-4 errors is not identified by the data, because the Fall 2023 and Spring 2025 administrations differ in multiple respects: different circuits (10 kV vs 100 V), different problems (three time periods vs a continuous demand range), different numbers of embedded errors (nearly a dozen vs three), different error salience (the paper itself says the GPT-4 errors were 'much more obvious'), different cohorts, and different instructors' rubrics. This is a between-semester, between-task comparison, not a controlled model comparison. The abstract and Section V.C should be revised to present the two semesters as separate descriptive case studies or to explicitly disclaim any model-version attribution, and the comparative language in the abstract and Section VI should be softened accordingly.
- [Sections V.A and V.B] The student-success measurements rest on unblinded instructor grading of student work with no reported inter-rater reliability. Because the same authors who designed the problems and know the intended answers graded the open-ended responses, the percentages in Tables III and IV (36.66%, 23.33%, 27%, etc.) could reflect grader expectations or leniency. The manuscript should report the rubric in full, state whether grading was blinded, and provide inter-rater reliability (e.g., Cohen's kappa) for at least a subsample, or explicitly acknowledge this as a limitation in the text.
- [Section V.A, Table IV, and Section VI] There are arithmetic inconsistencies in the reported percentages that undermine the quantitative claims: (i) the text states '66.33%' did not detect all errors, but 100% − 36.66% = 63.34%; (ii) Table IV's percentages sum to 102% (27+24+10+20+14+7); and (iii) the conclusion states that the share of the class unable to identify any error was larger than the share able to solve the entire problem correctly, but Table III shows 23.33% identified zero errors while Table IV's first category (27%) is larger, not smaller. These numbers should be recomputed and the conclusion reworded to match the tables.
minor comments (4)
- [Section III.A and Table I] The circuit diagram labels the switched branch with jXl and jXc, while the text and Table I use Xc = -208.3 Ω; the sign convention for capacitive reactance should be defined consistently to avoid confusion for readers.
- [Abstract] There is a grammatical typo in the abstract: 'the errors from the ChatGPT o1 version much more difficult' is missing the verb 'were'.
- [Figure 2] The subplots in Figure 2 would benefit from axis labels and a legend indicating the switch-open and switch-closed segments, as the four line segments are otherwise easy to misread.
- [Section V.C] The phrase 'students found the errors from the ChatGPT o1 version much more difficult to identify' implies a causal comparison; even after the major revision, the text should consistently use wording such as 'appeared more difficult' or 'were reported as more difficult' to match the retrospective design.
Circularity Check
No significant circularity: the technical correction is derived parameter-free from the circuit equations, and the empirical claims are observational rather than fitted predictions.
full rationale
I walked the paper's claimed derivation chain. The technical diagnosis of ChatGPT o1's error and the corrected solution (Section IV.B) are derived directly from the stated circuit relations Qs = Qd + V^2/Xl (switch open) and Qs = Qd + V^2/Xl + V^2/Xc (switch closed), combined with the minimax condition that the worst-case |Qs| is minimized when the four line segments in Fig. 2b have equal maximum height. This optimization does not take the ChatGPT output as an input; it recomputes Xl = 833.3 ohms and Xc = -208.3 ohms from the unity-power-factor points at Qd = -12 and Qd = 36 VAr, and its result (worst-case |Qs| = 24 VAr, pf_min = 0.9015) is independently confirmed by the parameter-free o3 grid search reported in Appendix C. The empirical student-success claims in Sections V.A and V.C are observational comparisons of two course offerings; no parameter is fitted to a subset of data and then relabeled as a prediction, and no load-bearing step relies on a self-citation or an imported uniqueness theorem. Concerns about the between-semester comparison (different problems, error counts, cohorts, and rubrics) are threats to statistical validity, not circularity. I find no circular step in the paper.
Assumptions & free parameters
assumptions (3)
- standard math The optimal power factor correction design places the two unity-power-factor points at the quarter and three-quarter positions of the reactive demand range, so all four segments of the |Qs| curve reach the same maximum height.
- domain assumption Conservation of power in a lossless reactive network gives Ps = Pd = 50 W for the real power supplied by the source.
- ad hoc to paper A single ChatGPT response per model and problem is treated as representative of that model version's behavior on this problem.
Cite this review
Pith. "Pith review of VArsity: Can Large Language Models Keep Power Engineering Students in Phase?." pith.science (2026). https://pith.science/paper/M2GTK7PW
@misc{pith2026250720995,
author = {Pith},
title = {Pith review of: VArsity: Can Large Language Models Keep Power Engineering Students in Phase?},
year = {2026},
howpublished = {\url{https://pith.science/paper/M2GTK7PW}},
note = {Machine review of arXiv:2507.20995}
}
read the original abstract
This paper provides an educational case study regarding our experience in deploying ChatGPT Large Language Models (LLMs) in the Spring 2025 and Fall 2023 offerings of ECE 4320: Power System Analysis and Control at Georgia Tech. As part of course assessments, students were tasked with identifying, explaining, and correcting errors in the ChatGPT outputs corresponding to power factor correction problems. While most students successfully identified the errors in the outputs from the GPT-4 version of ChatGPT used in Fall 2023, students found the errors from the ChatGPT o1 version much more difficult to identify in Spring 2025. As shown in this case study, the role of LLMs in pedagogy, assessment, and learning in power engineering classrooms is an important topic deserving further investigation.
Figures
Reference graph
Works this paper leans on
-
[1]
Exploring the Capabilities and Limitations of Large Language Models in the Electric Energy Sector,
S. Majumder, L. Dong, F. Doudi, Y . Cai, C. Tian, D. Kalathil, K. Ding, A. A. Thatte, N. Li, and L. Xie, “Exploring the Capabilities and Limitations of Large Language Models in the Electric Energy Sector,” Joule, vol. 8, pp. 1544–1549, June 2024
work page 2024
-
[2]
Large Language Models Challenge the Future of Higher Education,
S. Milano, J. A. McGrane, and S. Leonelli, “Large Language Models Challenge the Future of Higher Education,” Nature Machine Intelligence , vol. 5, pp. 333–334, Apr. 2023
work page 2023
-
[3]
ChatGPT for Good? On Opportunities and Challenges of Large Language Models for Education,
E. Kasneci, K. Sessler, S. K ¨uchemann, M. Bannert, D. Dementieva, F. Fischer, U. Gasser, G. Groh, S. G ¨unnemann, E. H ¨ullermeier, S. Krusche, G. Kutyniok, T. Michaeli, C. Nerdel, J. Pfeffer, O. Poquet, M. Sailer, A. Schmidt, T. Seidel, M. Stadler, J. Weller, J. Kuhn, and G. Kasneci, “ChatGPT for Good? On Opportunities and Challenges of Large Language M...
work page 2023
-
[4]
Examining the Potential and Pitfalls of ChatGPT in Science and Engineering Problem-Solving,
K. D. Wang, E. Burkholder, C. Wieman, S. Salehi, and N. Haber, “Examining the Potential and Pitfalls of ChatGPT in Science and Engineering Problem-Solving,” Frontiers in Education, vol. 8, Jan. 2024
work page 2024
-
[5]
DualSchool: How Reliable are LLMs for Optimization Education?,
M. Klamkin, A. Deza, S. Cheng, H. Zhao, and P. V . Hentenryck, “DualSchool: How Reliable are LLMs for Optimization Education?,” arXiv:2505.21775, 2025
arXiv 2025
-
[6]
Assessing GPT Performance in a Proof-Based University-Level Course Under Blind Grading
M. Ding, R. Kyng, F. Solda, and W. Yuan, “Assessing GPT Perfor- mance in a Proof-Based University-Level Course Under Blind Grading,” arXiv:2505.13664, 2025
work page Pith review arXiv 2025
-
[7]
A. Hickey, C. O. Faol ´ain, and P. Cuffe, “Large Language Models in Power Engineering Education: A Case Study on Solving Optimal Dispatch Coursework Problems,” in 21st International Conference on Information Technology Based Higher Education and Training (ITHET) , 2024. APPENDIX CONTRASTING WITH FALL 2023: T HE PITFALLS OF GPT-4 In the Fall 2023 semester,...
work page 2024
-
[8]
Specifically, the power factor is 10/ √ 105 + 52 = 0 .8944 lagging
The load’s power factor is not given during the morning time period, but we can infer this from the complex power consumption of 10 + j5 MV A. Specifically, the power factor is 10/ √ 105 + 52 = 0 .8944 lagging. There is no need to arbitrarily choose a power factor for the load during this or any other time period
Show all 39 references
-
[9]
With a complex power of 10 + j5 MV A, the apparent power is √ 102 + 52 = 11.1803 MV A
-
[10]
The correct expression is Z = |V |2/S∗
The formula Z = V 2/S is not correct. The correct expression is Z = |V |2/S∗. The load impedance is Z = (10×103)2/(10×106 +j5×106)∗ = 8+ j4 Ω
-
[11]
Additionally, we do not necessarily want the fixed capacitor to supply 5 MV Ar
The reactive power consumed by a shunt capacitance is Q = −|V |2/(Xc), so the capacitive reactance that would supply a 5 MV Ar is Xc,fixed = −|V |2/Q = −(10 × 103)2/5 × 106 = −20 Ω, not −13.2 Ω. Additionally, we do not necessarily want the fixed capacitor to supply 5 MV Ar. Wh...
-
[12]
A capacitor has a negative reactance, so Xc,fixed should be a negative value
-
[13]
We can choose values that result in a higher power factor during all time periods
The assertion that we should aim for a power factor of 0.9 lagging is not correct. We can choose values that result in a higher power factor during all time periods
-
[14]
The load impedance during the afternoon should be Z = (10 × 103)2/(32 × 106 + j24 × 106)∗ = 2 + j1.5 Ω
Same as in 4), the formula Z = V 2/S is not correct (and the value is also wrong even if the formula had been correct). The load impedance during the afternoon should be Z = (10 × 103)2/(32 × 106 + j24 × 106)∗ = 2 + j1.5 Ω
-
[15]
The value should be (10 × 103)2/ − 24 × 106 = −4.1667 Ω
This expression for the desired capacitive reactance has both a sign error and the value of 4167 Ω is not the output of this expression (off by three orders of magnitude). The value should be (10 × 103)2/ − 24 × 106 = −4.1667 Ω
-
[16]
However, it is not correct to directly subtract the reactances since the capacitors are connected in parallel
The switched capacitor’s reactance is based on the difference between the fixed capacitor’s reactance and the desired reactance. However, it is not correct to directly subtract the reactances since the capacitors are connected in parallel
-
[17]
The choice of the fixed capacitor’s reactance should be cognizant of the resulting deviation in the power factor during the evening period
This analysis of the evening load demand neglects the fact that the fixed capacitor is supplying reactive power and thus changing the power factor away from unity. The choice of the fixed capacitor’s reactance should be cognizant of the resulting deviation in the power factor ...
-
[18]
7 1⃝ 2⃝ 3⃝ P3+jQ3=−1.50−j0.75 R12+jX12= 0 +j0.20 R23+jX23= 0 +j0.10 R13+jX13= 0.10 +j0.20 V1∠θ1 P1+jQ1 V2∠θ2 P2+jQ2 Fig
The values in the summary are incorrect due to the errors earlier in the solution. 7 1⃝ 2⃝ 3⃝ P3+jQ3=−1.50−j0.75 R12+jX12= 0 +j0.20 R23+jX23= 0 +j0.10 R13+jX13= 0.10 +j0.20 V1∠θ1 P1+jQ1 V2∠θ2 P2+jQ2 Fig. 4: The one-line diagram presented to the Fall 2023 cohort of ECE 4320 stu...
2023
-
[19]
The voltage magnitudes at buses 1 and 2 are specified values, so these are not variables. The ChatGPT solution adds constraints to enforce V1 = 1 and V2 = 1, which is not wrong but is not necessary since these can just be eliminated from the problem by substituting in the corr...
-
[20]
The active and reactive power injections at bus 3 are specified values, so these are not variables
-
[21]
Rather, the voltage magnitude V2 is specified and the active power, while not specified explicitly, is equal to the active power at bus 1
Bus 2 is not a PQ bus since neither the active power nor the reactive power are specified. Rather, the voltage magnitude V2 is specified and the active power, while not specified explicitly, is equal to the active power at bus 1
-
[22]
First, these expressions are missing the terms corresponding to the power flowing on the line from bus 2 to bus 3
The expressions for active and reactive power balance at bus 2 have several errors. First, these expressions are missing the terms corresponding to the power flowing on the line from bus 2 to bus 3. Second, the expressions use the line resistance and reactance values instead o...
-
[23]
First, the expressions for active and reactive power balance at bus 3 are missing the specified values for P3 = −1.50 and Q3 = −0.75
The expressions for active and reactive power balance at bus 3 have several errors. First, the expressions for active and reactive power balance at bus 3 are missing the specified values for P3 = −1.50 and Q3 = −0.75. Second, these expressions are missing the terms correspondi...
-
[24]
The specified x vector and system of equations for f (x) = 0 inherit the errors above and also are clearly incorrect since the number of equations (seven) is inconsistent with the number of variables (eleven)
-
[25]
Actually, this whole portion of the response is not needed
The real and reactive power expressions are not multiplied by the impedance values. Actually, this whole portion of the response is not needed
-
[26]
Finally, the formulation (if appropriately written) should account for line losses
The problem does not have any transformers, so tap ratios do not need to be considered, and there are also no shunt elements. Finally, the formulation (if appropriately written) should account for line losses. 8 C. Can new LLM models do better? At the conclusion of the Spring ...
2025
-
[27]
Response
Power factor correction problem (Spring 2025): Below is the response provided by the o3 model for the power factor correction prompt in Section III-B. Response
2025
-
[28]
• Inductor: QL = 10000 XL , (positive → lagging) • Capacitor: QC = − 10000 XC , (QC < 0) (negative → leading)
Basic relations for a shunt element At the node where the 100 V-rms source is connected Q = V 2 rms X (Vrms = 100 V). • Inductor: QL = 10000 XL , (positive → lagging) • Capacitor: QC = − 10000 XC , (QC < 0) (negative → leading)
-
[29]
Because the inductor is always in circuit but the capacitor can be switched, the reactive power seen by the source is Qs = ( Qd + QL, S open, Qd + QL + QC , S closed
Two possible reactive-power states Let the load’s reactive demand vary in the interval Qd ∈ [−36, 60] V Ar. Because the inductor is always in circuit but the capacitor can be switched, the reactive power seen by the source is Qs = ( Qd + QL, S open, Qd + QL + QC , S closed. De...
-
[30]
Because the two straight-line functions are parallel, the optimal choice is obtained when
Min-max (Chebyshev) design of A and B We want to minimise the worst-case reactive mismatch max Qd∈[−36,60] min(|Qd + A|, |Qd + A − B|). Because the two straight-line functions are parallel, the optimal choice is obtained when
-
[31]
They cross exactly halfway between their two extreme errors, and
-
[32]
Carrying this out (or by a quick optimisation sweep) gives A = 12 V Ar, B = 48 V Ar
The magnitudes of the residuals at both ends of the interval are equal. Carrying this out (or by a quick optimisation sweep) gives A = 12 V Ar, B = 48 V Ar. Hence, QL = +12 V Ar, Q C = −48 V Ar. The worst residual reactive power is then only ±24 VAr everywhere in the interval
-
[33]
Reactances XL = V 2 rms QL = 10000 12 ≈ 833 Ω XC = V 2 rms |QC | = 10000 48 ≈ 208 Ω (The capacitor’s reactance is negative in the usual sign convention: XC ≈ −208 Ω.)
-
[34]
Measure (or estimate) the load’s instantaneous Qd
-
[35]
anywhere from -36 VAr to +12 VAr) leave the capacitor disconnected (switch open)
If Qd ≤ +12 VAr (i.e. anywhere from -36 VAr to +12 VAr) leave the capacitor disconnected (switch open)
-
[36]
from +12 VAr up to +60 VAr) connect the capacitor into the circuit (switch closed)
If Qd > +12 VAr (i.e. from +12 VAr up to +60 VAr) connect the capacitor into the circuit (switch closed). With this rule Qs = ( Qd + 12, Q d ≤ 12, Qd − 36, Q d > 12, =⇒ |Qs| ≤24 V Ar
-
[37]
Resulting power-factor band Since the real power drawn from the source is constant at Ps = 50 W, the minimum power factor occurs at |Qs| = 24 VAr: pfmin = 50√ 502 + 242 ≈ 0.90. Thus the source operates between 0.90 and 1.0 (leading or lagging) over the entire range of load rea...
-
[38]
For this problem, o3 took an interesting approach during extensive chain-of- thought reasoning, which lasted for approximately 2.5 minutes
Power factor correction problem (Fall 2023): Similarly to the Spring 2025 power factor correction problem, the o3 model also solved the power factor correction prompt described in Appendix A essentially correctly. For this problem, o3 took an interesting approach during extens...
2023
-
[39]
Namely, o3 was able to immediately formulate the correct nodal admittance matrix: Response
AC power flow formulation problem: For the AC power flow formulation problem in Appendix B, ChatGPT version o3 demonstrated a significantly improved solution approach compared to by GPT-4. Namely, o3 was able to immediately formulate the correct nodal admittance matrix: Respon...
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.