REVIEW 3 major objections 2 minor 48 references
Four widely available AI systems score 82–92% on a decade of AP Physics free-response questions, yet share the same visual and spatial failure modes.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-15 13:14 UTC pith:H6QJ4R64
load-bearing objection Useful, limited-scope AP Physics free-response LLM benchmark with a clear error taxonomy—but we only have the abstract; the supplied “full text” is a different security paper. the 3 major comments →
How Well Do AI Systems Solve AP Physics? A Comparative Evaluation of Large Language Models on Algebra-Based Free Response Questions
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
When four accessible AI systems are given standardized free-response AP Physics 1 and 2 questions spanning 2015–2025 and scored by experts with official College Board guidelines, they achieve mean scores of 82–92%, yet they exhibit the same recurring error patterns in diagram and graph interpretation, vector reasoning, circuit topology, qualitative explanation, and three-dimensional concepts such as the right-hand rule. Algebraic competence is therefore strong; spatial and visual competence is not.
What carries the argument
Standardized exam-style prompting of four models, followed by independent expert scoring against official College Board free-response rubrics, which yields both quantitative mean scores and a shared qualitative catalog of failure modes.
Load-bearing premise
The assumption that the particular exam-style prompts and the three experts’ application of College Board rubrics produce scores that fairly represent each model’s true free-response physics ability, without systematic bias from prompt wording or diagram-to-text encoding.
What would settle it
Re-run the identical 2015–2025 free-response set with systematically varied prompts (including different diagram encodings) and measure whether mean scores or the shared error catalog change by more than a few percentage points; or publish inter-rater reliability statistics showing the three experts disagree on more than a small fraction of points.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The abstract claims a systematic evaluation of four widely accessible LLMs (ChatGPT 4.1 mini, Gemini 2.5 Flash, Claude 4.0 Sonnet, DeepSeek R1) on AP Physics 1 and 2 free-response questions (2015–2025), with solutions elicited under standardized exam-style prompting and scored by three independent physics experts using official College Board guidelines. It reports high mean scores (82–92%), substantial year-to-year variability (especially AP Physics 1, with no consistent model hierarchy), statistically significant differences on AP Physics 2 favoring Gemini and DeepSeek over Claude, and a shared qualitative error catalog (diagram/graph misinterpretation, vector direction, circuit topology, partial qualitative explanations, right-hand rule / 3D concepts). The supplied full manuscript body, however, is an entirely different paper: a goal-driven system-level security risk framework for LLM-powered systems that combines data-flow modeling, Attack–Defense Trees, and CVSS v3.1 exploitability scoring, demonstrated on a healthcare assistant case study (goals G1–G3: procedure intervention, EHR leakage, availability disruption). The methods, data, statistics, and error analysis claimed for the AP Physics study are therefore not present in the review package.
Significance. If the AP Physics evaluation were fully documented and held up under scrutiny, it would be a useful empirical contribution to physics education research: multi-year free-response coverage, expert rubric scoring, and a concrete catalog of spatial/visual/conceptual failure modes would inform both AI-assisted instruction and assessment design. Those contributions cannot be credited or stress-tested from the materials provided, because the body of the manuscript does not match the title or abstract under review. The security manuscript that was supplied instead is a separate, methodologically ambitious risk-assessment workflow; it is not the paper announced by paper_id 2603.07457.
major comments (3)
- Manuscript identity mismatch (title/abstract vs. full text): The review package identifies arXiv:2603.07457 as an AP Physics LLM evaluation (physics.ed-ph), but the FULL MANUSCRIPT TEXT is the security paper “Where Do LLM-based Systems Break?…” (ADT/CVSS healthcare risk framework). No methods section, prompt template, diagram-encoding procedure, scoring protocol, inter-rater statistics, year-by-year score tables, or qualitative coding scheme for the AP Physics study appears in the supplied body. The central quantitative claims (82–92% means; year-to-year variability; AP1 lack of hierarchy; AP2 significant differences) and the qualitative error catalog are therefore unsupported by any inspectable evidence in this package.
- Load-bearing design choices uncheckable from abstract alone: Even granting the abstract’s design sketch (standardized exam-style prompting; three independent experts; official College Board guidelines), the abstract does not report inter-rater reliability, the exact prompt/decoding settings, or how diagrams and graphs were encoded into model inputs. Those choices are load-bearing for both the reported means and the claimed model hierarchy / non-hierarchy results; without the correct methods and data sections they cannot be evaluated.
- Statistical and sample claims cannot be verified: The abstract asserts “statistical testing” for year-to-year variability and for AP Physics 2 model differences, but supplies no test names, sample sizes per year/exam, effect sizes, or multiple-comparison corrections. In the absence of the matching manuscript body (tables/figures/results), these claims are not reviewable.
minor comments (2)
- The abstract itself is clearly written and states the intended design elements (standardized prompting; three expert scorers; College Board rubrics; qualitative error themes). Presentation of the abstract is not the issue; the issue is that the full text does not correspond to it.
- If the correct AP Physics manuscript is later supplied, the authors should ensure the abstract’s numerical claims are backed by explicit tables (per-model, per-year means and SDs), IRR metrics (e.g., ICC or percent exact agreement), and a reproducible prompt/diagram-encoding appendix.
Circularity Check
No circularity: empirical scoring study with external expert rubrics; no derivation that folds inputs into claimed predictions.
full rationale
The paper is an empirical evaluation of four LLMs on AP Physics 1/2 free-response questions (2015–2025). Model solutions are elicited under standardized exam-style prompting and scored by three independent physics experts using official College Board guidelines. Mean scores (82–92%), year-to-year variability, statistical comparisons, and a qualitative catalog of shared error modes (diagram/graph misinterpretation, vector direction, circuit topology, right-hand rule, etc.) are reported as observed outcomes of that external scoring process. There is no mathematical derivation chain, no fitted parameter re-presented as a prediction, no uniqueness theorem imported from the authors’ prior work, and no self-definitional loop. The supplied full-text block is a mismatched security manuscript (ADTrees/CVSS risk assessment) and therefore cannot introduce circular steps into the physics-education claims; those claims rest solely on the abstract’s empirical design. Circularity score is therefore 0.
Axiom & Free-Parameter Ledger
free parameters (2)
- exam-style prompt template and decoding settings
- diagram/graph encoding into model input
axioms (3)
- domain assumption Official College Board free-response scoring guidelines constitute an adequate ground-truth measure of physics problem-solving quality for model evaluation.
- domain assumption Three independent physics experts applying the same rubrics produce scores whose mean is a stable estimate of model performance (adequate inter-rater reliability).
- domain assumption Year-to-year exam difficulty and content sampling are comparable enough that score variance can be attributed primarily to model behavior rather than exam form effects.
read the original abstract
The rapid advancement of LLMs has generated growing interest in their potential role in physics education and assessment, yet a focused evaluation of their performance on multi-faceted, free-response physics problems remains underexplored. In this study, we systematically evaluate the performance of four widely accessible AI systems-ChatGPT 4.1 mini, Gemini 2.5 Flash, Claude 4.0 Sonnet, and DeepSeek R1-on AP Physics 1 and 2 free-response questions administered between 2015 and 2025. Model-generated solutions were produced under standardized exam-style prompting and evaluated by three independent physics experts using official College Board scoring guidelines. All models achieved relatively high mean scores (82-92%), indicating strong capability in structured algebraic problem solving. However, substantial year-to-year variability was observed, particularly for AP Physics 1, where statistical testing revealed no consistent performance hierarchy among models. In contrast, AP Physics 2 results showed statistically significant differences, with Gemini and DeepSeek demonstrating more consistent performance than Claude. A qualitative analysis revealed recurring error patterns across all models, including misinterpretation of diagrams and graphs, incorrect graph construction, incorrect reasoning about vector direction, circuit topology errors, partial and misleading qualitative explanations, and difficulties applying three-dimensional concepts such as the right-hand rule. These findings suggest that while contemporary AI systems can effectively support routine physics problem solving, they remain limited in tasks requiring spatial reasoning, visual interpretation, and conceptual integration. The results highlight both the instructional potential and current pedagogical limitations of AI-assisted learning tools in physics education.
Reference graph
Works this paper leans on
-
[1]
GPT-4 Technical Report,
OpenAI, J. Achiam, S. Adler, and et al., “GPT-4 Technical Report,” OpenAI, Tech. Rep., 2024, available at https://arxiv.org/abs/2303. 08774
2024
-
[2]
Potential of large language models in health care: Delphi study,
K. Denecke, R. May, LLMHealthGroup, and O. Rivera Romero, “Potential of large language models in health care: Delphi study,” Journal of Medical Internet Research, vol. 26, p. e52399, 2024. [Online]. Available: https://doi.org/10.2196/52399
-
[3]
Ethical considerations and fundamental principles of large language models in medical education: Viewpoint,
L. Zhui, L. Fenghe, W. Xuehu, F. Qining, and R. Wei, “Ethical considerations and fundamental principles of large language models in medical education: Viewpoint,”Journal of Medical Internet Research, vol. 26, p. e60083, 2024. [Online]. Available: https://www.jmir.org/2024/1/e60083
2024
-
[4]
(2025, Sep.) Microsoft security devel- opment lifecycle (sdl)
Microsoft Corporation. (2025, Sep.) Microsoft security devel- opment lifecycle (sdl). Microsoft Learn, Microsoft Corporation. Accessed January 19, 2026; Microsoft security assurance documentation on SDL processes and phases. [Online]. Available: https://learn.microsoft.com/en-us/compliance/assurance/ assurance-microsoft-security-development-lifecycle
2025
-
[5]
An early categorization of prompt injection attacks on large language models,
S. Rossi, A. M. Michel, R. R. Mukkamala, and J. B. Thatcher, “An early categorization of prompt injection attacks on large language models,” 2024. [Online]. Available: https://arxiv.org/abs/2402.00898
Pith/arXiv arXiv 2024
-
[6]
Comprehensive assessment of jailbreak attacks against llms,
J. Chu, Y . Liu, Z. Yang, X. Shen, M. Backes, and Y . Zhang, “Comprehensive assessment of jailbreak attacks against llms,” 2024. [Online]. Available: https://arxiv.org/abs/2402.05668
Pith/arXiv arXiv 2024
-
[7]
Shostack,Threat Modeling: Designing for Security
A. Shostack,Threat Modeling: Designing for Security. John Wiley & Sons, 2014
2014
-
[8]
[Online]
FIRST.Org, Inc.,Common Vulnerability Scoring System v3.1: Specification Document, FIRST.Org, Inc., 2019, version 3.1. [Online]. Available: https://www.first.org/cvss/v3-1/specification-document
2019
-
[9]
Cyber threat modeling of an llm-based healthcare system,
N. Nagaraja and H. Bahsi, “Cyber threat modeling of an llm-based healthcare system,” inProceedings of the 11th International Con- ference on Information Systems Security and Privacy - Volume 1: ICISSP, INSTICC. SciTePress, 2025, pp. 325–336
2025
-
[10]
Goal-driven risk assessment for llm-powered systems: A healthcare case study,
——, “Goal-driven risk assessment for llm-powered systems: A healthcare case study,” 2026. [Online]. Available: https: //arxiv.org/abs/2603.03633
arXiv 2026
-
[11]
A review of attack graph and attack tree visual syntax in cyber security,
H. S. Lallie, K. Debattista, and J. Bal, “A review of attack graph and attack tree visual syntax in cyber security,”Computer Science Review, vol. 35, p. 100219, 2020
2020
-
[12]
Threat modeling of indus- trial control systems: A systematic literature review,
S. M. Khalil, H. Bahsi, and T. Korotko, “Threat modeling of indus- trial control systems: A systematic literature review,”Computers & Security, vol. 136, p. 103543, 2024
2024
-
[13]
Threat modelling and risk analysis for large language model (llm)-powered applications,
S. B. Tete, “Threat modelling and risk analysis for large language model (llm)-powered applications,” 2024. [Online]. Available: https://arxiv.org/abs/2406.11007
Pith/arXiv arXiv 2024
-
[14]
Mapping llm security landscapes: A comprehensive stakeholder risk assessment proposal,
R. Pankajakshan, S. Biswal, Y . Govindarajulu, and G. Gressel, “Mapping llm security landscapes: A comprehensive stakeholder risk assessment proposal,” 2024. [Online]. Available: https://arxiv. org/abs/2403.13309
Pith/arXiv arXiv 2024
-
[15]
A review of large language models in healthcare: Taxonomy, threats, vulnerabilities, and framework,
R. Hamid and S. Brohi, “A review of large language models in healthcare: Taxonomy, threats, vulnerabilities, and framework,”Big Data and Cognitive Computing, vol. 8, no. 11, p. 161, 2024. [Online]. Available: https://doi.org/10.3390/bdcc8110161
-
[16]
Prompt injection attacks on large language models in oncology,
J. Clusmann, D. Ferber, I. C. Wiest, C. V . Schneider, T. J. Brinker, S. Foersch, D. Truhn, and J. N. Kather, “Prompt injection attacks on large language models in oncology,” 2024. [Online]. Available: https://arxiv.org/abs/2407.18981
Pith/arXiv arXiv 2024
-
[17]
M. Aijaz, M. Nazir, and M. N. A. Mohammad, “Threat modeling and assessment methods in the healthcare-it system: A critical review and systematic evaluation,”SN Computer Science, vol. 4, no. 6, September 2023. [Online]. Available: https://doi.org/10.1007/s42979-023-02221-1
-
[18]
Cybersecurity monitoring/mapping of usa healthcare (all hospitals): Magnified vulnerability due to shared it infrastructure, market concentration, and geographical distribution,
W. Yurcik, A. Schick, S. North, M. T. Gastner, F. R. de Miranda, R. d. S. Avelino, A. F. d. M. Batista, G. Pluta, and I. Brooks, “Cybersecurity monitoring/mapping of usa healthcare (all hospitals): Magnified vulnerability due to shared it infrastructure, market concentration, and geographical distribution,” inProceedings of the 2024 ACM Workshop on Cybers...
2024
-
[19]
Available: https://doi.org/10.1145/3689942.3694754
[Online]. Available: https://doi.org/10.1145/3689942.3694754
-
[20]
Threat modeling of internet of things health devices,
A. Omotosho, B. A. Haruna, and O. M. Olaniyi, “Threat modeling of internet of things health devices,”Journal of Applied Security Research, vol. 14, no. 1, pp. 1–16, April 2019. [Online]. Available: https://doi.org/10.1080/19361610.2019.1545278
-
[21]
Threat modeling and risk analysis for miniaturized wireless biomedical devices,
V . Vakhter, B. Soysal, P. Schaumont, and U. Guler, “Threat modeling and risk analysis for miniaturized wireless biomedical devices,”IEEE Internet of Things Journal, vol. PP, no. 99, pp. 1–1, August 2022. [Online]. Available: https://doi.org/10.1109/JIOT.2022.3144130
-
[22]
Medicalharm - a threat modeling de- signed for modern medical devices,
E. Kwarteng and M. Cebe, “Medicalharm - a threat modeling de- signed for modern medical devices,” in2023 IEEE 22nd Interna- tional Conference on Trust, Security and Privacy in Computing and Communications (TrustCom), 2023, pp. 1147–1156
2023
-
[23]
There are rabbit holes i want to go down that i’m not allowed to go down: An investigation of security expert threat modeling practices for medical device,
R. E. Thompsonet al., “There are rabbit holes i want to go down that i’m not allowed to go down: An investigation of security expert threat modeling practices for medical device,” inProc. USENIX Security 2024, 2024, pp. 4909–4926. [Online]. Available: https:// www.usenix.org/conference/usenixsecurity24/presentation/thompson
2024
-
[24]
Using attack-defense trees to analyze threats and countermeasures in an atm: A case study,
M. Fraile, M. Ford, O. Gadyatskaya, R. Trujillo-Rasuaet al., “Using attack-defense trees to analyze threats and countermeasures in an atm: A case study,” inProc. IFIP PoEM 2016, vol. 267, 2016, pp. 365–373. [Online]. Available: https://doi.org/10.1007/978-3-319-48393-1 24
-
[25]
Threat modeling ai/ml with the attack tree,
S. V . Hoseiniet al., “Threat modeling ai/ml with the attack tree,” pp. 1–1, January 2024, license: CC BY-NC-ND 4.0. [Online]. Available: https://doi.org/10.1109/ACCESS.2024.3497011
-
[26]
A limited technical background is sufficient for attack-defense tree acceptability,
N. D. Schiele and O. Gadyatskaya, “A limited technical background is sufficient for attack-defense tree acceptability,” 2025. [Online]. Available: https://arxiv.org/abs/2502.11920
Pith/arXiv arXiv 2025
-
[27]
On the validity of traditional vulnerability scoring systems for adversarial attacks against llms,
A. A. M. Bahar and A. S. Wazan, “On the validity of traditional vulnerability scoring systems for adversarial attacks against llms,”
-
[28]
Available: https://arxiv.org/abs/2412.20087
[Online]. Available: https://arxiv.org/abs/2412.20087
-
[29]
Atlas matrix,
MITRE, “Atlas matrix,” 2024. [Online]. Available: https://atlas.mitre. org/matrices/ATLAS
2024
-
[30]
Owasp top 10 for large language model applications,
OW ASP, “Owasp top 10 for large language model applications,” 2023-2024. [Online]. Available: https://genai.owasp.org/llm-top-10- 2023-24/
2023
-
[31]
Survey: Automatic generation of attack trees and attack graphs,
A.-M. Konsta, A. Lluch Lafuente, B. Spiga, and N. Dragoni, “Survey: Automatic generation of attack trees and attack graphs,” Computers & Security, vol. 137, p. 103602, 2024. [Online]. Available: https://doi.org/10.1016/j.cose.2023.103602
-
[32]
(2019) Common vulnerability scoring system (cvss) version 3.1 specification document
Forum of Incident Response and Security Teams (FIRST). (2019) Common vulnerability scoring system (cvss) version 3.1 specification document. FIRST.org. Accessed January 19, 2026; Official CVSS v3.1 specification from FIRST, providing the standard for scoring software vulnerability severity. [Online]. Available: https://www.first.org/cvss/v3-1/specificatio...
2019
-
[33]
(2023, Aug.) Threat modeling for drivers
Microsoft. (2023, Aug.) Threat modeling for drivers. Mi- crosoft Learn. Defines the DREAD risk model (Damage, Reproducibility, Exploitability, Affected users, Discoverability) and describes a 1–10 scoring approach. [Online]. Avail- able: https://learn.microsoft.com/en-us/windows-hardware/drivers/ driversecurity/threat-modeling-for-drivers
2023
-
[34]
(2023) Common vulnerability scoring system version 4.0
Forum of Incident Response and Security Teams (FIRST). (2023) Common vulnerability scoring system version 4.0. FIRST.org. Official CVSS v4.0 standard information; CVSS v4.0 was officially released on November 1, 2023. [Online]. Available: https://www.first.org/cvss/v4.0/
2023
-
[35]
N. Nagaraja, L. Zhang, Z. Wang, B. Zhang, and P. Patil, “Image- based prompt injection: Hijacking multimodal llms through visually embedded adversarial instructions,” in2025 3rd International Conference on Foundation and Large Language Models (FLLM). IEEE, Nov. 2025, p. 916–922. [Online]. Available: http://dx.doi.org/ 10.1109/FLLM67465.2025.11391218
-
[36]
K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not what you’ve signed up for: Compromising real- world llm-integrated applications with indirect prompt injection,” in Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security, ser. AISec ’23. New York, NY , USA: Association for Computing Machinery, 2023, p. 79–9...
-
[37]
Agentpoison: Red-teaming llm agents via poisoning memory or knowledge bases,
Z. Chen, Z. Xiang, C. Xiao, D. Song, and B. Li, “Agentpoison: Red-teaming llm agents via poisoning memory or knowledge bases,”
-
[38]
Available: https://arxiv.org/abs/2407.12784
[Online]. Available: https://arxiv.org/abs/2407.12784
-
[39]
Red-teaming llm multi-agent systems via communication attacks,
P. He, Y . Lin, S. Dong, H. Xu, Y . Xing, and H. Liu, “Red-teaming llm multi-agent systems via communication attacks,” 2025. [Online]. Available: https://arxiv.org/abs/2502.14847
Pith/arXiv arXiv 2025
-
[40]
Owasp top 10 for large language model applications, 2025,
OW ASP, “Owasp top 10 for large language model applications, 2025,” 2025. [Online]. Available: https://genai.owasp.org/llm-top-10/
2025
-
[41]
Stealing part of a production language model,
N. Carlini, D. Paleka, K. D. Dvijotham, T. Steinke, J. Hayase, A. F. Cooper, K. Lee, M. Jagielski, M. Nasr, A. Conmy, I. Yona, E. Wallace, D. Rolnick, and F. Tram `er, “Stealing part of a production language model,” 2024. [Online]. Available: https://arxiv.org/abs/2403.06634
Pith/arXiv arXiv 2024
-
[42]
March 20 chatgpt outage,
OpenAI, “March 20 chatgpt outage,” 2024, accessed December 15, 2025. [Online]. Available: https://openai.com/index/march-20- chatgpt-outage/
2024
-
[43]
I know what you asked: Prompt leakage via kv-cache sharing in multi-tenant llm serving,
G. Wu, Z. Zhang, Y . Zhang, W. Wang, J. Niu, Y . Wu, and Y . Zhang, “I know what you asked: Prompt leakage via kv-cache sharing in multi-tenant llm serving,” inNDSS, 2025. [Online]. Available: https: //www.ndss-symposium.org/ndss-paper/i-know-what-you-asked- prompt-leakage-via-kv-cache-sharing-in-multi-tenant-llm-serving/
2025
-
[44]
Extracting training data from large language models,
N. Carlini, F. Tram `er, E. Wallace, M. Jagielski, A. Herbert-V oss, K. Lee, A. Roberts, T. Brown, D. Song, ´U. Erlingsson, A. Oprea, and C. Raffel, “Extracting training data from large language models,” in30th USENIX Security Symposium (USENIX Security 21). USENIX Association, Aug. 2021, pp. 2633–2650. [Online]. Available: https://www.usenix.org/conferen...
2021
-
[45]
Membership inference attacks from first principles,
N. Carlini, S. Chien, M. Nasr, S. Song, A. Terzis, and F. Tramer, “Membership inference attacks from first principles,” 2022. [Online]. Available: https://arxiv.org/abs/2112.03570
Pith/arXiv arXiv 2022
-
[46]
To protect the llm agent against the prompt injection attack with polymorphic prompt,
Z. Wang, N. Nagaraja, L. Zhang, H. Bahsi, P. Patil, and P. Liu, “To protect the llm agent against the prompt injection attack with polymorphic prompt,” in2025 55th Annual IEEE/IFIP International Conference on Dependable Systems and Networks - Supplemental Volume (DSN-S), 2025, pp. 22–28
2025
-
[47]
UcedaVelez and M
T. UcedaVelez and M. M. Morana,Risk Centric Threat Modeling: process for attack simulation and threat analysis. John Wiley & Sons, 2015
2015
-
[48]
Cybersecurity threat modeling the genomic data sequencing workflow: An example threat model implementation for genomic data sequencing and analysis (draft),
R. Pulivarti, J. Wagner, J. Zook, B. Kreider, J. Snyder, K. Wilson, S. Ross, P. Whitlow, E. Alim, I. Brownet al., “Cybersecurity threat modeling the genomic data sequencing workflow: An example threat model implementation for genomic data sequencing and analysis (draft),” US Department of Commerce, Tech. Rep., 2024. Branch Precondition CVEs (ex- amples) E...
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.