Pith. sign in

REVIEW 1 cited by

Assessing GPT Performance in a Proof-Based University-Level Course Under Blind Grading

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.13664 v1 pith:3PLTF5AO submitted 2025-05-19 cs.CY cs.CL

classification cs.CYcs.CL
keywords performancecourseeducationgpt-4ogradingmodelso1-previewpassing
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As large language models (LLMs) advance, their role in higher education, particularly in free-response problem-solving, requires careful examination. This study assesses the performance of GPT-4o and o1-preview under realistic educational conditions in an undergraduate algorithms course. Anonymous GPT-generated solutions to take-home exams were graded by teaching assistants unaware of their origin. Our analysis examines both coarse-grained performance (scores) and fine-grained reasoning quality (error patterns). Results show that GPT-4o consistently struggles, failing to reach the passing threshold, while o1-preview performs significantly better, surpassing the passing score and even exceeding the student median in certain exercises. However, both models exhibit issues with unjustified claims and misleading arguments. These findings highlight the need for robust assessment strategies and AI-aware grading policies in education.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VArsity: Can Large Language Models Keep Power Engineering Students in Phase?

    cs.CY 2025-07 conditional novelty 5.0 of 10

    In two offerings of Georgia Tech's ECE 4320, 36.66% of students found all three reasoning errors in a ChatGPT o1 power factor solution, versus 75% finding all or nearly all of GPT-4's cruder errors, and the newest mod...

Pith tools