Pith. sign in

REVIEW 1 cited by

Evaluating GPT-4 at Grading Handwritten Solutions in Math Exams

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.05231 v2 pith:LFSRL2E6 submitted 2024-11-07 cs.CY cs.CLcs.LG

Evaluating GPT-4 at Grading Handwritten Solutions in Math Exams

classification cs.CY cs.CLcs.LG
keywords responsesgradinghandwrittenalignmentexamsgpt-4omathstudent
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Recent advances in generative artificial intelligence (AI) have shown promise in accurately grading open-ended student responses. However, few prior works have explored grading handwritten responses due to a lack of data and the challenge of combining visual and textual information. In this work, we leverage state-of-the-art multi-modal AI models, in particular GPT-4o, to automatically grade handwritten responses to college-level math exams. Using real student responses to questions in a probability theory exam, we evaluate GPT-4o's alignment with ground-truth scores from human graders using various prompting techniques. We find that while providing rubrics improves alignment, the model's overall accuracy is still too low for real-world settings, showing there is significant room for growth in this task.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. LLM Performance on a Real, Double-Marked GCSE Benchmark

    cs.CL 2026-06 unverdicted novelty 6.0

    Off-the-shelf LLMs match or exceed inter-examiner agreement on a new 32k-response double-marked GCSE dataset spanning five subjects and handwritten scripts.