Pith. sign in

REVIEW 1 cited by

CHECK-MAT: Checking Hand-Written Mathematical Answers for the Russian Unified State Exam

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2507.22958 v1 pith:65S5CVSV submitted 2025-07-29 cs.CV cs.AIcs.LG

CHECK-MAT: Checking Hand-Written Mathematical Answers for the Russian Unified State Exam

classification cs.CV cs.AIcs.LG
keywords solutionsmathematicalassessmentbenchmarkexamgradeshand-writtenrussian
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

This paper introduces a novel benchmark, EGE-Math Solutions Assessment Benchmark, for evaluating Vision-Language Models (VLMs) on their ability to assess hand-written mathematical solutions. Unlike existing benchmarks that focus on problem solving, our approach centres on understanding student solutions, identifying mistakes, and assigning grades according to fixed criteria. We compile 122 scanned solutions from the Russian Unified State Exam (EGE) together with official expert grades, and evaluate seven modern VLMs from Google, OpenAI, Arcee AI, and Alibaba Cloud in three inference modes. The results reveal current limitations in mathematical reasoning and human-rubric alignment, opening new research avenues in AI-assisted assessment. You can find code in https://github.com/Karifannaa/Auto-check-EGE-math

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. EDU-CIRCUIT-HW: Evaluating Multimodal Large Language Models on Real-World University-Level STEM Student Handwritten Solutions

    cs.CV 2026-01 unverdicted novelty 6.0

    EDU-CIRCUIT-HW reveals large latent recognition failures in MLLMs on real handwritten university STEM solutions, limiting auto-grading reliability, though hybrid human-AI routing of only 3.3% cases improves outcomes.