Pith. sign in

REVIEW 1 cited by

ChatGPT or Grammarly? Evaluating ChatGPT on Grammatical Error Correction Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.13648 v1 pith:RCLXNP2S submitted 2023-03-15 cs.CL

classification cs.CL
keywords chatgptevaluationgrammaticalautomaticbenchmarkcorrectionerrorfind
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

ChatGPT is a cutting-edge artificial intelligence language model developed by OpenAI, which has attracted a lot of attention due to its surprisingly strong ability in answering follow-up questions. In this report, we aim to evaluate ChatGPT on the Grammatical Error Correction(GEC) task, and compare it with commercial GEC product (e.g., Grammarly) and state-of-the-art models (e.g., GECToR). By testing on the CoNLL2014 benchmark dataset, we find that ChatGPT performs not as well as those baselines in terms of the automatic evaluation metrics (e.g., $F_{0.5}$ score), particularly on long sentences. We inspect the outputs and find that ChatGPT goes beyond one-by-one corrections. Specifically, it prefers to change the surface expression of certain phrases or sentence structure while maintaining grammatical correctness. Human evaluation quantitatively confirms this and suggests that ChatGPT produces less under-correction or mis-correction issues but more over-corrections. These results demonstrate that ChatGPT is severely under-estimated by the automatic evaluation metrics and could be a promising tool for GEC.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. KoGEC : Korean Grammatical Error Correction with Pre-trained Translation Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A 3.3B-parameter NLLB model fine-tuned on Korean social media data beats GPT-4o and HCX-3 on BLEU for Korean grammatical error correction.

Pith tools