Pith. sign in

REVIEW 2 cited by

CLEME: Debiasing Multi-reference Evaluation for Grammatical Error Correction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.10819 v2 pith:VYGGL5PK submitted 2023-05-18 cs.CL

classification cs.CL
keywords evaluationclememulti-referencecorrectiongrammaticalreferencessystemstask
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Evaluating the performance of Grammatical Error Correction (GEC) systems is a challenging task due to its subjectivity. Designing an evaluation metric that is as objective as possible is crucial to the development of GEC task. However, mainstream evaluation metrics, i.e., reference-based metrics, introduce bias into the multi-reference evaluation by extracting edits without considering the presence of multiple references. To overcome this issue, we propose Chunk-LEvel Multi-reference Evaluation (CLEME), designed to evaluate GEC systems in the multi-reference evaluation setting. CLEME builds chunk sequences with consistent boundaries for the source, the hypothesis and references, thus eliminating the bias caused by inconsistent edit boundaries. Furthermore, we observe the consistent boundary could also act as the boundary of grammatical errors, based on which the F$_{0.5}$ score is then computed following the correction independence assumption. We conduct experiments on six English reference sets based on the CoNLL-2014 shared task. Extensive experiments and detailed analyses demonstrate the correctness of our discovery and the effectiveness of CLEME. Further analysis reveals that CLEME is robust to evaluate GEC systems across reference sets with varying numbers of references and annotation style.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exploring the Implicit Semantic Ability of Multimodal Large Language Models: A Pilot Study on Entity Set Expansion

    cs.CL 2024-12 conditional novelty 5.0 of 10

    LUSAR applies listwise sampling and ranking to multimodal LLMs for entity set expansion and reports improved MESED scores, though the gains are confounded with supervised fine-tuning.

  2. Loss-Aware Curriculum Learning for Chinese Grammatical Error Correction

    cs.CL 2024-12 reject novelty 4.0 of 10

    A two-level curriculum, batch ordering by loss and instance/token reweighting by Monte Carlo dropout confidence, yields about 0.5 to 1.2 F0.5 gains for BART, mT5, and SynGEC on NLPCC and MuCGEC.

Pith tools