Pith. sign in

REVIEW 2 cited by

Linguistic Rules-Based Corpus Generation for Native Chinese Grammatical Error Correction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.10442 v1 pith:EM3GTVPI submitted 2022-10-19 cs.CL

classification cs.CL
keywords cgecchinesegrammaticalerrorsmodelsnativetrainingapplication
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Chinese Grammatical Error Correction (CGEC) is both a challenging NLP task and a common application in human daily life. Recently, many data-driven approaches are proposed for the development of CGEC research. However, there are two major limitations in the CGEC field: First, the lack of high-quality annotated training corpora prevents the performance of existing CGEC models from being significantly improved. Second, the grammatical errors in widely used test sets are not made by native Chinese speakers, resulting in a significant gap between the CGEC models and the real application. In this paper, we propose a linguistic rules-based approach to construct large-scale CGEC training corpora with automatically generated grammatical errors. Additionally, we present a challenging CGEC benchmark derived entirely from errors made by native Chinese speakers in real-world scenarios. Extensive experiments and detailed analyses not only demonstrate that the training data constructed by our method effectively improves the performance of CGEC models, but also reflect that our benchmark is an excellent resource for further development of the CGEC field.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exploring the Implicit Semantic Ability of Multimodal Large Language Models: A Pilot Study on Entity Set Expansion

    cs.CL 2024-12 conditional novelty 5.0 of 10

    LUSAR applies listwise sampling and ranking to multimodal LLMs for entity set expansion and reports improved MESED scores, though the gains are confounded with supervised fine-tuning.

  2. Loss-Aware Curriculum Learning for Chinese Grammatical Error Correction

    cs.CL 2024-12 reject novelty 4.0 of 10

    A two-level curriculum, batch ordering by loss and instance/token reweighting by Monte Carlo dropout confidence, yields about 0.5 to 1.2 F0.5 gains for BART, mT5, and SynGEC on NLPCC and MuCGEC.

Pith tools