Pith. sign in

REVIEW 4 cited by

On the (In)Effectiveness of Large Language Models for Chinese Text Correction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.09007 v2 pith:IVPQDTQN submitted 2023-07-18 cs.CL

classification cs.CL
keywords chinesellmscorrectiontextcapabilitieslanguagemodelsperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, the development and progress of Large Language Models (LLMs) have amazed the entire Artificial Intelligence community. Benefiting from their emergent abilities, LLMs have attracted more and more researchers to study their capabilities and performance on various downstream Natural Language Processing (NLP) tasks. While marveling at LLMs' incredible performance on all kinds of tasks, we notice that they also have excellent multilingual processing capabilities, such as Chinese. To explore the Chinese processing ability of LLMs, we focus on Chinese Text Correction, a fundamental and challenging Chinese NLP task. Specifically, we evaluate various representative LLMs on the Chinese Grammatical Error Correction (CGEC) and Chinese Spelling Check (CSC) tasks, which are two main Chinese Text Correction scenarios. Additionally, we also fine-tune LLMs for Chinese Text Correction to better observe the potential capabilities of LLMs. From extensive analyses and comparisons with previous state-of-the-art small models, we empirically find that the LLMs currently have both amazing performance and unsatisfactory behavior for Chinese Text Correction. We believe our findings will promote the landing and application of LLMs in the Chinese NLP community.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EXECUTE: A Multilingual Benchmark for LLM Token Understanding

    cs.CL 2025-05 conditional novelty 7.0 of 10

    A multilingual extension of the CUTE benchmark shows that LLM token-manipulation performance varies by language and script, with surprisingly strong results on low-resource languages and weak results on sub-character ...

  2. Mixture of Small and Large Models for Chinese Spelling Check

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A decoding-time mixture of a fine-tuned BERT classifier and a frozen LLM improves Chinese spelling correction across five benchmarks.

  3. Refine Knowledge of Large Language Models via Adaptive Contrastive Learning

    cs.CL 2025-02 conditional novelty 6.0 of 10

    An adaptive contrastive learning strategy that uses a model's own sampled response accuracy to create per-region positive and negative training pairs improves LLM truthful rate by up to 6.9% over IDK-SFT.

  4. From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI

    cs.AI 2026-06 conditional novelty 4.0 of 10

    Autonomous AI becomes dependable when tool use is embedded in persistent workspaces with reusable skills, shifting evaluation from answers to task closure.

Pith tools