Pith. sign in

REVIEW 2 cited by

PETCI: A Parallel English Translation Dataset of Chinese Idioms

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.09509 v1 pith:7NFCALNN submitted 2022-02-19 cs.CL cs.LG

classification cs.CLcs.LG
keywords translationidiomsmachinepetcichinesedatasetidiommodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Idioms are an important language phenomenon in Chinese, but idiom translation is notoriously hard. Current machine translation models perform poorly on idiom translation, while idioms are sparse in many translation datasets. We present PETCI, a parallel English translation dataset of Chinese idioms, aiming to improve idiom translation by both human and machine. The dataset is built by leveraging human and machine effort. Baseline generation models show unsatisfactory abilities to improve translation, but structure-aware classification models show good performance on distinguishing good translations. Furthermore, the size of PETCI can be easily increased without expertise. Overall, PETCI can be helpful to language learners and machine translation systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating LLMs on Chinese Idiom Translation

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    Across 900 annotated translation pairs from nine MT systems, the best system still mistranslates Chinese idioms in 28% of cases, and standard metrics miss these errors (Pearson correlation below 0.48).

  2. A Survey of Idiom Datasets for Psycholinguistic and Computational Research

    cs.CL 2025-08 conditional novelty 4.0 of 10

    A survey of 53 idiom datasets finds that psycholinguistic norming resources and computational idiom corpora have essentially no overlap or cross-usage.

Pith tools