Pith. sign in

REVIEW 1 cited by

Examining the Tip of the Iceberg: A Data Set for Idiom Translation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1802.04681 v1 pith:VNC6NIU2 submitted 2018-02-13 cs.CL

classification cs.CL
keywords translationdataidiomfirstidiomslanguagebettercorpus
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Neural Machine Translation (NMT) has been widely used in recent years with significant improvements for many language pairs. Although state-of-the-art NMT systems are generating progressively better translations, idiom translation remains one of the open challenges in this field. Idioms, a category of multiword expressions, are an interesting language phenomenon where the overall meaning of the expression cannot be composed from the meanings of its parts. A first important challenge is the lack of dedicated data sets for learning and evaluating idiom translation. In this paper we address this problem by creating the first large-scale data set for idiom translation. Our data set is automatically extracted from a widely used German-English translation corpus and includes, for each language direction, a targeted evaluation set where all sentences contain idioms and a regular training corpus where sentences including idioms are marked. We release this data set and use it to perform preliminary NMT experiments as the first step towards better idiom translation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating LLMs on Chinese Idiom Translation

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    Across 900 annotated translation pairs from nine MT systems, the best system still mistranslates Chinese idioms in 28% of cases, and standard metrics miss these errors (Pearson correlation below 0.48).

Pith tools