REVIEW 2 cited by
On the Implications of Verbose LLM Outputs: A Case Study in Translation Evaluation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper investigates the impact of verbose LLM translations on evaluation. We first demonstrate the prevalence of this behavior across several LLM outputs drawn from the WMT 2024 general shared task on machine translation. We then identify the primary triggers of verbosity, including safety, copyright concerns, and insufficient context in short input queries. Finally, we show that ignoring this behavior unfairly penalizes more verbose LLMs according to both automatic and human evaluations, highlighting the need to address this issue for more accurate future evaluations.
Forward citations
Cited by 2 Pith papers
-
How Important is `Perfect' English for Machine Translation Prompts?
For LLM machine translation, prompt choice affects output quality more than realistic user errors, with spelling errors hurting most and phrase-level errors often harmless.
-
GPL-SLAM: A Laser SLAM Framework with Gaussian Process Based Extended Landmarks
A laser SLAM framework that models each object as a Gaussian-process contour, updated recursively and inferred jointly with the robot pose in a Bayesian framework.
Discussion (0). Continue with ORCID to comment.