REVIEW 2 cited by
Preliminary WMT24 Ranking of General MT Systems and LLMs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This is the preliminary ranking of WMT24 General MT systems based on automatic metrics. The official ranking will be a human evaluation, which is superior to the automatic ranking and supersedes it. The purpose of this report is not to interpret any findings but only provide preliminary results to the participants of the General MT task that may be useful during the writing of the system submission.
Forward citations
Cited by 2 Pith papers
-
ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities
ConsistencyChecker ranks LLMs by how well they survive chains of reversible transformations, and those scores track WMT 2024 translation quality rankings (r > 0.7) without using WMT paired data.
-
How to Select Datapoints for Efficient Human Evaluation of NLG Models?
Selecting human-evaluation items by metric variance, metric consistency, output diversity, or IRT-based informativeness matches random-sampling ranking accuracy with roughly 70% of the annotation budget in WMT23 and SummEval.
Discussion (0). Continue with ORCID to comment.