{"id":"67e9106c-5bda-448f-b919-09f63fdd3509","arxiv_id":"2501.08621","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A shared-task report and dataset release for four Vietnamese-Chinese and Vietnamese-Lao translation directions, with human post-editing scores used for official rankings.","lead":"This paper reports the results of the VLSP 2022-2023 machine translation shared tasks for Vietnamese-Chinese and Vietnamese-Lao, including automatic BLEU scores and expert human rankings. It also releases a benchmark dataset of bilingual sentence pairs and human-corrected translations.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 9's own averaged human scores contradict the announced 2023 winner: MTA AI (56.17) beats Bluesky (52.83), so the ranking claim is internally inconsistent.","rationale":"The reader's rejection is well founded, and the stress-test finds a direct internal contradiction that does not depend on external assumptions. The paper itself defines the champion-selection procedure as averaging the two directions, yet the only table that should present those averages lacks the Final Score column and its rank structure is inconsistent with the announced winners. Recomputing the averages from the two per-direction human-evaluation tables yields a different champion. This is an internal contradiction, not a disagreement with the authors' choice of metric or with external consensus. It is load-bearing because the abstract and conclusions present the VLSP 2022-2023 rankings as the paper's deliverable; if the ranking table is unreliable, the central claim does not stand. The released dataset and the summaries of participant systems may still be useful, but neither rescues the ranking claim. A corrected table, an explicit scoring formula, and the underlying post-edit scores would be needed before the paper could serve as an authoritative record. The reader's weakest assumption about human post-editing reliability is related, but the controlling problem is the paper's own arithmetic contradiction, which needs no external data to expose. No ad hominem is intended; the issue is with the consistency of the reported results, not with the authors' intent.","tokens_in":13175,"tokens_out":5607,"duration_ms":48412,"concrete_test":"Independently reconstruct the 2023 official ranking from Tables 5 and 6 using the rule stated in the text: for each team, average the Human scores for Lo-Vi and Vi-Lo. Compare the resulting order with (a) the sentence 'the winning team is Bluesky, second is MTA AI' and (b) the rank column of Table 9. If the recomputed averages place MTA AI above Bluesky, the central ranking claim is contradicted. As a secondary check, inspect the released VLSP2023-MT leaderboard files, if present, for an actual Final Score column to determine which formula was used for the official standings.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is the official VLSP 2022-2023 rankings. Section 6 says systems were 'officially ranked according to human evaluation error calculated on all the collected postedits', and the text explicitly says the 2023 champion is chosen by averaging the two directions: 'The Final Score column of Table 9 shows that the winning team is Bluesky, second is MTA AI, third is BGSV AI.' Table 9, however, has no Final Score column, its rank column is malformed (Bluesky and MTA AI share rank 1; Faiz AIO and BGSV AI share rank 2; rank 3 is empty), and its one description ('SacreBLEU highest') points to a different metric from the human scores that are said to be decisive. Averaging the Human columns from Tables 5 and 6 gives MTA AI (51.03+61.31)/2 = 56.17, Bluesky (54.28+51.37)/2 = 52.83, Faiz AIO (47.83+46.51)/2 = 47.17, BGSV AI (27.51+31.56)/2 = 29.54. Under the stated averaging rule, MTA AI beats Bluesky, so the published champion contradicts the published data. Since the identity of the winning system is the most load-bearing factual assertion, this inconsistency cannot be dismissed as a mere typo in one column; it invalidates the advertised authoritative ranking unless a corrected table and explicit scoring formula are provided.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports on the VLSP 2022 and VLSP 2023 machine translation shared tasks for Vietnamese-Chinese and Vietnamese-Lao, respectively. It describes the released ViBidirectionMT-Eval dataset, the participating systems, and the automatic and human evaluations. The paper's main claims are the official rankings: SDS wins the 2022 Chinese-Vietnamese task and Bluesky wins the 2023 Lao-Vietnamese task, with MTA AI second and BGSV AI third. The paper also states that human post-editing evaluation was the decisive factor for the official rankings.","tokens_in":13443,"tokens_out":5728,"duration_ms":52372,"significance":"If the results and rankings were reliable, the paper would be a useful reference for low-resource machine translation between Vietnamese-Chinese and Vietnamese-Lao, and the released dataset on HuggingFace would be a community resource. The paper also contains strengths that should be acknowledged: the dataset covers four translation directions, uses 1,000 test sentences per direction, includes both automatic and human evaluation, and makes a distinction between constrained and unconstrained systems. However, the central ranking claims are undermined by internal inconsistencies in the reported scores and by a lack of detail in the human evaluation methodology, so the paper cannot currently serve as an authoritative record of the shared tasks.","major_comments":[{"comment":"The announced 2023 winner contradicts the data in the paper's own tables. The text states 'The Final Score column of Table 9 shows that the winning team is Bluesky, second is MTA AI, third is BGSV AI', but Table 9 contains no Final Score column. Averaging the Human columns of Tables 5 and 6 gives MTA AI (51.03+61.31)/2 = 56.17 versus Bluesky (54.28+51.37)/2 = 52.83; under the stated averaging rule, MTA AI would be the 2023 champion. The paper must provide the exact scoring formula and a corrected Table 9, or the central ranking claim is unsupported.","section":"§5, Tables 5, 6, and 9"},{"comment":"The official rankings depend entirely on human post-editing scores, but the paper never defines the scoring function referred to as 'human evaluation error', reports no inter-annotator agreement, and provides no significance testing. Section 6 also states that there were five submissions per task and five translators, yet Section 4 reports 5 official teams in 2022 and 7 in 2023, and Tables 5 and 6 list 7 and 5 teams respectively. These inconsistencies make it impossible to verify the final rankings.","section":"§6"},{"comment":"The sentence 'Due to various reasons, the S-NLP team could not complete the technical report, so we have removed the S-NLP team from the final standings' is duplicated from the 2022 discussion; S-NLP appears only in the 2022 tables. This duplication calls into question the reliability of the surrounding final-standings text and must be corrected.","section":"§5, VLSP 2023 paragraph"},{"comment":"The tables list SacreBLEU and 'Human' scores but do not define what the Human score represents, how it was computed, or whether higher values are better. The Description column in Table 9 labels Bluesky as 'SacreBLEU highest,' which refers to a different criterion from the human-based ranking claimed in the text; the relationship between the automatic and human metrics needs to be stated explicitly.","section":"§5, Tables 5 and 6"}],"minor_comments":[{"comment":"The manuscript contains numerous grammatical errors ('an results', 'submission were evaluated', 'this approarch') that should be corrected in a revision.","section":"Abstract and throughout"},{"comment":"The citation markers are malformed ('[ Papineni]', '[ Matt:2018]', '[ Cho2014LearningPR]'), and the reference list includes many entries that are not cited in the text (e.g., LAMBADA, TruthfulQA, BIG-Bench) while missing standard, correctly formatted references for BLEU and SacreBLEU.","section":"References"},{"comment":"Large blocks of text appear to be pasted from the participating teams' reports, including repeated figures and equations (e.g., Equations (1)-(3) in Section 4.2.3), which makes it difficult to distinguish the organizers' own evaluation description from the teams' internal reports.","section":"§4.1.1 and §4.2.3"},{"comment":"Team names are inconsistent across the paper ('F AIZ AIO' vs. 'Faiz AIO', 'BLUESKY' vs. 'Bluesky', 'TESTLA V100' vs. 'TeslaV100'), and 'ScareBLEU' appears instead of 'SacreBLEU' in the text.","section":"Tables 5 and 6"},{"comment":"The labels 'Result' and 'Final Score' are used without defining the unit or the direction of the scale; the text should explicitly state that these are percentages with higher values indicating better quality.","section":"Tables 7 and 8"}],"recommendation":"reject","confidential_remarks":"The manuscript appears to be a very rough draft with pasted sections from team reports. The contradiction between the announced winner and Table 9 is severe and load-bearing: the paper's main purpose is to report the official rankings, and the reported data do not support the stated winner. The lack of a defined human-evaluation scoring function and the inconsistent team counts further undermine the paper's reliability. Unless the authors can supply the raw scores and a corrected, fully specified evaluation description, the paper does not meet publication standards for a serious evaluation report."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe genuinely useful piece here is the artifact: a released benchmark for Vietnamese-Chinese and Vietnamese-Lao MT with human post-edited references. That fills a real gap for two low-resource pairs, and the raw 2022-2023 VLSP numbers are new. If you work on those pairs, the HuggingFace dataset is worth a look.\n\nThe paper itself, however, is not publishable in its current shape. The central ranking is internally inconsistent. The text says the 2023 champion is chosen by averaging the two human scores and declares Bluesky the winner, but Table 9 has no Final Score column, and averaging the Human columns from Tables 5 and 6 gives MTA AI (56.17) above Bluesky (52.83). The stress-test note is correct: this is a load-bearing contradiction, not a typo. The announced winner is contradicted by the paper's own numbers.\n\nThere are also integrity problems. Large blocks of prose appear copied from WMT shared-task papers—the subtitling paragraph in the Introduction and the post-editing rationale in Section 6 are essentially verbatim from earlier WMT reports. The same 'S-NLP removed' paragraph appears in both the 2022 and 2023 result sections. The reference list contains irrelevant entries (LAMBADA, MMLU, Vistral) that suggest the bibliography was assembled by copy-paste. And the human evaluation section never defines the scoring function, reports no inter-annotator agreement, and does no significance testing, yet all official rankings rest on those scores. That is a thin base even when the arithmetic is consistent.\n\nWhat the paper does well is modest: the automatic evaluation setup is standard SacreBLEU, and the team descriptions are serviceable summaries of standard techniques (mBART fine-tuning, back-translation, ensembling). The dataset release, if it is actually available and clean, is the real contribution.\n\nI would not send this to peer review as is. It needs the tables fixed, the scoring formula stated, the copied passages rewritten, and the duplicate paragraph removed. If the authors do that and the dataset checks out, a revised report would deserve a serious referee. For now, I would desk reject and invite a corrected submission.\n\nBest,","headline":"The released VLSP 2022-2023 MT benchmark is worth having, but this report's own tables contradict its announced winner, so it cannot be trusted as the official record.","tokens_in":13996,"tokens_out":3394,"would_cite":false,"duration_ms":30005,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The VLSP 2022–2023 MT shared tasks produced a public Vietnamese–Chinese and Vietnamese–Lao benchmark, with human post-editing scores naming SDS (2022) and Bluesky (2023) as winners.","keywords":["low-resource machine translation","Vietnamese-Chinese","Vietnamese-Lao","VLSP shared task","human post-editing evaluation","BLEU","SacreBLEU","benchmark dataset"],"falsifier":"Recompute the final Lao–Vietnamese standings from the Human/FinalScore column of Table 9: MTA AI's two-direction mean (61.31 + 51.03)/2 = 56.17 exceeds Bluesky's (54.28 + 51.37)/2 = 52.83, which contradicts the paper's announcement of Bluesky as champion; additionally, any release of individual translator scores that showed large disagreements would undermine the reliability of the human rankings.","tokens_in":12948,"feed_emoji":"🏆","tokens_out":5728,"duration_ms":49348,"temperature":0.7,"pith_summary":"This paper reports the organization and results of two machine translation shared tasks run under the VLSP evaluation campaign: Vietnamese–Chinese in 2022 and Vietnamese–Lao in 2023. It introduces the released evaluation dataset ViBidirectionMT-Eval, describes the participating systems, and presents both automatic scores and human post-editing scores for four translation directions. The paper's central claim is that the human post-editing evaluation, conducted by five professional translators per task, provides a reliable basis for the official system rankings, which name SDS as the 2022 champion and Bluesky as the 2023 champion. A sympathetic reader would care because the dataset and rankings are a reusable public benchmark for low-resource Southeast Asian language pairs.","feed_headline":"New Vietnamese MT benchmark crowns SDS and Bluesky","feed_subtitle":"Human post-editing scores rank five and seven systems for four translation directions.","key_machinery":"The load-bearing mechanism is the post-editing human evaluation framework: for each translation direction, outputs from all participating systems are randomly assigned in equal numbers to five professional translators, who post-edit them; the resulting edited texts become new reference translations, and systems are ranked by a human evaluation error score computed against all collected post-edits. This protocol is what turns subjective judgments into a ranked table, and it is the basis for the paper's official winner announcements.","core_discovery":"On its own terms, the paper establishes that the VLSP 2022–2023 machine translation tasks can be meaningfully evaluated by combining automatic metrics (BLEU, SacreBLEU) with a post-editing protocol in which five translators edit each system's output and the edited versions serve as additional references. Based on that protocol, the official final standings are: for Chinese–Vietnamese, SDS first, VBD-MT second, JNLP third, VC-Datamining fourth; for Lao–Vietnamese, Bluesky first, MTA AI second, BGSV AI third. The paper also contributes the ViBidirectionMT-Eval dataset, including human post-edits, as a reusable resource for these language pairs.","pith_inferences":["The announced 2023 winner does not follow from the paper's own numbers: averaging the two direction scores in Table 9 gives MTA AI an arithmetic mean of 56.17 versus Bluesky's 52.83, so the stated ranking appears inconsistent with the presented data.","Because no inter-annotator agreement or significance test is reported, the human scores should be treated as a descriptive summary rather than a statistically grounded comparison.","The same post-editing protocol could be extended to other under-resourced Southeast Asian language pairs, such as Vietnamese–Khmer, using the released dataset as a template.","Readers should verify the reproducibility of the human scoring by checking whether the five translators' individual post-edit scores are released alongside the aggregate."],"forward_implications":["Future MT systems for Vietnamese–Chinese and Vietnamese–Lao can be compared directly against the published scores and dataset.","The post-edited outputs provide additional reference translations that can be used for training and evaluation beyond the original test set.","The demonstrated success of data synthesis, back-translation, and mBART fine-tuning offers a template for other low-resource language pairs in the region."],"supporting_citations":[{"why":"Supplies the BLEU metric used for automatic evaluation of all system outputs.","marker":"[Papineni]"},{"why":"Supplies SacreBLEU, the recommended and reported automatic metric.","marker":"[Matt:2018]"},{"why":"Source of the mBART-50 model that the winning 2022 team (SDS) fine-tunes and that Bluesky adapts for Lao.","marker":"[Tang2020MultilingualTW]"},{"why":"Provides the Fairseq baseline used by the VBD-MT system for the 2022 Chinese–Vietnamese task.","marker":"[ott-etal-2019-fairseq]"},{"why":"Supplies the SentencePiece tokenization used by multiple systems, including Bluesky and BGSV AI.","marker":"[kudo-richardson-2018-sentencepiece]"}],"fun_headline_variants":["SDS and Bluesky top VLSP Vietnamese MT shared tasks","ViBidirectionMT-Eval: dataset and rankings for Vietnamese MT","Human post-edits improve MT evaluation for Vietnamese pairs","VLSP 2022-23 MT: SDS first for zh-vi, Bluesky for lo-vi","New MT benchmark combines BLEU and human edits for Vietnamese"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire ranking rests on the assumption that the human post-editing scores—averaged across five translators with no reported inter-annotator agreement or significance testing—accurately reflect translation quality.","fun_headline_variants_meta":{"raw":{"variants":["SDS and Bluesky top VLSP Vietnamese MT shared tasks","ViBidirectionMT-Eval: dataset and rankings for Vietnamese MT","Human post-edits improve MT evaluation for Vietnamese pairs","VLSP 2022-23 MT: SDS first for zh-vi, Bluesky for lo-vi","New MT benchmark combines BLEU and human edits for Vietnamese"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000301,"raw_usage":{"total_tokens":1682,"prompt_tokens":835,"completion_tokens":847,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":451,"completion_tokens_details":{"reasoning_tokens":751}},"tokens_in":451,"tokens_out":847,"duration_ms":8050,"temperature":1.0,"reasoning_tokens":751,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:21:39.522800+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the final Lao–Vietnamese standings from the Human/FinalScore column of Table 9: MTA AI's two-direction mean (61.31 + 51.03)/2 = 56.17 exceeds Bluesky's (54.28 + 51.37)/2 = 52.83, which contradicts the paper's announcement of Bluesky as champion; additionally, any release of individual translator scores that showed large disagreements would undermine the reliability of the human rankings.","supporting_citations":[],"review_version":1}