Pith. sign in

REVIEW 3 major objections 5 minor 69 references

Multi-Hypothesis Distillation of Multilingual Neural Translation Models for Low-Resource Languages

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Training a small translation model on several translations of each source sentence, rather than the teacher's single best beam-search output, improves quality for low-resource languages and also softens two known side effects of…

desk verdict A careful, reproducible empirical study showing that training students on multiple sampled teacher translations beats single-beam KD for low-resource pairs; the main claims hold, though the abstract oversells the corpus-size result and the hallucination analysis is thin. read the letter →

arxiv 2507.21568 v2 pith:XLIOGDKK submitted 2025-07-29 cs.CL

classification cs.CL
keywords multi-hypothesisdistillationsequence-levelknowledgelow-resourcemachinetranslationdecodingmethodsmultilingualneuralgenderbiashallucinationschrF++
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Large multilingual translation models are too big for many real uses, so they are often compacted by knowledge distillation: a small student is trained on translations that a large teacher produces. The standard recipe keeps only the teacher's single most-likely beam-search translation per sentence. This paper argues that this discards most of what the teacher knows and proposes Multi-Hypothesis Distillation (MHD), which generates several translations per source sentence and trains the student on all of them. Across seven low-resource directions involving Swahili, Igbo, and Bambara, MHD students match or beat the standard distilled students, even when the added translations are individually lower quality than the beam output. The same diversity also reduces the gender-bias amplification and the hallucinations that distillation typically brings.

What carries the argument

The central object is the teacher's output distribution, represented in MHD by $M$ decoded hypotheses per source instead of a single mode. The training signal is $$\mathcal{L}_{\mathrm{MHD}} = -\sum_{i=1}^{N}\sum_{m=1}^{M}\sum_{t=1}^{T_i} \log P(\tilde{y}_{i,m,t} \mid \tilde{y}_{i,m,<t}, x_i; \theta_S),$$ which is simply the standard sequence-level KD loss on a corpus where each source sentence appears $M$ times. The decoding method is what shapes that corpus: beam search and diverse beam search return ranked high-probability lists, giving low variability and, for poorly fitted languages, increasingly improbable continuations as $M$ grows; top-$p$ and top-$k$ sample independently, giving high lexical variability and stable probabilities; MBR reranks epsilon-sampled candidates by expected chrF, filtering out bad translations. The paper uses these properties to explain when MHD helps: sampling wins where the teacher is weak or the corpus is small, while high-quality ranked outputs regain the edge when monolingual data are abundant.

What would settle it

Have human translators rank the FLORES+ devtest outputs of the D1-BS, D10-top-p, and D10-MBR students for eng-ibo and bam-swh; if the sampling- or MBR-trained students do not beat the beam-KD student, the central claim fails.

Watch

Extended reading notes

Core claim

MHD is a sequence-level knowledge distillation method: a teacher translation model (NLLB-200 in the 1.3B and 3.3B sizes) decodes $M$ hypotheses $\tilde{y}_{i,1}, \ldots, \tilde{y}_{i,M}$ for each source sentence $x_i$ using one of five decoding strategies, and the student is trained with the standard cross-entropy objective on the repeated dataset, so every source appears $M$ times paired with a different target. The paper's central empirical claim is that, for low-resource directions, this beats the standard sequence-level KD baseline $D^1_{BS}$ in which each source has only its beam-search output. Gains are largest when the hypotheses are sampled (top-$p$ and top-$k$) rather than ranked (beam search and diverse beam search), with MBR decoding best for the least-resourced Bambara pairs. The paper also reports that multi-hypothesis training reduces gender-bias amplification as measured by contrastive conditioning, and reduces hallucinations in most settings, although multiple beam-search hypotheses from a weak teacher can increase them for Bambara.

Load-bearing premise

The load-bearing premise is that chrF++ on the FLORES+ devtest set judges translation quality faithfully for all seven language pairs; the paper validates against COMET only for English-Swahili and Swahili-English, so a metric artifact for Igbo or Bambara would undermine the ranking at the center of the claim.

Editorial extensions

If this is right

  • MHD reaches results comparable to standard sequence-level KD while using a much smaller monolingual corpus, so the method lowers the data requirement for distilling a competitive student.
  • For the lowest-resource directions (involving Bambara and the into-English pairs), sampling-based hypotheses give the strongest students; top-$p$ is nearly as effective as MBR and much faster.
  • With one million source sentences, the advantage of sampling over beam search narrows, and diverse beam search can even lead for Swahili-English; the right decoding choice depends on corpus size and teacher quality.
  • Where the teacher is poorly calibrated, multiple beam-search hypotheses degrade student quality, while multiple sampled hypotheses keep improving it, which is direct evidence that the mode is not a good summary of the teacher's distribution.
  • Training on $M=10$ hypotheses systematically mitigates gender-bias amplification compared to a single beam hypothesis, with sampling methods reducing it most.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: if the gains come from diversity rather than from the identity of any single good translation, then controlling diversity directly, for example by tuning top-$p$ while filtering repeated sentences, could further improve student models at lower cost than MBR.
  • Beyond the paper: the paper's vocabulary-coverage curves suggest a practical recipe: with scarce monolingual data, spend distillation effort on covering the target vocabulary with several sampled hypotheses first, then switch to high-quality single hypotheses once coverage saturates.
  • Beyond the paper: because MHD needs only monolingual text and access to the teacher's outputs, the same recipe could be applied to even lower-resource languages or to API-only teachers, and could be combined with back-translation to grow the source side as well.
  • Beyond the paper: the multi-hypothesis idea is not specific to translation; other sequence-generation tasks where the teacher's mode is unrepresentative might benefit from the same keep-several-outputs training signal, but that is an extrapolation.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes Multi-Hypothesis Distillation (MHD), a sequence-level knowledge distillation method that trains a compact student translation model on multiple teacher-generated translations per source sentence. The teacher is an NLLB-200 model and the students are 65M-parameter Transformers. Experiments cover seven low-resource directions (eng-swh, eng-ibo, eng-bam, swh-eng, ibo-eng, bam-eng, bam-swh), two teacher sizes (1.3B and 3.3B), several decoding methods (beam search, diverse beam search, top-p, top-k, MBR), sweeps over the number of hypotheses M, corpus-size sweeps, and analyses of gender bias, hallucinations, vocabulary coverage, and decoding-parameter sensitivity. The central claim is that MHD with M=10 hypotheses, especially with sampling-based decoding, improves student performance over standard single-hypothesis beam-search KD while also reducing gender-bias amplification and hallucinations.

Significance. If the claims hold, the paper makes a useful practical contribution: it shows that a black-box multilingual teacher accessed only through decoding can be distilled into a much smaller bilingual student using monolingual data, and that sampling-based hypothesis generation can beat beam-search KD in low-resource settings. The study is unusually thorough: 482 trained models, two teacher scales, significance testing with paired approximate randomization, and public code. The vocabulary-coverage and corpus-size analyses (Figures 7-9) are informative and could guide practitioners. The claims are falsifiable and the experimental protocol is mostly reproducible from the description.

major comments (3)
  1. [Section 5.1, Eq. (3), Fig. 2] The comparison between D^10_Z and D^1_BS confounds the number of hypotheses per source with the total number of training examples and optimization steps. A 100k-source D^10_Z corpus contains 1M target sentences, while D^1_BS contains 100k target sentences. The 'best translation per source' control reported in Section 5.1 rules out a single lucky translation but does not control for data quantity. To attribute the gains to diversity rather than to tenfold more training data, please add a control in which the single D^1_BS translation is repeated ten times per source sentence, or otherwise match the total number of target sentences across conditions.
  2. [Section 5.1, Tables 8-9] The zero-shot direction bam-swh shows a clear metric disagreement: BLEU ranks D10_BS below D1_BS (1.2 vs 2.1), while chrF++ ranks it above (20.4 vs 8.7). The conclusion that MHD with beam search helps in the zero-shot scenario therefore rests entirely on chrF++, which is validated with COMET only for eng-swh and swh-eng in Section 4.3. Please add a neural metric or human evaluation for at least bam-swh, or restrict the claim to sampling-based MHD, for which BLEU and chrF++ agree in direction.
  3. [Section 5.4, Table 4] The gender-bias reductions reported in Table 4 are small (e.g., eng-swh D10_BS 51.0 vs D1_BS 49.2; eng-bam D10_BS 50.3 vs 50.8) and no significance testing, confidence intervals, or run-level variance is reported. Since bias mitigation is part of the headline contribution, please provide uncertainty estimates or significance tests for these differences, or soften the corresponding conclusion.
minor comments (5)
  1. [Table 1 caption] The caption reads 'ChrfF++ scores'; the metric should be written 'chrF++'.
  2. [Section 2.1] The text spells 'Kullback-Leiber'; the correct name is 'Kullback-Leibler'.
  3. [Figure 8 caption] The caption contains an unmatched parenthesis in 'the swh-eng) training corpus'; please fix the parenthetical.
  4. [Appendix D.4, Tables 8-9] The captions refer to underlined and bolded values, but these visual markers are not described in the text; please state explicitly in the caption which comparison each marker refers to and ensure the markers are visible in the published PDF.
  5. [Template front matter] The received/revised dates in the JAIR template ('Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009') appear to be leftover template text and should be updated or removed.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the MHD objective is a standard cross-entropy loss over teacher-generated hypotheses, and the central comparison against single-beam KD is evaluated on external FLORES+ benchmarks without fitting to the test set.

full rationale

The paper's derivation chain is self-contained and empirical. Equation 3 defines MHD simply as sequence-level cross-entropy over M teacher-generated translations per source sentence; no parameter of the method is fitted to FLORES+ devtest, and the student models are evaluated with external metrics (chrF++, BLEU, and COMET for eng-swh and swh-eng) against held-out references. The central claim that sampling-based MHD with M=10 outperforms standard D1_BS sequence-level KD is not defined in terms of its own output: the teacher's synthetic corpora are produced by fixed decoding methods (beam search, diverse beam search, top-p, top-k, MBR) with standard hyperparameters, and student quality is measured independently on FLORES+. The paper even runs a control experiment selecting only the best COMET-without-reference translation per source for eng-swh D10_top-p, obtaining performance similar to D1_top-p, which directly addresses the concern that a single good translation drives the result. The only metric-alignment caveat is the acknowledged use of fastChrF as the MBR utility function while chrF++ is the primary evaluation metric; the authors explicitly state 'we cannot rule out a metric bias introduced by using the same type of metric to rank the MBR candidates and for evaluation', and they report BLEU rankings showing MBR is not most effective under BLEU. This is an honest limitation affecting the MBR-specific comparison, not the main sampling-versus-beam-search claim, and it is not a hidden reduction of the method to its evaluation. Self-citations [17] and [18] are descriptive references to the authors' prior NAACL Findings paper and related work on low-resource NMT; they are not invoked as uniqueness theorems or as substitutes for the experimental evidence, so they are not load-bearing circularity. No fitted parameter is renamed as a prediction, no known result is repackaged under new coordinates, and no claim reduces by construction to an input of the paper.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The paper does not introduce new theoretical entities, forces, or conserved quantities. Its empirical claims rest on standard machine-learning assumptions and on prior-work hypotheses about decoding distributions, evaluation metrics, and bias measurement.

free parameters (5)
  • M (number of hypotheses per source sentence) = 10
    The number of translations generated per source sentence, fixed to 10 for the corpus-size experiments even though the optimal M varies by language pair (e.g., for bam-swh, D3_BS beats D5_BS and D10_BS).
  • p (top-p sampling threshold) = 0.7
    Nucleus sampling parameter chosen following Eikema and Aziz [12]; sensitivity analysis in Section 5.3 shows student performance is stable across p values.
  • k (top-k sampling size) = 10
    Top-k parameter from Fan et al. [15], also used in back-translation work [66]; sensitivity analysis in Section 5.3.
  • epsilon (MBR candidate sampling threshold) = 0.02
    Epsilon sampling threshold for generating 256 MBR candidates, following Finkelstein and Freitag [16].
  • beam width and DBS diversity penalty = n=10, lambda=0.5
    Beam size for BS and DBS, and diversity penalty for DBS, following Vijayakumar et al. [57].
assumptions (5)
  • domain assumption The teacher model's output distribution beyond the mode contains transferable knowledge for the student.
    This is the motivating hypothesis for MHD, adopted from Eikema and Aziz [11] 'the inadequacy of the mode', not established within this paper.
  • domain assumption chrF++ on FLORES+ devtest is a valid proxy for translation quality in all seven language pairs.
    Section 4.3 adopts chrF++ as the primary metric; COMET is only reported for eng-swh and swh-eng, so evaluation of Igbo and Bambara relies on chrF++ alone.
  • domain assumption Contrastive conditioning with NLLB-200 1.3B as evaluator measures gender bias amplification in the target languages.
    Section 5.4 uses the English-only WinoMT dataset and NLLB-200 1.3B as the evaluator; the validity of this proxy for Swahili, Igbo, and Bambara is assumed from Vamvas and Sennrich [54].
  • domain assumption SONAR embedding cosine similarity identifies hallucinations.
    Section 5.5 relies on prior findings [6,62] that cross-lingual sentence embeddings outperform COMET for hallucination detection.
  • domain assumption The chosen monolingual corpora (OSCAR, ParaCrawl, bayelemabaga, etc.) are representative of the target languages.
    Section 4.2 and Appendix C; noisy or unrepresentative corpora could change the diversity-quality trade-off and the corpus-size conclusions.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Multi-Hypothesis Distillation of Multilingual Neural Translation Models for Low-Resource Languages." pith.science (2026). https://pith.science/paper/XLIOGDKK

@misc{pith2026250721568,
  author       = {Pith},
  title        = {Pith review of: Multi-Hypothesis Distillation of Multilingual Neural Translation Models for Low-Resource Languages},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XLIOGDKK}},
  note         = {Machine review of arXiv:2507.21568}
}
abstract

This paper explores sequence-level knowledge distillation (KD) of multilingual pre-trained encoder-decoder translation models. We argue that the teacher model's output distribution holds valuable insights for the student, beyond the approximated mode obtained through beam search (the standard decoding method), and present Multi-Hypothesis Distillation (MHD), a sequence-level KD method that generates multiple translations for each source sentence. This provides a larger representation of the teacher model distribution and exposes the student model to a wider range of target-side prefixes. We leverage $n$-best lists from beam search to guide the student's learning and examine alternative decoding methods to address issues like low variability and the under-representation of infrequent tokens. For low-resource languages, our research shows that while sampling methods may slightly compromise translation quality compared to beam search based approaches, they enhance the generated corpora with greater variability and lexical richness. This ultimately improves student model performance and mitigates the gender bias amplification often associated with KD.

Figures

Figures reproduced from arXiv: 2507.21568 by the authors.

Figure 1
Figure 1. Multi-Hypothesis Knowledge Distillation (MHD) approach. The teacher model generates [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Average chrF++ score obtained by student models trained on M samples generated with different decoding methods [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Similarity among 10 generated translations per source sentence as evaluated by self-BLEU (y-axis). [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figures from the paper (12 more)
Figure 4
Figure 4. Figure 4: Zipf’s distribution over Swahili corpora. Similar patterns were observed for the other languages. [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]
Figure 5
Figure 5. Figure 5: Probabilities normalised by length of 10 swh-eng translation hypotheses generated with NLLB-200 1.3B for each [PITH_FULL_IMAGE:figures/full_fig_p011_5.png]
Figure 6
Figure 6. Figure 6: Percentage of translations (y-axis) for the same source with a probability lower than the median of the 256 translations [PITH_FULL_IMAGE:figures/full_fig_p012_6.png]
Figure 7
Figure 7. Figure 7: Scores attained for different corpus sizes in number of sentences (x-axis). [PITH_FULL_IMAGE:figures/full_fig_p013_7.png]
Figure 8
Figure 8. Figure 8: Relationship between the vocabulary size of the swh-eng) training corpus and the chrF++ of the student models. [PITH_FULL_IMAGE:figures/full_fig_p013_8.png]
Figure 9
Figure 9. Figure 9: Effect of vocabulary coverage and teacher translation quality. X-axis shows decoding methods ranked by variability. [PITH_FULL_IMAGE:figures/full_fig_p014_9.png]
Figure 10
Figure 10. Figure 10: Relationship between the teacher translation quality and variability and the student models score for eng-swh. [PITH_FULL_IMAGE:figures/full_fig_p014_10.png]
Figure 11
Figure 11. Figure 11: Scheme of contrastive conditioning. The evaluator model calculates the probability of the generated translation for [PITH_FULL_IMAGE:figures/full_fig_p015_11.png]
Figure 12
Figure 12. Figure 12: Kernel density estimations (bandwidth=1.0) for SONAR-based cosine similarities between the output produced by [PITH_FULL_IMAGE:figures/full_fig_p017_12.png]
Figure 13
Figure 13. Figure 13: Average BLEU score obtained by student models trained on M samples generated with different decoding methods [PITH_FULL_IMAGE:figures/full_fig_p026_13.png]
Figure 14
Figure 14. Figure 14: Average chrF++ score obtained by student models trained on M samples generated by NLLB-200 3.3B with different [PITH_FULL_IMAGE:figures/full_fig_p026_14.png]
Figure 15
Figure 15. Figure 15: Average BLEU score obtained by student models trained on M samples generated by NLLB-200 3.3B with different [PITH_FULL_IMAGE:figures/full_fig_p027_15.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

69 extracted references · 32 canonical work pages

  1. [1]

    David Adelani, Md Mahfuz Ibn Alam, Antonios Anastasopoulos, Akshita Bhagia, Marta R. Costa-jussà, Jesse Dodge, Fahim Faisal, Christian Federmann, Natalia Fedorova, Francisco Guzmán, Sergey Koshelev, Jean Maillard, Vukosi Marivate, Jonathan Mbuya, Alexandre Mourachko, Safiyyah Saleem, Holger Schwenk, and Guillaume Wenzek. 2022. Findings of the WMT’22 Share...

  2. [2]

    Rishabh Agarwal, Nino Vieillard, Yongchao Zhou, Piotr Stanczyk, Sabela Ramos Garea, Matthieu Geist, and Olivier Bachem. 2024. On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes. InThe Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=3zKtaqxLhW

  3. [3]

    Jaimeen Ahn, Hwaran Lee, Jinhwa Kim, and Alice Oh. 2022. Why Knowledge Distillation Amplifies Gender Bias and How to Mitigate from the Perspective of DistilBERT. InProceedings of the 4th Workshop on Gender Bias in Natural Language Processing (GeBNLP). Association for Computational Linguistics, Seattle, Washington, 266–272. https://doi.org/10.18653/v1/2022...

  4. [4]

    Varshney

    Sourya Basu, Govardana Sachitanandam Ramachandran, Nitish Shirish Keskar, and Lav R. Varshney. 2021. Mirostat: A Neural Text Decoding Algorithm that Directly Controls Perplexity. arXiv:2007.14966 [cs.CL]

  5. [5]

    Laurie Burchell, Alexandra Birch, and Kenneth Heafield. 2022. Exploring diversity in back translation for low-resource machine translation. InProceedings of the Third Workshop on Deep Learning for Low-Resource Natural Language Processing. Association for Computational Linguistics, Hybrid, 67–79. https://doi.org/10.18653/v1/2022.deeplo-1.8

  6. [6]

    Costa-jussà

    David Dale, Elena Voita, Loic Barrault, and Marta R. Costa-jussà. 2023. Detecting and Mitigating Hallucinations in Machine Translation: Model Internal Workings Alone Do Well, Sentence Similarity Even Better. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Anna Rogers, Jordan Boyd-Graber, a...

  7. [7]

    Ona De Gibert, Raúl Vázquez, Mikko Aulamo, Yves Scherrer, Sami Virpioja, and Jörg Tiedemann. 2023. Four Approaches to Low-Resource Multilingual NMT: The Helsinki Submission to the AmericasNLP 2023 Shared Task. InProceedings of the Workshop on Natural Language Processing for Indigenous Languages of the Americas (AmericasNLP), Manuel Mager, Abteen Ebrahimi,...

  8. [8]

    Alexandra DeLucia, Aaron Mueller, Xiang Lisa Li, and João Sedoc. 2021. Decoding Methods for Neural Narrative Generation. In Proceedings of the 1st Workshop on Natural Language Generation, Evaluation, and Metrics (GEM 2021). Association for Computational Linguistics, Online, 166–185. https://doi.org/10.18653/v1/2021.gem-1.16 Submited to JAIR on July 2025. ...

Show all 69 references
  1. [9]

    Heejin Do and Gary Geunbae Lee. 2023. Target-Oriented Knowledge Distillation with Language-Family-Based Grouping for Multilingual NMT.ACM Trans. Asian Low-Resour. Lang. Inf. Process.22, 2, Article 42 (mar 2023), 18 pages. https://doi.org/10.1145/3546067

  2. [10]

    Paul-Ambroise Duquenne, Holger Schwenk, and Benoît Sagot. 2023. SONAR: Sentence-Level Multimodal and Language-Agnostic Representations. arXiv:2308.11466 [cs.CL] https://arxiv.org/abs/2308.11466

  3. [11]

    Bryan Eikema and Wilker Aziz. 2020. Is MAP Decoding All You Need? The Inadequacy of the Mode in Neural Machine Translation. In Proceedings of the 28th International Conference on Computational Linguistics. International Committee on Computational Linguistics, Barcelona, Spain ...

  4. [12]

    Bryan Eikema and Wilker Aziz. 2022. Sampling-Based Approximations to Minimum Bayes Risk Decoding for Neural Machine Translation. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). As...

  5. [13]

    Maxim Enis and Mark Hopkins. 2024. From LLM to NMT: Advancing Low-Resource Machine Translation with Claude. arXiv:2404.13813 [cs.CL] https://arxiv.org/abs/2404.13813

  6. [14]

    Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, Naman Goyal, Tom Birch, Vitaliy Liptchinsky, Sergey Edunov, Michael Auli, and Armand Joulin. 2021. Beyond English-Centric M...

  7. [15]

    Angela Fan, Mike Lewis, and Yann Dauphin. 2018. Hierarchical Neural Story Generation. arXiv:1805.04833 [cs.CL]

  8. [16]

    Mara Finkelstein and Markus Freitag. 2024. MBR and QE Finetuning: Training-time Distillation of the Best and Most Expensive Decoding Methods. InThe Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=bkNx3O0sND

  9. [17]

    Sánchez-Cartagena

    Aarón Galiano-Jiménez, Juan Antonio Pérez-Ortiz, Felipe Sánchez-Martínez, and Víctor M. Sánchez-Cartagena. 2025. Beyond the Mode: Sequence-Level Distillation of Multilingual Translation Models for Low-Resource Language Pairs. InFindings of the Association for Computational Lin...

  10. [18]

    Sánchez-Cartagena, and Juan Antonio Pérez-Ortiz

    Aarón Galiano-Jiménez, Felipe Sánchez-Martínez, Víctor M. Sánchez-Cartagena, and Juan Antonio Pérez-Ortiz. 2023. Exploiting large pre-trained models for low-resource neural machine translation. InProceedings of the 24th Annual Conference of the European Association for Machine...

  11. [19]

    Vikrant Goyal, Sourav Kumar, and Dipti Misra Sharma. 2020. Efficient Neural Machine Translation for Low-Resource Languages via Exploiting Related Languages. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop. As...

  12. [20]

    Miguel Graça, Yunsu Kim, Julian Schamper, Shahram Khadivi, and Hermann Ney. 2019. Generalizing Back-Translation in Neural Machine Translation. InProceedings of the Fourth Conference on Machine Translation (Volume 1: Research Papers). Association for Computational Linguistics, ...

  13. [21]

    Alex Graves. 2012. Sequence Transduction with Recurrent Neural Networks. arXiv:1211.3711 [cs.NE]

  14. [22]

    Guerreiro, Elena Voita, and André Martins

    Nuno M. Guerreiro, Elena Voita, and André Martins. 2023. Looking for a Needle in a Haystack: A Comprehensive Study of Hallucinations in Neural Machine Translation. InProceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, An...

  15. [23]

    Varun Gumma, Raj Dabre, and Pratyush Kumar. 2023. An Empirical Study of Leveraging Knowledge Distillation for Compressing Multilingual Neural Machine Translation Models. InProceedings of the 24th Annual Conference of the European Association for Ma- chine Translation, Mary Nur...

  16. [24]

    John Hewitt, Christopher Manning, and Percy Liang. 2022. Truncation Sampling as Language Model Desmoothing. InFindings of the Association for Computational Linguistics: EMNLP 2022, Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang (Eds.). Association for Computational Linguistic...

  17. [25]

    Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015. Distilling the Knowledge in a Neural Network. arXiv:1503.02531 [stat.ML]

  18. [26]

    Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The Curious Case of Neural Text Degeneration. arXiv:1904.09751 [cs.CL]

  19. [27]

    Vivek Iyer, Bhavitvya Malik, Pavel Stepachev, Pinzhen Chen, Barry Haddow, and Alexandra Birch. 2024. Quality or Quantity? On Data Scale and Diversity in Adapting Large Language Models for Low-Resource Translation. InProceedings of the Ninth Conference on Submited to JAIR on Ju...

  20. [28]

    Yoon Kim and Alexander M. Rush. 2016. Sequence-Level Knowledge Distillation. InProceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Austin, Texas, 1317–1327. https://doi.org/10.18653/v1/D16- 1139

  21. [29]

    Kingma and Jimmy Ba

    Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Optimization. In3rd International Conference on Learning Representations, ICLR 2015, Conference Track Proc.http://arxiv.org/abs/1412.6980

  22. [30]

    Tom Kocmi, Eleftherios Avramidis, Rachel Bawden, Ondřej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Markus Freitag, Thamme Gowda, Roman Grundkiewicz, Barry Haddow, Marzena Karpinska, Philipp Koehn, Benjamin Marie, Christof Monz, Kenton Murray, Masaaki Nagata, Ma...

  23. [31]

    Geza Kovacs, Daniel Deutsch, and Markus Freitag. 2024. Mitigating Metric Bias in Minimum Bayes Risk Decoding. InProceedings of the Ninth Conference on Machine Translation, Barry Haddow, Tom Kocmi, Philipp Koehn, and Christof Monz (Eds.). Association for Computational Linguisti...

  24. [32]

    Taku Kudo and John Richardson. 2018. SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing. InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. Association for ...

  25. [33]

    Choquette-Choo, Katherine Lee, Derrick Xin, Aditya Kusupati, Romi Stella, Ankur Bapna, and Orhan Firat

    Sneha Kudugunta, Isaac Caswell, Biao Zhang, Xavier Garcia, Christopher A. Choquette-Choo, Katherine Lee, Derrick Xin, Aditya Kusupati, Romi Stella, Ankur Bapna, and Orhan Firat. 2023. MADLAD-400: A Multilingual And Document-Level Large Audited Dataset. arXiv:2309.04662 [cs.CL]

  26. [34]

    Ilia Kulikov, Alexander Miller, Kyunghyun Cho, and Jason Weston. 2019. Importance of Search and Evaluation Strategies in Neural Dialogue Modeling. InProceedings of the 12th International Conference on Natural Language Generation. Association for Computational Linguistics, Toky...

  27. [35]

    Solomon Kullback and Richard A Leibler. 1951. On Information and Sufficiency.The Annals of Mathematical Statistics22, 1 (1951), 79–86

  28. [36]

    Shankar Kumar and William Byrne. 2004. Minimum Bayes-Risk Decoding for Statistical Machine Translation. InProceedings of the Human Language Technology Conference of the North American Chapter of the Association for Computational Linguistics: HLT-NAACL

  29. [37]

    Wen Lai, Jindřich Libovický, and Alexander Fraser. 2021. The LMU Munich System for the WMT 2021 Large-Scale Multilingual Machine Translation Shared Task. InProceedings of the Sixth Conference on Machine Translation. Association for Computational Linguistics, Online, 412–417. h...

  30. [38]

    Senyu Li, Jiayi Wang, Felermino D. M. A. Ali, Colin Cherry, Daniel Deutsch, Eleftheria Briakou, Rui Sousa-Silva, Henrique Lopes Cardoso, Pontus Stenetorp, and David Ifeoluwa Adelani. 2025. SSA-COMET: Do LLMs Outperform Learned Metrics in Evaluating MT for Under-Resourced Afric...

  31. [39]

    Mathias Müller and Rico Sennrich. 2021. Understanding the Properties of Minimum Bayes Risk Decoding in Neural Machine Translation. InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural L...

  32. [40]

    NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prang...

  33. [41]

    Krishna Pillutla, Swabha Swayamdipta, Rowan Zellers, John Thickstun, Sean Welleck, Yejin Choi, and Zaid Harchaoui. 2021. MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence Frontiers. arXiv:2102.01454 [cs.CL]

  34. [42]

    Maja Popović. 2017. chrF++: words helping character n-grams. InProceedings of the Second Conference on Machine Translation. Association for Computational Linguistics, Copenhagen, Denmark, 612–618. https://doi.org/10.18653/v1/W17-4770

  35. [43]

    Marc’Aurelio Ranzato, Sumit Chopra, Michael Auli, and Wojciech Zaremba. 2015. Sequence Level Training with Recurrent Neural Networks.CoRRabs/1511.06732 (2015). https://api.semanticscholar.org/CorpusID:7147309

  36. [44]

    Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020. COMET: A Neural Framework for MT Evaluation. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics, Online, 2685–2702. https:/...

  37. [45]

    Stefan Riezler and John T. Maxwell. 2005. On Some Pitfalls in Automatic Evaluation and Significance Testing for MT. InProceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization. Association for Computational Ling...

  38. [46]

    Sánchez-Cartagena, Marta Bañón, Sergio Ortiz-Rojas, and Gema Ramírez-Sánchez

    Víctor M. Sánchez-Cartagena, Marta Bañón, Sergio Ortiz-Rojas, and Gema Ramírez-Sánchez. 2018. Prompsit’s submission to WMT 2018 Parallel Corpus Filtering shared task. InProceedings of the Third Conference on Machine Translation, Volume 2: Shared Task Papers. Association for Co...

  39. [47]

    Barbara Scalvini, Iben Nyholm Debess, Annika Simonsen, and Hafsteinn Einarsson. 2025. Rethinking Low-Resource MT: The Surprising Effectiveness of Fine-Tuned Multilingual Models in the LLM Age. InProceedings of the Joint 25th Nordic Conference on Computational Linguistics and 1...

  40. [48]

    Chufan Shi, Haoran Yang, Deng Cai, Zhisong Zhang, Yifan Wang, Yujiu Yang, and Wai Lam. 2024. A Thorough Examination of Decoding Methods in the Era of LLMs. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal,...

  41. [49]

    Yewei Song, Saad Ezzini, Jacques Klein, Tegawende Bissyande, Clément Lefebvre, and Anne Goujon. 2023. Letz Translate: Low-Resource Machine Translation for Luxembourgish. arXiv:2303.01347 [cs.CL]

  42. [50]

    Smith, and Luke Zettlemoyer

    Gabriel Stanovsky, Noah A. Smith, and Luke Zettlemoyer. 2019. Evaluating Gender Bias in Machine Translation. InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Florence, Italy, 1679–1684. https:...

  43. [51]

    Yixuan Su, Tian Lan, Yan Wang, Dani Yogatama, Lingpeng Kong, and Nigel Collier. 2022. A Contrastive Framework for Neural Text Generation. arXiv:2202.06417 [cs.CL]

  44. [52]

    Xu Tan, Yi Ren, Di He, Tao Qin, and Tie-Yan Liu. 2019. Multilingual Neural Machine Translation with Knowledge Distillation. InSeventh International Conference on Learning Representations. https://openreview.net/forum?id=S1gUsoR9YX

  45. [53]

    Chau Tran, Shruti Bhosale, James Cross, Philipp Koehn, Sergey Edunov, and Angela Fan. 2021. Facebook AI WMT21 News Translation Task Submission. InProc. of the Sixth Conference on Machine Translation (WMT). 205–215

  46. [54]

    Jannis Vamvas and Rico Sennrich. 2021. Contrastive Conditioning for Assessing Disambiguation in MT: A Case Study of Distilled Bias. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Online and P...

  47. [55]

    Jannis Vamvas and Rico Sennrich. 2024. Linear-time Minimum Bayes Risk Decoding with Reference Aggregation. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). ...

  48. [56]

    Gomez, Łukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. InProceedings of the 31st International Conference on Neural Information Processing Systems(Long Beach, California, US...

  49. [57]

    Ashwin Vijayakumar, Michael Cogswell, Ramprasaath Selvaraju, Qing Sun, Stefan Lee, David Crandall, and Dhruv Batra. 2018. Diverse Beam Search for Improved Description of Complex Scenes.Proceedings of the AAAI Conference on Artificial Intelligence32, 1 (Apr. 2018). https://doi....

  50. [58]

    Jiayi Wang, David Ifeoluwa Adelani, Sweta Agrawal, Marek Masiak, Ricardo Rei, Eleftheria Briakou, Marine Carpuat, Xuanli He, Sofia Bourhim, Andiswa Bukula, Muhidin Mohamed, Temitayo Olatoye, Tosin Adewumi, Hamam Mokayed, Christine Mwase, Wangui Kimotho, Foutse Yuehgoh, Anuoluw...

  51. [59]

    Jun Wang, Eleftheria Briakou, Hamid Dadkhahi, Rishabh Agarwal, Colin Cherry, and Trevor Cohn. 2024. Don’t Throw Away Data: Better Sequence Knowledge Distillation. arXiv:2407.10456 [cs.CL] https://arxiv.org/abs/2407.10456

  52. [60]

    Gian Wiher, Clara Meister, and Ryan Cotterell. 2022. On Decoding Strategies for Neural Text Generators. arXiv:2203.15721 [cs.CL]

  53. [61]

    Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mari...

  54. [62]

    Martindale, and Marine Carpuat

    Weijia Xu, Sweta Agrawal, Eleftheria Briakou, Marianna J. Martindale, and Marine Carpuat. 2023. Understanding and Detecting Hallucinations in Neural Machine Translation via Model Introspection.Transactions of the Association for Computational Linguistics11 (2023), 546–564. htt...

  55. [63]

    Zhengzhe Yu, Daimeng Wei, Zongyao Li, Hengchao Shang, Xiaoyu Chen, Zhanglin Wu, Jiaxin Guo, Minghan Wang, Lizhi Lei, Min Zhang, Hao Yang, and Ying Qin. 2021. HW-TSC’s Participation in the WMT 2021 Large-Scale Multilingual Translation Task. InProceedings of the Sixth Conference...

  56. [64]

    Hugh Zhang, Daniel Duckworth, Daphne Ippolito, and Arvind Neelakantan. 2021. Trading Off Diversity and Quality in Natural Language Generation. InProceedings of the Workshop on Human Evaluation of NLP Systems (HumEval). Association for Computational Linguistics, Online, 25–33. ...

  57. [65]

    Songming Zhang, Yunlong Liang, Shuaibo Wang, Yufeng Chen, Wenjuan Han, Jian Liu, and Jinan Xu. 2023. Towards Understanding and Improving Knowledge Distillation for Neural Machine Translation. InProceedings of the 61st Annual Meeting of the Association for Computational Linguis...

  58. [66]

    Yuhao Zhang, Ziyang Wang, Runzhe Cao, Binghao Wei, Weiqiao Shan, Shuhan Zhou, Abudurexiti Reheman, Tao Zhou, Xin Zeng, Laohu Wang, et al. 2020. The niutrans machine translation systems for wmt20. InProceedings of the Fifth Conference on Machine Translation. 338–345

  59. [67]

    Wenhao Zhu, Hongyi Liu, Qingxiu Dong, Jingjing Xu, Shujian Huang, Lingpeng Kong, Jiajun Chen, and Lei Li. 2024. Multilingual Machine Translation with Large Language Models: Empirical Results and Analysis. InFindings of the Association for Computational Linguistics: NAACL 2024,...

  60. [68]

    Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu. 2018. Texygen: A Benchmarking Platform for Text Generation Models. InThe 41st International ACM SIGIR Conference on Research & Development in Information Retrieval(Ann Arbor, MI, USA)(SIGIR ’18)...

  61. [2004]

    https://aclanthology.org/N04-1022

    Association for Computational Linguistics, Boston, Massachusetts, USA, 169–176. https://aclanthology.org/N04-1022

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.