REVIEW 4 major objections 6 minor 79 references
Are We Paying Attention to Her? Investigating Gender Disambiguation and Attention in Machine Translation
T0 review · 4 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read English-to-Italian translation models ignore explicit gender pronouns in most sentence pairs, defaulting to stereotypical gender instead.
desk verdict A clean new minimal-pair metric for measuring gender-cue reliance, with a robust male-default asymmetry result; the attention analysis is exploratory and the alignment pipeline needs validation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Minimal Pair Accuracy (MPA) is the central artifact: two English sentences identical except for the pronoun (he vs. she) referring to the same profession noun are translated, and a pair counts as correct only if the model produces the matching grammatical gender in the Italian target noun in both directions. The second piece is an attention-weight analysis on the encoder: for the correctly gendered pairs, the authors extract the average self-attention weight from the profession noun to the gender cue across layers and heads, treating weights above the uniform baseline as evidence of cue integration. These two instruments together let the paper separate 'happens to produce the right gender' from 'consistently uses the cue to decide the gender.'
What would settle it
Run the MPA calculation on a human-verified subset of the same sentences, manually checking alignment and target gender, and see whether the 6.12%, 30.24%, and 38.45% numbers move; if they change materially, the automatic pipeline is responsible for the reported cue-ignoring behavior. Alternatively, mask or remove the pronoun from the source and measure whether MPA drops; if it does not, the cue is not doing the work attributed to it.
Extended reading notes
Core claim
The paper's central claim is that gender disambiguation in neural machine translation is driven more by learned statistical associations than by the explicit gendered pronoun in the source. Concretely, MPA—the percentage of minimal pairs in which a model correctly adapts the grammatical gender of the profession noun in both the pro-stereotypical and anti-stereotypical versions of a sentence—is 6.12% for OPUS-MT, 30.24% for NLLB-200, and 38.45% for mBART on the English–Italian challenge set. Among the pairs that are correctly disambiguated, the majority involve professions stereotypically associated with women receiving a masculine cue: 82.29%, 69.10%, and 61.90% respectively, while the reverse—feminine cues applied to male-stereotyped professions—ranges from 17.71% to 38.10%. The paper further claims that encoder self-attention between the profession noun and the pronoun shows gender-specific patterns: feminine pronouns produce concentrated, specialized attention, masculine pronouns produce weaker, more diffuse attention, and the models with more distributed attention (NLLB-200, mBART) are also the ones with higher MPA.
Load-bearing premise
The results assume that the automatic word alignment and morphological analysis used to pinpoint the profession noun and its grammatical gender in each translated sentence are accurate on this dataset, and that encoder attention weights are a meaningful proxy for how much the model actually uses the gender cue.
Editorial extensions
If this is right
- A model can show respectable gender accuracy while almost never using pronouns as cues; MPA exposes that separation.
- Masculine defaults are asymmetric: overriding a female-stereotyped profession with a male pronoun is common, while overriding a male-stereotyped profession with a female pronoun is rare.
- Attention patterns differ by cue gender: feminine cues are encoded by specialized, localized heads, while masculine cues are diffuse; more diffuse, multi-layer encoding is associated with higher MPA.
- Because the framework needs only a gendered target language, MPA can be applied to other language pairs and to any encoder-decoder or decoder-only model.
- The observed attention results are correlational; the paper itself cautions that they do not prove a causal link, which motivates intervention-based follow-ups.
Reading between the lines
- If MPA becomes a standard companion to accuracy metrics, model rankings could shift: a system optimized for surface gender correctness would no longer be credited with context sensitivity, and smaller models that genuinely track cues might be recognized.
- The masculine-default asymmetry may not be specific to Italian morphology; the same MPA design could be applied to languages with different agreement systems or to gender-neutral cue forms to test whether the default-to-masculine pattern is grammatical or social in origin.
- The attention result suggests a testable mechanism: if distributed encoding of a cue is what enables disambiguation, then fine-tuning or constraining attention in early layers of a single-head model like OPUS-MT should improve MPA; that hypothesis goes beyond what the paper proves.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Minimal Pair Accuracy (MPA), a metric for measuring whether NMT models use gendered pronouns in English as contextual cues for disambiguating the grammatical gender of profession nouns in Italian translations. Applying MPA to the WinoMT challenge set for English-to-Italian translation, the authors report values of 6.12% for OPUS-MT, 30.24% for NLLB-200, and 38.45% for mBART (Table 2), and a breakdown showing that correctly disambiguated minimal pairs are predominantly associated with female-stereotyped professions (Table 3). They also analyze encoder self-attention weights between the gender cue and the profession noun on accurately gendered minimal pairs, concluding that masculine cues elicit more diffuse attention while feminine cues elicit more concentrated attention, with differences across models. The paper includes standard WinoMT accuracy results, an exploratory cross-attention analysis, a limitations section, and publicly released code.
Significance. If the MPA results are reliable, they provide a useful evaluation dimension beyond surface-level gender accuracy, directly targeting whether models adapt their translations to contextual gender cues rather than defaulting to stereotypes. The Pro-F/Pro-M asymmetry is an interesting empirical finding consistent with the male-as-norm bias and is worth reporting. The attention analysis is exploratory and the authors are appropriately cautious about causal claims; the release of the evaluation code is a concrete strength. However, the central quantitative claims depend on an unvalidated alignment and morphological-analysis pipeline for English-Italian, and the attention conclusions rest on a vaguely defined relevance threshold and visual inspection of averaged heatmaps. Both need to be addressed before the results can be fully credited.
major comments (4)
- [§5.1–5.2, Tables 2–3] The MPA results are computed using WinoMT's automatic pipeline to locate the profession noun in the Italian translation and extract its grammatical gender, together with fast_align to map source and target token indices. No validation or error analysis is reported for this English–Italian subset. Italian realizes gender on articles and adjectives as well as noun endings, and there are numerous epicene nouns (e.g., 'cantante', 'pianista') whose form alone is ambiguous; if the alignment or morphological tagging is wrong for even a modest fraction of sentences, the MPA percentages and the Pro-F/Pro-M split are systematically biased. Please report alignment quality and morphological-analyzer accuracy on a manually inspected sample, or otherwise demonstrate that the pipeline is accurate for this language pair.
- [§5.2, Tables 2–3] The headline differences and asymmetries are reported as point estimates without confidence intervals or significance tests. The differences between models (6.12% vs. 30.24% vs. 38.45%) and the Pro-F vs. Pro-M asymmetry (82.29% vs. 17.71% for OPUS-MT) should be accompanied by a bootstrap or McNemar test to establish that they are not due to sampling variability, especially because the number of correctly disambiguated minimal pairs is much smaller than the total number of pairs.
- [§6.2, Figures 4–5] The attention analysis identifies 'relevant' heads by comparing average attention weights to a uniform baseline of approximately 1/13, but the 'notable margin' above this baseline is never defined. The conclusions that masculine cues elicit 'weaker, more dispersed' attention and feminine cues elicit 'more localized, concentrated' attention are based on visual inspection of averaged heatmaps. Please define a quantitative concentration measure (e.g., entropy, variance, or maximum minus baseline), specify the threshold in advance, and report per-condition statistics rather than only aggregate heatmaps.
- [Abstract and §5.2] The abstract claims that the models 'ignore available gender cues in most cases in favour of (statistical) stereotypical gender interpretation,' but this is not directly supported by MPA as defined. MPA is the proportion of minimal pairs in which both the pro-stereotypical and anti-stereotypical sentences receive the grammatically correct target gender; a low MPA could result from inconsistent cue use, morphological errors, or systematic defaulting. Without a baseline (e.g., a no-cue model or a chance-level expectation), the phrase 'ignore available gender cues' overinterprets the metric. Please either qualify the claim or provide a comparison baseline that would support the causal reading.
minor comments (6)
- [Abstract] The sentence 'We evaluate a number of NMT models using this metric, we show that they ignore available gender cues in most cases in favour of (statistical) stereotypical gender interpretation' contains a comma splice and should be rewritten for grammatical completeness.
- [Figure 3] In the mBART ANTI-S example, 'Il analista' is ungrammatical Italian; the correct form is 'L'analista'. If this is verbatim model output, please indicate so; otherwise correct the translation in the figure.
- [Table 3] The caption defines Pro-F and Pro-M as percentages of correctly disambiguated minimal pairs where the profession is stereotypically associated with women or men, respectively, but since each minimal pair contains both a pro- and an anti-stereotypical sentence, the relation to the pronoun should be clarified to avoid confusion.
- [§6.1] The description of extracting attention weights should specify the direction of attention (from the profession noun to the pronoun, or vice versa) and whether averaging over subword tokens is performed over query or key positions; this matters for interpreting the heatmaps.
- [References] The reference 'Savoldi et al. (2024) A decade of gender bias in machine translation' is marked '[under review]'; references should be complete or moved to footnote, and the entry for 'Bentivogli Luisa et al. (2020)' reverses given and family names.
- [Figures 4–7] The heatmaps use a 'standardized colormap' but do not include a colorbar, making it impossible for readers to map colors to numerical attention values; please add a color scale.
Circularity Check
No significant circularity: MPA is a direct external measurement against WinoMT, and the attention analysis is explicitly exploratory with caveats rather than a prediction derived from its own inputs.
full rationale
The paper's central quantitative claims rest on Minimal Pair Accuracy (MPA), computed from the external WinoMT challenge set (Stanovsky et al., 2019). MPA is defined as the proportion of pro/anti minimal pairs, which differ only in the gendered pronoun, where the model produces the correct target gender in both sentences. This is a direct measurement against gold labels: no parameter is fitted, no quantity is predicted from the metric itself, and the metric is not defined in terms of the conclusion it supports. The low MPA values and the Pro-F/Pro-M asymmetry in Tables 2 and 3 are re-aggregations of the same external per-sentence correctness labels, not self-referential derivations. The attention analysis in Section 6 selects accurately gendered minimal pairs and averages encoder self-attention weights between the profession noun and the pronoun. This is an observational interpretation of model internals, not a fitted parameter renamed as a prediction. The authors explicitly disclaim causality ('they should not be taken as definitive explanations of model decision-making as no causal relationship between gender cue integration and translation outputs is established') and acknowledge the gender-composition imbalance and its confound with pro/anti-stereotypical contexts in Sections 6.2 and 7.3. Those are validity caveats, not circular reductions. Citations to the authors' own prior work (e.g., Vanmassenhove et al., 2018) appear only as background context and are not load-bearing for the MPA or attention findings. No uniqueness theorem, no ansatz smuggled via self-citation, and no renaming of a known result as a new derivation are present. The derivation chain is therefore self-contained against an external benchmark, and the appropriate circularity score is 0.
Assumptions & free parameters
free parameters (1)
- Attention relevance threshold =
~0.08 (1/13)
assumptions (3)
- domain assumption The WinoMT evaluation pipeline (automatic word alignment plus morphological analysis) correctly identifies the grammatical gender of the primary entity in Italian translations.
- domain assumption The profession nouns in WinoMT and their Italian translations are lexically gender-ambiguous in a way that requires contextual cues for correct disambiguation.
- domain assumption Encoder self-attention weights between pronoun and profession noun are a meaningful indicator of gender cue integration.
Cite this review
Pith. "Pith review of Are We Paying Attention to Her? Investigating Gender Disambiguation and Attention in Machine Translation." pith.science (2026). https://pith.science/paper/3YBJMF7L
@misc{pith2026250508546,
author = {Pith},
title = {Pith review of: Are We Paying Attention to Her? Investigating Gender Disambiguation and Attention in Machine Translation},
year = {2026},
howpublished = {\url{https://pith.science/paper/3YBJMF7L}},
note = {Machine review of arXiv:2505.08546}
}
read the original abstract
While gender bias in modern Neural Machine Translation (NMT) systems has received much attention, traditional evaluation metrics do not to fully capture the extent to which these systems integrate contextual gender cues. We propose a novel evaluation metric called Minimal Pair Accuracy (MPA), which measures the reliance of models on gender cues for gender disambiguation. MPA is designed to go beyond surface-level gender accuracy metrics by focusing on whether models adapt to gender cues in minimal pairs -- sentence pairs that differ solely in the gendered pronoun, namely the explicit indicator of the target's entity gender in the source language (EN). We evaluate a number of NMT models on the English-Italian (EN--IT) language pair using this metric, we show that they ignore available gender cues in most cases in favor of (statistical) stereotypical gender interpretation. We further show that in anti-stereotypical cases, these models tend to more consistently take masculine gender cues into account while ignoring the feminine cues. Furthermore, we analyze the attention head weights in the encoder component and show that while all models encode gender information to some extent, masculine cues elicit a more diffused response compared to the more concentrated and specialized responses to feminine gender cues.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Lauren Ackerman. 2019. https://doi.org/10.5334/gjgl.721 Syntactic and cognitive issues in investigating gendered coreference . Glossa: a journal of general linguistics, 4(1)
-
[2]
Rajas Bansal. 2022. A survey on bias and fairness in natural language processing. arXiv preprint arXiv:2204.09591
arXiv 2022
-
[3]
Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2018. https://arxiv.org/abs/1811.01157 Identifying and controlling important neurons in neural machine translation . In Proceedings of the Seventh International Conference on Learning Representations (ICLR)
arXiv 2018
-
[4]
Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, and James Glass. 2017 a . https://doi.org/10.18653/v1/P17-1080 What do neural machine translation models learn about morphology? In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 861--872, Vancouver, Canada. Association for ...
-
[5]
Yonatan Belinkov, Llu \'i s M \`a rquez, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2017 b . https://aclanthology.org/I17-1001/ Evaluating layers of representation in neural machine translation on part-of-speech and semantic tagging tasks . In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume ...
work page 2017
-
[6]
Adrien Bibal, R \'e mi Cardon, David Alfter, Rodrigo Wilkens, Xiaoou Wang, Thomas Fran c ois, and Patrick Watrin. 2022. https://doi.org/10.18653/v1/2022.acl-long.269 Is attention explanation? an introduction to the debate . In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3889--3900,...
-
[7]
Su Lin Blodgett, Solon Barocas, Hal Daum \'e III, and Hanna Wallach. 2020. https://doi.org/10.18653/v1/2020.acl-main.485 Language (technology) is power: A critical survey of bias in NLP . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5454--5476, Online. Association for Computational Linguistics
-
[8]
Franck Burlot and Fran c ois Yvon. 2017. https://doi.org/10.18653/v1/W17-4705 Evaluating the morphological competence of machine translation systems . In Proceedings of the Second Conference on Machine Translation, pages 43--55, Copenhagen, Denmark. Association for Computational Linguistics
Show all 79 references
-
[9]
Yang Trista Cao and Hal Daum \'e III. 2020. https://doi.org/10.18653/v1/2020.acl-main.418 Toward gender-inclusive coreference resolution . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4568--4595. Association for Computationa...
2020 doi
-
[10]
Prafulla Kumar Choubey, Anna Currey, Prashant Mathur, and Georgiana Dinu. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.123 GFST : G ender-filtered self-training for more accurate gender in translation . In Proceedings of the 2021 Conference on Empirical Methods in Natural...
2021 doi
-
[11]
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019. https://doi.org/10.18653/v1/W19-4828 What does BERT look at? an analysis of BERT `s attention . In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP...
2019 doi
-
[12]
Alexis Conneau, German Kruszewski, Guillaume Lample, Lo \"i c Barrault, and Marco Baroni. 2018. https://doi.org/10.18653/v1/P18-1198 What you can cram into a single \ & ! \# * vector: Probing sentence embeddings for linguistic properties . In Proceedings of the 56th Annual Mee...
2018 doi
-
[13]
Marta R Costa-juss \`a . 2019. An analysis of Gender Bias studies in Natural Language Processing . Nature Machine Intelligence, 1(11):495--496
2019
-
[14]
Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansan...
2022 arXiv
-
[15]
Costa-jussà, Carlos Escolano, Christine Basta, Javier Ferrando, Roser Batlle, and Ksenia Kharitonova
Marta R. Costa-jussà, Carlos Escolano, Christine Basta, Javier Ferrando, Roser Batlle, and Ksenia Kharitonova. 2020. https://arxiv.org/abs/2012.13176 Gender bias in multilingual neural machine translation: The architecture matters . Preprint, arXiv:2012.13176
2020 arXiv
-
[16]
Marcel Danesi. 2014. https://doi.org/10.4324/9781315705194 Dictionary of Media and Communications . Routledge
2014 doi
-
[17]
G \'a llego, and Marta R
Javier Ferrando, Gerard I. G \'a llego, and Marta R. Costa-juss \`a . 2022. https://doi.org/10.18653/v1/2022.emnlp-main.595 Measuring the mixing of contextual information in the transformer . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Proces...
2022 doi
-
[18]
Joel Escud \'e Font and Marta R Costa-juss \`a . 2019. Equalizing gender bias in neural machine translation with word embeddings techniques. In Proceedings of the First Workshop on Gender Bias in Natural Language Processing, pages 147--154
2019
-
[19]
Akshay Goindani and Manish Shrivastava. 2021. https://aclanthology.org/2021.ranlp-1.52/ A dynamic head importance computation mechanism for neural machine translation . In Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP 2021...
2021
-
[20]
Nizar Habash, Houda Bouamor, and Christine Chung. 2019. Automatic gender identification and reinflection in arabic. In Proceedings of the First Workshop on Gender Bias in Natural Language Processing, pages 155--165
2019
-
[21]
Stefan Heimersheim and Neel Nanda. 2024. https://arxiv.org/abs/2404.15255 How to use and interpret activation patching . Preprint, arXiv:2404.15255
2024 arXiv
-
[22]
Tosho Hirasawa and Mamoru Komachi. 2019. Debiasing word embeddings improves multimodal machine translation. In Proceedings of Machine Translation Summit XVII: Research Track, pages 32--42
2019
-
[23]
Sarthak Jain and Byron C. Wallace. 2019. https://doi.org/10.18653/v1/N19-1357 A ttention is not E xplanation . In Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and...
2019 doi
-
[24]
Jae-young Jo and Sung-Hyon Myaeng. 2020. https://doi.org/10.18653/v1/2020.acl-main.311 Roles and utilization of attention heads in transformer-based neural language models . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3404-...
2020 doi
-
[25]
Jaap Jumelet, Willem Zuidema, and Dieuwke Hupkes. 2019. https://doi.org/10.18653/v1/K19-1001 Analysing neural language models: Contextual decomposition reveals default reasoning in number and gender assignment . In Proceedings of the 23rd Conference on Computational Natural La...
2019 doi
-
[26]
Yunsu Kim, Duc Thanh Tran, and Hermann Ney. 2019. https://doi.org/10.18653/v1/D19-6503 When and why is document-level context useful in neural machine translation? In Proceedings of the Fourth Workshop on Discourse in Machine Translation (DiscoMT 2019), pages 24--34, Hong Kong...
2019 doi
-
[27]
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.574 Attention is not only a weight: Analyzing transformers with vector norms . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Pro...
2020 doi
-
[28]
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.373 I ncorporating R esidual and N ormalization L ayers into A nalysis of M asked L anguage M odels . In Proceedings of the 2021 Conference on Empirical Methods ...
2021 doi
-
[29]
Tom Kocmi, Tomasz Limisiewicz, and Gabriel Stanovsky. 2020. https://aclanthology.org/2020.wmt-1.39/ Gender coreference and bias evaluation at WMT 2020 . In Proceedings of the Fifth Conference on Machine Translation, pages 357--364, Online. Association for Computational Linguistics
2020
-
[30]
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019. https://doi.org/10.18653/v1/D19-1445 Revealing the dark secrets of BERT . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference ...
2019 doi
-
[31]
Jaesong Lee, Joong-Hwi Shin, and Jun-Seok Kim. 2017. https://doi.org/10.18653/v1/D17-2021 Interactive visualization and manipulation of attention-based neural machine translation . In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing: Syste...
2017 doi
-
[32]
Yongjie Lin, Yi Chern Tan, and Robert Frank. 2019. https://doi.org/10.18653/v1/W19-4825 Open sesame: Getting inside BERT `s linguistic knowledge . In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 241--253, Florence,...
2019 doi
-
[33]
Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. https://doi.org/10.1162/tacl_a_00343 Multilingual denoising pre-training for neural machine translation . Transactions of the Association for Computational...
2020 doi
-
[34]
Bentivogli Luisa, Beatrice Savoldi, Negri Matteo, A Di Gangi Mattia, Cattoni Roldano, Turchi Marco, et al. 2020. Gender in danger? evaluating speech translation technology on the must-she corpus. In Proceedings of the 58th Annual Meeting of the Association for Computational Li...
2020
-
[35]
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. https://proceedings.neurips.cc/paper_files/paper/2022/file/6f1d43d5a82a37e89b0665b33bf3a182-Paper-Conference.pdf Locating and editing factual associations in gpt . In Advances in Neural Information Processing Sy...
2022
-
[36]
Ali Modarressi, Mohsen Fayyaz, Yadollah Yaghoobzadeh, and Mohammad Taher Pilehvar. 2022. https://doi.org/10.18653/v1/2022.naacl-main.19 G lob E nc: Quantifying global token attribution by incorporating the whole encoder layer in transformers . In Proceedings of the 2022 Confer...
2022 doi
-
[37]
Wafaa Mohammed and Vlad Niculae. 2024. https://aclanthology.org/2024.findings-eacl.113/ On measuring context utilization in document-level MT systems . In Findings of the Association for Computational Linguistics: EACL 2024, pages 1633--1643, St. Julian ' s, Malta. Association...
2024
-
[38]
Hosein Mohebbi, Grzegorz Chrupa a, Willem Zuidema, and Afra Alishahi. 2023 a . https://doi.org/10.18653/v1/2023.emnlp-main.513 Homophone disambiguation reveals patterns of context mixing in speech transformers . In Proceedings of the 2023 Conference on Empirical Methods in Nat...
2023 doi
-
[39]
Hosein Mohebbi, Willem Zuidema, Grzegorz Chrupa a, and Afra Alishahi. 2023 b . https://doi.org/10.18653/v1/2023.eacl-main.245 Quantifying context mixing in transformers . In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguis...
2023 doi
-
[40]
Amit Moryossef, Roee Aharoni, and Yoav Goldberg. 2019. Filling gender & number gaps in neural machine translation with black-box context injection. In Proceedings of the First Workshop on Gender Bias in Natural Language Processing, pages 49--54
2019
-
[41]
Krithika Ramesh, Gauri Gupta, and Sanjay Singh. 2021. Evaluating gender bias in hindi-english machine translation. In Proceedings of the 3rd Workshop on Gender Bias in Natural Language Processing, pages 16--23
2021
-
[42]
Viegas, Andy Coenen, Adam Pearce, and Been Kim
Emily Reif, Ann Yuan, Martin Wattenberg, Fernanda B. Viegas, Andy Coenen, Adam Pearce, and Been Kim. 2019. Visualizing and measuring the geometry of bert. In Advances in Neural Information Processing Systems (NeurIPS), pages 8594--8603
2019
-
[43]
Argentina Anna Rescigno, Eva Vanmassenhove, Johanna Monti, and Andy Way. 2020. A case study of natural gender phenomena in translation a comparison of google translate, bing microsoft translator and deepl for english to italian, french and spanish. Computational Linguistics CL...
2020
-
[44]
Annette Rios Gonzales, Laura Mascarell, and Rico Sennrich. 2017. https://doi.org/10.18653/v1/W17-4702 Improving word sense disambiguation in neural machine translation with sense embeddings . In Proceedings of the Second Conference on Machine Translation, pages 11--19, Copenha...
2017 doi
-
[45]
Tim Rockt \"a schel, Edward Grefenstette, Karl Moritz Hermann, Tom \'a s Ko c isk \'y , and Phil Blunsom. 2016. Reasoning about entailment with neural attention. In International Conference on Learning Representations (ICLR)
2016
-
[46]
Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. 2018. https://doi.org/10.18653/v1/N18-2002 Gender bias in coreference resolution . In Proceedings of the 2018 Conference of the North A merican Chapter of the Association for Computational Linguistics: H...
2018 doi
-
[47]
Gabriele Sarti, Grzegorz Chrupała, Malvina Nissim, and Arianna Bisazza. 2024. https://arxiv.org/abs/2310.01188 Quantifying the plausibility of context reliance in neural machine translation . In Proceedings of the Twelfth International Conference on Learning Representations (ICLR)
2024 arXiv
-
[48]
Danielle Saunders and Bill Byrne. 2020. Reducing gender bias in neural machine translation as a domain adaptation problem. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7724--7736
2020
-
[49]
Beatrice Savoldi, Jasmijn Bastings, Lucia Bentivogli, and Eva Vanmassenhove. 2024. A decade of gender bias in machine translation. [under review]
2024
-
[50]
Beatrice Savoldi, Marco Gaido, Luisa Bentivogli, Matteo Negri, and Marco Turchi. 2021. Gender bias in machine translation. Transactions of the Association for Computational Linguistics, 9:845--874
2021
-
[51]
Rico Sennrich. 2017. https://aclanthology.org/E17-2060/ How grammatical is character-level neural machine translation? assessing MT quality with contrastive translation pairs . In Proceedings of the 15th Conference of the E uropean Chapter of the Association for Computational ...
2017
-
[52]
Dagmar Stahlberg, Friederike Braun, Lisa Irmen, and Sabine Sczesny. 2007. http://psycnet.apa.org/psycinfo/2007-01308-006 Representation of the sexes in language , page 163–187
2007
-
[53]
Karolina Stanczak and Isabelle Augenstein. 2021. A survey on gender bias in natural language processing . arXiv preprint arXiv:2112.14168
2021 arXiv
-
[54]
Smith, and Luke Zettlemoyer
Gabriel Stanovsky, Noah A. Smith, and Luke Zettlemoyer. 2019. https://doi.org/10.18653/v1/p19-1164 Evaluating gender bias in machine translation . Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics
2019 doi
-
[55]
Tony Sun, Andrew Gaut, Shirlyn Tang, Yuxin Huang, Mai ElSherief, Jieyu Zhao, Diba Mirza, Elizabeth Belding, Kai-Wei Chang, and William Yang Wang. 2019. https://doi.org/10.18653/v1/P19-1159 Mitigating gender bias in natural language processing: Literature review . In Proceeding...
2019 doi
-
[56]
Tony Sun, Kellie Webster, Apu Shah, William Yang Wang, and Melvin Johnson. 2021. They, them, theirs: Rewriting with gender-neutral english. arXiv preprint arXiv:2102.06788
2021 arXiv
-
[57]
o rg Tiedemann, Mikko Aulamo, Daria Bakshandaeva, Michele Boggia, Stig Arne Gr \
J \"o rg Tiedemann, Mikko Aulamo, Daria Bakshandaeva, Michele Boggia, Stig Arne Gr \"o nroos, Tommi Nieminen, Alessandro Raganato, Yves Scherrer, Ra \'u l V \'a zquez, and Sami Virpioja. 2023. https://doi.org/10.1007/s10579-023-09704-w Democratizing neural machine translation ...
2023 doi
-
[58]
Jörg Tiedemann. 2012. Parallel data, tools and interfaces in OPUS . In Proceedings of the Eight International Conference on Language Resources and Evaluation (LREC'12), Istanbul, Turkey. European Language Resources Association (ELRA)
2012
-
[59]
European Union. 2024. https://eur-lex.europa.eu/eli/reg/2024/1689/oj Regulation (eu) 2024/1689 of the european parliament and of the council of 13 june 2024 laying down harmonised rules on artificial intelligence
2024
-
[60]
Jannis Vamvas and Rico Sennrich. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.803 Contrastive conditioning for assessing disambiguation in MT : A case study of distilled bias . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, page...
2021 doi
-
[61]
Jannis Vamvas and Rico Sennrich. 2022. https://doi.org/10.18653/v1/2022.acl-short.53 As little as possible, as much as necessary: Detecting over- and undertranslations with contrastive conditioning . In Proceedings of the 60th Annual Meeting of the Association for Computationa...
2022 doi
-
[62]
Oskar Van Der Wal, Jaap Jumelet, Katrin Schulz, and Willem Zuidema. 2022. https://aclanthology.org/2022.gebnlp-1.8 The birth of bias: A case study on the evolution of gender bias in an E nglish language model . In Proceedings of the 4th Workshop on Gender Bias in Natural Langu...
2022
-
[63]
Eva Vanmassenhove. 2024. Gender bias in machine translation and the era of large language models. Gendered Technology in Translation and Interpreting: Centering Rights in the Development of Language Technology, page 225
2024
-
[64]
Eva Vanmassenhove, Chris Emmery, and Dimitar Shterionov. 2021 a . Neutral rewriter: A rule-based and neural approach to automatic rewriting into gender-neutral alternatives. arXiv preprint arXiv:2109.06105
2021 arXiv
-
[65]
Eva Vanmassenhove, Christian Hardmeier, and Andy Way. 2018. https://doi.org/10.18653/v1/D18-1334 Getting gender right in neural machine translation . In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3003--3008, Brussels, Belgium....
2018 doi
-
[66]
Eva Vanmassenhove, Dimitar Shterionov, and Matthew Gwilliam. 2021 b . Machine translationese: Effects of algorithmic bias on linguistic complexity in machine translation. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguis...
2021
-
[67]
Eva Vanmassenhove, Dimitar Shterionov, and Andy Way. 2019. Lost in translation: Loss and decay of linguistic richness in machine translation. In Proceedings of Machine Translation Summit XVII: Research Track, pages 222--232
2019
-
[68]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30. Curran Associates, Inc
2017
-
[69]
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Simas Sakenis, Jason Huang, Yaron Singer, and Stuart Shieber. 2020. https://arxiv.org/abs/2004.12265 Causal mediation analysis for interpreting neural nlp: The case of gender bias . Preprint, arXiv:2004.12265
2020 arXiv
-
[70]
Elena Voita, Rico Sennrich, and Ivan Titov. 2021. https://doi.org/10.18653/v1/2021.acl-long.91 Analyzing the source and target contributions to predictions in neural machine translation . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistic...
2021 doi
-
[71]
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019. https://doi.org/10.18653/v1/P19-1580 Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned . In Proceedings of the 57th Annual Meeting of the Associatio...
2019 doi
-
[72]
Yequan Wang, Minlie Huang, Li Zhao, et al. 2016. Attention-based lstm for aspect-level sentiment classification. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 606--615
2016
-
[73]
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015. Show, attend and tell: Neural image caption generation with visual attention. In Proceedings of the International Conference on Machine Learning (ICML), p...
2015
-
[74]
Kayo Yin, Patrick Fernandes, Danish Pruthi, Aditi Chaudhary, Andr \'e F. T. Martins, and Graham Neubig. 2021. https://doi.org/10.18653/v1/2021.acl-long.65 Do context-aware translation models pay the right attention? In Proceedings of the 59th Annual Meeting of the Association ...
2021 doi
-
[75]
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018. https://doi.org/10.18653/v1/N18-2003 Gender bias in coreference resolution: Evaluation and debiasing methods . In Proceedings of the 2018 Conference of the North A merican Chapter of the Associati...
2018 doi
-
[76]
Jinman Zhao, Yitian Ding, Chen Jia, Yining Wang, and Zifan Qian. 2024. https://arxiv.org/abs/2403.00277 Gender bias in large language models across multiple languages . arXiv
2024 arXiv
-
[77]
Ran Zmigrod, Sabrina J Mielke, Hanna Wallach, and Ryan Cotterell. 2019. Counterfactual data augmentation for mitigating gender stereotypes in languages with rich morphology. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 1651--1661
2019
-
[78]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[79]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.