{"paper":{"title":"Missing Information, Unresponsive Authors, Experimental Flaws: The Impossibility of Assessing the Reproducibility of Previous Human Evaluations in NLP","license":"http://creativecommons.org/licenses/by-nc-nd/4.0/","headline":"","cross_cats":[],"primary_cat":"cs.CL","authors_text":"Albert Gatt, Anouck Braggaar, Antonio Toral, Anya Belz, Chris van der Lee, Craig Thomson, Dimitra Gkatzia, Dirk Hovy, Diyi Yang, Ehud Reiter, Elizabeth Clark, Emiel Krahmer, Emiel van Miltenburg, Filip Klubicka, Gavin Abercrombie, Huiyuan Lai, Javier Gonz\\'alez-Corbelle, Jie Ruan, Joel Tetreault, John D. Kelleher, Jose M. Alonso-Moral, Kees van Deemter, Leo Wanner, Lewis Watson, Malvina Nissim, Manuela H\\\"urlimann, Margot Mieskes, Mark Cieliebak, Mingqi Gao, Mohammad Arvan, Natalie Parde, Ond\\v{r}ej Du\\v{s}ek, Ond\\v{r}ej Pl\\'atek, Pablo Mosteiro, Qixiang Fang, Saad Mahamood, Steffen Eger, Takumi Ito, Tanvi Dinkar, Verena Rieser, Xiaojun Wan, Yiru Li","submitted_at":"2023-05-02T17:46:12Z","abstract_excerpt":"We report our efforts in identifying a set of previous human evaluations in NLP that would be suitable for a coordinated study examining what makes human evaluations in NLP more/less reproducible. We present our results and findings, which include that just 13\\% of papers had (i) sufficiently low barriers to reproduction, and (ii) enough obtainable information, to be considered for reproduction, and that all but one of the experiments we selected for reproduction was discovered to have flaws that made the meaningfulness of conducting a reproduction questionable. As a result, we had to change o"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2305.01633","kind":"arxiv","version":2},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2305.01633/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"}