REVIEW 4 major objections 4 minor 45 references
BESSTIE: A Benchmark for Sentiment and Sarcasm Classification for Varieties of English
T0 review · 4 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read BESSTIE: a benchmark for sentiment and sarcasm in three varieties of English, showing models lag on Indian English.
desk verdict BESSTIE fills a real gap as the first labeled sentiment/sarcasm benchmark for national English varieties, but its headline inner-circle advantage may be an artifact of label imbalance, and the variety validation is too thin to bear the paper's claims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the BESSTIE dataset itself: roughly 15,000 annotated texts assembled by two complementary collection methods (location-based Google Places reviews and topic-based Reddit subreddits selected by native speakers), then filtered to ambiguous star ratings (2 or 4 stars) to avoid trivial polarity. Variety representation is checked by a two-step validation: manual annotation by en-IN and en-UK speakers (inter-annotator agreement κ=0.26) and an automatic variety predictor fine-tuned on ICE-Australia and ICE-India. Sentiment and sarcasm labels come from single native-speaker annotators per variety, with a secondary annotation pass reporting κ between 0.47 and 0.79 for sentiment and 0.47 to 0.63 for sarcasm. The evaluation protocol fine-tunes nine models on the benchmark and compares in-variety and cross-variety performance, which is the mechanism that produces the inner-circle vs. outer-circle gap.
What would settle it
Re-annotate a fresh random sample of the same sources with a panel of multiple native speakers per variety; if those annotations show the en-IN texts are not reliably distinguishable from en-AU or en-UK, or if a model trained on the re-annotated labels no longer shows a performance gap for en-IN, the paper's central bias claim would fail. Alternatively, a human evaluation showing BESSTIE's variety labels are no more accurate than chance would invalidate the benchmark.
Extended reading notes
Core claim
The paper claims BESSTIE is the first benchmark to provide sentiment and sarcasm labels for natural, user-generated text across multiple varieties of English. Across nine fine-tuned LLMs (six encoders and three decoders), models average F-scores of 0.78 for en-AU, 0.74 for en-UK, and 0.63 for en-IN across domain–task pairs, with sarcasm classification on Reddit comments particularly low (overall 0.59). The authors interpret this as evidence that current LLMs are biased against outer-circle English varieties, and that sarcasm specifically resists cross-variety generalisation because it depends on local cultural and contextual knowledge. They also report that monolingual models marginally outperform multilingual ones, suggesting that broad pre-training language coverage does not automatically extend to varieties of a single language.
Load-bearing premise
The entire comparison rests on the assumption that the collected texts genuinely represent the three national varieties, which depends on location- and topic-based filtering supported by a variety-validation step with low annotator agreement (κ=0.26) and a predictor trained only on ICE-Australia and ICE-India yet reporting en-UK scores.
Editorial extensions
If this is right
- BESSTIE gives future work a fixed, public testbed for measuring how well models handle Australian, Indian, and British English in sentiment and sarcasm tasks.
- Sarcasm classification across varieties is far from solved; the best average F-score is 0.68 and cross-variety fine-tuning hurts out-of-variety sarcasm performance.
- The consistent en-IN deficit suggests that models need variety-specific training data or debiasing methods tuned to outer-circle English, not just more multilingual pre-training.
- Monolingual encoders outperform multilingual decoders on these tasks on average, so architectural choices matter when evaluating variety robustness.
Reading between the lines
- If the variety labels hold up under wider scrutiny, BESSTIE could be extended to other outer-circle varieties (e.g., Nigerian or Singaporean English) to test whether the en-IN gap generalises or is specific to that variety.
- The low annotator agreement on variety identification (κ=0.26) suggests that national variety is often ambiguous in short web text; future benchmarks might combine location signals with author self-identification or expert adjudication to strengthen labels.
- A natural testable extension is to evaluate instruction-tuned LLMs zero-shot on BESSTIE without fine-tuning, which would isolate pre-training bias from adaptation effects.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces BESSTIE, a labelled benchmark for sentiment and sarcasm classification in three national varieties of English: Australian (en-AU), Indian (en-IN), and British (en-UK). Data are collected from Google Places reviews (location-based filtering) and Reddit comments (topic-based filtering), then annotated by native speakers. The authors validate variety representation through manual annotation and an automatic variety predictor, fine-tune nine language models on the benchmark, and report that models consistently perform better on the inner-circle varieties (en-AU, en-UK) than on en-IN, especially for sarcasm classification. The paper also reports cross-variety and cross-domain experiments, an error analysis of misclassified examples, and releases the dataset publicly.
Significance. If the dataset and the empirical claims hold, BESSTIE would be a useful resource for evaluating LLM bias across national varieties of English, an area where labelled benchmarks are scarce. The paper's strengths include the public release of the dataset, the inclusion of two tasks (sentiment and sarcasm) and two domains, the evaluation of nine diverse models, and a manual error analysis that identifies concrete dialectal and colloquial features. The dataset is likely to be reused by the community for bias measurement and cross-variety generalization studies, provided the validity of the variety labels and the robustness of the headline performance comparison are established.
major comments (4)
- [Section 4.1, Table 5] The sentence 'due to the absence of sarcasm labels in GOOGLE subset (as shown in Table 5)' conflicts with Table 5, which reports % Pos. Sarc. values of 7, 1, and 0 for the GOOGLE subsets of en-AU, en-IN, and en-UK. If sarcasm labels were collected but not used, the paper should explain why they are excluded from evaluation; if they were not collected, Table 5 should not report them. As written, the scope of the benchmark's sarcasm component is unclear and the reader cannot determine whether the GOOGLE subset does or does not contain sarcasm labels.
- [Section 4.1, Table 5] The headline claim that models perform worse on en-IN 'particularly for sarcasm classification' is confounded by label prevalence. Table 5 shows REDDIT sarcasm positive rates of 42% for en-AU, 13% for en-IN, and 22% for en-UK, and REDDIT sentiment positive rates of 32%, 25%, and 12% for en-AU, en-IN, and en-UK. Because all reported metrics are macro-averaged F1, the score for a rare positive class is strongly affected by the class prior: a model with comparable per-class behavior will obtain a lower macro-F1 when the positive class is 13% than when it is 42%. The observed en-IN deficit, and especially the 'particularly for sarcasm' part, may therefore reflect prior differences rather than variety-specific model difficulty. Please report per-class F1, balanced accuracy, or re-evaluate on prevalence-matched test sets before drawing the central conclusion.
- [Section 2.2, Table 3] The automated variety validation does not validate en-UK. The DISTIL predictor is fine-tuned only on ICE-Australia and ICE-India for a binary inner-circle versus outer-circle classification task, yet Table 3 reports en-UK P(v) and F-scores. Since the predictor has never seen en-UK training data, high probabilities for en-UK texts likely reflect the binary decision boundary rather than evidence that the texts represent British English. In addition, the manual annotation agreement is low (inter-annotator kappa = 0.26; agreement with the true label is 0.41 and 0.34 for the en-IN and en-UK annotators), which weakens the claim that the subsets 'collectively form a good representative sample.' The validation section should either be redone with an en-UK-aware predictor and higher-agreement annotation, or the limitations of both validation steps should be acknowledged explicitly in the conclusion of Section 2.2.
- [Section 2.3, Table 4] The reliability of the sentiment and sarcasm labels is asserted as 'high degree of reliability' based on agreement between the original annotator and one independent annotator on only 50 instances per variety. For sarcasm, the reported kappas are 0.47 (en-AU), 0.51 (en-IN), and 0.63 (en-UK); these are moderate levels of agreement, not high reliability. Since the full dataset is annotated by a single annotator per variety, and two of the three original annotators are also authors, the annotation reliability claim should be tempered. Reporting per-label disagreement rates, confidence scores, or a larger independent annotation sample would substantially strengthen the benchmark.
minor comments (4)
- [Section 3] There is a typo in the prompt description: 'we the prompt the model with' should read 'we prompt the model with.'
- [Appendix D] The sentence 'Both models perform better for REDDIT → GOOGLE , suggesting that models may be better at transferring sentiment from a relatively informal domain .' has a stray space before the period and the reasoning is incomplete; the directionality claim would benefit from a reference or a brief explanation.
- [Figures 3 and 4] The figures report single F1 values without variance or significance information. Given the small performance differences between models, standard deviations across multiple random seeds would help the reader assess whether the reported trends are stable.
- [Section 2.2] The lack of access to ICE-GB is mentioned in a footnote but should also be stated in the main text or limitations, since it directly affects the validity of the en-UK variety validation.
Circularity Check
No circularity: BESSTIE's benchmark construction and model evaluation are independent; self-citations are background or calibration, and the DISTIL-based filtering is a dataset-design choice, not a fitted prediction.
full rationale
BESSTIE's central claims are (i) that the collected texts represent en-AU, en-IN, and en-UK, and (ii) that fine-tuned LLMs score higher on inner-circle varieties than on en-IN, especially for sarcasm. Neither claim is derived from the benchmark's own inputs by construction. Variety labels come from external collection signals (Google Places location and topic-selected subreddits) and are checked against an independent ICE-trained predictor; sentiment and sarcasm labels come from manual annotation, not from star ratings or model outputs. The DISTIL experiment in Section 2.1 is used to decide which star ratings to retain (2 and 4 rather than 1 and 5), which shapes the dataset, but the final evaluation reports held-out F-scores of separately fine-tuned models on manually annotated test splits, so no fitted parameter is renamed as a prediction. Self-citations, including Srirag et al. (2025), Joshi et al. (2016), Joshi et al. (2025), and Painter et al. (2022), are used for background motivation or agreement calibration and are not load-bearing premises. The paper's own Limitations section candidly acknowledges single-annotator bias and regional variation, further supporting that no derivation is hidden. The main caveat is statistical rather than circular: Table 5 shows very different sarcasm positive rates across varieties (42% for en-AU vs 13% for en-IN), and macro-F1 is prevalence-sensitive, so the headline inner/outer gap may be confounded; this is a correctness and robustness concern, not a circularity concern.
Assumptions & free parameters
free parameters (5)
- Star-rating filter =
2 and 4 stars
- fastText English probability threshold =
0.98
- City population thresholds =
en-AU: 20K; en-IN: 100K; en-UK: 50K
- Reddit sampling size =
3,000 comments per variety
- Learning rate grid =
1e-5, 2e-5, 3e-5
assumptions (5)
- domain assumption Location-based filtering of Google Places reviews approximates the national variety of the reviewer.
- domain assumption Topic-based filtering on selected subreddits approximates the national variety of commenters.
- domain assumption Manual annotation by one native speaker per variety yields reliable sentiment and sarcasm labels.
- domain assumption The ICE corpora (ICE-Australia and ICE-India) are a valid training source for automatic variety prediction.
- domain assumption National varieties are homogeneous enough to treat as single categories.
Cite this review
Pith. "Pith review of BESSTIE: A Benchmark for Sentiment and Sarcasm Classification for Varieties of English." pith.science (2026). https://pith.science/paper/V5BQNTSG
@misc{pith2026241204726,
author = {Pith},
title = {Pith review of: BESSTIE: A Benchmark for Sentiment and Sarcasm Classification for Varieties of English},
year = {2026},
howpublished = {\url{https://pith.science/paper/V5BQNTSG}},
note = {Machine review of arXiv:2412.04726}
}
read the original abstract
Despite large language models (LLMs) being known to exhibit bias against non-standard language varieties, there are no known labelled datasets for sentiment analysis of English. To address this gap, we introduce BESSTIE, a benchmark for sentiment and sarcasm classification for three varieties of English: Australian (en-AU), Indian (en-IN), and British (en-UK). We collect datasets for these language varieties using two methods: location-based for Google Places reviews, and topic-based filtering for Reddit comments. To assess whether the dataset accurately represents these varieties, we conduct two validation steps: (a) manual annotation of language varieties and (b) automatic language variety prediction. Native speakers of the language varieties manually annotate the datasets with sentiment and sarcasm labels. We perform an additional annotation exercise to validate the reliance of the annotated labels. Subsequently, we fine-tune nine LLMs (representing a range of encoder/decoder and mono/multilingual models) on these datasets, and evaluate their performance on the two tasks. Our results show that the models consistently perform better on inner-circle varieties (i.e., en-AU and en-UK), in comparison with en-IN, particularly for sarcasm classification. We also report challenges in cross-variety generalisation, highlighting the need for language variety-specific datasets such as ours. BESSTIE promises to be a useful evaluative benchmark for future research in equitable LLMs, specifically in terms of language varieties. The BESSTIE dataset is publicly available at: https://huggingface.co/ datasets/unswnlporg/BESSTIE.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Gavin Abercrombie and Dirk Hovy. 2016. https://doi.org/10.18653/v1/P16-3016 Putting sarcasm detection into context: The effects of class imbalance and manual labelling on supervised machine classification of T witter conversations . In Proceedings of the ACL 2016 Student Research Workshop , pages 107--113, Berlin, Germany. Association for Computational Li...
-
[4]
Ibrahim Abu Farha, Wajdi Zaghouani, and Walid Magdy. 2021. https://aclanthology.org/2021.wanlp-1.36 Overview of the WANLP 2021 shared task on sarcasm and sentiment detection in A rabic . In Proceedings of the Sixth Arabic Natural Language Processing Workshop, pages 296--305, Kyiv, Ukraine (Virtual). Association for Computational Linguistics
work page 2021
-
[5]
Basma Alharbi, Hind Alamro, Manal Alshehri, Zuhair Khayyat, Manal Kalkatawi, Inji Ibrahim Jaber, and Xiangliang Zhang. 2020. Asad: A twitter-based benchmark arabic sentiment analysis dataset. arXiv preprint arXiv:2011.00578
arXiv 2020
-
[6]
Su Lin Blodgett, Lisa Green, and Brendan O ' Connor. 2016. https://doi.org/10.18653/v1/D16-1120 Demographic dialectal variation in social media: A case study of A frican- A merican E nglish . In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 1119--1130, Austin, Texas. Association for Computational Linguistics
-
[7]
Dushyant Singh Chauhan, SR Dhanush, Asif Ekbal, and Pushpak Bhattacharyya. 2020. Sentiment and emotion help sarcasm? a multi-task learning framework for multi-modal sarcasm, sentiment and emotion analysis. In Proceedings of the 58th annual meeting of the association for computational linguistics, pages 4351--4360
work page 2020
-
[8]
Mark Cieliebak, Jan Milan Deriu, Dominic Egger, and Fatih Uzdilli. 2017. https://doi.org/10.18653/v1/W17-1106 A T witter corpus and benchmark resources for G erman sentiment analysis . In Proceedings of the Fifth International Workshop on Natural Language Processing for Social Media, pages 45--51, Valencia, Spain. Association for Computational Linguistics
Show all 45 references
-
[9]
Nicholas Deas, Jessica Grieser, Shana Kleiner, Desmond Patton, Elsbeth Turcan, and Kathleen McKeown. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.421 Evaluation of A frican A merican language bias in natural language generation . In Proceedings of the 2023 Conference on E...
2023 doi
-
[10]
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. https://openreview.net/forum?id=OUIFPHEgJU QL o RA : Efficient finetuning of quantized LLM s . In Thirty-seventh Conference on Neural Information Processing Systems
2023
-
[11]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423 BERT : Pre-training of deep bidirectional transformers for language understanding . In Proceedings of the 2019 Conference of the North A merican Chapter of the Associat...
2019 doi
-
[12]
AbdelRahim Elmadany, ElMoatez Billah Nagoudi, and Muhammad Abdul-Mageed. 2023. https://doi.org/10.18653/v1/2023.findings-acl.609 ORCA : A challenging benchmark for A rabic language understanding . In Findings of the Association for Computational Linguistics: ACL 2023, pages 95...
2023 doi
-
[13]
Fahim Faisal, Orevaoghene Ahia, Aarohi Srivastava, Kabir Ahuja, David Chiang, Yulia Tsvetkov, and Antonios Anastasopoulos. 2024. https://doi.org/10.18653/v1/2024.acl-long.777 DIALECTBENCH : An NLP benchmark for dialects, varieties, and closely-related languages . In Proceeding...
2024 doi
-
[14]
Elena Filatova. 2012. http://www.lrec-conf.org/proceedings/lrec2012/pdf/661_Paper.pdf Irony and sarcasm: Corpus generation and analysis using crowdsourcing . In Proceedings of the Eighth International Conference on Language Resources and Evaluation ( LREC '12) , pages 392--398...
2012
-
[15]
Gemma Team , Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, Johan Ferret, Peter Liu, Pouya Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charline Le La...
2024 arXiv
-
[16]
Naman Goyal, Jingfei Du, Myle Ott, Giri Anantharaman, and Alexis Conneau. 2021. https://doi.org/10.18653/v1/2021.repl4nlp-1.4 Larger-scale transformers for multilingual masked language modeling . In Proceedings of the 6th Workshop on Representation Learning for NLP (RepL4NLP-2...
2021 doi
-
[17]
Edouard Grave, Piotr Bojanowski, Prakhar Gupta, Armand Joulin, and Tomas Mikolov. 2018. Learning word vectors for 157 languages. In Proceedings of the International Conference on Language Resources and Evaluation (LREC 2018)
2018
-
[18]
Sidney Greenbaum and Gerald Nelson. 1996. https://doi.org/10.1111/j.1467-971X.1996.tb00088.x The international corpus of english (ice) project . World Englishes, 15(1):3--15
1996 doi
-
[19]
Nihal Jain, Dejiao Zhang, Wasi Uddin Ahmad, Zijian Wang, Feng Nan, Xiaopeng Li, Ming Tan, Ramesh Nallapati, Baishakhi Ray, Parminder Bhatia, Xiaofei Ma, and Bing Xiang. 2023. https://doi.org/10.18653/v1/2023.acl-long.355 C ontra CLM : Contrastive learning for causal language m...
2023 doi
-
[20]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas ...
2023 arXiv
-
[21]
Aditya Joshi, Raj Dabre, Diptesh Kanojia, Zhuang Li, Haolan Zhan, Gholamreza Haffari, and Doris Dippold. 2025. https://doi.org/10.1145/3712060 Natural language processing for dialects of a language: A survey . ACM Comput. Surv., 57(6)
2025 doi
-
[22]
Aditya Joshi, Vaibhav Tripathi, Pushpak Bhattacharyya, and Mark J. Carman. 2016. https://doi.org/10.18653/v1/K16-1015 Harnessing sequence labeling for sarcasm detection in dialogue from TV series F riends' . In Proceedings of the 20th SIGNLL Conference on Computational Natural...
2016 doi
-
[23]
Mikhail Khodak, Nikunj Saunshi, and Kiran Vodrahalli. 2018. https://aclanthology.org/L18-1102 A large self-annotated corpus for sarcasm . In Proceedings of the Eleventh International Conference on Language Resources and Evaluation ( LREC 2018) , Miyazaki, Japan. European Langu...
2018
-
[24]
Douwe Kiela, Max Bartolo, Yixin Nie, Divyansh Kaushik, Atticus Geiger, Zhengxuan Wu, Bertie Vidgen, Grusha Prasad, Amanpreet Singh, Pratik Ringshia, Zhiyi Ma, Tristan Thrush, Sebastian Riedel, Zeerak Waseem, Pontus Stenetorp, Robin Jia, Mohit Bansal, Christopher Potts, and Adi...
2021 doi
-
[25]
Bernd Kortmann, Kerstin Lunkenheimer, and Katharina Ehret, editors. 2020. https://ewave-atlas.org/ eWAVE
2020
-
[26]
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut. 2020. https://openreview.net/forum?id=H1eA7AEtvS Albert: A lite bert for self-supervised learning of language representations . In International Conference on Learning Representations
2020
-
[27]
Heather Lent, Kushal Tatariya, Raj Dabre, Yiyi Chen, Marcell Fekete, Esther Ploeger, Li Zhou, Ruth-Ann Armstrong, Abee Eijansantos, Catriona Malau, Hans Erik Heje, Ernests Lavrinovics, Diptesh Kanojia, Paul Belony, Marcel Bollmann, Loïc Grobol, Miryam de Lhoneux, Daniel Hershc...
2024 doi
-
[28]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. https://arxiv.org/abs/1907.11692 Roberta: A robustly optimized bert pretraining approach . Preprint, arXiv:1907.11692
2019 arXiv
-
[29]
Niklas Muennighoff, Nouamane Tazi, Loic Magne, and Nils Reimers. 2023. https://doi.org/10.18653/v1/2023.eacl-main.148 MTEB : Massive text embedding benchmark . In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages...
2023 doi
-
[30]
Shamsuddeen Hassan Muhammad, Idris Abdulmumin, Abinew Ali Ayele, Nedjma Ousidhoum, David Ifeoluwa Adelani, Seid Muhie Yimam, Ibrahim Sa'id Ahmad, Meriem Beloucif, Saif M. Mohammad, Sebastian Ruder, Oumaima Hourrane, Pavel Brazdil, Felermino Dário Mário António Ali, Davis David...
2023 arXiv
-
[31]
Dong Nguyen. 2021. 10 dialect variation on social media. Similar Languages, Varieties, and Dialects: A Computational Perspective, page 204
2021
-
[32]
Silviu Oprea and Walid Magdy. 2020. https://doi.org/10.18653/v1/2020.acl-main.118 i S arcasm: A dataset of intended sarcasm . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1279--1289, Online. Association for Computational Linguistics
2020 doi
-
[33]
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and...
2022
-
[34]
Jordan Painter, Helen Treharne, and Diptesh Kanojia. 2022. https://doi.org/10.18653/v1/2022.nlpcss-1.22 Utilizing weak supervision to create S 3 D : A sarcasm annotated dataset . In Proceedings of the Fifth Workshop on Natural Language Processing and Computational Social Scien...
2022 doi
-
[35]
Barbara Plank. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.731 The `` problem '' of human label variation: On ground truth in data, modeling and evaluation . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 10671--10682, Ab...
2022 doi
-
[36]
Tom \'a s Pt \'a c ek, Ivan Habernal, and Jun Hong. 2014. https://aclanthology.org/C14-1022 Sarcasm detection on C zech and E nglish T witter . In Proceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Papers , pages 213--223, Dub...
2014
-
[37]
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2020. https://arxiv.org/abs/1910.01108 Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter . Preprint, arXiv:1910.01108
2020 arXiv
-
[38]
Adam Smith and Pam Peters. 2023. https://doi.org/10.25949/24769173.v1 International Corpus of English (ICE)
2023 doi
-
[39]
Manning, Andrew Ng, and Christopher Potts
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013. https://aclanthology.org/D13-1170 Recursive deep models for semantic compositionality over a sentiment treebank . In Proceedings of the 2013 Conference on Emp...
2013
-
[40]
Dipankar Srirag, Nihar Ranjan Sahoo, and Aditya Joshi. 2025. https://aclanthology.org/2025.sumeval-2.3/ Evaluating dialect robustness of language models via conversation understanding . In Proceedings of the Second Workshop on Scaling Up Multilingual & Multi-Cultural Evaluatio...
2025
-
[41]
Jiao Sun, Thibault Sellam, Elizabeth Clark, Tu Vu, Timothy Dozat, Dan Garrette, Aditya Siddhant, Jacob Eisenstein, and Sebastian Gehrmann. 2023. https://doi.org/10.18653/v1/2023.acl-long.331 Dialect-robust evaluation of generated text . In Proceedings of the 61st Annual Meetin...
2023 doi
-
[42]
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2019. https://proceedings.neurips.cc/paper_files/paper/2019/file/4496bf24afe7fab6f046bf4923da8de6-Paper.pdf Superglue: A stickier benchmark for general-purp...
2019
-
[43]
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018. https://doi.org/10.18653/v1/W18-5446 GLUE : A multi-task benchmark and analysis platform for natural language understanding . In Proceedings of the 2018 EMNLP Workshop B lackbox NLP : A...
2018 doi
-
[44]
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jianxin Yang, Jin Xu, Jingren Zhou, Jinze...
2024 arXiv
-
[45]
Caleb Ziems, William Held, Jingfeng Yang, Jwala Dhamala, Rahul Gupta, and Diyi Yang. 2023. https://doi.org/10.18653/v1/2023.acl-long.44 Multi- VALUE : A framework for cross-dialectal E nglish NLP . In Proceedings of the 61st Annual Meeting of the Association for Computational ...
2023 doi
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.