{"id":"90819448-a0fe-48c3-ae39-bb2d5da736a3","arxiv_id":"2412.04726","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"BESSTIE is the first sentiment and sarcasm benchmark for three varieties of English, and models trained on it score lowest on Indian English.","lead":"This paper introduces BESSTIE, a labeled dataset for sentiment and sarcasm in Australian, Indian, and British English. It fine-tunes nine language models and finds they perform worse on Indian English, especially for sarcasm.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline en-IN sarcasm deficit is confounded by label prevalence: en-IN REDDIT has 13% sarcastic positives vs 42% en-AU, and macro-F1 is prevalence-dependent, so the comparison may not isolate variety.","rationale":"The reader's weakest assumption was that collected texts truly represent the intended national varieties, citing low annotator agreement (kappa=0.26) and the automatic predictor trained without ICE-GB. That remains a real concern for the benchmark's validity as a variety resource. However, the most load-bearing problem for the paper's empirical conclusion is the confound between label prevalence and macro-F1 in the sarcasm comparison, which is concrete, quantitative, and directly testable. Both issues are fixable and do not by themselves invalidate the dataset release, so I would not move the verdict away from CONDITIONAL. I set agreement_with_reader to partial because the reader identified variety-label reliability rather than the class-prior confound; my proposed check targets the latter. The check I propose—balanced subsets plus AUC—would settle whether the reported en-IN sarcasm deficit reflects genuine variety-specific difficulty or merely differences in how often sarcasm appears in each subset. This is a stronger and more decisive test than re-doing the variety annotation alone, because even perfect variety labels would not remove the class-prior confound from the current F-score comparison.","tokens_in":16140,"tokens_out":5629,"duration_ms":63091,"concrete_test":"Recompute the REDDIT-sarcasm evaluation on class-balanced test sets, e.g., subsample negative instances so that each variety has equal positive and negative counts, and report positive-class F1, macro-F1, and AUC with bootstrap confidence intervals. If the en-AU/en-UK > en-IN gap persists under balanced priors and prevalence-independent metrics, the variety effect is supported; if the gap shrinks or reverses, the headline result is an artifact of label distribution.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central empirical claim—models 'consistently perform better on inner-circle varieties (i.e., en-AU and en-UK), in comparison with en-IN, particularly for sarcasm classification' (Section 4.1)—rests on macro-averaged F-score comparisons across test sets with very different positive-class rates. Table 5 reports REDDIT sarcasm positive rates of 42% for en-AU, 13% for en-IN, and 22% for en-UK. Macro-F1 is not prevalence-neutral: for a fixed per-class error pattern, positive-class precision, and therefore F1, drops as the positive class becomes rarer. Thus the en-AU > en-IN gap in sarcasm can arise from class-prior differences alone, without any variety-specific model difficulty. The same confound affects REDDIT sentiment (positive rates 32%, 25%, 12% for en-AU, en-IN, en-UK) and to a lesser extent GOOGLE sentiment, though GOOGLE sentiment priors are similar across varieties (73–75%). The paper's use of weighted cross-entropy for encoder models does not resolve this, since the reported metric is still macro-F1, and decoder models are fine-tuned with unweighted maximum likelihood. Because the 'particularly for sarcasm' claim is the most prominent evidence of inner/outer-circle bias, the comparison must be shown to be robust to label prevalence before the headline conclusion is accepted.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces BESSTIE, a labelled benchmark for sentiment and sarcasm classification in three national varieties of English: Australian (en-AU), Indian (en-IN), and British (en-UK). Data are collected from Google Places reviews (location-based filtering) and Reddit comments (topic-based filtering), then annotated by native speakers. The authors validate variety representation through manual annotation and an automatic variety predictor, fine-tune nine language models on the benchmark, and report that models consistently perform better on the inner-circle varieties (en-AU, en-UK) than on en-IN, especially for sarcasm classification. The paper also reports cross-variety and cross-domain experiments, an error analysis of misclassified examples, and releases the dataset publicly.","tokens_in":16407,"tokens_out":3673,"duration_ms":37471,"significance":"If the dataset and the empirical claims hold, BESSTIE would be a useful resource for evaluating LLM bias across national varieties of English, an area where labelled benchmarks are scarce. The paper's strengths include the public release of the dataset, the inclusion of two tasks (sentiment and sarcasm) and two domains, the evaluation of nine diverse models, and a manual error analysis that identifies concrete dialectal and colloquial features. The dataset is likely to be reused by the community for bias measurement and cross-variety generalization studies, provided the validity of the variety labels and the robustness of the headline performance comparison are established.","major_comments":[{"comment":"The sentence 'due to the absence of sarcasm labels in GOOGLE subset (as shown in Table 5)' conflicts with Table 5, which reports % Pos. Sarc. values of 7, 1, and 0 for the GOOGLE subsets of en-AU, en-IN, and en-UK. If sarcasm labels were collected but not used, the paper should explain why they are excluded from evaluation; if they were not collected, Table 5 should not report them. As written, the scope of the benchmark's sarcasm component is unclear and the reader cannot determine whether the GOOGLE subset does or does not contain sarcasm labels.","section":"Section 4.1, Table 5"},{"comment":"The headline claim that models perform worse on en-IN 'particularly for sarcasm classification' is confounded by label prevalence. Table 5 shows REDDIT sarcasm positive rates of 42% for en-AU, 13% for en-IN, and 22% for en-UK, and REDDIT sentiment positive rates of 32%, 25%, and 12% for en-AU, en-IN, and en-UK. Because all reported metrics are macro-averaged F1, the score for a rare positive class is strongly affected by the class prior: a model with comparable per-class behavior will obtain a lower macro-F1 when the positive class is 13% than when it is 42%. The observed en-IN deficit, and especially the 'particularly for sarcasm' part, may therefore reflect prior differences rather than variety-specific model difficulty. Please report per-class F1, balanced accuracy, or re-evaluate on prevalence-matched test sets before drawing the central conclusion.","section":"Section 4.1, Table 5"},{"comment":"The automated variety validation does not validate en-UK. The DISTIL predictor is fine-tuned only on ICE-Australia and ICE-India for a binary inner-circle versus outer-circle classification task, yet Table 3 reports en-UK P(v) and F-scores. Since the predictor has never seen en-UK training data, high probabilities for en-UK texts likely reflect the binary decision boundary rather than evidence that the texts represent British English. In addition, the manual annotation agreement is low (inter-annotator kappa = 0.26; agreement with the true label is 0.41 and 0.34 for the en-IN and en-UK annotators), which weakens the claim that the subsets 'collectively form a good representative sample.' The validation section should either be redone with an en-UK-aware predictor and higher-agreement annotation, or the limitations of both validation steps should be acknowledged explicitly in the conclusion of Section 2.2.","section":"Section 2.2, Table 3"},{"comment":"The reliability of the sentiment and sarcasm labels is asserted as 'high degree of reliability' based on agreement between the original annotator and one independent annotator on only 50 instances per variety. For sarcasm, the reported kappas are 0.47 (en-AU), 0.51 (en-IN), and 0.63 (en-UK); these are moderate levels of agreement, not high reliability. Since the full dataset is annotated by a single annotator per variety, and two of the three original annotators are also authors, the annotation reliability claim should be tempered. Reporting per-label disagreement rates, confidence scores, or a larger independent annotation sample would substantially strengthen the benchmark.","section":"Section 2.3, Table 4"}],"minor_comments":[{"comment":"There is a typo in the prompt description: 'we the prompt the model with' should read 'we prompt the model with.'","section":"Section 3"},{"comment":"The sentence 'Both models perform better for REDDIT → GOOGLE , suggesting that models may be better at transferring sentiment from a relatively informal domain .' has a stray space before the period and the reasoning is incomplete; the directionality claim would benefit from a reference or a brief explanation.","section":"Appendix D"},{"comment":"The figures report single F1 values without variance or significance information. Given the small performance differences between models, standard deviations across multiple random seeds would help the reader assess whether the reported trends are stable.","section":"Figures 3 and 4"},{"comment":"The lack of access to ICE-GB is mentioned in a footnote but should also be stated in the main text or limitations, since it directly affects the validity of the en-UK variety validation.","section":"Section 2.2"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a timely and useful gap, and the resource itself is likely to be valuable. However, the central empirical claim is currently under-supported by the prevalence confound and the variety-validation gaps. These issues are fixable with additional analysis and revised reporting, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know about BESSTIE: it is genuinely the first labeled sentiment and sarcasm dataset for Australian, Indian, and British English, and the dataset is public. That alone makes it worth a look if you work on variety robustness or fairness. The authors collected Google reviews and Reddit comments, annotated them with native speakers, and ran a reasonable sweep of nine encoder/decoder models. The error analysis is a nice touch, and the cross-variety experiments are useful even if preliminary.\n\nThe soft spots are not fatal to the resource, but they do undercut the main empirical claim. The headline result—models consistently do worse on en-IN, especially sarcasm—is confounded by label prevalence. On Reddit sarcasm, en-AU has 42% positive examples, en-IN has 13%, en-UK 22%. Macro-F1 is not prevalence-neutral; a model that performs identically on all three will show a lower macro-F1 on the rarer-positive dataset. The paper does not control for this, so the 'particularly for sarcasm' conclusion is not supported as written. Weighted training loss does not fix the evaluation metric. This needs a prevalence-matched or precision-recall balanced analysis before the claim is taken seriously.\n\nThere is also an internal inconsistency: Section 4.1 says sarcasm labels are absent from the GOOGLE subset, but Table 5 lists 7% for en-AU and 1% for en-IN. It looks like only en-UK GOOGLE has zero sarcasm. Either the table or the text is wrong, and it should be corrected.\n\nThe variety validation is weaker than the narrative suggests. Inter-annotator agreement is κ=0.26, which is poor. The automatic predictor is trained on ICE-Australia and ICE-India only, yet reports en-UK scores; that is not a validation of the en-UK subset. The en-IN Reddit F-score of 0.69 means a quarter of that data may not actually be Indian English, or at least the predictor cannot recognise it. The authors acknowledge some of this in the limitations, but the central comparison between varieties rests on labels that are shaky.\n\nBottom line: the dataset is a contribution and the release is honest in terms of numbers. But the paper's headline finding needs a prevalence-controlled re-analysis and the validation needs more work. I'd send it to review—the resource is worth referee time—but I would not accept it in its current form. The authors need to either add a matched analysis or substantially soften the model-bias claims.","headline":"BESSTIE fills a real gap as the first labeled sentiment/sarcasm benchmark for national English varieties, but its headline inner-circle advantage may be an artifact of label imbalance, and the variety validation is too thin to bear the paper's claims.","tokens_in":16956,"tokens_out":2966,"would_cite":true,"duration_ms":26827,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"BESSTIE: a benchmark for sentiment and sarcasm in three varieties of English, showing models lag on Indian English.","keywords":["sentiment analysis","sarcasm detection","English varieties","Indian English","Australian English","British English","LLM bias","benchmark dataset"],"falsifier":"Re-annotate a fresh random sample of the same sources with a panel of multiple native speakers per variety; if those annotations show the en-IN texts are not reliably distinguishable from en-AU or en-UK, or if a model trained on the re-annotated labels no longer shows a performance gap for en-IN, the paper's central bias claim would fail. Alternatively, a human evaluation showing BESSTIE's variety labels are no more accurate than chance would invalidate the benchmark.","tokens_in":15899,"feed_emoji":"🗣️","tokens_out":4383,"duration_ms":38483,"temperature":0.7,"pith_summary":"BESSTIE is the first labelled benchmark for sentiment and sarcasm classification built specifically for national varieties of English, covering Australian (en-AU), Indian (en-IN), and British (en-UK) English. The paper's central finding is that fine-tuned large language models consistently perform better on the inner-circle varieties (en-AU and en-UK) than on en-IN, with the gap widest for sarcasm detection. The authors argue the benchmark provides a reusable resource for measuring and eventually mitigating LLM bias against non-standard English varieties. The dataset combines Google Places reviews and Reddit comments, and each text is annotated by native speakers for positive/negative sentiment and sarcasm.","feed_headline":"Benchmark shows LLMs lag on Indian English sentiment and sarcasm","feed_subtitle":"BESSTIE is the first labelled dataset for measuring how well models handle Australian, Indian, and British English.","key_machinery":"The load-bearing object is the BESSTIE dataset itself: roughly 15,000 annotated texts assembled by two complementary collection methods (location-based Google Places reviews and topic-based Reddit subreddits selected by native speakers), then filtered to ambiguous star ratings (2 or 4 stars) to avoid trivial polarity. Variety representation is checked by a two-step validation: manual annotation by en-IN and en-UK speakers (inter-annotator agreement κ=0.26) and an automatic variety predictor fine-tuned on ICE-Australia and ICE-India. Sentiment and sarcasm labels come from single native-speaker annotators per variety, with a secondary annotation pass reporting κ between 0.47 and 0.79 for sentiment and 0.47 to 0.63 for sarcasm. The evaluation protocol fine-tunes nine models on the benchmark and compares in-variety and cross-variety performance, which is the mechanism that produces the inner-circle vs. outer-circle gap.","core_discovery":"The paper claims BESSTIE is the first benchmark to provide sentiment and sarcasm labels for natural, user-generated text across multiple varieties of English. Across nine fine-tuned LLMs (six encoders and three decoders), models average F-scores of 0.78 for en-AU, 0.74 for en-UK, and 0.63 for en-IN across domain–task pairs, with sarcasm classification on Reddit comments particularly low (overall 0.59). The authors interpret this as evidence that current LLMs are biased against outer-circle English varieties, and that sarcasm specifically resists cross-variety generalisation because it depends on local cultural and contextual knowledge. They also report that monolingual models marginally outperform multilingual ones, suggesting that broad pre-training language coverage does not automatically extend to varieties of a single language.","pith_inferences":["If the variety labels hold up under wider scrutiny, BESSTIE could be extended to other outer-circle varieties (e.g., Nigerian or Singaporean English) to test whether the en-IN gap generalises or is specific to that variety.","The low annotator agreement on variety identification (κ=0.26) suggests that national variety is often ambiguous in short web text; future benchmarks might combine location signals with author self-identification or expert adjudication to strengthen labels.","A natural testable extension is to evaluate instruction-tuned LLMs zero-shot on BESSTIE without fine-tuning, which would isolate pre-training bias from adaptation effects."],"forward_implications":["BESSTIE gives future work a fixed, public testbed for measuring how well models handle Australian, Indian, and British English in sentiment and sarcasm tasks.","Sarcasm classification across varieties is far from solved; the best average F-score is 0.68 and cross-variety fine-tuning hurts out-of-variety sarcasm performance.","The consistent en-IN deficit suggests that models need variety-specific training data or debiasing methods tuned to outer-circle English, not just more multilingual pre-training.","Monolingual encoders outperform multilingual decoders on these tasks on average, so architectural choices matter when evaluating variety robustness."],"supporting_citations":[{"why":"Provides DialectBench, the recent dialectal NLP benchmark that BESSTIE extends by adding English sentiment and sarcasm datasets.","marker":"Faisal et al. (2024)"},{"why":"Multi-VALUE supplies a synthetic cross-dialect English benchmark; BESSTIE positions itself as the natural-text alternative that better represents real variety usage.","marker":"Ziems et al. (2023)"},{"why":"Demonstrates LLM bias against African-American English, which motivates the benchmark's goal of measuring bias against other non-standard English varieties.","marker":"Deas et al. (2023)"},{"why":"Shows LLM bias against Indian English, directly motivating the inclusion of en-IN in the benchmark.","marker":"Srirag et al. (2025)"},{"why":"The ICE corpora provide the training data for the automatic variety predictor used to validate BESSTIE's variety labels.","marker":"Greenbaum and Nelson (1996)"},{"why":"Supplies a reference sarcasm annotation agreement (κ=0.44) that the paper uses to argue BESSTIE's sarcasm labels are comparably reliable.","marker":"Joshi et al. (2016)"}],"fun_headline_variants":["New benchmark exposes LLM bias against Indian English sarcasm","BESSTIE: first benchmark for sentiment and sarcasm in English varieties","LLMs struggle with Indian English sarcasm in new benchmark","Benchmark reveals LLMs favor Australian, British over Indian English"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire comparison rests on the assumption that the collected texts genuinely represent the three national varieties, which depends on location- and topic-based filtering supported by a variety-validation step with low annotator agreement (κ=0.26) and a predictor trained only on ICE-Australia and ICE-India yet reporting en-UK scores.","fun_headline_variants_meta":{"raw":{"variants":["New benchmark exposes LLM bias against Indian English sarcasm","BESSTIE: first benchmark for sentiment and sarcasm in English varieties","LLMs struggle with Indian English sarcasm in new benchmark","Benchmark reveals LLMs favor Australian, British over Indian English"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000199,"raw_usage":{"total_tokens":1410,"prompt_tokens":1020,"completion_tokens":390,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":636,"completion_tokens_details":{"reasoning_tokens":318}},"tokens_in":636,"tokens_out":390,"duration_ms":4290,"temperature":1.0,"reasoning_tokens":318,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T21:18:54.461342+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate a fresh random sample of the same sources with a panel of multiple native speakers per variety; if those annotations show the en-IN texts are not reliably distinguishable from en-AU or en-UK, or if a model trained on the re-annotated labels no longer shows a performance gap for en-IN, the paper's central bias claim would fail. Alternatively, a human evaluation showing BESSTIE's variety labels are no more accurate than chance would invalidate the benchmark.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows LLM bias against Indian English, directly motivating the inclusion of en-IN in the benchmark."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies a reference sarcasm annotation agreement (κ=0.44) that the paper uses to argue BESSTIE's sarcasm labels are comparably reliable."}],"review_version":1}