REVIEW 3 major objections 5 minor 1 cited by
Can Prompting LLMs Unlock Hate Speech Detection across Languages? A Zero-shot and Few-shot Study
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Prompted multilingual language models are worse at recognizing hate speech in real-world social media posts than fine-tuned encoders, yet they generalize better on controlled functional tests of hate speech detection across eight…
desk verdict A broad, useful prompt-taxonomy sweep for multilingual hate speech detection whose headline split is probably real, but whose Table 2 numbers are picked on the test set and need a held-out selection protocol before being quoted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The argument is carried by two evaluation instruments paired with a wide prompt-design search. The real-world test sets are 2,000-sample subsets of eight existing hate speech datasets (1,000 for Arabic, 1,500 for French) drawn from actual social media conversation. The functional instrument is the HateCheck benchmark and its multilingual extension, which supply controlled test cases for capabilities such as detecting implicit hate, handling negation, and not flagging non-hateful uses of slurs. The other mechanism is the prompt zoo: nine or more zero-shot templates (vanilla, classification, definition, chain-of-thought, NLI, role-play, multilingual, translate, distinction) plus few-shot variants with 1, 3, or 5 examples per class, applied to four instruction-tuned LLMs. The study's key comparison is the best per-language prompting result against two fine-tuned encoders, XLM-T and mDeBERTa, on both test types.
What would settle it
Re-run the same evaluation but select each language's prompt on a held-out validation set (or on a separate development portion of HateCheck) and then apply the chosen prompts to the real-world test sets and the held-out functional tests; if prompted LLMs no longer beat fine-tuned encoders on the functional tests, the reported generalization advantage is an artifact of test-set prompt selection.
Extended reading notes
Core claim
On the paper's own terms, the discovery is a performance split that depends on the evaluation distribution, not on which approach is universally better. On real-world datasets, the best prompted LLM result trails the best fine-tuned encoder in most languages—for example, Turkish real-world F1 is 81.76 for few-shot prompting versus 88.32 for XLM-T, and German 77.55 versus 79.18—though prompting already beats one or both encoders in languages like Arabic and French. On HateCheck functional tests, the ordering reverses across all eight languages: Hindi few-shot prompting reaches 65.93 versus 23.26 for XLM-T, Arabic 71.88 versus 25.47, Spanish 87.40 versus 67.93. The paper attributes the reversal to stronger generalization by instruction-tuned LLMs in controlled settings and shows that few-shot examples (typically five) with language-appropriate prompt templates raise functional-test scores further. The conclusion is not that prompting replaces fine-tuning, but that the choice should be guided by data availability and by whether the target is real-world distribution or functional robustness.
Load-bearing premise
The comparison assumes that choosing each language's best prompt from the test results themselves, as Table 2 does, gives a fair estimate of prompting performance rather than an optimistically selected one.
Editorial extensions
If this is right
- In low-resource settings, a team with only a few hundred labeled examples can use zero- or few-shot prompting to match or exceed a fine-tuned encoder, lowering the data threshold for deploying hate speech detection in a new language.
- For robustness evaluations, prompted LLMs are a better starting point than fine-tuned encoders, since they generalize better on controlled functional tests such as HateCheck.
- Few-shot prompting, usually five examples per class, should be part of functional-test deployments because it consistently raises HateCheck scores in most languages.
- Because the winning prompt varies by language, model, and test type, practical systems should search over prompt templates rather than assume one prompt transfers.
- With abundant training data, fine-tuning an encoder remains the more effective route for real-world distributions; prompting is not a substitute there.
Reading between the lines
- Because the best-prompt numbers are selected on the test set, the magnitude of the prompting advantage on HateCheck may be overstated; a held-out prompt-selection protocol could shrink the gap.
- HateCheck cases are generated from templates, so part of the LLMs' functional-test edge may come from having seen similar template patterns during pretraining; an adversarially perturbed functional set could narrow the gap.
- The paper's language-specific prompt rankings suggest a testable controller: use a small labeled validation set to pick among prompt templates per language, then measure real-world and functional F1; this could turn prompt search into a deployable decision rule.
- The data-threshold points (roughly 100-200 examples for Spanish, 300-400 for Hindi, 600-700 for German) imply a practical annotation-budget rule: below these budgets, prompt; above them, fine-tune—though the paper presents these numbers as observations, not as a rule.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a comparative study of zero-shot and few-shot prompting of four multilingual instruction-tuned LLMs (LLaMA-3.1-8B-Instruct, Qwen2.5-7B-Instruct, Aya-101, BloomZ-7B1) against fine-tuned encoder classifiers (mDeBERTa, XLM-T) for hate speech detection in eight non-English languages. Using many prompt templates and 1/3/5-shot variants, it evaluates on real-world tweet test splits and the multilingual HateCheck functional benchmark. The central claim is that fine-tuned encoders generally outperform prompted LLMs on real-world data, while prompting (especially few-shot) achieves higher F1-macro on functional tests, and that prompt design is language-dependent.
Significance. The paper has strong practical scope: it covers four LLMs, eight languages, a wide prompt grid, and both naturalistic and functional test sets, and it includes a data-curve analysis that is relevant for low-resource deployment. If the comparative claim survives closer scrutiny, it offers a useful trade-off rule for practitioners. The reproducibility-oriented release of code and prompts is a further strength. The main weakness is that the headline comparison is based on per-language best prompts selected on the test set, so the quantitative size of the prompting advantage on functional tests is not yet established.
major comments (3)
- [Section 6, Table 2 and Appendix C] The prompting scores in Table 2 are 'best zero- and few-shot prompting results', obtained by taking the maximum F1-macro over four LLMs and the full prompt grid on the same test sets used for evaluation. No held-out validation-based prompt selection is described. Appendix D shows that the selected prompt changes by language and by condition (e.g., Spanish functional Llama3: vanilla 86.37 vs. CoT 33.59), so the selection penalty can be large. This makes the exact magnitude of the claimed functional-test advantage over encoders uncertain. Please either use a fixed prompt per condition chosen on a validation split, report distributions over prompts, or report the sensitivity of the headline comparison to the selection procedure.
- [Section 6, Figure 1 and Appendix C] The data-curve comparison, including the statements that prompting becomes competitive with 100–200 examples in Spanish, 300–400 in Hindi, and 600–700 in German, compares an oracle-selected prompting score against fine-tuned XLM-T without error bars or significance tests. The crossover points are therefore not established with any stated uncertainty. Please add confidence intervals or seed-level variability, and state explicitly whether the same best-prompt selection criterion is applied on the encoder side or on a separate validation split.
- [Section 6 and Conclusion] The wording 'stronger generalization ability' is stronger than what is measured: the LLMs are not trained, and the functional-test advantage is a performance comparison after test-set prompt selection, not a controlled measure of generalization. Please qualify the claim (for example, 'on held-out functional test suites, prompted LLMs achieve higher F1-macro') and avoid implying a general capability ordering beyond the evaluated conditions.
minor comments (5)
- [Table 1] The model header 'Qwan' is a typo and should read 'Qwen'.
- [Table 1 and Appendix B] The prompt labels 'general' and 'distiction' appear in Table 1 but do not match any template in Tables 3–5; please add the missing templates or rename the labels for consistency.
- [Appendix D] Full per-prompt results are shown only for Spanish and Portuguese; since the headline comparison depends on the full prompt grid, please make complete tables available for all languages in the supplementary material.
- [Section 3] The description of random test-set construction should state whether the split is stratified by class and whether a fixed random seed is used, for reproducibility.
- [Section 5 and Appendix B] The few-shot naming convention is slightly confusing: a 'five-shot' condition includes ten examples (five per class). Please clarify this at first use.
Circularity Check
No circularity: empirical comparison against external benchmarks; test-set prompt selection is a methodology concern, not a circular derivation.
full rationale
This paper is an empirical benchmark study with no formal derivation or fitted model whose outputs are defined by its inputs. The central comparison—prompted LLMs versus fine-tuned encoder models on real-world and HateCheck functional test sets—is evaluated against external benchmarks (Sections 3–6). Encoder baselines are fine-tuned on training splits; LLM prompting is run in inference mode with no parameter updates (Appendix A). The only selection step is the per-language, per-condition choice of the best prompt reported in Table 2 as 'best zero- and few-shot prompting results.' This is an oracle-style model selection on the test set and may overstate prompting performance, especially on functional tests, but it is a statistical or evaluation concern rather than a circularity: the reported F1 scores are computed from model outputs on test items, not constructed from the selection criterion. No load-bearing claim is justified by self-citation: the cited prior work by overlapping authors (Masud et al. 2024; Zhang et al. 2025) is used only as related work and does not supply the paper's assumptions or conclusions. Therefore no step reduces to its own inputs by definition, and the circularity score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption The selected datasets follow the protected-group definition of hate speech and their binary labels are comparable across languages.
- domain assumption Multilingual HateCheck functional tests are a valid measure of real-world generalization.
- domain assumption Instruction-tuned LLMs in inference mode produce usable binary classifications in the target languages without decoding modifications.
Cite this review
Pith. "Pith review of Can Prompting LLMs Unlock Hate Speech Detection across Languages? A Zero-shot and Few-shot Study." pith.science (2026). https://pith.science/paper/TTMVFDKW
@misc{pith2026250506149,
author = {Pith},
title = {Pith review of: Can Prompting LLMs Unlock Hate Speech Detection across Languages? A Zero-shot and Few-shot Study},
year = {2026},
howpublished = {\url{https://pith.science/paper/TTMVFDKW}},
note = {Machine review of arXiv:2505.06149}
}
read the original abstract
Despite growing interest in automated hate speech detection, most existing approaches overlook the linguistic diversity of online content. Multilingual instruction-tuned large language models such as LLaMA, Aya, Qwen, and BloomZ offer promising capabilities across languages, but their effectiveness in identifying hate speech through zero-shot and few-shot prompting remains underexplored. This work evaluates LLM prompting-based detection across eight non-English languages, utilizing several prompting techniques and comparing them to fine-tuned encoder models. We show that while zero-shot and few-shot prompting lag behind fine-tuned encoder models on most of the real-world evaluation sets, they achieve better generalization on functional tests for hate speech detection. Our study also reveals that prompt design plays a critical role, with each language often requiring customized prompting techniques to maximize performance.
Figures
Forward citations
Cited by 1 Pith paper
-
LLM Harms: A Taxonomy and Discussion
This paper proposes a taxonomy of LLM harms in five categories and suggests mitigation strategies plus a dynamic auditing system for responsible development.
Reference graph
Works this paper leans on
-
[1]
Mistral AI. 2024. Un ministral, des ministraux. https://mistral.ai/news/ministraux. Accessed: 2025-04-19
work page 2024
-
[2]
Mehdi Ali, Michael Fromm, Klaudia Thellmann, Jan Ebert, Alexander Arno Weber, Richard Rutmann, Charvi Jain, Max Lübbering, Daniel Steinigen, Johannes Leveling, Katrin Klug, Jasper Schulze Buschhoff, Lena Jurkschat, Hammam Abdelwahab, Benny Jörg Stein, Karl-Heinz Sylla, Pavel Denisov, Nicolo' Brandizzi, Qasid Saleem, Anirban Bhowmick, Lennard Helmer, Chels...
arXiv 2024
-
[3]
Francesco Barbieri, Luis Espinosa Anke, and Jose Camacho-Collados. 2022. https://aclanthology.org/2022.lrec-1.27/ XLM - T : Multilingual language models in T witter for sentiment analysis and beyond . In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 258--266, Marseille, France. European Language Resources Association
work page 2022
-
[4]
Valerio Basile, Cristina Bosco, Elisabetta Fersini, Debora Nozza, Viviana Patti, Francisco Manuel Rangel Pardo, Paolo Rosso, and Manuela Sanguinetti. 2019. https://doi.org/10.18653/v1/S19-2007 S em E val-2019 task 5: Multilingual detection of hate speech against immigrants and women in T witter . In Proceedings of the 13th International Workshop on Semant...
-
[5]
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzm \'a n, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. https://doi.org/10.18653/v1/2020.acl-main.747 Unsupervised cross-lingual representation learning at scale . In Proceedings of the 58th Annual Meeting of the Association for Comp...
-
[6]
Arid Hasan, Imran Razzak, and Usman Naseem
Krishno Dey, Prerona Tarannum, Md. Arid Hasan, Imran Razzak, and Usman Naseem. 2024. http://arxiv.org/abs/2410.13153 Better to ask in english: Evaluation of large language models on english, low-resource and cross-lingual settings
arXiv 2024
-
[7]
Fatema Tuj Johora Faria, Laith H. Baniata, and Sangwoo Kang. 2024. https://doi.org/10.3390/math12233687 Investigating the predominance of large language models in low-resource bangla language over transformer models for hate speech detection: A comparative analysis . Mathematics, 12(23)
-
[8]
Paula Fortuna, Jo \ a o Rocha da Silva, Juan Soler-Company, Leo Wanner, and S \'e rgio Nunes. 2019. https://doi.org/10.18653/v1/W19-3510 A hierarchically-labeled P ortuguese hate speech dataset . In Proceedings of the Third Workshop on Abusive Language Online, pages 94--104, Florence, Italy. Association for Computational Linguistics
Show all 44 references
-
[9]
Janis Goldzycher, Paul R \"o ttger, and Gerold Schneider. 2024. https://doi.org/10.18653/v1/2024.naacl-long.248 Improving adversarial data collection by supporting annotators: Lessons from GAHD , a G erman hate speech dataset . In Proceedings of the 2024 Conference of the Nort...
2024 doi
-
[10]
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Art...
2024 arXiv
-
[11]
Keyan Guo, Alexander Hu, Jaden Mu, Ziheng Shi, Ziming Zhao, Nishant Vishwamitra, and Hongxin Hu. 2023. https://ieeexplore.ieee.org/abstract/document/10459901 An investigation of large language models for real-world hate speech detection . In 2023 International Conference on Ma...
2023
-
[12]
Lawrence Han and Hao Tang. 2022. https://ieeexplore.ieee.org/document/10216739/ Designing of prompts for hate speech recognition with in-context learning . In 2022 International Conference on Computational Science and Computational Intelligence (CSCI), pages 319--320. IEEE
2022
-
[13]
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021. https://openreview.net/forum?id=XP9wzJg9dG Deberta: Decoding-enhanced bert with disentangled attention . In International Conference on Learning Representations (ICLR)
2021
-
[14]
Fan Huang, Haewoon Kwak, and Jisun An. 2023. https://dl.acm.org/doi/10.1145/3543873.3587368 Is chatgpt better than human annotators? potential and limitations of chatgpt in explaining implicit hate speech . In Companion proceedings of the ACM web conference 2023, pages 294--297
2023
-
[15]
Lingyao Li, Lizhou Fan, Shubham Atreja, and Libby Hemphill. 2024. https://dl.acm.org/doi/10.1145/3643829 “hot” chatgpt: The promise of chatgpt in detecting and discriminating hateful, offensive, and toxic comments on social media . ACM Transactions on the Web, 18(2):1--36
2024 doi
-
[16]
Sean MacAvaney, Hao-Ren Yao, Eugene Yang, Katina Russell, Nazli Goharian, and Ophir Frieder. 2019. https://doi.org/https://doi.org/10.1371/journal.pone.0221152 Hate speech detection: Challenges and solutions . PloS one, 14(8):e0221152
2019 doi
-
[17]
Sarah Masud, Sahajpreet Singh, Viktor Hangya, Alexander Fraser, and Tanmoy Chakraborty. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.886 Hate personified: Investigating the role of LLM s in content moderation . In Proceedings of the 2024 Conference on Empirical Methods in...
2024 doi
-
[18]
Sandip Modha, Thomas Mandl, Gautam Kishore Shahi, Hiren Madhu, Shrey Satapara, Tharindu Ranasinghe, and Marcos Zampieri. 2021. https://dl.acm.org/doi/10.1145/3503162.3503176 Overview of the hasoc subtrack at fire 2021: Hate speech and offensive content identification in englis...
2021
-
[19]
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward ...
2023 doi
-
[20]
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward ...
2023 arXiv
-
[21]
Nedjma Ousidhoum, Zizheng Lin, Hongming Zhang, Yangqiu Song, and Dit-Yan Yeung. 2019. https://doi.org/10.18653/v1/D19-1474 Multilingual and multi-aspect hate speech analysis . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th...
2019 doi
-
[22]
Filippo Pedrazzini. 2025. https://blog.premai.io/multilingual-llms-progress-challenges-and-future-directions/ Multilingual llms: Progress, challenges, and future directions . Accessed: 2025-04-14
2025
-
[23]
Hao Peng, Xiaozhi Wang, Jianhui Chen, Weikai Li, Yunjia Qi, Zimu Wang, Zhili Wu, Kaisheng Zeng, Bin Xu, Lei Hou, and Juanzi Li. 2023. http://arxiv.org/abs/2311.08993 When does in-context learning fall short and why? a study on specification-heavy tasks
2023 arXiv
-
[24]
Qwen, :, An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, ...
2025 arXiv
-
[25]
Paul R \"o ttger, Debora Nozza, Federico Bianchi, and Dirk Hovy. 2022 a . https://doi.org/10.18653/v1/2022.emnlp-main.383 Data-efficient strategies for expanding hate speech detection into under-resourced languages . In Proceedings of the 2022 Conference on Empirical Methods i...
2022 doi
-
[26]
Paul R \"o ttger, Haitham Seelawi, Debora Nozza, Zeerak Talat, and Bertie Vidgen. 2022 b . https://doi.org/10.18653/v1/2022.woah-1.15 Multilingual H ate C heck: Functional tests for multilingual hate speech detection models . In Proceedings of the Sixth Workshop on Online Abus...
2022 doi
-
[27]
Paul R \"o ttger, Bertie Vidgen, Dong Nguyen, Zeerak Waseem, Helen Margetts, and Janet Pierrehumbert. 2021. https://doi.org/10.18653/v1/2021.acl-long.4 H ate C heck: Functional tests for hate speech detection models . In Proceedings of the 59th Annual Meeting of the Associatio...
2021 doi
-
[28]
Sarthak Roy, Ashish Harshvardhan, Animesh Mukherjee, and Punyajoy Saha. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.407 Probing LLM s for hate speech detection: strengths and vulnerabilities . In Findings of the Association for Computational Linguistics: EMNLP 2023, ...
2023 doi
-
[29]
Manuela Sanguinetti, Gloria Comandini, Elisa Di Nuovo, Simona Frenda, Marco Stranisci, Cristina Bosco, Tommaso Caselli, Viviana Patti, and Irene Russo. 2020. https://ceur-ws.org/Vol-2765/paper162.pdf Haspeede 2@ evalita2020: Overview of the evalita 2020 hate speech detection t...
2020
-
[30]
Uri Shaham, Jonathan Herzig, Roee Aharoni, Idan Szpektor, Reut Tsarfaty, and Matan Eyal. 2024. https://doi.org/10.18653/v1/2024.findings-acl.136 Multilingual instruction tuning with just a pinch of multilinguality . In Findings of the Association for Computational Linguistics:...
2024 doi
-
[31]
Michał Skibicki. 2025. https://www.stxnext.com/blog/large-language-models-functionality-and-impact-on-everyday-applications Large language models: Functionality and impact on everyday applications . Accessed: 2025-04-14
2025
-
[32]
Daniela Stockmann, Sophia Schlosser, and Paxia Ksatryo. 2023. https://onlinelibrary.wiley.com/doi/full/10.1002/poi3.348 Social media governance and strategies to combat online hatespeech in germany . Policy & Internet, 15(4):627--645
2023 doi
-
[33]
Kurt Thomas, Devdatta Akhawe, Michael Bailey, Dan Boneh, Elie Bursztein, Sunny Consolvo, Nicola Dell, Zakir Durumeric, Patrick Gage Kelley, Deepak Kumar, Damon McCoy, Sarah Meiklejohn, Thomas Ristenpart, and Gianluca Stringhini. 2021. https://doi.org/10.1109/SP40001.2021.00028...
2021
-
[34]
Hale, Samuel P
Manuel Tonneau, Diyi Liu, Niyati Malhotra, Scott A. Hale, Samuel P. Fraiberger, Victor Orozco-Olvera, and Paul Röttger. 2024. http://arxiv.org/abs/2411.15462 Hateday: Insights from a global hate speech dataset representative of a day on twitter
2024 arXiv
-
[35]
Cagri Toraman, Furkan S ahinu c , and Eyup Yilmaz. 2022. https://aclanthology.org/2022.lrec-1.238/ Large-scale hate speech detection with cross-domain transfer . In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 2215--2225, Marseille, France. ELRA
2022
-
[36]
Ahmet \"U st \"u n, Viraat Aryabumi, Zheng Yong, Wei-Yin Ko, Daniel D ' souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh, Hui-Lee Ooi, Amr Kayid, Freddie Vargus, Phil Blunsom, Shayne Longpre, Niklas Muennighoff, Marzieh Fadaee, Julia Kreutzer, and Sara Hooker. 2024. ht...
2024 doi
-
[37]
Janikke Solstad Vedeler, Terje Olsen, and John Eriksen. 2019. https://www.tandfonline.com/doi/abs/10.1080/09687599.2018.1515723 Hate speech harms: A social justice discussion of disabled norwegians’ experiences . Disability & Society, 34(3):368--383
2019
- [38]
-
[39]
Anwar Hossain Zahid, Monoshi Kumar Roy, and Swarna Das. 2025. http://arxiv.org/abs/2502.19612 Evaluation of hate speech detection using large language models and geographical contextualization
2025 arXiv
-
[40]
Shengyu Zhang, Linfeng Dong, Xiaoya Li, Sen Zhang, Xiaofei Sun, Shuhe Wang, Jiwei Li, Runyi Hu, Tianwei Zhang, Fei Wu, and Guoyin Wang. 2024. http://arxiv.org/abs/2308.10792 Instruction tuning for large language models: A survey
2024
-
[41]
Yaqi Zhang, Viktor Hangya, and Alexander Fraser. 2025. https://aclanthology.org/2025.coling-main.188/ LLM sensitivity challenges in abusive language detection: Instruction-tuned vs. human feedback . In Proceedings of the 31st International Conference on Computational Linguisti...
2025
-
[42]
Yiming Zhu, Peixian Zhang, Ehsan-Ul Haq, Pan Hui, and Gareth Tyson. 2025. https://link.springer.com/chapter/10.1007/978-3-031-78548-1_2 Exploring the capability of chatgpt to reproduce human labels for social computing tasks . In Social Networks Analysis and Mining, pages 13--...
2025 doi
-
[43]
URL: " 'urlintro :=
ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before...
-
[44]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.