REVIEW 4 major objections 5 minor 78 references
Multilingual and Explainable Text Detoxification with Parallel Corpora
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This paper claims that manually curated parallel corpora in German, Chinese, Arabic, Hindi, and Amharic, combined with a cluster-aware Chain-of-Thought prompt, make text detoxification both more multilingual and more explainable.
desk verdict The five-language parallel detoxification corpora are a real resource worth citing; the CoT method claim is not supported by the current evaluation and needs per-language validation of the toxicity classifier, significance testing, and held-out cluster selection. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the cluster-conditioned Chain-of-Thought (CoT) prompt. First, GPT-4 extracts five descriptive features (toxicity level, tone, language type, implied sentiment, negative connotations) for toxic and detoxified sentences; the toxic sentences' feature vectors are validated against native speakers, then one-hot encoded and clustered by k-means into three groups per language. The three clusters correspond to repair strategies—removing profanities, rephrasing condescending or biased language, and lightly adjusting informal text—and each cluster carries a human-readable explanation and a representative parallel pair. When a new toxic sentence arrives, the prompt asks the LLM to estimate its features, assign it to a cluster, and detoxify it using that cluster's explanation and example. This turns the explainability analysis into a conditioning signal for generation, which the paper says reduces hallucination and makes the edit more targeted.
What would settle it
Give native speakers a blind sample of 200 GPT-4 CoT outputs and 200 GPT-4 few-shot outputs per language and ask them to judge toxicity and meaning preservation. If human raters do not rate CoT outputs as non-toxic more often, or if the per-language ranking changes, the reported STA-based Joint advantage is an artifact of the toxicity checker. A direct pilot: measure the checker's agreement with human toxicity judgments on the human detoxification references for Chinese and Amharic, where the checker's scores diverge most.
Extended reading notes
Core claim
The paper's central claim is that supervised text detoxification can be extended to German, Chinese, Arabic, Hindi, and Amharic by manually curating parallel toxic-to-neutral pairs, and that explaining toxicity with GPT-4 can be turned into a better detoxification prompt. The authors collect 400 training and 600 test sentence pairs for each new language, then ask GPT-4 to label sentences across the nine languages with toxicity level, tone, language style, implied sentiment, and negative connotations, with native-speaker validation reported at 98% agreement. The toxic sentences' labels are one-hot encoded and clustered by k-means into three per-language detoxification strategies, and a Chain-of-Thought prompt first assigns a new toxic sentence to a cluster and then rewrites it using the cluster's explanation and a representative example. The paper reports that this method achieves the highest Style Transfer Accuracy (the share of outputs judged non-toxic by the toxicity classifier) among compared approaches and the highest average Joint score, an aggregate of non-toxicity, content preservation, and fluency; it interprets this as evidence that cluster knowledge reduces hallucination and yields more precise edits.
Load-bearing premise
The load-bearing assumption is that the automatic toxicity checker used to score outputs is a fair judge of non-toxicity in all nine languages; the paper's own Table 9 shows Chinese human detoxification references receiving only 0.266 on that check, so if the checker is biased, the reported rankings of methods, including the CoT advantage, would not be trustworthy.
Editorial extensions
If this is right
- The new 400/600 train/test splits for German, Hindi, Amharic, Arabic, and Chinese give subsequent work a standard benchmark for supervised detoxification in languages with no prior parallel corpus.
- The cluster-conditioned CoT prompt can be applied without fine-tuning: given any toxic sentence, the model diagnoses the edit type from the cluster and rewrites accordingly, reducing hallucinated content.
- The Delete baseline's strong showing in Chinese, Arabic, and Amharic implies that for some languages, removing toxic tokens is a competitive fallback when no good paraphrase model exists.
- The descriptive feature analysis provides a reusable map of each language's toxic lexicon, for instance animal insults in Hindi and Amharic and refugee-related wordplay in German, that can inform language-specific moderation and generation.
Reading between the lines
- Editorial inference: because the STA checker disagrees sharply with human detoxification references in some languages (Chinese human references get STA 0.266), the reported Joint-score gaps could be re-ranked by a human-validated toxicity measure; the CoT method's advantage should be treated as pending that check.
- Editorial inference: the describe-and-cluster recipe is not specific to toxicity; applying feature extraction plus k-means over repair strategies to formality transfer or sentiment transfer could produce the same kind of conditioning signal, and existing parallel datasets would allow that test.
- Editorial inference: the cluster explanations in the paper are English-only, so varying the language and phrasing of the cluster instructions, or the number of clusters, is an untested dimension that could matter more in lower-resource languages such as Amharic.
- Editorial inference: a direct human evaluation of fluency and content preservation on the test outputs would settle whether the STA advantage reflects genuinely better detoxifications or merely a shift toward the classifier's notion of non-toxicity.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper extends parallel text detoxification corpora to German, Chinese, Arabic, Hindi, and Amharic, reports a detailed annotation pipeline for each language, and uses GPT-4 to annotate descriptive features of toxic and non-toxic sentences across nine languages. On the basis of this feature analysis, the authors build per-language K-means clusters and propose a Chain-of-Thought (CoT) prompting method that identifies a test sentence's cluster and supplies a representative detoxification example in the prompt. The central empirical claim is that this cluster-conditioned CoT method achieves the best average joint score J = STA x SIM x ChrF1 and the best STA scores against several unsupervised, supervised, and few-shot baselines.
Significance. If the claims hold, the resource contribution is genuinely valuable: manually curated parallel detoxification data for five non-European languages, with public data and code, extends the reach of supervised detoxification substantially. The descriptive feature analysis across nine languages is also a useful reference for cross-lingual work on toxicity. However, the empirical comparison that supports the CoT method is currently not trustworthy. The STA toxicity classifier used in the joint metric is never validated per language, and internal tables show it disagrees sharply with human gold labels for Chinese and Amharic. In addition, the clusters and prompt examples used in the CoT method are derived from the full 1,000-pair set, including the 600 test instances, so the evaluation is not a clean test of generalization. These issues are fixable with additional experiments and reporting, but they are load-bearing for the paper's central claim.
major comments (4)
- [Section 5, Tables 4, 9, 10] The STA component of the joint score is computed by an XLM-R-large classifier fine-tuned on 5,000 subsampled examples per language, but the paper reports no per-language validation, accuracy, F1, or calibration of this classifier. The internal results show that the classifier is not consistently measuring non-toxicity: in Table 9, Chinese human detoxified references receive STA 0.266, meaning most gold non-toxic rewrites are deemed toxic, while GPT-4 CoT on the same language receives STA 0.716; in Table 10, the Duplicate baseline for Amharic receives STA 0.426, so unchanged toxic inputs are called non-toxic nearly half the time. Because J is multiplicative, such a biased STA directly distorts every comparison in Table 4, including the claimed CoT advantage over few-shot prompting. The 98% expert agreement reported in Section 4.1 concerns GPT-4's descriptive feature annotations and not the validity of the STA classifier, so it does not mitigate this concern.
- [Sections 3.6, 4.1, 4.5, Appendix A.3] The CoT method is evaluated on test instances that were used to construct the method. Feature extraction and K-means clustering are run on all 1,000 pairs per language, and the 600 test sentences are a subset of these 1,000 pairs. The cluster definitions, the choice of K=3, and the representative cluster examples embedded in the CoT prompt are therefore informed by the same inputs and references that are later scored in Table 4. This is a form of test-set leakage: the comparison does not measure how the method would perform on unseen inputs. The authors should derive clusters from the 400 training pairs only, or otherwise exclude the test set from all prompt and cluster construction, and then re-run the evaluation.
- [Section 7, Table 4, Appendix C.5] The reported advantage of GPT-4 CoT over few-shot prompting is very small on average (0.331 vs 0.324) and negative for English (0.326 vs 0.475), yet no significance tests, confidence intervals, or multiple-run variance are reported. Appendix C.5 states that GPT-4 was run with default hyperparameters including temperature=1.0, and no number of repeated inferences is given. With a single stochastic run per input, the observed differences may be sampling noise. Paired bootstrap tests or multiple runs with confidence intervals are needed before claiming that CoT 'achieved the highest scores across all approaches'.
- [Section 4.1, Table 7, Appendix E] The feature-extraction analysis relies on GPT-4 annotations for the full 1,000-pair set, and the paper states that experts agreed with GPT-4 in 98% of cases. However, the size of the expert-reviewed sample, the number of annotators per language, and the agreement measure are not reported. Since the method's clusters and prompts are built directly on these annotations, the reliability of this 98% figure matters for the validity of the CoT approach. The authors should specify the validation protocol and report per-language agreement.
minor comments (5)
- [Section 3, Annotators Compensations] The text reports 'C20 per hour' and 'C7.65 above the minimum wage'; the 'C' appears to be a rendering error for the euro sign, which should be corrected.
- [Section 3.3.2, Amharic Annotation Process] 'Two annotators ... were evolved in the main annotation' should read 'were involved in the main annotation'.
- [Section 3.1.2, German Annotation Process] The text says each sample was transcribed by only one annotator, while Table 1 lists two annotators per sentence for German; the apparent discrepancy between the prose and the table should be clarified.
- [Appendix C.5] The phrase 'top_k=0.0' is not a standard GPT-4 API parameter in the same sense as temperature and may confuse readers; the paper should either explain the sampling configuration precisely or omit the unsupported parameter.
- [Section 2, Related Work] The related-work section would benefit from a short paragraph positioning the new corpora against the MultiParaDetox and TextDetox CLEF-2024 shared task, since the data are already described as the basis of that task in footnote 3.
Circularity Check
No significant circularity; the corpus construction, metric computation, and CoT comparison are not equivalent to their inputs by construction.
full rationale
The paper's central contributions—new manually curated parallel detoxification corpora (Sections 3.1–3.5), GPT-4-based descriptive feature analysis with native-speaker validation (Section 4), and the Chain-of-Thoughts prompting method (Section 4.5)—do not reduce, by the paper's own equations or definitions, to their inputs. The corpora are curated from external toxic-language datasets and are publicly released; the STA classifier is fine-tuned on 5,000 subsampled toxicity-classification examples per language explicitly 'not used for ParaDetox data collection' (Section 5), so the J-score comparisons in Tables 4, 8–10 are not fitted from the detoxified outputs being scored. The CoT method uses GPT-4-derived cluster prompts, but the detoxified sentences are generated by GPT-4 and are not computed from the cluster fit; the cluster label is a conditioning input, and the claimed improvement is an empirical comparison against few-shot prompting on the same test set. The self-citations (Dementieva et al. 2024a for the toxicity definition; Logacheva et al. 2022 for the evaluation pipeline; Dementieva et al. 2024b for the shared task) point to published, externally available resources and are not load-bearing. Concerns raised by reviewers—the unvalidated per-language STA signal (e.g., Table 9 gives Chinese human references STA 0.266) and the use of all 1,000 pairs per language, including test pairs, to build the CoT clusters (Sections 4.1 and 4.5)—are correctness, calibration, and potential test-set-leakage risks, not by-construction equivalences between claimed predictions and fitted inputs. The paper itself acknowledges limitations, including reliance on closed-source GPT-4 and English-only cluster explanations. Therefore no significant circularity is present.
Assumptions & free parameters
free parameters (3)
- Number of clusters K =
3 per language
- Chinese toxic-score filtering threshold =
0.978
- Chinese curation thresholds =
1 to 5 toxic words, toxic word ratio < 0.5, length 3 to 50 words
assumptions (3)
- domain assumption GPT-4 feature annotations accurately characterize toxicity and detoxification in all nine languages.
- domain assumption The STA classifier trained on 5,000 subsampled examples per language reliably measures style transfer accuracy.
- domain assumption Toxicity is defined as vulgar or profane language excluding deep insults and hate, following Dementieva et al. (2024a).
Cite this review
Pith. "Pith review of Multilingual and Explainable Text Detoxification with Parallel Corpora." pith.science (2026). https://pith.science/paper/2CIZYQW5
@misc{pith2026241211691,
author = {Pith},
title = {Pith review of: Multilingual and Explainable Text Detoxification with Parallel Corpora},
year = {2026},
howpublished = {\url{https://pith.science/paper/2CIZYQW5}},
note = {Machine review of arXiv:2412.11691}
}
read the original abstract
Even with various regulations in place across countries and social media platforms (Government of India, 2021; European Parliament and Council of the European Union, 2022, digital abusive speech remains a significant issue. One potential approach to address this challenge is automatic text detoxification, a text style transfer (TST) approach that transforms toxic language into a more neutral or non-toxic form. To date, the availability of parallel corpora for the text detoxification task (Logachevavet al., 2022; Atwell et al., 2022; Dementievavet al., 2024a) has proven to be crucial for state-of-the-art approaches. With this work, we extend parallel text detoxification corpus to new languages -- German, Chinese, Arabic, Hindi, and Amharic -- testing in the extensive multilingual setup TST baselines. Next, we conduct the first of its kind an automated, explainable analysis of the descriptive features of both toxic and non-toxic sentences, diving deeply into the nuances, similarities, and differences of toxicity and detoxification across 9 languages. Finally, based on the obtained insights, we experiment with a novel text detoxification method inspired by the Chain-of-Thoughts reasoning approach, enhancing the prompting process through clustering on relevant descriptive attributes.
Figures
Figures from the paper (11 more)
Reference graph
Works this paper leans on
-
[1]
Katherine Atwell, Sabit Hassan, and Malihe Alikhani. 2022. https://aclanthology.org/2022.coling-1.530 APPDIA: A discourse-aware transformer-based style transfer model for offensive social media conversations . In Proceedings of the 29th International Conference on Computational Linguistics, COLING 2022, Gyeongju, Republic of Korea, October 12-17, 2022 , p...
work page 2022
-
[2]
Abinew Ali Ayele, Skadi Dinter, Tadesse Destaw Belay, Tesfa Tegegne Asfaw, Seid Muhie Yimam, and Chris Biemann. 2022. https://ieeexplore.ieee.org/document/9971189 The 5Js in Ethiopia: Amharic hate speech data annotation using Toloka Crowdsourcing Platform . In Proceedings of the 4th International Conference on Information and Communication Technology for ...
-
[3]
Abinew Ali Ayele, Seid Muhie Yimam, Tadesse Destaw Belay, Tesfa Asfaw, and Chris Biemann. 2023. https://aclanthology.org/2023.ranlp-1.6 Exploring A mharic hate speech data collection and classification approaches . In Proceedings of the 14th International Conference on Recent Advances in Natural Language Processing, pages 49--59, Varna, Bulgaria. INCOMA L...
work page 2023
-
[4]
Anatoly Belchikov. 2019. Russian language toxic comments. https://www.kaggle.com/blackmoon/russian-language-toxic-comments https://www.kaggle.com/blackmoon/russian-language-toxic-comments. Accessed: 2023-12-14
work page 2019
-
[5]
Kateryna Bobrovnyk. 2019 a . https://ena.lpnu.ua:8443/server/api/core/bitstreams/c4c645c1-f465-4895-98dd-765f862cf186/content Automated building and analysis of ukrainian twitter corpus for toxic text detection . In COLINS 2019. Volume II: Workshop
work page 2019
-
[6]
Kateryna Bobrovnyk. 2019 b . The dictionary of ukrainian obscene words. https://github.com/saganoren/obscene-ukr https://github.com/saganoren/obscene-ukr. Accessed: 2024-12-12
work page 2019
-
[7]
Aditya Bohra, Deepanshu Vijay, Vinay Singh, Syed Sarfaraz Akhtar, and Manish Shrivastava. 2018. https://doi.org/10.18653/v1/W18-1105 A dataset of H indi- E nglish code-mixed social media text for hate speech detection . In Proceedings of the Second Workshop on Computational Modeling of People ' s Opinions, Personality, and Emotions in Social Media , pages...
-
[8]
Eleftheria Briakou, Di Lu, Ke Zhang, and Joel Tetreault. 2021. https://doi.org/10.18653/v1/2021.naacl-main.256 Ol \'a , bonjour, salve! XFORMAL : A benchmark for multilingual formality style transfer . In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 31...
Show all 78 references
-
[9]
Keith Carlson, Allen Riddell, and Daniel Rockmore. 2018. https://royalsocietypublishing.org/doi/10.1098/rsos.171920 Evaluating prose style transfer with the bible . Royal Society open science, 5(10):171920
2018 doi
-
[10]
Jennifer Cobbe. 2021. Algorithmic censorship by social platforms: Power and resistance. Philosophy & Technology, 34(4):739--766
2021
-
[11]
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzm \' a n, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. https://doi.org/10.18653/V1/2020.ACL-MAIN.747 Unsupervised cross-lingual representation learning...
2020 doi
-
[12]
Costa - juss \` a , James Cross, Onur C elebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Y
Marta R. Costa - juss \` a , James Cross, Onur C elebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Y. Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Lo \" c Barrault, Gabriel Mejia Gonzalez, Pr...
-
[13]
Costa - juss \` a , Mariano Coria Meglioli, Pierre Andrews, David Dale, Prangthip Hansanti, Elahe Kalbassi, Alexandre Mourachko, Christophe Ropers, and Carleigh Wood
Marta R. Costa - juss \` a , Mariano Coria Meglioli, Pierre Andrews, David Dale, Prangthip Hansanti, Elahe Kalbassi, Alexandre Mourachko, Christophe Ropers, and Carleigh Wood. 2024. https://doi.org/10.48550/ARXIV.2401.05060 Mutox: Universal multilingual audio-based toxicity da...
-
[14]
David Dale, Anton Voronov, Daryna Dementieva, Varvara Logacheva, Olga Kozlova, Nikita Semenov, and Alexander Panchenko. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.629 Text detoxification using large pre-trained neural models . In Proceedings of the 2021 Conference on Em...
2021 doi
-
[15]
Daryna Dementieva, Nikolay Babakov, and Alexander Panchenko. 2024 a . https://doi.org/10.18653/v1/2024.naacl-short.12 M ulti P ara D etox: Extending text detoxification with parallel data to new languages . In Proceedings of the 2024 Conference of the North American Chapter of...
2024 doi
-
[16]
Krotova, Nikita Semenov, Tatiana Shavrina, and Alexander Panchenko
Daryna Dementieva, Varvara Logacheva, Irina Nikishina, Alena Fenogenova, David Dale, I. Krotova, Nikita Semenov, Tatiana Shavrina, and Alexander Panchenko. 2022. https://api.semanticscholar.org/CorpusID:253169495 RUSSE-2022: Findings of the First Russian Detoxification Shared ...
2022
-
[17]
Daryna Dementieva, Daniil Moskovskiy, Nikolay Babakov, Abinew Ali Ayele, Naquee Rizwan, Frolian Schneider, Xintog Wang, Seid Muhie Yimam, Dmitry Ustalov, Elisei Stakovskii, Alisa Smirnova, Ashraf Elnagar, Animesh Mukherjee, and Alexander Panchenko. 2024 b . https://ceur-ws.org...
2024
-
[18]
Daryna Dementieva, Daniil Moskovskiy, David Dale, and Alexander Panchenko. 2023. https://doi.org/10.18653/v1/2023.ijcnlp-main.70 Exploring methods for cross-lingual text style transfer: The case of text detoxification . In Proceedings of the 13th International Joint Conference...
2023 doi
-
[19]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423 BERT : Pre-training of deep bidirectional transformers for language understanding . In Proceedings of the 2019 Conference of the North A merican Chapter of the Associat...
2019 doi
-
[20]
European Parliament and Council of the European Union . 2022. https://eur-lex.europa.eu/eli/reg/2022/2065/oj Regulation (eu) 2022/2065 of the european parliament and of the council of 19 october 2022 on a single market for digital services (digital services act) and amending d...
2022
-
[21]
Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. 2022. https://doi.org/10.18653/V1/2022.ACL-LONG.62 Language-agnostic BERT sentence embedding . In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long...
2022 doi
-
[22]
Griffin Floto, Mohammad Mahdi Abdollah Pour, Parsa Farinneya, Zhenwei Tang, Ali Pesaranghader, Manasa Bharadwaj, and Scott Sanner. 2023. https://doi.org/10.18653/v1/2023.findings-acl.478 D iffu D etox: A mixed diffusion model for text detoxification . In Findings of the Associ...
2023 doi
-
[23]
Robert James Gabriel. 2023. English full list of bad words and top swear words banned by google. https://github.com/coffee-and-fun/google-profanity-words/blob/main/data/en.txt https://github.com/coffee-and-fun/google-profanity-words/blob/main/data/en.txt. Accessed: 2024-12-12
2023
-
[24]
Gongane, Mousami V
Vaishali U. Gongane, Mousami V. Munot, and Alwin D. Anuse. 2024. https://doi.org/10.1007/S42001-024-00248-9 A survey of explainable AI techniques for detection of fake news and hate speech on social media platforms . J. Comput. Soc. Sci., 7(1):587--623
2024 doi
-
[25]
Government of India . 2021. https://www.meity.gov.in/writereaddata/files/Intermediary_Guidelines_and_Digital_Media_Ethics_Code_Rules-2021.pdf Information technology (intermediary guidelines and digital media ethics code) rules, 2021 . Ministry of Electronics and Information Te...
2021
-
[26]
Hatem Haddad, Hala Mulki, and Asma Oueslati. 2019. https://doi.org/10.1007/978-3-030-32959-4\_18 T-HSAB: A tunisian hate speech and abusive dataset . In Arabic Language Processing: From Theory to Practice - 7th International Conference, ICALP 2019, Nancy, France, October 16-17...
2019 doi
-
[27]
Skyler Hallinan, Alisa Liu, Yejin Choi, and Maarten Sap. 2023. https://doi.org/10.18653/v1/2023.acl-short.21 Detoxifying text with M a RC o: Controllable revision with experts and anti-experts . In Proceedings of the 61st Annual Meeting of the Association for Computational Lin...
2023 doi
-
[28]
Nhat Hoang, Xuan Long Do, Duc Anh Do, Duc Anh Vu, and Anh Tuan Luu. 2024. https://doi.org/10.18653/v1/2024.naacl-long.359 T o XCL : A unified framework for toxic speech detection and explanation . In Proceedings of the 2024 Conference of the North American Chapter of the Assoc...
2024 doi
-
[29]
Zachary Horvitz, Ajay Patel, Chris Callison - Burch, Zhou Yu, and Kathleen R. McKeown. 2024. https://doi.org/10.1609/AAAI.V38I16.29780 Paraguide: Guided diffusion paraphrasers for plug-and-play textual style transfer . In Thirty-Eighth AAAI Conference on Artificial Intelligenc...
2024 doi
-
[30]
Imbwaga, Nagaratna B
Joan L. Imbwaga, Nagaratna B. Chittaragi, and Shashidhar G. Koolagudi. 2024. https://doi.org/10.1007/S10772-024-10135-3 Explainable hate speech detection using LIME . Int. J. Speech Technol., 27(3):793--815
2024 doi
-
[31]
Aiqi Jiang, Xiaohan Yang, Yang Liu, and Arkaitz Zubiaga. 2022. https://doi.org/10.1016/J.OSNEM.2021.100182 SWSR: A chinese dataset and lexicon for online sexism detection . Online Soc. Networks Media, 27:100182
2022
-
[32]
Jigsaw. 2017. Toxic comment classification challenge. https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge. Accessed: 2024-03-18
2017
-
[33]
Di Jin, Zhijing Jin, Zhiting Hu, Olga Vechtomova, and Rada Mihalcea. 2022. https://doi.org/10.1162/coli_a_00426 Deep learning for text style transfer: A survey . Computational Linguistics, 48(1):155--205
2022 doi
- [34]
-
[35]
Enes Kulenovi \'c . 2023. https://link.springer.com/article/10.1007/s10677-022-10336-2 Should democracies ban hate speech? hate speech laws and counterspeech . Ethical Theory and Moral Practice, 26(4):511--532
2023 doi
-
[36]
Teyun Kwon and Anandha Gopalan. 2021. https://arxiv.org/abs/2112.00819 CO-STAR: conceptualisation of stereotypes for analysis and reasoning . CoRR, abs/2112.00819
2021 arXiv
-
[37]
Chak Tou Leong, Yi Cheng, Jiashuo Wang, Jian Wang, and Wenjie Li. 2023. https://doi.org/10.18653/V1/2023.EMNLP-MAIN.269 Self-detoxifying language models via toxification reversal . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP...
2023 doi
-
[38]
Juncen Li, Robin Jia, He He, and Percy Liang. 2018. https://doi.org/10.18653/V1/N18-1169 Delete, retrieve, generate: a simple approach to sentiment and style transfer . In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Lin...
2018 doi
-
[39]
Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020. https://doi.org/10.1162/TACL\_A\_00343 Multilingual denoising pre-training for neural machine translation . Trans. Assoc. Comput. Linguistics, 8:726--742
2020 doi
-
[40]
Varvara Logacheva, Daryna Dementieva, Sergey Ustyantsev, Daniil Moskovskiy, David Dale, Irina Krotova, Nikita Semenov, and Alexander Panchenko. 2022. https://doi.org/10.18653/v1/2022.acl-long.469 P ara D etox: Detoxification with parallel data . In Proceedings of the 60th Annu...
2022 doi
-
[41]
Junyu Lu, Bo Xu, Xiaokun Zhang, Changrong Min, Liang Yang, and Hongfei Lin. 2023. https://aclanthology.org/2023.acl-long.898 Facilitating fine-grained detection of C hinese toxic language: Hierarchical taxonomy, resources, and benchmarks . In Proceedings of the 61st Annual Mee...
2023
-
[42]
Lundberg and Su - In Lee
Scott M. Lundberg and Su - In Lee. 2017. https://proceedings.neurips.cc/paper/2017/hash/8a20a8621978632d76c43dfd28b67767-Abstract.html A unified approach to interpreting model predictions . In Advances in Neural Information Processing Systems 30: Annual Conference on Neural In...
2017
-
[43]
Thomas Mandl, Sandip Modha, Prasenjit Majumder, Daksh Patel, Mohana Dave, Chintak Mandlia, and Aditya Patel. 2019. https://doi.org/10.1145/3368567.3368584 Overview of the hasoc track at fire 2019: Hate speech and offensive content identification in indo-european languages . In...
2019
-
[44]
Binny Mathew, Punyajoy Saha, Seid Muhie Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee. 2021. https://doi.org/10.1609/AAAI.V35I17.17745 Hatexplain: A benchmark dataset for explainable hate speech detection . In Thirty-Fifth AAAI Conference on Artificial Intelligence,...
2021 doi
-
[45]
Jos \' e Mar \' a Molero, Jorge P \' e rez - Mart \' n, \' A lvaro Rodrigo, and Anselmo Pe \ n as. 2023. https://doi.org/10.1109/ACCESS.2023.3310244 Offensive language detection in spanish social media: Testing from bag-of-words to transformers models . IEEE Access , 11:95639--95652
2023
-
[46]
Edoardo Mosca, Daryna Dementieva, Tohid Ebrahim Ajdari, Maximilian Kummeth, Kirill Gringauz, Yutong Zhou, and Georg Groh. 2023. https://doi.org/10.18653/v1/2023.ijcnlp-demo.7 IFAN : An explainability-focused interaction framework for humans and NLP models . In Proceedings of t...
2023 doi
-
[47]
Daniil Moskovskiy, Daryna Dementieva, and Alexander Panchenko. 2022. https://doi.org/10.18653/v1/2022.acl-srw.26 Exploring cross-lingual text detoxification with large multilingual language models. In Proceedings of the 60th Annual Meeting of the Association for Computational ...
2022 doi
-
[48]
Hamdy Mubarak, Kareem Darwish, Walid Magdy, Tamer Elsayed, and Hend Al-Khalifa. 2020. https://aclanthology.org/2020.osact-1.7 Overview of OSACT 4 A rabic offensive language detection shared task . In Proceedings of the 4th Workshop on Open-Source Arabic Corpora and Processing ...
2020
-
[49]
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M. Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward...
2023 doi
-
[50]
Ojha, and Ond r ej Du s ek
Sourabrata Mukherjee, Akanksha Bansal, Pritha Majumdar, Atul Kr. Ojha, and Ond r ej Du s ek. 2023. https://doi.org/10.18653/v1/2023.banglalp-1.5 Low-resource text style transfer for B angla: Data & models . In Proceedings of the First Workshop on Bangla Language Processing (BL...
2023 doi
-
[51]
Ojha, Akanksha Bansal, Deepak Alok, John P
Sourabrata Mukherjee, Atul Kr. Ojha, Akanksha Bansal, Deepak Alok, John P. McCrae, and Ondrej Dusek. 2024 a . https://doi.org/10.48550/ARXIV.2405.20805 Multilingual text style transfer: Datasets & models for indian languages . CoRR, abs/2405.20805
- [52]
-
[53]
Hala Mulki and Bilal Ghanem. 2021. https://aclanthology.org/2021.wanlp-1.16 Let-mi: An A rabic L evantine T witter dataset for misogynistic language . In Proceedings of the Sixth Arabic Natural Language Processing Workshop, pages 154--163, Kyiv, Ukraine (Virtual). Association ...
2021
-
[54]
Hala Mulki, Hatem Haddad, Chedi Bechikh Ali, and Halima Alshabani. 2019. https://doi.org/10.18653/v1/W19-3512 L - HSAB : A L evantine T witter dataset for hate speech and abusive language . In Proceedings of the Third Workshop on Abusive Language Online, pages 111--118, Floren...
2019 doi
-
[55]
Cicero Nogueira dos Santos, Igor Melnyk, and Inkit Padhi. 2018. https://doi.org/10.18653/v1/P18-2031 Fighting offensive language on social media with unsupervised text style transfer . In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (...
2018 doi
-
[56]
OpenAI. 2022. https://openai.com/blog/chatgpt Chatgpt: Optimizing language models for dialogue . Accessed: 2024-05-31
2022
-
[57]
Juan Carlos Pereira - Kohatsu, Lara Quijano S \' a nchez, Federico Liberatore, and Miguel Camacho - Collados. 2019. https://doi.org/10.3390/S19214654 Detecting and monitoring hate speech in twitter . Sensors, 19(21):4654
2019 doi
-
[58]
Juan Manuel P \'e rez, Dami \'a n Ariel Furman, Laura Alonso Alemany, and Franco M. Luque. 2022. https://aclanthology.org/2022.lrec-1.785 R o BERT uito: a pre-trained language model for social media text in S panish . In Proceedings of the Thirteenth Language Resources and Eva...
2022
-
[59]
Matt Post. 2018. https://doi.org/10.18653/V1/W18-6319 A call for clarity in reporting BLEU scores . In Proceedings of the Third Conference on Machine Translation: Research Papers, WMT 2018, Belgium, Brussels, October 31 - November 1, 2018 , pages 186--191. Association for Comp...
2018 doi
-
[60]
Shuofei Qiao, Yixin Ou, Ningyu Zhang, Xiang Chen, Yunzhi Yao, Shumin Deng, Chuanqi Tan, Fei Huang, and Huajun Chen. 2023. https://doi.org/10.18653/v1/2023.acl-long.294 Reasoning with language model prompting: A survey . In Proceedings of the 61st Annual Meeting of the Associat...
2023 doi
-
[61]
Sudha Rao and Joel Tetreault. 2018. https://doi.org/10.18653/v1/N18-1012 Dear sir or madam, may I introduce the GYAFC dataset: Corpus, benchmarks and metrics for formality style transfer . In Proceedings of the 2018 Conference of the North A merican Chapter of the Association ...
2018 doi
-
[62]
why should I trust you?
Marco T \' u lio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016. https://doi.org/10.1145/2939672.2939778 "why should I trust you?": Explaining the predictions of any classifier . In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data M...
2016
-
[63]
Julian Risch, Anke Stoll, Lena Wilms, and Michael Wiegand. 2021. https://aclanthology.org/2021.germeval-1.1 Overview of the germeval 2021 shared task on the identification of toxic, engaging, and fact-claiming comments . In Proceedings of the GermEval 2021 Shared Task on the I...
2021
-
[64]
Bj \"o rn Ross, Michael Rist, Guillermo Carbonell, Benjamin Cabrera, Nils Kurowsky, and Michael Wojatzki. 2016. https://d-nb.info/1119886848/34#page=12 Measuring the Reliability of Hate Speech Annotations: The Case of the European Refugee Crisis . In Proceedings of NLP4CMC III...
2016
-
[65]
Sarthak Roy, Ashish Harshvardhan, Animesh Mukherjee, and Punyajoy Saha. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.407 Probing LLM s for hate speech detection: strengths and vulnerabilities . In Findings of the Association for Computational Linguistics: EMNLP 2023, ...
2023 doi
-
[66]
Aleksandr Semiletov. 2020. Toxic Russian Comments: Labelled comments from the popular Russian social network . https://www.kaggle.com/alexandersemiletov/toxic-russian-comments https://www.kaggle.com/alexandersemiletov/toxic-russian-comments. Accessed: 2023-12-14
2020
-
[67]
Zheyuan Ryan Shi, Claire Wang, and Fei Fang. 2020. https://arxiv.org/abs/2001.01818 Artificial intelligence for social good: A survey . CoRR, abs/2001.01818
2020 arXiv
-
[68]
Inc Shutterstock. 2020. List of dirty, naughty, obscene, and otherwise bad words. https://github.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words https://github.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words. Accessed: 2024-12-12
2020
- [69]
-
[70]
Yuqing Tang, Chau Tran, Xian Li, Peng - Jen Chen, Naman Goyal, Vishrav Chaudhary, Jiatao Gu, and Angela Fan. 2020. https://arxiv.org/abs/2008.00401 Multilingual translation with extensible multilingual pretraining and finetuning . CoRR, abs/2008.00401
2020 arXiv
-
[71]
Mariona Taul \' e , Montserrat Nofre, V \' ctor Bargiela, and Xavier Bonet Casals. 2024. https://doi.org/10.1007/S10579-023-09711-X Newscom-tox: a corpus of comments on news articles annotated for toxicity in spanish . Lang. Resour. Evaluation, 58(4):1115--1155
2024 doi
-
[72]
Rachel Ung. 2023. https://waseda.repo.nii.ac.jp/record/2000931/files/t5121FG17.pdf Formality Style Transfer between Japanese and English . Ph.D. thesis, Waseda University
2023
-
[73]
Shang Wang, Tianqing Zhu, Bo Liu, Ming Ding, Xu Guo, Dayong Ye, Wanlei Zhou, and Philip S. Yu. 2024. https://doi.org/10.48550/ARXIV.2406.07973 Unique security and privacy threats of large language model: A comprehensive survey . CoRR, abs/2406.07973
2024 doi
-
[74]
Michael Wiegand, Melanie Siegel, and Josef Ruppenhofer. 2018. https://epub.oeaw.ac.at/?arp=0x003a10d2 Overview of the GermEval 2018 Shared Task on the Identification of Offensive Language . In Proceedings of GermEval 2018, 14th Conference on Natural Language Processing (KONVEN...
2018
- [75]
-
[76]
Chiyu Zhang, Honglong Cai, Yuezhang Li, Yuexin Wu, Le Hou, and Muhammad Abdul-Mageed. 2024. https://doi.org/10.18653/v1/2024.naacl-srw.21 Distilling text style transfer with self-explanation from LLM s . In Proceedings of the 2024 Conference of the North American Chapter of th...
2024 doi
-
[77]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[78]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.