REVIEW 3 major objections 2 minor 6 cited by
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
T0 review · 3 major / 2 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read This submission's abstract promises an LLM-safety survey, but the full text is a theorem about iterates of $x^d+c$.
desk verdict The abstract promises an LLM safety survey, but the supplied full text is a number-theory paper by a different author; nothing here is reviewable as submitted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the unicritical polynomial family $f_{d,c}(x)=x^d+c$ and its iterates $f^n_{d,c}(x)$; the claim is that $\tau(d)$, the divisor-counting function of the degree $d$, bounds the number of irreducible factors of $f^n(x)-\alpha$ for all $n$. The proof mechanism is the abc-conjecture together with height conditions ($h(c)>0$, $h(c-\alpha)>0$) that keep $\alpha$ away from periodic-point behavior, the known source of unbounded factor growth. Height theory converts the arithmetic of the iterates into factor-count control, which then feeds the orbit-density and integral-point applications.
What would settle it
Open the full text and search for the topics advertised in the abstract: if 'toxicity', 'jailbreak', 'RLHF', or 'safety alignment' never appear and the body instead proves a factorization theorem about $x^d+c$, the survey claim is settled as unsupported by this artifact. For the embedded theorem, compute the number of irreducible factors of $f^n(x)-\alpha$ for any admissible small pair $(d,c)$ with $n=2$: a count above $\tau(d)$ would falsify Theorem 1.1.
Extended reading notes
Core claim
The mathematical body proves Theorem 1.1: for a number field $K$ over which the abc-conjecture holds, fixed $\alpha\in K$, and $f(x)=x^d+c$ with $d\ge2$, $c\in K$, if $\varphi(d)$ exceeds a constant fraction of $d$, every prime divisor of $d$ exceeds a constant depending on $\alpha$ and $K$, $\alpha$ is not a fixed point of $f$, and both $h(c)$ and $h(c-\alpha)$ are positive, then $f^n(x)-\alpha$ has at most $\tau(d)$ irreducible factors in $K[x]$ for all $n\ge1$. The author presents this as the first eventual-stability result for a large class of unicritical polynomials of non-prime-powered degree with nonzero basepoint, and draws consequences for the density of prime divisors in forward or
Load-bearing premise
The load-bearing premise for the advertised survey is that the attached body is the survey announced in the title and abstract; separately, the embedded mathematics stands on the abc-conjecture, and if that conjecture fails the theorem's conclusion is unsupported.
Editorial extensions
If this is right
- For a positive-density set of degrees $d$---including all sufficiently large prime powers---the iterates $f^n(x)$ have at most $\tau(d)$ factors, so the pair $(f,\alpha)$ is eventually stable.
- The factorization bound yields a controlled density of prime divisors in forward orbits for many unicritical polynomials, conditional on the abc-conjecture.
- The finiteness of integral points in backward orbits follows in the covered cases.
- The result extends known eventual-stability results beyond prime-powered degree, a step that prior literature had left open.
Reading between the lines
- Editorial inference: the most plausible explanation for the mismatch is an upload error---the wrong full text was attached to the survey's front matter; nothing in the supplied body supports the abstract's claims about toxicity, jailbreaks, or RLHF.
- Editorial inference: a reader who wants the survey should treat this submission as evidence of a pipeline failure, not as an instance of LLM harm or a taxonomy of defenses.
- Editorial inference: for the embedded mathematics, a testable next step is to compute the exact factor counts for small admissible pairs $(d,c)$ to see whether the $\tau(d)$ bound is sharp or merely an upper bound.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The submission is presented as a survey of LLM safety, titled 'Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM'. The abstract promises a systematic review of unintentional toxicity, adversarial jailbreak attacks, and mitigation strategies, and proposes a unified taxonomy of LLM-related harms and defenses. However, the supplied full text is the opening page of an unrelated mathematics paper, 'On the Factorization of Iterates of x^d + c in Large Degree' by Wade Hindes (arXiv:2508.05795v1, math.NT). That text states Theorem 1.1 about factorization of iterates conditional on the abc-conjecture and contains no LLM safety content whatsoever.
Significance. If the survey described in the abstract existed as claimed, it could be a useful systematic review with a unified taxonomy, potentially valuable to the LLM safety community. However, the submitted manuscript does not contain that survey. The paper's central claim—that it systematically reviews recent studies and proposes a taxonomy—is entirely unsupported by the provided body, which is a different paper in a different field. The artifact is internally inconsistent, and the claims in the abstract are unverifiable against the text. Because the core content is absent, the potential significance cannot be assessed.
major comments (3)
- [Abstract vs. Full Text] The abstract states: 'In this survey, we systematically review recent studies encompassing unintentional toxicity, adversarial jailbreak attacks, and comprehensive mitigation strategies' and 'we propose a unified taxonomy of LLM-related harms and defenses.' The full text, however, is the opening page of an arithmetic dynamics paper by a different author with a different title and subject classification. No survey content, taxonomy, RLHF discussion, or jailbreak analysis appears anywhere in the provided body. This is a load-bearing internal inconsistency: the central claim of the paper is unsupported by the manuscript text.
- [Theorem 1.1] The only substantive mathematical content in the provided text is Theorem 1.1, concerning the number of factors of iterates of x^d + c over number fields conditional on the abc-conjecture. This theorem is unrelated to the announced topic of LLM safety. If the body is the intended content, then the title and abstract misdescribe it; if the surveyed content is genuinely absent, the abstract's claims are unverifiable. Either way, the submission cannot be evaluated as a coherent paper.
- [Missing proof for Theorem 1.1] Even taking the body on its own terms, the displayed theorem is stated without proof. The introduction references prior work but provides no argument for the claimed result. The paper does not deliver the comprehensive review announced in the abstract, nor does it provide a self-contained mathematical contribution. This further undermines the manuscript's claims.
minor comments (2)
- [Bibliographic metadata] The submission's title and author do not match the full text's title ('On the Factorization of Iterates of x^d + c in Large Degree') and author (Wade Hindes). The arXiv identifier embedded in the text (2508.05795v1, math.NT) also differs from the cover identifier of this submission. This metadata mismatch should be resolved before any resubmission.
- [Internal consistency] The abstract uses terminology such as 'toxicity', 'jailbreak attacks', and 'RLHF' that does not appear anywhere in the full text. The body uses mathematical notation (e.g., φ(·), τ(·), h(·)) with no connection to the abstract. Ensuring that the title, abstract, and body describe the same work is a basic prerequisite for review.
Circularity Check
No circular derivation found; abstract/body mismatch is a verifiability issue, not circularity.
full rationale
The mathematical portion of the manuscript proves Theorem 1.1 conditionally on the abc-conjecture. Its inputs are explicit hypotheses (conditions (1)-(4) on d, c, and alpha) and previously cited results; the conclusion about factor counts of iterates follows from those stated assumptions. There is no fitted parameter relabeled as a prediction, no target quantity defined in terms of its own output, and no author-overlapping uniqueness theorem invoked to force the choice of framework. The abstract promises a survey of LLM safety with a unified taxonomy, while the body is an arithmetic-dynamics paper on factorization of iterates of x^d + c. That mismatch makes the abstract's central claim unverifiable against the supplied text, and it is a serious artifact-level inconsistency, but it is not a circular step: the body's derivation chain does not reduce to its own inputs. Accordingly, no circularity is present and the score is 0.
Assumptions & free parameters
assumptions (2)
- ad hoc to paper The body of this submission is the survey described by the title and abstract
- domain assumption The abc-conjecture holds over number fields
Cite this review
Pith. "Pith review of Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM." pith.science (2026). https://pith.science/paper/DN4LQ5YA
@misc{pith2026250805775,
author = {Pith},
title = {Pith review of: Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM},
year = {2026},
howpublished = {\url{https://pith.science/paper/DN4LQ5YA}},
note = {Machine review of arXiv:2508.05775}
}
read the original abstract
Large Language Models (LLMs) have revolutionized content creation across digital platforms, offering unprecedented capabilities in natural language generation and understanding. Meanwhile, they pose risks by inadvertently producing toxic, offensive, or biased content. This dual role of LLMs, both as powerful tools for text generation and as potential sources of harmful language, presents a pressing sociotechnical challenge. In this survey, we systematically review recent studies encompassing unintentional toxicity, adversarial jailbreak attacks, and comprehensive mitigation strategies. We explore LLMs' dual role as both generators of harm and enablers of safety through detection, classification, content moderation, and prevention. We propose a unified taxonomy of LLM-related harms and defenses, analyze emerging multimodal and LLM-assisted jailbreak strategies, and assess mitigation efforts, including reinforcement learning with human feedback (RLHF), prompt engineering, and safety alignment. Our review highlights the evolving landscape of LLM safety and identifies limitations in current evaluation methodologies. Ultimately, our review outlines future research directions to guide the development of robust and ethically aligned language technologies.
Forward citations
Cited by 6 Pith papers
-
Do Coding Agents Understand Least-Privilege Authorization?
Coding agents struggle to infer least-privilege file permissions by omitting needed accesses while granting unused or sensitive ones, but Sufficiency-Tightness Decomposition improves sensitive-task success by up to 15...
-
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue
TurnGate uses a new multi-turn intent dataset to detect the harm-enabling closure point in dialogues, outperforming baselines with low over-refusal and generalizing across domains.
-
One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue
TurnGate identifies the critical turn in multi-turn dialogues where a response would complete hidden malicious intent, outperforming baselines on the new MTID dataset while keeping over-refusal low.
-
Why Do Large Language Models Generate Harmful Content?
Causal mediation analysis shows harmful LLM outputs arise in late layers from MLP failures and gating neurons, with early layers handling harm context detection and signal propagation.
-
BarrierSteer: LLM Safety via Learning Barrier Steering
BarrierSteer applies control barrier functions to LLM latent states for constraint-guided steering that reduces unsafe generations while preserving utility.
-
Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows
A survey of 49 LLM fraud and trust-and-safety papers finds that fraud work reports almost no per-decision latency, cost, or calibration evidence, while moderation work reports more.
Reference graph
Works this paper leans on
-
[1]
arXiv preprint arXiv:2204.05862 , year=
Training a helpful and harmless assistant with reinforcement learning from human feedback , author=. arXiv preprint arXiv:2204.05862 , year=
-
[2]
arXiv preprint arXiv:2309.02705 , year=
Certifying llm safety against adversarial prompting , author=. arXiv preprint arXiv:2309.02705 , year=
-
[3]
and Arora, Simran and Mazeika, Manias and Hendrycks, Dan and Lin, Zinan and Cheng, Yu and Koyejo, Sanmi and Song, Dawn and Li, Bo , title =
Wang, Boxin and Chen, Weixin and Pei, Hengzhi and Xie, Chulin and Kang, Mintong and Zhang, Chenhui and Xu, Chejian and Xiong, Zidi and Dutta, Ritik and Schaeffer, Rylan and Truong, Sang T. and Arora, Simran and Mazeika, Manias and Hendrycks, Dan and Lin, Zinan and Cheng, Yu and Koyejo, Sanmi and Song, Dawn and Li, Bo , title =. Proceedings of the 37th Int...
2023
-
[4]
arXiv preprint arXiv:2306.11507 , year=
Trustgpt: A benchmark for trustworthy and responsible large language models , author=. arXiv preprint arXiv:2306.11507 , year=
-
[5]
Discourse & Society , volume=
ChatGPT-4 as a journalist: Whose perspectives is it reproducing? , author=. Discourse & Society , volume=. 2024 , publisher=
2024
-
[6]
Advances in Neural Information Processing Systems , volume=
Are aligned neural networks adversarially aligned? , author=. Advances in Neural Information Processing Systems , volume=
-
[7]
Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and Defenses , pages=
Understanding multi-turn toxic behaviors in open-domain chatbots , author=. Proceedings of the 26th International Symposium on Research in Attacks, Intrusions and Defenses , pages=
-
[8]
Findings of the association for computational linguistics: EMNLP 2023 , pages=
Toxicity in chatgpt: Analyzing persona-assigned language models , author=. Findings of the association for computational linguistics: EMNLP 2023 , pages=
2023
Show all 251 references
-
[9]
IEEE Transactions on Cognitive and Developmental Systems , year=
The inadequacy of reinforcement learning from human feedback-radicalizing large language models via semantic vulnerabilities , author=. IEEE Transactions on Cognitive and Developmental Systems , year=
-
[10]
Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security , pages=
Why so toxic? measuring and triggering toxic behavior in open-domain chatbots , author=. Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security , pages=
2022
-
[11]
Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=
Unveiling the implicit toxicity in large language models , author=. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=
2023
-
[12]
arXiv preprint arXiv:2402.13926 , year=
Large Language Models are Vulnerable to Bait-and-Switch Attacks for Generating Harmful Content , author=. arXiv preprint arXiv:2402.13926 , year=
-
[13]
arXiv preprint arXiv:2410.11459 , year=
Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models , author=. arXiv preprint arXiv:2410.11459 , year=
-
[14]
Security and Communication Networks , volume=
Adversarial Attacks on Large Language Model-Based System and Mitigating Strategies: A Case Study on ChatGPT , author=. Security and Communication Networks , volume=. 2023 , publisher=
2023
-
[15]
Proceedings of the First Workshop on Social Influence in Conversations (SICon 2023) , pages=
Detoxifying online discourse: A guided response generation approach for reducing toxicity in user-generated text , author=. Proceedings of the First Workshop on Social Influence in Conversations (SICon 2023) , pages=
2023
-
[16]
arXiv preprint arXiv:2311.08487 , year=
Alignment is not sufficient to prevent large language models from generating harmful information: A psychoanalytic perspective , author=. arXiv preprint arXiv:2311.08487 , year=
-
[17]
arXiv preprint arXiv:2302.14003 , year=
Systematic rectification of language models via dead-end analysis , author=. arXiv preprint arXiv:2302.14003 , year=
-
[18]
arXiv preprint arXiv:2311.04921 , year=
Successor Features for Efficient Multisubject Controlled Text Generation , author=. arXiv preprint arXiv:2311.04921 , year=
-
[19]
arXiv preprint arXiv:2405.19299 , year=
Expert-Guided Extinction of Toxic Tokens for Debiased Generation , author=. arXiv preprint arXiv:2405.19299 , year=
-
[20]
arXiv preprint arXiv:2405.12900 , year=
Adversarial DPO: Harnessing Harmful Data for Reducing Toxicity with Minimal Impact on Coherence and Evasiveness in Dialogue Agents , author=. arXiv preprint arXiv:2405.12900 , year=
-
[21]
arXiv preprint arXiv:2404.05143 , year=
Plug and Play with Prompts: A Prompt Tuning Approach for Controlling Text Generation , author=. arXiv preprint arXiv:2404.05143 , year=
-
[22]
2024 IEEE Security and Privacy Workshops (SPW) , pages=
Exploiting programmatic behavior of llms: Dual-use through standard security attacks , author=. 2024 IEEE Security and Privacy Workshops (SPW) , pages=. 2024 , organization=
2024
-
[23]
I’m fully who I am
“I’m fully who I am”: Towards Centering Transgender and Non-Binary Voices to Measure Biases in Open Language Generation , author=. Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency , pages=
2023
-
[24]
arXiv preprint arXiv:2305.15336 , year=
From text to mitre techniques: Exploring the malicious use of large language models for generating cyber attack payloads , author=. arXiv preprint arXiv:2305.15336 , year=
-
[25]
Scientific Reports , volume=
Bias of AI-generated content: an examination of news produced by large language models , author=. Scientific Reports , volume=. 2024 , publisher=
2024
-
[26]
Otolaryngology--Head and Neck Surgery , year=
Gender Bias in Artificial Intelligence-Written Letters of Reference , author=. Otolaryngology--Head and Neck Surgery , year=
-
[27]
The 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=
A Robot Walks into a Bar: Can Language Models Serve as Creativity SupportTools for Comedy? An Evaluation of LLMs’ Humour Alignment with Comedians , author=. The 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=
2024
-
[28]
The 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=
Auditing GPT's Content Moderation Guardrails: Can ChatGPT Write Your Favorite TV Show? , author=. The 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=
2024
-
[29]
arXiv preprint arXiv:2406.19497 , year=
Inclusivity in large language models: Personality traits and gender bias in scientific abstracts , author=. arXiv preprint arXiv:2406.19497 , year=
-
[30]
arXiv preprint arXiv:2407.01270 , year=
The African Woman is Rhythmic and Soulful: An Investigation of Implicit Biases in LLM Open-ended Text Generation , author=. arXiv preprint arXiv:2407.01270 , year=
-
[31]
TATuP-Zeitschrift f
Misuse of large language models: Exploiting weaknesses for target-specific outputs , author=. TATuP-Zeitschrift f. 2024 , publisher=
2024
-
[32]
Gender Bias in Decision-Making with Large Language Models: A Study of Relationship Conflicts
Levy, Sharon and Adler, William and Karver, Tahilin Sanchez and Dredze, Mark and Kaufman, Michelle R. Gender Bias in Decision-Making with Large Language Models: A Study of Relationship Conflicts. Findings of the Association for Computational Linguistics: EMNLP 2024. 2024
2024
-
[33]
Advances in Neural Information Processing Systems , volume=
Tree of attacks: Jailbreaking black-box llms automatically , author=. Advances in Neural Information Processing Systems , volume=
-
[34]
arXiv preprint arXiv:2310.12505 , year=
Attack prompt generation for red teaming and defending large language models , author=. arXiv preprint arXiv:2310.12505 , year=
-
[35]
arXiv preprint arXiv:2402.09132 , year=
Exploring the Adversarial Capabilities of Large Language Models , author=. arXiv preprint arXiv:2402.09132 , year=
-
[36]
arXiv preprint arXiv:2312.06924 , year=
Safety Alignment in NLP Tasks: Weakly Aligned Summarization as an In-Context Attack , author=. arXiv preprint arXiv:2312.06924 , year=
-
[37]
arXiv preprint arXiv:2409.00787 , year=
The dark side of human feedback: Poisoning large language models via user inputs , author=. arXiv preprint arXiv:2409.00787 , year=
-
[38]
arXiv preprint arXiv:2410.08776 , year=
F2A: An Innovative Approach for Prompt Injection by Utilizing Feign Security Detection Agents , author=. arXiv preprint arXiv:2410.08776 , year=
-
[39]
Generation of Korean Offensive Language by Leveraging Large Language Models via Prompt Design , author=. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computationa...
-
[40]
NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following , year=
Automatic Construction of a Korean Toxic Instruction Dataset for Ethical Tuning of Large Language Models , author=. NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following , year=
2023
-
[41]
Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=
People make better edits: Measuring the efficacy of LLM-generated counterfactually augmented data for harmful language detection , author=. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=
2023
-
[42]
Fifty Shades of Bias
" Fifty Shades of Bias": Normative Ratings of Gender Bias in GPT Generated English Text , author=. The 2023 Conference on Empirical Methods in Natural Language Processing , year =
2023
-
[43]
K o C o S a: K orean Context-aware Sarcasm Detection Dataset
Kim, Yumin and Suh, Heejae and Kim, Mingi and Won, Dongyeon and Lee, Hwanhee. K o C o S a: K orean Context-aware Sarcasm Detection Dataset. Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). 2024
2024
-
[44]
arXiv preprint arXiv:2404.12010 , year=
ParaFusion: A Large-Scale LLM-Driven English Paraphrase Dataset Infused with High-Quality Lexical and Syntactic Diversity , author=. arXiv preprint arXiv:2404.12010 , year=
-
[45]
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Indicllmsuite: A blueprint for creating pre-training and fine-tuning datasets for indian languages , author=. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
-
[46]
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pages=
SeeGULL Multilingual: a Dataset of Geo-Culturally Situated Stereotypes , author=. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pages=
-
[47]
arXiv preprint arXiv:2310.14429 , year=
Text generation for dataset augmentation in security classification tasks , author=. arXiv preprint arXiv:2310.14429 , year=
-
[48]
arXiv preprint arXiv:2306.09442 , year=
Explore, establish, exploit: Red teaming language models from scratch , author=. arXiv preprint arXiv:2306.09442 , year=
-
[49]
International Conference on Machine Learning , pages=
Automatically auditing large language models via discrete optimization , author=. International Conference on Machine Learning , pages=. 2023 , organization=
2023
-
[50]
arXiv preprint arXiv:2308.13768 , year=
Adversarial Fine-Tuning of Language Models: An Iterative Optimisation Approach for the Generation and Detection of Problematic Content , author=. arXiv preprint arXiv:2308.13768 , year=
-
[51]
arXiv preprint arXiv:2403.00829 , year=
TroubleLLM: Align to Red Team Expert , author=. arXiv preprint arXiv:2403.00829 , year=
-
[52]
arXiv preprint arXiv:2405.18540 , year=
Learning diverse attacks on large language models for robust red-teaming and safety tuning , author=. arXiv preprint arXiv:2405.18540 , year=
-
[53]
Outcome-Constrained Large Language Models for Countering Hate Speech
Hong, Lingzi and Luo, Pengcheng and Blanco, Eduardo and Song, Xiaoying. Outcome-Constrained Large Language Models for Countering Hate Speech. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024
2024
-
[54]
Proceedings of the CHI Conference on Human Factors in Computing Systems , pages=
A Piece of Theatre: Investigating How Teachers Design LLM Chatbots to Assist Adolescent Cyberbullying Education , author=. Proceedings of the CHI Conference on Human Factors in Computing Systems , pages=
-
[55]
Discgen: A framework for discourse-informed counterspeech generation , author=. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics , pages=
-
[56]
Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=
On zero-shot counterspeech generation by LLMs , author=. Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=
2024
-
[57]
arXiv preprint arXiv:2306.03097 , year=
Seeing seeds beyond weeds: Green teaming generative ai for beneficial uses , author=. arXiv preprint arXiv:2306.03097 , year=
-
[58]
Intrinsic Self-correction for Enhanced Morality: An Analysis of Internal Mechanisms and the Superficial Hypothesis
Liu, Guangliang and Mao, Haitao and Tang, Jiliang and Johnson, Kristen. Intrinsic Self-correction for Enhanced Morality: An Analysis of Internal Mechanisms and the Superficial Hypothesis. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024
2024
-
[59]
arXiv preprint arXiv:2408.10668 , year=
Probing the safety response boundary of large language models via unsafe decoding path generation , author=. arXiv preprint arXiv:2408.10668 , year=
-
[60]
arXiv preprint arXiv:2406.05587 , year=
Creativity has left the chat: The price of debiasing language models , author=. arXiv preprint arXiv:2406.05587 , year=
-
[61]
Controlled Text Generation for Large Language Model with Dynamic Attribute Graphs
Liang, Xun and Wang, Hanyu and Song, Shichao and Hu, Mengting and Wang, Xunzhi and Li, Zhiyu and Xiong, Feiyu and Tang, Bo. Controlled Text Generation for Large Language Model with Dynamic Attribute Graphs. Findings of the Association for Computational Linguistics: ACL 2024. 2024
2024
-
[62]
Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=
Gpt-hatecheck: Can llms write better functional tests for hate speech detection? , author=. Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages=
2024
-
[63]
Yale JL & Tech
The virtues of moderation , author=. Yale JL & Tech. , volume=. 2015 , publisher=
2015
-
[64]
Engineering, Technology & Applied Science Research , volume=
Towards optimal NLP solutions: analyzing GPT and LLaMA-2 models across model scale, dataset size, and task diversity , author=. Engineering, Technology & Applied Science Research , volume=
-
[65]
Machine-Generated Tweets , author=
Unmasking the Imposters: In-Domain Detection of Human vs. Machine-Generated Tweets , author=. arXiv e-prints , pages=
-
[66]
Natural Language Processing , pages=
Focal inferential infusion coupled with tractable density discrimination for implicit hate detection , author=. Natural Language Processing , pages=. 2023 , publisher =
2023
-
[67]
International Conference on Computational Science and Its Applications , pages=
LLMs and finetuning: benchmarking cross-domain performance for hate speech detection , author=. International Conference on Computational Science and Its Applications , pages=. 2025 , organization=
2025
-
[68]
arXiv preprint arXiv:2311.00203 , year=
Modeling subjectivity (by Mimicking Annotator Annotation) in toxic comment identification across diverse communities , author=. arXiv preprint arXiv:2311.00203 , year=
-
[69]
THOS: A Benchmark Dataset for Targeted Hate and Offensive Speech , author=. Proc. of Data-centric Machine Learning Research (DMLR) Workshop at ICML 2023 , year=
2023
-
[70]
Regulating Hate Speech Created by Generative AI , pages=
Generative ai for hate speech detection: Evaluation and findings , author=. Regulating Hate Speech Created by Generative AI , pages=. 2024 , publisher=
2024
-
[71]
arXiv preprint arXiv:2212.05613 , year=
A Study of Slang Representation Methods , author=. arXiv preprint arXiv:2212.05613 , year=
-
[72]
arXiv preprint arXiv:2212.10440 , year=
Perplexed by quality: A perplexity-based method for adult and harmful content detection in multilingual heterogeneous web data , author=. arXiv preprint arXiv:2212.10440 , year=
-
[73]
arXiv preprint arXiv:2212.10154 , year=
Human-guided fair classification for natural language processing , author=. arXiv preprint arXiv:2212.10154 , year=
-
[74]
arXiv preprint arXiv:2301.12534 , year=
Vicarious offense and noise audit of offensive speech classifiers: unifying human and machine disagreement on what is offensive , author=. arXiv preprint arXiv:2301.12534 , year=
-
[75]
International Conference on Advances in Social Networks Analysis and Mining , pages=
Non-binary gender expression in online interactions , author=. International Conference on Advances in Social Networks Analysis and Mining , pages=. 2024 , organization=
2024
-
[76]
Explicit Toxicity Detection Models with Interactive Visualization , year=
ToxVis: Enabling Interpretability of Implicit vs , author=. Explicit Toxicity Detection Models with Interactive Visualization , year=
-
[77]
arXiv preprint arXiv:2306.02978 , year=
Which Argumentative Aspects of Hate Speech in Social Media can be reliably identified? , author=. arXiv preprint arXiv:2306.02978 , year=
-
[78]
Proceedings of the Second Workshop on NLP for Positive Impact (NLP4PI) , pages=
Critical perspectives: A benchmark revealing pitfalls in PerspectiveAPI , author=. Proceedings of the Second Workshop on NLP for Positive Impact (NLP4PI) , pages=
-
[79]
How good is chatgpt for detecting hate speech in portuguese? , author=. Simp. 2023 , organization=
2023
-
[80]
Topological data mapping of online hate speech, misinformation, and general mental health: A large language model based study
Alexander, Andrew and Wang, Hongbin. Topological data mapping of online hate speech, misinformation, and general mental health: A large language model based study. 2309.13098
-
[81]
arXiv preprint arXiv:2310.10707 , year=
Demonstrations are all you need: Advancing offensive content paraphrasing using in-context learning , author=. arXiv preprint arXiv:2310.10707 , year=
-
[82]
Beyond plain toxic: building datasets for detection of flammable topics and inappropriate statements
Babakov, Nikolay and Logacheva, Varvara and Panchenko, Alexander. Beyond plain toxic: building datasets for detection of flammable topics and inappropriate statements. Lang. Resour. Eval
-
[83]
Evaluation of ChatGPT and BERT-based models for Turkish hate speech detection
C am, Nur Bengisu and \"O zg \"u r, Arzucan. Evaluation of ChatGPT and BERT-based models for Turkish hate speech detection. 2023 8th International Conference on Computer Science and Engineering ( UBMK )
2023
-
[84]
arXiv preprint arXiv:2310.13985 , year=
Haterephrase: Zero-and few-shot reduction of hate intensity in online posts using large language models , author=. arXiv preprint arXiv:2310.13985 , year=
-
[85]
arXiv preprint arXiv:2311.18580 , year=
Fft: Towards harmlessness evaluation and analysis for llms with factuality, fairness, toxicity , author=. arXiv preprint arXiv:2311.18580 , year=
-
[86]
Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=
ROBBIE: Robust Bias Evaluation of Large Generative Language Models , author=. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing , pages=
2023
-
[87]
arXiv preprint arXiv:2312.06674 , year=
Automated Toxicity Detection and Mitigation in Online Discussions Using LLMs: A Case Study with ChatGPT , author=. arXiv preprint arXiv:2312.06674 , year=
-
[88]
2024 , isbn =
Liu, Yi and Yu, Junzhe and Sun, Huijia and Shi, Ling and Deng, Gelei and Chen, Yuqi and Liu, Yang , title =. 2024 , isbn =. doi:10.1145/3691620.3695018 , booktitle =
2024
-
[89]
arXiv preprint arXiv:2402.14258 , year=
Toxicity Detection in User-Generated Content: A Prompt-Based Approach , author=. arXiv preprint arXiv:2402.14258 , year=
-
[90]
Proceedings of the first workshop on language technology for equality, diversity and inclusion , pages=
Cross-lingual transfer learning for hate speech detection , author=. Proceedings of the first workshop on language technology for equality, diversity and inclusion , pages=
-
[91]
Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=
Data-Efficient Hate Speech Detection via Cross-Lingual Nearest Neighbor Retrieval with Limited Labeled Data , author=. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=
2025
-
[92]
M isgender M ender: A Community-Informed Approach to Interventions for Misgendering
Hossain, Tamanna and Dev, Sunipa and Singh, Sameer. M isgender M ender: A Community-Informed Approach to Interventions for Misgendering. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologi...
2024
-
[93]
arXiv preprint arXiv:2405.01577 , year=
HateTinyLLM: Hate Speech Detection Using Tiny Large Language Models , author=. arXiv preprint arXiv:2405.01577 , year=
-
[94]
arXiv preprint arXiv:2404.17841 , year=
Toxicity Classification in Ukrainian , author=. arXiv preprint arXiv:2404.17841 , year=
-
[95]
Proceedings of the International AAAI Conference on Web and Social Media , volume=
The peripatetic hater: predicting movement among hate subreddits , author=. Proceedings of the International AAAI Conference on Web and Social Media , volume=
-
[96]
Proceedings of the International AAAI Conference on Web and Social Media , volume=
LLM-Based Semantic Augmentation for Harmful Content Detection , author=. Proceedings of the International AAAI Conference on Web and Social Media , volume=
-
[97]
Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop , pages=
Align before attend: Aligning visual and textual features for multimodal hateful content detection , author=. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Student Research Workshop , pages=
-
[98]
arXiv preprint arXiv:2407.15227 , year=
Toxicity Detection in Multilingual Settings Using Aligned Representations in LLMs , author=. arXiv preprint arXiv:2407.15227 , year=
-
[99]
and Saha, Sriparna and Pasupa, Kitsuchart , title =
Maity, Krishanu and Poornash, A.S. and Saha, Sriparna and Pasupa, Kitsuchart , title =. 2024 , isbn =. doi:10.1145/3627673.3680004 , booktitle =
2024
-
[100]
arXiv preprint arXiv:2410.10414 , year=
Detecting Deceptive Toxicity in Online Debates Using Instruction-tuned LLMs , author=. arXiv preprint arXiv:2410.10414 , year=
-
[101]
ACM Computing Surveys , volume=
Handling bias in toxic speech detection: A survey , author=. ACM Computing Surveys , volume=. 2023 , publisher=
2023
-
[102]
Proceedings of the 32nd ACM International Conference on Information and Knowledge Management , year=
Detecting Harmful Language in Code-Switched Social Media Text , author=. Proceedings of the 32nd ACM International Conference on Information and Knowledge Management , year=
-
[103]
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=
Toxicity detection is NOT all you need: measuring the gaps to supporting volunteer content moderators through a user-centric method , author=. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=
2024
-
[104]
Findings of the Association for Computational Linguistics: EMNLP 2023 , year=
HARE: Explainable Hate Speech Detection with Self-Rationale Supervision , author=. Findings of the Association for Computational Linguistics: EMNLP 2023 , year=
2023
-
[105]
Proceedings of the 32nd ACM International Conference on Information and Knowledge Management , year=
Toxicity Detection in Online Communities: Combining Social and Linguistic Features , author=. Proceedings of the 32nd ACM International Conference on Information and Knowledge Management , year=
-
[106]
IEEE Access , year=
Hate speech detection using large language models: A comprehensive review , author=. IEEE Access , year=
-
[107]
and Saha, Sriparna and Bhattacharyya, Pushpak
Maity, Krishanu and Poornash, A.S. and Saha, Sriparna and Bhattacharyya, Pushpak. T ox V id LM : A Multimodal Framework for Toxicity Detection in Code-Mixed Videos. Findings of the Association for Computational Linguistics: ACL 2024. 2024. doi:10.18653/v1/2024.findings-acl.663
2024 doi
-
[108]
International Conference on Learning Representations , volume=
More rlhf, more trust? on the impact of preference alignment on trustworthiness , author=. International Conference on Learning Representations , volume=
-
[109]
arXiv preprint arXiv:2405.17410 , year=
Mitigating Hate Speech in Social Media through Adaptive Prompt Tuning , author=. arXiv preprint arXiv:2405.17410 , year=
-
[110]
Don’t Go To Extremes: Revealing the Excessive Sensitivity and Calibration Limitations of LLMs in Implicit Hate Speech Detection , url =
Zhang, Min and He, Jianfeng and Ji, Taoran and Lu, Chang-Tien , year =. Don’t Go To Extremes: Revealing the Excessive Sensitivity and Calibration Limitations of LLMs in Implicit Hate Speech Detection , url =. doi:10.18653/v1/2024.acl-long.652 , booktitle =
2024 doi
-
[111]
Realistic Evaluation of Toxicity in Large Language Models , url =
Luong, Tinh and Le, Thanh-Thien and Ngo, Linh and Nguyen, Thien , year =. Realistic Evaluation of Toxicity in Large Language Models , url =. doi:10.18653/v1/2024.findings-acl.61 , booktitle =
2024 doi
-
[112]
Beyond Binary Gender Labels: Revealing Gender Bias in LLMs through Gender-Neutral Name Predictions , url =
You, Zhiwen and Lee, HaeJin and Mishra, Shubhanshu and Jeoung, Sullam and Mishra, Apratim and Kim, Jinseok and Diesner, Jana , year =. Beyond Binary Gender Labels: Revealing Gender Bias in LLMs through Gender-Neutral Name Predictions , url =. doi:10.18653/v1/2024.gebnlp-1.16 ,...
2024 doi
-
[113]
Hate Speech Detection using CoT and Post-hoc Explanation through Instruction-based Fine Tuning in Large Language Models , url =
Gupta, Palak and Jain, Nikhilesh and Bhat, Aruna , year =. Hate Speech Detection using CoT and Post-hoc Explanation through Instruction-based Fine Tuning in Large Language Models , url =. doi:10.1109/icaaic60222.2024.10575336 , booktitle =
2024
-
[114]
Sujitha and J, Anitha
Vakayil, Sonia and Juliet, D. Sujitha and J, Anitha. and Vakayil, Sunil , year =. RAG-Based LLM Chatbot Using Llama-2 , url =. doi:10.1109/icdcs59278.2024.10561020 , booktitle =
2024
-
[115]
Assessing Gender and Racial Bias in Large Language Model‐Powered Virtual Reference , volume =
Liu, Jieli and Wang, Haining , year =. Assessing Gender and Racial Bias in Large Language Model‐Powered Virtual Reference , volume =. Proceedings of the Association for Information Science and Technology , publisher =. doi:10.1002/pra2.1061 , number =
-
[116]
Efficient Toxic Content Detection by Bootstrapping and Distilling Large Language Models , volume =
Zhang, Jiang and Wu, Qiong and Xu, Yiming and Cao, Cheng and Du, Zheng and Psounis, Konstantinos , year =. Efficient Toxic Content Detection by Bootstrapping and Distilling Large Language Models , volume =. Proceedings of the AAAI Conference on Artificial Intelligence , publis...
-
[117]
OffMix-3L: A Novel Code-Mixed Dataset in Bangla-English-Hindi for Offensive Language Identification , publisher =
Goswami, Dhiman and Raihan, Md Nishat and Mahmud, Antara and Anastasopoulos, Antonios and Zampieri, Marcos , keywords =. OffMix-3L: A Novel Code-Mixed Dataset in Bangla-English-Hindi for Offensive Language Identification , publisher =. 2023 , copyright =. doi:10.48550/ARXIV.23...
-
[118]
Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=
Text classification via large language models , author=. Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=
2023
-
[119]
Proceedings of the 2024 conference on empirical methods in natural language processing , pages=
A survey on in-context learning , author=. Proceedings of the 2024 conference on empirical methods in natural language processing , pages=
2024
-
[120]
Tricking LLM s into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks
Rao, Abhinav Sukumar and Naik, Atharva Roshan and Vashistha, Sachin and Aditya, Somak and Choudhury, Monojit. Tricking LLM s into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks. Proceedings of the 2024 Joint International Conference on Computational Linguistics...
2024
-
[121]
33rd USENIX Security Symposium (USENIX Security 24) , pages=
Don't listen to me: Understanding and exploring jailbreak prompts of large language models , author=. 33rd USENIX Security Symposium (USENIX Security 24) , pages=
-
[122]
The Twelfth International Conference on Learning Representations , year=
Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models , author=. The Twelfth International Conference on Learning Representations , year=
-
[123]
arXiv preprint arXiv:2308.07308 , year=
Llm self defense: By self examination, llms know they are being tricked , author=. arXiv preprint arXiv:2308.07308 , year=
-
[124]
arXiv preprint arXiv:2308.13387 , year=
Do-not-answer: A dataset for evaluating safeguards in llms , author=. arXiv preprint arXiv:2308.13387 , year=
-
[125]
Sensitivity, Performance, Robustness: Deconstructing the Effect of Sociodemographic Prompting
Beck, Tilman and Schuff, Hendrik and Lauscher, Anne and Gurevych, Iryna. Sensitivity, Performance, Robustness: Deconstructing the Effect of Sociodemographic Prompting. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (...
2024 doi
-
[126]
ACM Transactions on Autonomous and Adaptive Systems , year=
Deceiving LLM through compositional instruction with hidden attacks , author=. ACM Transactions on Autonomous and Adaptive Systems , year=
-
[127]
arXiv preprint arXiv:2311.00172 , year=
Robust safety classifier for large language models: Adversarial prompt shield , author=. arXiv preprint arXiv:2311.00172 , year=
-
[128]
Proceedings of the AAAI Conference on Artificial Intelligence , volume=
Figstep: Jailbreaking large vision-language models via typographic visual prompts , author=. Proceedings of the AAAI Conference on Artificial Intelligence , volume=
-
[129]
Jailbreaker in jail: Moving target defense for large language models
Chen, Bocheng and Paliwal, Advait and Yan, Qiben. Jailbreaker in jail: Moving target defense for large language models. Proceedings of the 10th ACM Workshop on Moving Target Defense
-
[130]
Can LLMs deeply detect complex malicious queries? A framework for jailbreaking via obfuscating intent
Shang, Shang and Zhao, Xinqiang and Yao, Zhongjiang and Yao, Yepeng and Su, Liya and Fan, Zijing and Zhang, Xiaodan and Jiang, Zhengwei. Can LLMs deeply detect complex malicious queries? A framework for jailbreaking via obfuscating intent. 2405.03654
-
[131]
Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=
How alignment and jailbreak work: Explain llm safety through intermediate hidden states , author=. Findings of the Association for Computational Linguistics: EMNLP 2024 , pages=
2024
-
[132]
arXiv preprint arXiv:2406.08725 , year=
Rl-jack: Reinforcement learning-powered black-box jailbreaking attack against llms , author=. arXiv preprint arXiv:2406.08725 , year=
-
[133]
Self-deception: Reverse penetrating the semantic firewall of large language models
Wang, Zhenhua and Xie, Wei and Chen, Kai and Wang, Baosheng and Gui, Zhiwen and Wang, Enze. Self-deception: Reverse penetrating the semantic firewall of large language models. 2308.11521
-
[134]
A wolf in sheep's clothing: Generalized nested jailbreak prompts can fool large Language Models easily
Ding, Peng and Kuang, Jun and Ma, Dan and Cao, Xuezhi and Xian, Yunsen and Chen, Jiajun and Huang, Shujian. A wolf in sheep's clothing: Generalized nested jailbreak prompts can fool large Language Models easily. 2311.08268
-
[135]
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=
ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings , author=. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=
2024
-
[136]
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Prp: Propagating universal perturbations to attack large language model guard-rails , author=. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
-
[137]
33rd USENIX Security Symposium (USENIX Security 24) , pages=
Making them ask and answer: Jailbreaking large language models in few queries via disguise and reconstruction , author=. 33rd USENIX Security Symposium (USENIX Security 24) , pages=
-
[138]
D r A ttack: Prompt Decomposition and Reconstruction Makes Powerful LLM s Jailbreakers
Li, Xirui and Wang, Ruochen and Cheng, Minhao and Zhou, Tianyi and Hsieh, Cho-Jui. D r A ttack: Prompt Decomposition and Reconstruction Makes Powerful LLM s Jailbreakers. Findings of the Association for Computational Linguistics: EMNLP 2024. 2024
2024
-
[139]
Virtual Context: Enhancing jailbreak attacks with special token injection
Zhou, Yuqi and Lu, Lin and Sun, Hanchi and Zhou, Pan and Sun, Lichao. Virtual Context: Enhancing jailbreak attacks with special token injection. 2406.19845
-
[140]
Model Surgery: Modulating LLM ' s Behavior Via Simple Parameter Editing
Wang, Huanqian and Yue, Yang and Lu, Rui and Shi, Jingxin and Zhao, Andrew and Wang, Shenzhi and Song, Shiji and Huang, Gao. Model Surgery: Modulating LLM ' s Behavior Via Simple Parameter Editing. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of th...
2025 doi
-
[141]
DistillSeq : A framework for safety alignment testing in large Language Models using knowledge distillation
Yang, Mingke and Chen, Yuqi and Liu, Yi and Shi, Ling. DistillSeq : A framework for safety alignment testing in large Language Models using knowledge distillation. arXiv [cs.SE]
-
[142]
arXiv e-prints , pages=
Figure it out: Analyzing-based jailbreak attack on large language models , author=. arXiv e-prints , pages=
-
[143]
Jailbreak Antidote: Runtime Safety-Utility Balance via Sparse Representation Adjustment in Large Language Models , author=
-
[144]
Attack prompt generation for red teaming and defending large language models
Deng, Boyi and Wang, Wenjie and Feng, Fuli and Deng, Yang and Wang, Qifan and He, Xiangnan. Attack prompt generation for red teaming and defending large language models. Findings of the Association for Computational Linguistics: EMNLP 2023
2023
-
[145]
Alan Turing Institute , volume=
How much online abuse is there , author=. Alan Turing Institute , volume=
-
[146]
arXiv preprint arXiv:2405.15604 , year=
Text generation: A systematic literature review of tasks, evaluation, and challenges , author=. arXiv preprint arXiv:2405.15604 , year=
-
[147]
All You Need To Know About LLM Text Generation , year =
-
[148]
ACM Computing Surveys , volume=
Pre-trained language models for text generation: A survey , author=. ACM Computing Surveys , volume=. 2024 , publisher=
2024
-
[149]
Proceedings of the International AAAI Conference on Web and Social Media , volume=
Human and LLM biases in hate speech annotations: A socio-demographic analysis of annotators and targets , author=. Proceedings of the International AAAI Conference on Web and Social Media , volume=
-
[150]
and Susnjak, Teo and Liu, Tong and Watters, Paul and Halgamuge, Malka N
McIntosh, Timothy R. and Susnjak, Teo and Liu, Tong and Watters, Paul and Halgamuge, Malka N. , journal=. The Inadequacy of Reinforcement Learning From Human Feedback—Radicalizing Large Language Models via Semantic Vulnerabilities , year=
-
[151]
2023 , eprint=
Are Personalized Stochastic Parrots More Dangerous? Evaluating Persona Biases in Dialogue Systems , author=. 2023 , eprint=
2023
-
[152]
2025 , eprint=
HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate Campaigns , author=. 2025 , eprint=
2025
-
[153]
2024 , eprint=
A Comprehensive Study on NLP Data Augmentation for Hate Speech Detection: Legacy Methods, BERT, and LLMs , author=. 2024 , eprint=
2024
-
[154]
arXiv preprint arXiv:2304.10611 , year=
Joint Repetition Suppression and Content Moderation of Large Language Models , author=. arXiv preprint arXiv:2304.10611 , year=
-
[155]
Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=
Legally Enforceable Hate Speech Detection for Public Forums , author=. Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=
2023
-
[156]
Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence , pages=
Evaluating GPT-3 generated explanations for hateful content moderation , author=. Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence , pages=
-
[157]
Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=
Exploring Cross-Cultural Differences in English Hate Speech Annotations: From Dataset Construction to Analysis , author=. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: L...
2024
-
[158]
The Twelfth International Conference on Learning Representations , year =
Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions , author=. The Twelfth International Conference on Learning Representations , year =
-
[159]
Model Surgery: Modulating LLM’s Behavior Via Simple Parameter Editing , author=. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=
2025
-
[160]
arXiv preprint arXiv:2310.03400 , year=
Adapting large language models for content moderation: Pitfalls in data engineering and supervised fine-tuning , author=. arXiv preprint arXiv:2310.03400 , year=
-
[161]
Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=
Self-guard: Empower the llm to safeguard itself , author=. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=
2024
-
[162]
arXiv preprint arXiv:2312.03813 , year=
Improving activation steering in language models with mean-centring , author=. arXiv preprint arXiv:2312.03813 , year=
-
[163]
arXiv preprint arXiv:2312.04782 , year=
Make them spill the beans! coercive knowledge extraction from (production) llms , author=. arXiv preprint arXiv:2312.04782 , year=
-
[164]
GTA : Gated Toxicity Avoidance for LM Performance Preservation
Kim, Heegyu and Cho, Hyunsouk. GTA : Gated Toxicity Avoidance for LM Performance Preservation. Findings of the Association for Computational Linguistics: EMNLP 2023. 2023
2023
-
[165]
Proceedings of the 41st International Conference on Machine Learning , pages=
Learning and forgetting unsafe examples in large language models , author=. Proceedings of the 41st International Conference on Machine Learning , pages=
-
[166]
NAACL-Findings , year =
Are you talking to ['xem'] or ['x', 'em']? On Tokenization and Addressing Misgendering in LLMs with Pronoun Tokenization Parity , author =. NAACL-Findings , year =
-
[167]
arXiv preprint arXiv:2401.08491 , year=
Contrastive perplexity for controlled generation: An application in detoxifying large language models , author=. arXiv preprint arXiv:2401.08491 , year=
-
[168]
Proceedings of the International AAAI Conference on Web and Social Media , volume=
Watch your language: Investigating content moderation with large language models , author=. Proceedings of the International AAAI Conference on Web and Social Media , volume=
-
[169]
Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=
Lora-guard: Parameter-efficient guardrail adaptation for content moderation of large language models , author=. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=
2024
-
[170]
Preference Tuning For Toxicity Mitigation Generalizes Across Languages
Li, Xiaochen and Yong, Zheng Xin and Bach, Stephen. Preference Tuning For Toxicity Mitigation Generalizes Across Languages. Findings of the Association for Computational Linguistics: EMNLP 2024. 2024
2024
-
[171]
Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Genderalign: An alignment dataset for mitigating gender bias in large language models , author=. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
-
[172]
arXiv preprint arXiv:2406.13748 , year=
Every language counts: Learn and unlearn in multilingual LLMs , author=. arXiv preprint arXiv:2406.13748 , year=
-
[173]
Natural Language Processing Journal , volume=
Contrastive adversarial gender debiasing , author=. Natural Language Processing Journal , volume=. 2024 , publisher=
2024
-
[174]
ACL (Findings) , year=
Towards Tracing Trustworthiness Dynamics: Revisiting Pre-training Period of Large Language Models , author=. ACL (Findings) , year=
-
[175]
arXiv preprint arXiv:2407.21772 , year=
Shieldgemma: Generative ai content moderation based on gemma , author=. arXiv preprint arXiv:2407.21772 , year=
-
[176]
2024 IEEE Symposium on Security and Privacy (SP) , pages=
You only prompt once: On the capabilities of prompt learning on large language models to tackle toxic content , author=. 2024 IEEE Symposium on Security and Privacy (SP) , pages=. 2024 , organization=
2024
-
[177]
2024 IEEE Symposium on Security and Privacy (SP) , pages=
Moderating new waves of online hate with chain-of-thought reasoning in large language models , author=. 2024 IEEE Symposium on Security and Privacy (SP) , pages=. 2024 , organization=
2024
-
[178]
arXiv preprint arXiv:2410.04155 , year=
Toxic Subword Pruning for Dialogue Response Generation on Large Language Models , author=. arXiv preprint arXiv:2410.04155 , year=
-
[179]
Neurips Safe Generative AI Workshop 2024 , year=
Simplicity Prevails: Rethinking Negative Preference Optimization for LLM Unlearning , author=. Neurips Safe Generative AI Workshop 2024 , year=
2024
-
[180]
arXiv preprint arXiv:2410.03466 , year=
Is Safer Better? The Impact of Guardrails on the Argumentative Strength of LLMs in Hate Speech Countering , author=. arXiv preprint arXiv:2410.03466 , year=
-
[181]
arXiv preprint arXiv:2402.17262 , year=
Speak out of turn: Safety vulnerability of large language models in multi-turn dialogue , author=. arXiv preprint arXiv:2402.17262 , year=
-
[182]
The Thirteenth International Conference on Learning Representations , year=
FairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMs , author=. The Thirteenth International Conference on Learning Representations , year=
-
[183]
2025 IEEE International Conference on Multimedia and Expo (ICME) , pages=
Pico: Jailbreaking multimodal large language models via pictorial code contextualization , author=. 2025 IEEE International Conference on Multimedia and Expo (ICME) , pages=. 2025 , organization=
2025
-
[184]
arXiv preprint arXiv:2405.20773 , year=
Visual-roleplay: Universal jailbreak attack on multimodal large language models via role-playing image character , author=. arXiv preprint arXiv:2405.20773 , year=
-
[185]
arXiv preprint arXiv:2009.11462 , year=
Realtoxicityprompts: Evaluating neural toxic degeneration in language models , author=. arXiv preprint arXiv:2009.11462 , year=
2009 arXiv
-
[186]
arXiv preprint arXiv:2205.01833 , year=
OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts , author=. arXiv preprint arXiv:2205.01833 , year=
-
[187]
arXiv preprint arXiv:2306.13651 , year=
Bring your own data! self-supervised evaluation for large language models , author=. arXiv preprint arXiv:2306.13651 , year=
-
[188]
Advances in Neural Information Processing Systems , volume=
Beavertails: Towards improved safety alignment of llm via a human-preference dataset , author=. Advances in Neural Information Processing Systems , volume=
-
[189]
Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
An empirical analysis of parameter-efficient methods for debiasing pre-trained language models , author=. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
-
[190]
Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023) , pages=
Enabling Classifiers to Make Judgements Explicitly Aligned with Human Values , author=. Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023) , pages=
2023
-
[191]
SQ u AR e: A Large-Scale Dataset of Sensitive Questions and Acceptable Responses Created through Human-Machine Collaboration
Lee, Hwaran and Hong, Seokhee and Park, Joonsuk and Kim, Takyoung and Cha, Meeyoung and Choi, Yejin and Kim, Byoungpil and Kim, Gunhee and Lee, Eun-Ju and Lim, Yong and Oh, Alice and Park, Sangchul and Ha, Jung-Woo. SQ u AR e: A Large-Scale Dataset of Sensitive Questions and A...
2023
-
[192]
arXiv preprint arXiv:2401.15585 , year=
Evaluating gender bias in large language models via chain-of-thought prompting , author=. arXiv preprint arXiv:2401.15585 , year=
-
[193]
Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=
Steerlm: Attribute conditioned sft as an (user-steerable) alternative to rlhf , author=. Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=
2023
-
[194]
ECAI 2024 , pages=
REFINE-LM: Mitigating Language Model Stereotypes via Reinforcement Learning , author=. ECAI 2024 , pages=. 2024 , publisher=
2024
-
[195]
arXiv preprint arXiv:2408.09600 , year=
Antidote: Post-fine-tuning safety alignment for large language models against harmful fine-tuning , author=. arXiv preprint arXiv:2408.09600 , year=
-
[196]
HateCheck: Functional Tests for Hate Speech Detection Models , author =. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) , pages =. 2021 ,...
2021
-
[197]
arXiv preprint arXiv:2206.09917 , year =
Multilingual HateCheck: Functional Tests for Multilingual Hate Speech Detection Models , author =. arXiv preprint arXiv:2206.09917 , year =
-
[198]
Proceedings of the NAACL Student Research Workshop , pages =
Hateful Symbols or Hateful People? Predictive Features for Hate Speech Detection on Twitter , author =. Proceedings of the NAACL Student Research Workshop , pages =. 2016 , publisher =
2016
-
[199]
arXiv preprint arXiv:2005.05921 , year =
Intersectional Bias in Hate Speech and Abusive Language Datasets , author =. arXiv preprint arXiv:2005.05921 , year =
2005 arXiv
-
[200]
arXiv preprint arXiv:2205.06621 , year =
Analyzing Hate Speech Data along Racial, Gender and Intersectional Axes , author =. arXiv preprint arXiv:2205.06621 , year =
-
[201]
International Conference on Human-Computer Interaction , pages=
Interpretable and high-performance hate and offensive speech detection , author=. International Conference on Human-Computer Interaction , pages=. 2022 , organization=
2022
-
[202]
2024 International Conference on Content-Based Multimedia Indexing (CBMI) , pages=
Exploring the Plausibility of Hate and Counter Speech Detectors with Explainable AI , author=. 2024 International Conference on Content-Based Multimedia Indexing (CBMI) , pages=. 2024 , organization=
2024
-
[203]
Towards Interpretable Hate Speech Detection using Large Language Model-extracted Rationales
Nirmal, Ayushi and Bhattacharjee, Amrita and Sheth, Paras and Liu, Huan. Towards Interpretable Hate Speech Detection using Large Language Model-extracted Rationales. Proceedings of the 8th Workshop on Online Abuse and Harms (WOAH 2024). 2024. doi:10.18653/v1/2024.woah-1.17
2024 doi
-
[204]
Proceedings of the 31st International Conference on Computational Linguistics , pages =
HateBRXplain: A Benchmark Dataset with Human-Annotated Rationales for Explainable Hate Speech Detection in Brazilian Portuguese , author =. Proceedings of the 31st International Conference on Computational Linguistics , pages =. 2025 , publisher =
2025
-
[205]
Neurips Safe Generative AI Workshop 2024 , year=
Investigating annotator bias in large language models for hate speech detection , author=. Neurips Safe Generative AI Workshop 2024 , year=
2024
-
[206]
2023 IEEE/ACM 20th International Conference on Mining Software Repositories (MSR) , pages=
She elicits requirements and he tests: Software engineering gender bias in large language models , author=. 2023 IEEE/ACM 20th International Conference on Mining Software Repositories (MSR) , pages=. 2023 , organization=
2023
-
[207]
arXiv preprint arXiv:2307.10337 , year=
Are you in a masquerade? exploring the behavior and impact of large language model driven social bots in online social networks , author=. arXiv preprint arXiv:2307.10337 , year=
-
[208]
2024 , eprint =
OpenEval: Benchmarking Chinese LLMs across Capability, Alignment and Safety , author =. 2024 , eprint =
2024
-
[209]
Beyond Perplexity: Multi-dimensional Safety Evaluation of
Xu, Zhichao and Gupta, Ashim and Li, Tao and Bentham, Oliver and Srikumar, Vivek , editor =. Beyond Perplexity: Multi-dimensional Safety Evaluation of. Findings of the Association for Computational Linguistics: EMNLP 2024 , month = nov, year =. doi:10.18653/v1/2024.findings-em...
2024 doi
-
[210]
arXiv preprint arXiv:2305.11262 , year=
Chbias: Bias evaluation and mitigation of chinese conversational language models , author=. arXiv preprint arXiv:2305.11262 , year=
-
[211]
arXiv preprint arXiv:2311.10266 , year=
Diagnosing and debiasing corpus-based political bias and insults in GPT2 , author=. arXiv preprint arXiv:2311.10266 , year=
-
[212]
arXiv preprint arXiv:2505.02009 , year=
Towards safer pretraining: Analyzing and filtering harmful content in webscale datasets for responsible llms , author=. arXiv preprint arXiv:2505.02009 , year=
-
[213]
arXiv preprint arXiv:2502.12566 , year=
Exploring the impact of personality traits on llm bias and toxicity , author=. arXiv preprint arXiv:2502.12566 , year=
-
[214]
IEEE Access , year=
OffensiveLang: A Community Based Implicit Offensive Language Dataset , author=. IEEE Access , year=
-
[215]
arXiv preprint arXiv:2302.13971 , year=
Llama: Open and efficient foundation language models , author=. arXiv preprint arXiv:2302.13971 , year=
-
[216]
arXiv preprint arXiv:2403.05530 , year=
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context , author=. arXiv preprint arXiv:2403.05530 , year=
-
[217]
arXiv preprint arXiv:2312.11805 , year=
Gemini: a family of highly capable multimodal models , author=. arXiv preprint arXiv:2312.11805 , year=
-
[218]
arXiv preprint arXiv:1907.11692 , year=
Roberta: A robustly optimized bert pretraining approach , author=. arXiv preprint arXiv:1907.11692 , year=
1907 arXiv
-
[219]
Bert: Pre-training of deep bidirectional transformers for language understanding , author=. Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers) , pages=
2019
-
[220]
arXiv preprint arXiv:2410.08565 , year=
Baichuan-omni technical report , author=. arXiv preprint arXiv:2410.08565 , year=
-
[221]
arXiv preprint arXiv:2309.16609 , year=
Qwen technical report , author=. arXiv preprint arXiv:2309.16609 , year=
-
[222]
arXiv preprint arXiv:2505.10846 , year=
AutoRAN: Weak-to-Strong Jailbreaking of Large Reasoning Models , author=. arXiv preprint arXiv:2505.10846 , year=
-
[223]
arXiv preprint arXiv:1607.06520 , year=
Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings , author=. arXiv preprint arXiv:1607.06520 , year=
-
[224]
arXiv preprint arXiv:2212.03827 , year=
Discovering Latent Knowledge in Language Models Without Supervision , author=. arXiv preprint arXiv:2212.03827 , year=
-
[225]
Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Steering Llama 2 via Contrastive Activation Addition , author=. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
-
[226]
arXiv preprint arXiv:2310.01405 , year=
Representation Engineering: A Top-Down Approach to AI Transparency , author=. arXiv preprint arXiv:2310.01405 , year=
-
[227]
International Conference on Learning Representations (ICLR) , year=
Improving Instruction-Following in Language Models Through Activation Steering , author=. International Conference on Learning Representations (ICLR) , year=
-
[228]
Advances in Neural Information Processing Systems (NeurIPS) , year=
Refusal in Language Models Is Mediated by a Single Direction , author=. Advances in Neural Information Processing Systems (NeurIPS) , year=
-
[229]
Proceedings of the 41st International Conference on Machine Learning , volume=
Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications , author=. Proceedings of the 41st International Conference on Machine Learning , volume=
-
[230]
Proceedings of the 31st International Conference on Computational Linguistics , pages=
Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective , author=. Proceedings of the 31st International Conference on Computational Linguistics , pages=
-
[231]
arXiv preprint arXiv:2411.11296 , year=
Steering Language Model Refusal with Sparse Autoencoders , author=. arXiv preprint arXiv:2411.11296 , year=
-
[232]
arXiv preprint arXiv:2412.08201 , year=
Model-editing-based jailbreak against safety-aligned large language models , author=. arXiv preprint arXiv:2412.08201 , year=
-
[233]
Findings of the Association for Computational Linguistics: ACL 2025 , pages=
Delman: Dynamic defense against large language model jailbreaking with model editing , author=. Findings of the Association for Computational Linguistics: ACL 2025 , pages=
2025
-
[234]
arXiv preprint arXiv:2510.13901 , year=
RAID: Refusal-Aware and Integrated Decoding for Jailbreaking LLMs , author=. arXiv preprint arXiv:2510.13901 , year=
-
[235]
arXiv preprint arXiv:2601.04034 , year=
HoneyTrap: Deceiving Large Language Model Attackers to Honeypot Traps with Resilient Multi-Agent Defense , author=. arXiv preprint arXiv:2601.04034 , year=
-
[236]
arXiv preprint arXiv:2303.08774 , year=
Gpt-4 technical report , author=. arXiv preprint arXiv:2303.08774 , year=
-
[237]
arXiv preprint arXiv:2404.01954 , year=
Hyperclova x technical report , author=. arXiv preprint arXiv:2404.01954 , year=
-
[238]
Advances in neural information processing systems , volume=
Language models are few-shot learners , author=. Advances in neural information processing systems , volume=
-
[239]
The Thirteenth International Conference on Learning Representations (ICLR) , year =
Programming Refusal with Conditional Activation Steering , author =. The Thirteenth International Conference on Learning Representations (ICLR) , year =
-
[240]
Adaptive Detoxification: Safeguarding General Capabilities of
Lu, Yifan and Li, Jing and Zhou, Yigeng and Zhang, Yihui and Wang, Wenya and Li, Xiucheng and Zhang, Meishan and Liu, Fangming and Yu, Jun and Zhang, Min , booktitle =. Adaptive Detoxification: Safeguarding General Capabilities of
-
[241]
Controlled
Pynadath, Patrick and Zhang, Ruqi , journal =. Controlled
-
[242]
2017 , howpublished =
Perspective. 2017 , howpublished =
2017
-
[243]
Corpus Pragmatics , year =
Kolhatkar, Varada and Wu, Hanhan and Cavasso, Luca and Francis, Emilie and Shukla, Kavan and Taboada, Maite , title =. Corpus Pragmatics , year =
-
[244]
Royal Society Open Science , volume=
Generalization bias in large language model summarization of scientific research , author=. Royal Society Open Science , volume=. 2025 , publisher=
2025
-
[245]
Proceedings of the National Academy of Sciences , volume=
Explicitly unbiased large language models still form biased associations , author=. Proceedings of the National Academy of Sciences , volume=. 2025 , publisher=
2025
-
[246]
ICWSM , year =
Davidson, Thomas and Warmsley, Dana and Macy, Michael and Weber, Ingmar , title =. ICWSM , year =
-
[247]
Hate Speech , year =
-
[248]
Hateful Conduct , year =
-
[249]
, title =
Salminen, Joni and Hopf, Maximilian and Chowdhury, Shammur Absar and Jung, Soon-gyo and Almerekhi, Hind and Jansen, Bernard J. , title =. Human-centric Computing and Information Sciences , year =
-
[250]
Proceedings of the 13th International Workshop on Semantic Evaluation , year =
Zampieri, Marcos and Malmasi, Shervin and Nakov, Preslav and Rosenthal, Sara and Farra, Noura and Kumar, Ritesh , title =. Proceedings of the 13th International Workshop on Semantic Evaluation , year =
-
[251]
Proceedings of
Wiegand, Michael and Siegel, Melanie and Ruppendorfer, Josef , title =. Proceedings of. 2018 , publisher =
2018
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.