REVIEW 11 cited by
ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Despite remarkable advances that large language models have achieved in chatbots, maintaining a non-toxic user-AI interactive environment has become increasingly critical nowadays. However, previous efforts in toxicity detection have been mostly based on benchmarks derived from social media content, leaving the unique challenges inherent to real-world user-AI interactions insufficiently explored. In this work, we introduce ToxicChat, a novel benchmark based on real user queries from an open-source chatbot. This benchmark contains the rich, nuanced phenomena that can be tricky for current toxicity detection models to identify, revealing a significant domain difference compared to social media content. Our systematic evaluation of models trained on existing toxicity datasets has shown their shortcomings when applied to this unique domain of ToxicChat. Our work illuminates the potentially overlooked challenges of toxicity detection in real-world user-AI conversations. In the future, ToxicChat can be a valuable resource to drive further advancements toward building a safe and healthy environment for user-AI interactions.
Forward citations
Cited by 11 Pith papers
-
LLMs Encode Harmfulness and Refusal Separately
LLMs encode a separate internal harmfulness direction, distinct from the refusal direction, which is more robust to jailbreaks and adversarial finetuning.
-
Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders
Coverage of sparse-autoencoder-identified task features predicts post-training performance and can guide synthesis of small, high-impact datasets (2,000 vs. 300,000 samples).
-
Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments
A TEE-based protocol for cryptographically verifiable AI safety benchmark results, demonstrated on Llama-3.1 with AWS Nitro Enclaves.
-
Large Language Models Often Know When They Are Being Evaluated
Frontier language models distinguish evaluation transcripts from deployment transcripts with AUC up to 0.83, below the authors' human baseline of 0.92.
-
A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard)
C-Guard uses a constitution grid and a per-cell learnability score to aim RL training data, cutting over-refusal by 9.6 points while exposing a hidden 0.06 rise in adversarial under-refusal.
-
Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B
A 184M-parameter DeBERTa-v3 fine-tuned model is claimed to beat Llama-Guard-3-8B on all tested prompt-injection benchmarks while adding BFSI regulatory labels, but a leaked training/eval overlap undermines the zero-FPR claim.
-
GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio
A reasoning-based guardrail model trained on 148k text/image/video samples is claimed to beat prior content-safety moderators, though the abstract and body conflict on modalities and model sizes.
-
From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law
Across eight LLMs, most explicitly IHL-violating prompts are refused, and a single system-level safety prompt raises explanatory refusal rates in six of eight models, though the benchmark is not publicly released.
-
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use
A new cross-lingual benchmark shows large language models comply with explicit requests to use swear words far more often in Indic languages than in English, revealing a safety alignment gap.
-
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs
A structured survey of LLM safety evaluation that proposes a why/what/where/how taxonomy and catalogs metrics, datasets, benchmarks, evaluators, and frameworks.
-
Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences
A guardrail pipeline combining detection, retrieval grounding, rule-based wrappers, and a repair model is reported to match OpenAI moderation and fix 80.7 percent of hallucinated HaluEval answers.
Discussion (0). Continue with ORCID to comment.