Pith. sign in

REVIEW 10 cited by

ToxicChat: Unveiling Hidden Challenges of Toxicity Detection in Real-World User-AI Conversation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.17389 v1 pith:LILDZUIX submitted 2023-10-26 cs.CL cs.AI

classification cs.CLcs.AI
keywords toxicityuser-aidetectiontoxicchatchallengesmodelsreal-worldbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite remarkable advances that large language models have achieved in chatbots, maintaining a non-toxic user-AI interactive environment has become increasingly critical nowadays. However, previous efforts in toxicity detection have been mostly based on benchmarks derived from social media content, leaving the unique challenges inherent to real-world user-AI interactions insufficiently explored. In this work, we introduce ToxicChat, a novel benchmark based on real user queries from an open-source chatbot. This benchmark contains the rich, nuanced phenomena that can be tricky for current toxicity detection models to identify, revealing a significant domain difference compared to social media content. Our systematic evaluation of models trained on existing toxicity datasets has shown their shortcomings when applied to this unique domain of ToxicChat. Our work illuminates the potentially overlooked challenges of toxicity detection in real-world user-AI conversations. In the future, ToxicChat can be a valuable resource to drive further advancements toward building a safe and healthy environment for user-AI interactions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLMs Encode Harmfulness and Refusal Separately

    cs.CL 2025-07 conditional novelty 7.0 of 10

    LLMs encode a separate internal harmfulness direction, distinct from the refusal direction, which is more robust to jailbreaks and adversarial finetuning.

  2. Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders

    cs.CL 2026-02 conditional novelty 6.0 of 10

    Coverage of sparse-autoencoder-identified task features predicts post-training performance and can guide synthesis of small, high-impact datasets (2,000 vs. 300,000 samples).

  3. Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A TEE-based protocol for cryptographically verifiable AI safety benchmark results, demonstrated on Llama-3.1 with AWS Nitro Enclaves.

  4. Large Language Models Often Know When They Are Being Evaluated

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Frontier language models distinguish evaluation transcripts from deployment transcripts with AUC up to 0.83, below the authors' human baseline of 0.92.

  5. A Constitution-Grid Instrument for Data-Efficient RL Alignment (C-Guard)

    cs.CL 2026-07 conditional novelty 5.0 of 10

    C-Guard uses a constitution grid and a per-cell learnability score to aim RL training data, cutting over-refusal by 9.6 points while exposing a hidden 0.06 rise in adversarial under-refusal.

  6. Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B

    cs.LG 2026-05 reject novelty 5.0 of 10

    A 184M-parameter DeBERTa-v3 fine-tuned model is claimed to beat Llama-Guard-3-8B on all tested prompt-injection benchmarks while adding BFSI regulatory labels, but a leaked training/eval overlap undermines the zero-FPR claim.

  7. GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio

    cs.CR 2026-02 reject novelty 4.0 of 10

    A reasoning-based guardrail model trained on 148k text/image/video samples is claimed to beat prior content-safety moderators, though the abstract and body conflict on modalities and model sizes.

  8. From Rogue to Safe AI: The Role of Explicit Refusals in Aligning LLMs with International Humanitarian Law

    cs.CY 2025-06 conditional novelty 4.0 of 10

    Across eight LLMs, most explicitly IHL-violating prompts are refused, and a single system-level safety prompt raises explanatory refusal rates in six of eight models, though the benchmark is not publicly released.

  9. SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use

    cs.CL 2025-05 conditional novelty 4.0 of 10

    A new cross-lingual benchmark shows large language models comply with explicit requests to use swear words far more often in Indic languages than in English, revealing a safety alignment gap.

  10. The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs

    cs.CL 2025-06 conditional novelty 3.0 of 10

    A structured survey of LLM safety evaluation that proposes a why/what/where/how taxonomy and catalogs metrics, datasets, benchmarks, evaluators, and frameworks.

Pith tools