Pith. sign in

REVIEW 6 cited by

COLD: A Benchmark for Chinese Offensive Language Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2201.06025 v2 pith:IBKT6ECB submitted 2022-01-16 cs.CL cs.AI

classification cs.CLcs.AI
keywords offensivelanguagechinesemodelsbenchmarkcolddetectioncoldetector
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Offensive language detection is increasingly crucial for maintaining a civilized social media platform and deploying pre-trained language models. However, this task in Chinese is still under exploration due to the scarcity of reliable datasets. To this end, we propose a benchmark --COLD for Chinese offensive language analysis, including a Chinese Offensive Language Dataset --COLDATASET and a baseline detector --COLDETECTOR which is trained on the dataset. We show that the COLD benchmark contributes to Chinese offensive language detection which is challenging for existing resources. We then deploy the COLDETECTOR and conduct detailed analyses on popular Chinese pre-trained language models. We first analyze the offensiveness of existing generative models and show that these models inevitably expose varying degrees of offensive issues. Furthermore, we investigate the factors that influence the offensive generations, and we find that anti-bias contents and keywords referring to certain groups or revealing negative attitudes trigger offensive outputs easier.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fine-Grained Chinese Hate Speech Understanding: Span-Level Resources, Coded Term Lexicon, and Enhanced Detection Frameworks

    cs.CL 2025-07 reject novelty 7.0 of 10

    The paper creates a span-level Chinese hate speech dataset and a 830-term coded hate lexicon, but its two-stage training method's reported superiority is contradicted by the paper's own COLD results.

  2. Exploring Multimodal Challenges in Toxic Chinese Detection: Taxonomy, Benchmark, and Findings

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new taxonomy and dataset of 8 types of perturbed toxic Chinese show nine top LLMs often miss these obfuscated insults, and small-sample ICL or fine-tuning causes overcorrection.

  3. Breaking the Cloak! Unveiling Chinese Cloaked Toxicity with Homophone Graph and Toxic Lexicon

    cs.CL 2025-05 conditional novelty 6.0 of 10

    C2TU combines a Chinese pronunciation graph, a toxic lexicon, and language-model probability checking to find and correct homophone-cloaked toxic words without any training.

  4. Chinese Toxic Language Mitigation via Sentiment Polarity Consistent Rewrites

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new Chinese detoxification dataset and 17-model benchmark show that LLMs can remove toxic words but often distort emotional tone, especially for emoji, homophone, and dialogue-based toxicity.

  5. Cascading versus Joint Modeling for Hierarchical Offensive Language Detection

    cs.CL 2026-07 conditional novelty 4.0 of 10

    A controlled comparison shows a three-model cascade beats a shared-encoder multi-task model on hierarchical offensive language detection by up to 7.1 macro-F1 points, at three times the parameter count and 1.67 times ...

  6. Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges

    cs.AI 2025-07 reject novelty 1.0 of 10

    A broad survey of LLM alignment that catalogs objectives, benchmarks, SFT/RLHF/DPO methods, and safety challenges, without contributing new experimental or theoretical results.

Pith tools