Pith. sign in

REVIEW 6 cited by

A Holistic Approach to Undesired Content Detection in the Real World

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.03274 v2 pith:VZBCVPYR submitted 2022-08-05 cs.CL cs.LG

classification cs.CLcs.LG
keywords contentapproachsystemholisticincludingmoderationrobusttaxonomies
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a holistic approach to building a robust and useful natural language classification system for real-world content moderation. The success of such a system relies on a chain of carefully designed and executed steps, including the design of content taxonomies and labeling instructions, data quality control, an active learning pipeline to capture rare events, and a variety of methods to make the model robust and to avoid overfitting. Our moderation system is trained to detect a broad set of categories of undesired content, including sexual content, hateful content, violence, self-harm, and harassment. This approach generalizes to a wide range of different content taxonomies and can be used to create high-quality content classifiers that outperform off-the-shelf models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AI Native Games: A Survey and Roadmap

    cs.AI 2026-07 unverdicted novelty 6.0 of 10

    The paper proposes a counterfactual definition of AI-native games, screens 53 examples, introduces a G/N taxonomy, and outlines a research roadmap for the field.

  2. Predict, Don't React: Value-Based Safety Forecasting for LLM Streaming

    cs.CL 2026-04 conditional novelty 6.0 of 10

    Forecasting expected future harmfulness from prefixes via Monte Carlo rollouts yields stronger streaming LLM moderation than boundary detection, without exact onset labels.

  3. Learning from Mistakes: Can LLM Self-Recover after Misalignment?

    cs.CY 2026-03 conditional novelty 6.0 of 10

    LLMs sometimes regain safe behavior after multi-turn jailbreaks, and this recovery can be measured with turn-level safety trajectories and metrics such as misalignment length and recovery duration.

  4. Guard Vector: Beyond English LLM Guardrails with Task-Vector Composition and Streaming-Aware Prefix SFT

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A task vector from an English guard model transfers safety classification to Korean, Chinese, and Japanese models, and a prefix-SFT variant maintains accuracy under streaming with a single-token classifier.

  5. SEALGuard: Safeguarding the Multilingual Conversations in Southeast Asian Languages for LLM Software Systems

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A LoRA-adapted SeaLLM model detects unsafe and jailbreak prompts in nine Southeast Asian languages with 97% recall and 98% F1 on the authors' new SEALSBench benchmark, far above zero-shot LlamaGuard.

  6. GuardReasoner-Omni: A Reasoning-based Multi-modal Guardrail for Text, Image, Video, and Audio

    cs.CR 2026-02 reject novelty 4.0 of 10

    A reasoning-based guardrail model trained on 148k text/image/video samples is claimed to beat prior content-safety moderators, though the abstract and body conflict on modalities and model sizes.

Pith tools