Pith. sign in

REVIEW 3 cited by

Content Moderation by LLM: From Accuracy to Legitimacy

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.03219 v2 pith:PAK2ICMI submitted 2024-09-05 cs.CY cs.AIcs.ETcs.HCcs.LG

classification cs.CYcs.AIcs.ETcs.HCcs.LG
keywords moderationaccuracycasescontentllmsarticleapplicationdecisions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

One trending application of LLM (large language model) is to use it for content moderation in online platforms. Most current studies on this application have focused on the metric of accuracy -- the extent to which LLMs make correct decisions about content. This article argues that accuracy is insufficient and misleading because it fails to grasp the distinction between easy cases and hard cases, as well as the inevitable trade-offs in achieving higher accuracy. Closer examination reveals that content moderation is a constitutive part of platform governance, the key of which is to gain and enhance legitimacy. Instead of making moderation decisions correct, the chief goal of LLMs is to make them legitimate. In this regard, this article proposes a paradigm shift from the single benchmark of accuracy towards a legitimacy-based framework for evaluating the performance of LLM moderators. The framework suggests that for easy cases, the key is to ensure accuracy, speed, and transparency, while for hard cases, what matters is reasoned justification and user participation. Examined under this framework, LLMs' real potential in moderation is not accuracy improvement. Rather, LLMs can better contribute in four other aspects: to conduct screening of hard cases from easy cases, to provide quality explanations for moderation decisions, to assist human reviewers in getting more contextual information, and to facilitate user participation in a more interactive way. To realize these contributions, this article proposes a workflow for incorporating LLMs into the content moderation system. Using normative theories from law and social sciences to critically assess the new technological application, this article seeks to redefine LLMs' role in content moderation and redirect relevant research in this field.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows

    cs.CR 2026-07 conditional novelty 5.0 of 10

    A survey of 49 LLM fraud and trust-and-safety papers finds that fraud work reports almost no per-decision latency, cost, or calibration evidence, while moderation work reports more.

  2. Socio-Culturally Aware Evaluation Framework for LLM-Based Content Moderation

    cs.CL 2024-12 reject novelty 5.0 of 10

    A persona-based generation pipeline creates culturally varied content moderation test sets, but its central 'greater challenge' claim depends on unvalidated synthetic labels and unreleased data.

  3. AI and the Future of Digital Public Squares

    cs.CY 2024-12 unverdicted novelty 3.0 of 10

    A multi-stakeholder agenda argues that LLM-enabled collective dialogue, bridging, moderation, and proof-of-humanity tools can strengthen digital public squares if paired with research and safeguards.

Pith tools