Pith. sign in

REVIEW 1 cited by

Directions in Abusive Language Training Data: Garbage In, Garbage Out

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2004.01670 v3 pith:EIDBUXB3 submitted 2020-04-03 cs.CL

classification cs.CL
keywords abusivedatalanguagecontentgarbageanalysiscataloguingcollection
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Data-driven analysis and detection of abusive online content covers many different tasks, phenomena, contexts, and methodologies. This paper systematically reviews abusive language dataset creation and content in conjunction with an open website for cataloguing abusive language data. This collection of knowledge leads to a synthesis providing evidence-based recommendations for practitioners working with this complex and highly diverse data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CODEOFCONDUCT at Multilingual Counterspeech Generation: A Context-Aware Model for Robust Counterspeech Generation in Low-Resource Languages

    cs.CL 2025-01 reject novelty 4.0 of 10

    A simulated-annealing pipeline that generates and ranks counterspeech candidates with a language-model judge placed first for Basque and in the top three for English, Italian, and Spanish in the MCG-COLING-2025 shared task.

Pith tools